LoongCollector Ops
Turn natural-language requests into executable, verifiable, rollbackable workflows for users operating their own LoongCollector and SLS resources — from install through collection to query.
Architecture: Install (ECS/self-host/ACK/self-k8s) + SLS Project + Logstore + Index + MachineGroup + Pipeline (API or ClusterAliyunPipelineConfig) + binding + SLS Lens
Scope. Covers:
- Install/upgrade Linux collector: ECS
aliyun ecs run-command, self-host SSH, ACK addonloongcollector, self-k8s custom package. Must continue to collection + query; process/Addon ready is only a stage gate. - ACK first-use:
open-ack-service --type propayasgo+ CS service roles (scripts/ensure_ack_prereq.sh).create-clusteronly when the user asked to create a cluster. Eval hooks may pre-create a fixture cluster; that is not a production default. - Cloud onboarding: Project / Logstore / Index / MachineGroup / Pipeline Config / binding.
- K8s collection: default SLS Pipeline API (
create-logtail-pipeline-config+ bind official group). CRD apply is opt-in only when the user asks for GitOps/CRD and a reachable kube-apiserver exists (references/crd-pipeline.md). - Config management: create, modify, apply, remove, data acceptance (U1-U6).
- Machine group management: IP / user-defined identity, members, heartbeat, version.
- SLS Lens: run-log query (
get-logs-v2), topic/field contracts, version routing, degradation. - Basic troubleshooting: no-data, heartbeat abnormal.
Language and HITL Delivery Contract
Hard language rule: use the user's primary language for every user-facing message. This includes plans, clarification questions, confirmation questions and their answer options, reports, and error guidance. Product names, identifiers, CLI commands, JSON fields, error codes, and fixed status tags may remain in their original form. Never switch the surrounding prose to another language.
Canonical user-facing message and marker catalog: every value below is literal. Emit the selected value verbatim; never translate, paraphrase, or combine it with another question.
messages:
missing_task_scope: "请补充要执行的具体操作目标、地域和 SLS Project。"
missing_lens_parameters: "请补充业务 Project、地域和查询时间范围。"
machine_group_identity: "请选择机器组标识类型:IP 或 userdefined。"
r2_update: "是否确认执行上述变更计划?请选择:确认执行或取消。"
r2_create_bind: "是否确认创建上述资源并完成绑定?请选择:确认执行或取消。"
r3_unbind: "是否确认将上述旧配置从机器组解绑?请选择:确认解绑或取消。"
permission_recovery: "是否已完成所需 RAM 授权并允许重试?请选择:已授权或未授权。"
permission_recovery_short: "是否已完成所需 RAM 授权并允许重试?"
lens_entry: "请提供 SLS Lens 服务日志的 Project 和 Logstore。"
ecs_install: "是否确认在上述 ECS 上安装 LoongCollector?"
self_host_install: "是否确认在上述主机上安装 LoongCollector?"
ack_install: "是否确认在上述 ACK 集群安装 loongcollector 组件?"
self_k8s_install: "是否确认在上述 Kubernetes 集群安装 LoongCollector?"
kubeconfig: "请提供可用的 kubectl 与目标集群 context。"
ssh: "请提供已配置的 SSH(alias 或主机),不要在对话中发送私钥。"
collector_version: "请提供采集器版本(例如 3.3.9)。"
markers:
ownership_error: ["不属于当前账号", "项目不属于你"]
permission_decline: ["未授权", "停止"]
permission_grant: ["已授权"]
cancel: ["取消"]
private_ip: ["私网 IP"]
routing_intent: ["日志采集", "安装采集器"]
end: ["结束"]
collector_deployed_without_version: ["LoongCollector 已部署"]
binding_acceptance: ["完成绑定与验收"]
approval: ["确认", "确认执行", "确认解绑"]
install_intent: ["允许安装", "请安装", "直接执行", "已授权操作", "任务已预授权"]
install_only_status: ["仅安装完成、采集未接入"]
deferral: ["还没想好", "等会儿再说", "暂不确认", "第二次等待", "第N次暂不确认", "先放一放", "已达到上限", "请阻塞"]
data_incomplete: ["无法完成数据面验收"]
data_empty: ["无数据"]
reason: ["原因"]
forbidden_empty_success: ["采集成功", "所有验收标准均已满足", "全链路验收通过", "通过"]
data_arrived: ["数据到达"]
root_cause_located: ["根因已定位"]
not_exists: ["不存在"]
pending_read: ["未执行待办", "待办"]
Pair machine_group_identity with [AWAITING: MACHINE_GROUP_TYPE] (never R2_CONFIRMATION). Pair permission_recovery with [AWAITING: PERMISSION_CONFIRMATION]; lens_entry with [AWAITING: LENS_ENTRY]; every install message with [AWAITING: INSTALL_CONFIRMATION]; kubeconfig with [AWAITING: KUBECONFIG]; ssh with [AWAITING: SSH]; and collector_version with [AWAITING: COLLECTOR_VERSION]. The self_host_install message is allowed only after a real SSH probe succeeds.
Do not replace these with long English prose, bilingual tables, or newly invented status labels. Whenever you re-ask, reproduce the same short Chinese question verbatim before the required [AWAITING: ...] tag. Last-line hard rule: the matching tag immediately follows the question on the next line and is the last line of the turn — no blank line between question and tag, no blank line after it, no punctuation, and no extra sentence. The turn that emits a HITL tag must not copy any [AWAITING: ...] literal into a tool call, code block, outputs/*, or ran_scripts/*; duplicate tags break automatic matching. Install confirmation ends the turn with [AWAITING: INSTALL_CONFIRMATION] — never reuse R2_CONFIRMATION for install or for machine-group identity. Collection/create-bind confirmation tags MUST include the ask counter on the last line: first ask [AWAITING: R2_CONFIRMATION] ask=1; each deferral re-ask increments the counter. Lens-entry fallback ends the turn with [AWAITING: LENS_ENTRY]. Missing kubectl ends the turn with [AWAITING: KUBECONFIG]. Missing SSH ends the turn with [AWAITING: SSH]. Missing collector version ends the turn with [AWAITING: COLLECTOR_VERSION]. Machine-group identity ends with [AWAITING: MACHINE_GROUP_TYPE]. RAM recovery ends with [AWAITING: PERMISSION_CONFIRMATION].
Fixed English tokens (must appear verbatim; surrounding prose stays Chinese): [BLOCKED: …] / [CANCELLED: …] / [AWAITING: …] / ask=1 / ask=2 / ask=3 / [Error: permission|throttling|internal|parameter] / [RECOVERED: …] / resource_status: Resource not found / [Query: Incomplete] / INCOMPLETE. Rejection and confirmation-timeout turns: the sole content of that turn is the short tag — no English long sentence, no prefix or suffix.
Out of scope. Windows; Sidecar; uninstall/rollback/restart-as-lifecycle; creating ECS; OOS/ChatOps; writing AliyunLogConfig / NamespaceAliyunPipelineConfig; advanced troubleshooting (delay, duplicate, parse failure, container filter, data loss/truncation). kubectl exec and docker exec are forbidden. If the user asks for an out-of-scope lifecycle action, say so and stop that branch.
1. Prerequisites
Pre-check: Aliyun CLI >= 3.3.3 required
[MUST] Verify:
aliyun version— must be >= 3.3.3 (>= 3.3.5 recommended).
- First install or major upgrade:
/bin/bash -c "$(curl -fsSL --connect-timeout 10 --max-time 120 https://aliyuncli.alicdn.com/setup.sh)"- Routine update (CLI >= 3.3.5):
aliyun upgrade.- See
references/cli-installation-guide.md.
Pre-check: SLS plugin required
[MUST]
aliyun configure set --auto-plugin-install truethenaliyun plugin install --names aliyun-cli-slsandaliyun plugin update. Collection subcommands are provided by thealiyun-cli-slsplugin (hyphenated subcommands such asaliyun sls get-logs-v2). Verify withaliyun sls --help.
Pre-check: Alibaba Cloud Credentials Required
Security Rules:
- NEVER read, echo, or print AK/SK values (e.g.,
echo $ALIBABA_CLOUD_ACCESS_KEY_IDis FORBIDDEN)- NEVER use
cat,less,head,tail,grep,open,json.load, or any file-reading command on credential files (e.g.,~/.aliyun/config.json,~/.aws/credentials). To check file existence uselsonly — never display contents. Printing plaintext secrets is an immediate task failure and security incident.- NEVER install or import
aliyun-log-python-sdk/aliyun.log/LogClient, or any other SLS SDK, to bypass CLI.pip installof a cloud SDK is a task failure.- NEVER print,
cat, or paste kubeconfig / client certificates / tokens into the conversation. For opt-in CRD only: writedescribe-cluster-user-kubeconfigoutput to a0600tempfile.- NEVER ask the user to input AK/SK directly in the conversation or command line
- NEVER use
aliyun configure setwith literal credential values- ONLY use
aliyun configure listto check credential status.scripts/preflight.shalready does this.aliyun configure listCheck the output for a valid profile (AK, STS, or OAuth identity).
If no valid profile exists, STOP here.
- Obtain credentials from Alibaba Cloud Console
- Configure credentials outside of this session (via
aliyun configurein terminal or environment variables in shell profile)- Return and re-run after
aliyun configure listshows a valid profile
Run bash scripts/preflight.sh to check CLI version, plugin, credential presence, and scope in one step. preflight.sh already invokes aliyun configure list internally; running it is a valid credential check — do not cat CLI config files, and do not add a standalone aliyun configure list just to satisfy a checklist. Full gate details: references/prerequisites.md.
Environment Variables
| Variable | Required | Description |
|---|---|---|
| (none for credentials) | — | Credentials come from aliyun configure profiles; never introduce AK/SK env vars in-session |
| SKILL_SESSION_ID | Injected at script run | Same 32-hex session id as the session/{session-id} UserAgent token; set inline when invoking bundled scripts (see §4) |
2. RAM Policy
This skill uses the user's own identity and only touches resources they are authorized for. Permissions are layered ReadOnly / Operator / Destructive. Per-workflow RAM Actions are in references/ram-policies.md — do not default to broad AliyunLogFullAccess.
[MUST] Permission Failure Handling: When any command or API call fails due to permission errors at any point during execution, follow this process:
- Read
references/ram-policies.mdto get the full list of permissions required by this SKILL- Use
ram-permission-diagnoseskill to guide the user through requesting the necessary permissions- Pause and wait until the user confirms that the required permissions have been granted
Runtime detail (same gate, do not skip the three steps above):
- Report the missing RAM Action and
requestID; output[Error: permission]. Readingreferences/ram-policies.mdalone is not a successful diagnose call. - Try
ram-permission-diagnosewith the missing Actions andrequestID. FALLBACK: if it is unavailable, output Action/requestID/RAM-console guide manually, then ask catalog messagepermission_recoverywith last line[AWAITING: PERMISSION_CONFIRMATION]and pause. - Do not retry the affected write (including
--cli-dry-run) before confirmation. - READ-PATH HARD STOP: On 401/403/
Unauthorized/AccessDeniedforget-project/get-machine-group/list-machines/get-log-store, first read the message. If it is ownership (English ownership text or any catalogownership_errormarker) → this is not a RAM gate: emit[BLOCKED: RESOURCE_RESOLUTION_FAILED]and stop; do not askpermission_recovery_short, do not create the officialk8s-log-*name, and do not retry. Otherwise emit[Error: permission]with Action/requestID, then in the same turn ask exactlypermission_recoverywith last line[AWAITING: PERMISSION_CONFIRMATION], and issue zero furtheraliyun slscalls that turn — includingget-machine-group,list-machines,get-applied-configs, andget-log-store. Those unread calls are catalogpending_readitems, not queried conclusions. - After the user's permission answer (same gate for read-path and write/dry-run):
- Any catalog
permission_declinemarker or equivalent decline → zero tools that turn (nowrite_file, noaliyun sls). Explicitly state that execution is terminated, list every not-yet-run read as catalogpending_readwith its RAM Action, then put[BLOCKED: PERMISSION_REQUIRED]on the final line. If the task asked for machine-group heartbeat, the pending items must includeget-machine-group→log:GetMachineGroupandlist-machines→log:ListMachines(name the group). Never present an unrun heartbeat as a queried conclusion. - A catalog
permission_grantmarker → same turn, retry the identical failed command (if the failure was a dry-run, retry that dry-run first) and emit[RECOVERED: permission_granted]in the user-facing text immediately. Explicitly state that the disposition is human intervention followed by retry, so the recovery action is unambiguous.
- Any catalog
On Unauthorized/AccessDenied from a core write or its dry-run: stop the current write, enter the §6 permission-recovery branch, and never switch account/profile or widen scope.
3. Parameter Confirmation
IMPORTANT: Parameter Confirmation — Before executing any command or API call, ALL user-customizable parameters (e.g., RegionId, instance names, CIDR blocks, passwords, domain names, resource specifications, etc.) MUST be confirmed with the user. Do NOT assume or use default values without explicit user approval.
| Parameter | Required/Optional | Description | Default |
|---|---|---|---|
| region | Required | SLS region id, e.g. cn-hangzhou | none |
| project | Required | Target SLS project | none |
| logstore | Conditional | Target logstore (create/onboarding/query) | none |
| machine_group | Conditional | Machine group name (bind/heartbeat) | none |
| config_name | Conditional | Logtail pipeline config name | none |
| machine_identify_type | Conditional | ip or userdefined | none |
| machine_list | Conditional | IP list or user-defined identifiers | none |
| scenario | Conditional | host / docker / k8s / host_agentsight | none |
| environment | Conditional | ecs / self_host / ack / self_k8s for install.deploy | none |
| instance_id | Conditional | ECS instance id (Workbench) | none |
| cluster_id | Conditional | ACK / self-k8s cluster id | none |
| lens_project / lens_logstore | Optional | SLS Lens entry when auto-discovery unavailable | none |
Never substitute placeholder/example values for a missing real resource. Missing region, project, or a scope-changing target → STOP and ask.
For a Chinese request with a missing machine_identify_type, ask exactly catalog message machine_group_identity using Chinese option descriptions, and end the turn with last line [AWAITING: MACHINE_GROUP_TYPE] — never [AWAITING: R2_CONFIRMATION]. After the user chooses, run scripts/normalize_diff.py and only then ask the create-and-bind question with [AWAITING: R2_CONFIRMATION] ask=1. Do not emit a custom English status label or an English selection table. install.deploy exception: do not ask machine-group identity until the collection create-bind gate has been confirmed. If that gate receives a catalog cancel marker or equivalent rejection, never ask. If the user already named IP, userdefined, or a catalog private_ip marker, never ask.
4. Observability (MUST follow for every aliyun command)
Upon loading this skill, generate a random session ID (32-char lowercase hex string) once for the entire session. Use it as {session-id} below.
Rule: Every aliyun CLI command that calls a cloud API MUST include the --user-agent flag.
This covers aliyun sls, aliyun cs, aliyun ecs, and any other aliyun <product> cloud call this skill issues, including --cli-dry-run. Local utility commands (e.g. configure, plugin, version) do not support this flag and should be excluded. kubectl / Workbench / SSH / local validators are not Alibaba Cloud APIs and do not send this flag.
Use two space-separated product tokens (quote the whole value; the space is required):
--user-agent "AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops session/{session-id}"
| Token | Example | Query use |
|---|---|---|
| Skill identity | AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops | All traffic from this skill |
| Session | session/{session-id} | One session |
Never glue the session id onto the skill token (.../ops/{session-id} is forbidden). Never omit quotes. Never skip, alter, or drop either token.
Example (assuming session-id is a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6):
aliyun sls list-machines --project my-proj --machine-group my-group --region cn-hangzhou --user-agent "AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops session/a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6"
References that write --user-agent <ua> mean this exact quoted two-token string.
Script / Terraform execution: When running Python SDK scripts or Terraform commands or bash scripts, inject the session-id via inline environment variable so the code can read it at runtime:
# Local validator (no cloud call)
SKILL_SESSION_ID={session-id} python3 scripts/validate_pipeline.py --file rendered.json
# Bundled script that itself calls a cloud API
SKILL_SESSION_ID={session-id} bash scripts/wait_cs_task.sh --cluster-id c-xxx --region cn-shanghai
# Terraform
SKILL_SESSION_ID={session-id} terraform apply
Scripts and Terraform configs should read SKILL_SESSION_ID from the environment (default to empty string if absent). Any bundled script that itself invokes aliyun against a cloud API MUST send the same two-token UserAgent (currently scripts/wait_cs_task.sh).
Domain extension — ATOMIC CLOUD-CALL RULE (HARD): Every tool invocation that calls SLS must contain exactly one direct aliyun sls ... command with literal, fully expanded parameter values, and the command must start with aliyun sls. Do not hide a cloud call behind shell variables, environment assignments, functions, aliases, wrapper scripts, loops, command substitutions, eval, pipes, redirections (including 2>&1), or compound commands (;, &&, ||). Use only lowercase hyphenated SLS plugin subcommands such as get-project; never use a PascalCase OpenAPI alias, because it bypasses the validated command/mocking contract. Compute timestamps or JSON in a separate local step, then place the resulting literals in the cloud command. This applies equally to verification and acceptance reads: no loops, no cd … prefix, no $VAR or $(…) substitution, including in --from/--to and JSON bodies. Never write cloud calls into a .sh file and run it; the command record is plain-text notes, not a runnable script. Local validators must receive what the command actually returned — save the real stdout to a file and pass that file; retyping or echo-ing an expected response is fabricated evidence. Generate the session ID with python3 -c 'import secrets; print(secrets.token_hex(16))' and validate ^[0-9a-f]{32}$; never copy the example ID above into live commands.
5. Capability Router
Classify the request into exactly one capability, then load its references/navigation.md entry before acting. Do not load the whole knowledge base into context.
| Capability | Trigger | Required inputs | Adapters | Success state |
|---|---|---|---|---|
| install.deploy | Install/upgrade then collect and query | region, environment, instance or cluster | ecs run-command / ssh / aliyun_cs / kubectl / aliyun_sls | install gate + U1-U6 + query |
| config.modify | Change an existing config / parse / fields | region, project, config | aliyun_sls, local validator | config + index + data verified |
| config.create | Base resources exist, only create a config | region, project, logstore, machine_group, scenario | aliyun_sls (default); kubectl CRD only if user asked | config exists, bound, has data |
| onboarding.cloud | Collector running, wire up cloud side (API) | region, project, logstore, machine_group, source | aliyun_sls, local validator | U1-U6 pass |
| machine_group.manage | Create/modify group, members, binding | region, project, group | aliyun_sls | object + relations match target |
| lens.query | Query collection alarms/status/metrics | business project, time range, lens entry | aliyun_sls | query complete with context |
| troubleshoot.basic | No data / heartbeat abnormal | region, project, optional logstore/config/group | aliyun_sls, Lens | root cause or single blocker |
Full router spec (when_to_use / out_of_scope / entry_signals / success|blocked|failure_state): references/navigation.md. Track multi-step work with the unified task object in references/task-model.yaml.
6. Execution State Machine
Classify → Preflight → Observe → Plan → Approve → Execute → Verify → (Rollback)
install.deploy inserts an extra install Approve/Execute before collection Observe/Plan. Do not end after the process/Addon stage gate.
-
Classify: pick capability, scenario (
host/docker/k8s/host_agentsight), environment (ecs|self_host|ack|self_k8s), management plane (api|crd). Defaultapifor host and K8s collection. Setcrdonly when the user explicitly asks forClusterAliyunPipelineConfig/ GitOps /kubectl applyCR and a reachable kube-apiserver is proven (references/crd-pipeline.md). ACK collection must not stop on[AWAITING: KUBECONFIG]. Ask for scope-changing inputs; never guess. If the user names SLS, Log Service, LoongCollector, Logtail, or any catalogrouting_intentmarker without a concrete operation, stay in this skill: clarify region, project/cluster/host, and the intended operation before any cloud call; never improvise outside scope.Agentloop / AgentSight: Agentloop, AgentSight,
input_agentsight, eBPF Runtime, orebpf-event→config.create(oronboarding.cloudif the group must be created) with scenariohost_agentsight. Loadreferences/agentsight-agentloop.mdandreferences/input-agentsight.md. Names are product-fixed (runtime-ebpf-agentsight-config→ebpf-event). Lock forbids overwrite, not Plan: even whenget-logtail-pipeline-configalready returns the object, still runscripts/render_pipeline.py+scripts/validate_pipeline.py+scripts/normalize_diff.pyand ask catalog messager2_create_bind. After confirm, still issuecreate-logtail-pipeline-config(andapply-config-to-machine-groupif unbound).AlreadyExist/ already-bound → record the lock[Idempotent-Skip]and do not update. Host Linux, kernel>=5.10, collector>=3.3.9. Not OBI/OTLP.INTENT / PARAM CLARIFICATION STOP (hard): Ask at most one clarifying question, in Chinese, using the applicable exact catalog message:
missing_task_scopefor general scope ormissing_lens_parametersfor Lens. If the user still gives no concrete values — undecided, wants only the checklist, or asks you not to run commands — do not ask again. In that same turn output a minimal declarative checklist (region, Project/resource locator, operation goal; for Lens-only asks: business project or Lens entry + time range), put[BLOCKED: MISSING_REQUIRED_INPUT]on the final line, and end the turn — no further questions, question marks, invitations to provide data, cloud calls, Preflight, or Observe. In particular, do not repeat either fixed clarification subject after the checklist. Resume only when the user supplies concrete values.Project locator: exact user prefix/handoff only — at most one
list-project --project-name <prefix>, thenget-projecton the resolved full name; alist-projecthit alone never proves the target nor authorizes any next resource read. Theget-projectcall is mandatory for read-only config views and Idempotent-Skip checks too: never jump directly from EVAL_ACCOUNT_ID/name resolution toget-log-storeorget-logtail-pipeline-config. Never broaden/synthesize names. Zero matches →[BLOCKED: RESOURCE_RESOLUTION_FAILED] …; multiple → ask user to choose. -
Preflight:
scripts/preflight.sh(CLI/plugin/credential/scope). Add--need-ecs(ECS) /--need-cs(ACK) /--need-kubectl(self_k8s install only, or opt-in CRD). Hard-gate only (CLI / SLS plugin / credentials / ACK--need-cs): output[BLOCKED: PREFLIGHT_FAILED] gate=<gate>; <reason>and stop. Adapters are not hard gates:- ECS:
run-commandis the only host channel.--need-ecsis a warn, not a hard fail. Still runscripts/render_loongcollector_install_cmd.py, and still ask catalog messageecs_installwith last line[AWAITING: INSTALL_CONFIRMATION]. Never emit[BLOCKED: PREFLIGHT_FAILED] gate=workbench/gate=ecs. After confirm, use exactlyaliyun ecs run-command --biz-region-id <region> ... --instance-id <id>and poll withaliyun ecs describe-invocation-results --biz-region-id <region> --invoke-id <id>; never use--instance-id.1,--region-id, or--regionfor these plugin commands. Do not useworkbench exec/aliyun ecs-workbench/ OOS. ECS post-install fast path: After a successful invocation result, reuse every user-supplied Project, Logstore, machine-group, and config name verbatim. Do not read helper source, call--helpfor mapped commands, reinstall/update plugins, re-run Preflight, or writeoutputs/*/ran_scripts/*before collection approval. Run only the required Observe reads, render/validate/diff calls, then immediately emit the collection approval question. If invocation status is still running, wait and poll as separate tool calls; never combinesleepand the poll command. - self_k8s kubectl missing: this is the first stop — before Observe, Plan, values render, or any create-bind question. Ask exactly catalog message
kubeconfig+[AWAITING: KUBECONFIG]and end the turn. Never[BLOCKED: PREFLIGHT_FAILED] gate=kubectl. Cloud writes /create-logtail-pipeline-configare forbidden until kubectl is provided. A refusal or catalogendmarker → stop; do not continue to R2 create-bind. - self_host SSH: before any
create-*/apply-config-*/ install /INSTALL_CONFIRMATION, run exactly one direct probe:ssh -o BatchMode=yes -o ConnectTimeout=8 <alias> -- true. Use the user's existing SSH configuration; never addStrictHostKeyChecking=no,UserKnownHostsFile=/dev/null, or any option that weakens host-key verification. Do not wrap the probe in a pipe,tee, redirection, compound command, or logging helper that can mask its exit code. If the alias is missing, unresolvable, or the probe fails → immediately ask exactly catalog messagesshwith last line[AWAITING: SSH]and stop; after the failed probe, do not callwrite_file, update a plan, or narrate a report before emitting that two-line gate. Do not ask the install-confirmation sentence on a failed probe. Do not create/edit~/.ssh/config,/etc/hosts,authorized_keys, or rewrite the alias to127.0.0.1/ localhost to fake a working channel. Zero cloud writes that turn. Prompt-supplied alias that does not work is still “no usable SSH”. A catalogendmarker → output only[CANCELLED: SSH_REQUIRED]; do not call tools or continue to install. - ACK collection uses SLS API and must not ask for kubeconfig /
--need-kubectl. - ACK first-use:
--need-cshard-fails only on a missing CS plugin. Unopened ACK (ErrorNotEnabled/cskpro) or missingAliyunCSDefaultRoleis not[BLOCKED: PREFLIGHT_FAILED]. After install confirm, runbash scripts/ensure_ack_prereq.sh --region <r>, then retry the failed CS write once.create-clusteronly if the user asked:--biz-profile Default(never--profile), specack.standardthenack.pro.small. Do not inventsls-eval-loop-ackin production.
- ECS:
-
Observe (read-only): Get current objects + bindings + heartbeat; save a snapshot. Read collector version (
list-machines.binary) before choosing plugins. If the user message already states a version (e.g.3.2.6/3.3.9/LoongCollector 3.2.6), use that version and do not ask. If version is still unknown: ask exactly catalog messagecollector_version+[AWAITING: COLLECTOR_VERSION]and stop. Do not assume 3.x, do not silently pickprocessor_json, and do not ask Lens only to learn the version. A catalogcollector_deployed_without_versionmarker without a version string is still unknown. MANDATORY VERIFICATION COMMANDS: existence is proven only byget-project,get-machine-group, andget-log-store(the last before any create/bind on that logstore). Observe must start withget-project --project <full-name>(orlist-projectthenget-projecton the resolved name). Forconfig.create/ create-and-bind, checkget-machine-group(orlist-machines) immediately after Project resolution and beforeget-log-store, so the machine-group precheck is unambiguously recorded before config planning. Concatenating a prefix withEVAL_ACCOUNT_IDis not existence proof and does not authorize skippingget-projectto jump toget-log-store. Theget-projectresult is the only runtime project name for later calls.list-machines/get-applied-configs/list-log-storesprove heartbeat or binding, never existence. Enter create only after aget-*returns ResourceNotExist. Even onProjectNotExist, still issue the remaining independent gets once each (includingget-applied-configs). -
Plan: build target objects. A from-zero
onboarding.cloudplan includes Project, Logstore, Index, MachineGroup, PipelineConfig, and binding unless the user explicitly opts out of indexing; never infer that Index is optional merely because the prompt contains a catalogbinding_acceptancemarker. MANDATORY CHECKPOINT: every planned R2/R3/R4 resource or relation change (including Project, Logstore, Index, MachineGroup, binding, and unbinding) MUST have a target JSON and an executedscripts/normalize_diff.pyresult; use--kind autofor non-config/index objects. Never substitute rawdiff, visual inspection, or a handwritten diff. Exit code3means "valid diff contains changes", not failure. Config/index coupling uses this fixed order: snapshot config and index → validate the full target config withscripts/validate_pipeline.py→ runscripts/normalize_diff.py --kind config→ runscripts/normalize_diff.py --kind index. Exit code1from validation blocks the write. Do not enter Approve until every applicable mandatory script has executed successfully. Include impact, risk, rollback, and verification.mode=planMUST NOT call write commands. SCRIPT EVIDENCE MUST BE VISIBLE: invoke each required Skill script as its own shell/tool call and keep its JSON/status on that call's stdout. It is fine to useteeto persist the same stdout, but do not redirect all script output only intooutputs/*/ran_scripts/*and leave the tool result with just an exit code. A later file read or a handwritten execution record does not replace visible script-call evidence. VALIDATION FAILURE HARD STOP: ifscripts/validate_pipeline.pyreturns exit code1orstatus=invalid, immediately output[BLOCKED: VALIDATION_FAILED]and end the turn. Do not run--cli-dry-run; do not execute any create/update/apply/remove/delete command; do not treat a server-side 4xx from an actual write as validation evidence. This rule overrides user approval and every later Execute step. -
Approve: HARD GATE. For every R2 operation (create resource, update config, apply/bind, create/update index) you MUST, before Execute, explicitly output the normalized diff, ask the user to confirm, and end the turn. Ask at most one confirmation question per turn. For a Chinese request it must be the applicable exact subject from the Language and HITL Delivery Contract, with the catalog
approvalandcanceloptions, never English ones. Only an explicit positive answer authorizes a write. R3: explicit impact confirmation. R4: second confirmation, restate resources.SEPARATE INSTALL THEN COLLECTION GATES: For
install.deploy, first ask only the environment-specific install catalog message and end the turn with[AWAITING: INSTALL_CONFIRMATION]— never[AWAITING: R2_CONFIRMATION]. Cataloginstall_intentmarkers in the original request describe intent but are not the separate confirmation reply; only a fresh user message received after the install question authorizesrun-command/ SSH install / addon install. After install succeeds, if a new CR or API config/binding is required, ask in a new turn catalog messager2_create_bindwith last line[AWAITING: R2_CONFIRMATION] ask=1before those writes. Do not ask machine-group identity or collector version between the two gates. On a catalogcancelmarker or equivalent rejection:[CANCELLED: R2_CONFIRMATION_REJECTED], report cataloginstall_only_status, and issue zero collection writes. If Observe already matches the target (get proves logstore/group/config/binding), emit[Idempotent-Skip]and do not re-issue create/apply. Pure reuse of ACK default collection is the same skip.SECOND-GATE TAG UNIQUENESS: In the post-install turn that asks for collection approval, the complete
[AWAITING: R2_CONFIRMATION] ask=1token may appear exactly once: as the final response line. Before that final response, refer to the pending step only as "collection approval"; never put the token in a Plan/Todo item, tool description, tool argument, code block, output file, execution record, or narration. Do not write an execution record or update a plan after the final validation/diff call; emit the fixed question and tag immediately.SEPARATE UNBIND GATE: Create-and-bind and unbind are two independent confirmations. First ask only catalog message
r2_create_bind. After the user confirms and those writes (or exact Idempotent-Skip) finish, ask in a new turn catalog messager3_unbind. Never merge unbind into the create-and-bind question, and never runremove-config-from-machine-group(including--cli-dry-run) on the create-and-bind approval.DO NOT SKIP CONFIRMATION: Automation, urgency, complete parameters, and the original task wording never waive this gate. While the answer is outstanding, emit the Chinese question and
[AWAITING: R2_CONFIRMATION] ask=1on the first ask, end the turn, and wait for the user's next message. Same-turn writes are a gate failure: after the question, do not--cli-dry-run,create-*,apply-config-*, orupdate-*until the next user message contains an explicit catalogapprovalmarker.HARD GATE CHECKLIST (Approve → Execute): 0. Ask only once the plan is real:
scripts/validate_pipeline.pyhas passed on any config payload andscripts/normalize_diff.pyhas run for every planned write. Asking approval for a plan you have not validated and diffed is a gate failure.- Nothing you produce yourself is an answer. If the turn ends without the user having stated a decision, output
[AWAITING: R2_CONFIRMATION] ask=<n>with the identical question and wait. - Never treat the original task wording, “the task explicitly requires”, “parameters are complete”, or an already-rendered plan table as approval.
- Enter Execute only after the user explicitly answers yes/confirm/approve or an equivalent catalog
approvalmarker. --cli-dry-runfor any R2/R3 write is part of Execute: it is forbidden before that explicit approval. Showing a plan without asking, then dry-running, is a gate failure.- Violating this gate is a task failure.
NON-ANSWER RULE (hard stop): A blank reply, "later", "not sure yet", any catalog
deferralmarker, or any equivalent deferral is not approval. Maintain ask counternstarting at1on the first confirmation turn. After each deferral, restate in one line which resources and operation are still waiting, re-ask the identical Chinese confirmation subject exactly once, and put[AWAITING: R2_CONFIRMATION] ask=<n+1>on the last line (counter goes 1 → 2 → 3). Do not print the question or the AWAITING tag twice. Do not add a blank line after the tag. A bare repeated question withoutask=<n>, or a "take your time" soft-close with no question, both fail this rule. After the user replies to the third ask with another catalogdeferralmarker, the next turn's sole content is[BLOCKED: R2_CONFIRMATION_TIMEOUT]— do not ask a fourth time, do not write, do not dry-run. Explicit reject/cancel → sole content[CANCELLED: R2_CONFIRMATION_REJECTED]. Do not append English prose such asUser rejected the proposed plan…. Full semantics:references/risk-and-approval.md.TERMINAL-STATE HARD STOP: After any
[BLOCKED: …]/[CANCELLED: …]tag, run no further tools (includingwrite_file), dry-runs, writes, or Verify. The collection-gate rejection after a successful install is the one reporting exception: output[CANCELLED: R2_CONFIRMATION_REJECTED]and exact catalog statusinstall_only_status, then end the turn with zero tools. A permission rejection may also list the required not-yet-run reads from §2, but must explicitly state that execution has terminated. All other terminal turns contain only the short tag. Resume only with a fresh Plan + confirmation after the user re-opens the work. - Nothing you produce yourself is an answer. If the turn ends without the user having stated a decision, output
-
Execute: approved commands only, with
--user-agentand §4 atomic rule. Always run--cli-dry-runas its own call before the real write, so a rejected request surfaces before anything mutates state. Idempotency: get before create/apply; if state matches, skip both dry-run and write, verify via get/list, and emit exact[Idempotent-Skip] <create/apply-command> skipped; verified via <get/list-command> that state matches expectation.inChanges. If get already proves target Project/Logstore shard/TTL, anycreate-project/create-log-store(incl. dry-run) is a task failure. OnAlreadyExist, get+compare → matching Idempotent-Skip, or[BLOCKED: EXISTING_RESOURCE_CONFLICT]if mismatched.config.modifyexception: an explicit user request to change a config/index (rename field, change parse, sync index) must still runscripts/normalize_diff.py, ask catalog messager2_update, then after confirm issueupdate-logtail-pipeline-configand the coupledupdate-index(each with its own--cli-dry-runfirst). Do not skip those two writes just because the snapshot already matches — overwrite is the requested change. Create/apply Idempotent-Skip still applies. AgentSight lock: still Plan + HITL, thencreate-logtail-pipeline-config;AlreadyExist→[Idempotent-Skip], do not update. Logstore Idempotent-Skip output: when shardCount/TTL already match, the user-facing final answer must include the skipped command name, for example[Idempotent-Skip] create-log-store skipped; verified via get-log-store that state matches expectation.Error recovery (≤3 retries / 4 total; keeperrorCode+requestID): Preferpython3 scripts/classify_sls_error.pyand emit itserror_tagbefore narration. Mapping: 400/ParameterInvalid→[Error: parameter]then fix+retry same API →[RECOVERED: parameter_fixed]; 429/WriteQuotaExceed→[Error: throttling]+ backoff →[RECOVERED: throttling_retry]; 500→[Error: internal]+ same-command retry →[RECOVERED: internal_retry]; 401/403 whose message is ownership (does not belong to you) →[BLOCKED: RESOURCE_RESOLUTION_FAILED](not a RAM HITL); other 401/403→[Error: permission]then §2 Permission Failure Handling (ask catalog messagepermission_recovery) → user-facing[RECOVERED: permission_granted]on catalogpermission_grant(retry the identical command the same turn) or sole-content[BLOCKED: PERMISSION_REQUIRED]on catalogpermission_decline. Dry-run failures use the same branch; real write only after dry-run succeeds. -
When
get-logs-v2returnsmeta.progress=Incomplete, that is an incomplete query (not a transport success to treat as final): output[Query: Incomplete] attempt=<n>/4, then retry the identical request per §9. Whenmeta.progress=Complete, reportComplete— do not emitINCOMPLETEor fabricate Incomplete retries. -
Verify: acceptance evidence must be re-read after Execute — every created or changed object through its own
get-*(get-log-store,get-machine-group,get-logtail-pipeline-config,get-index), both binding directions (get-applied-configsandget-applied-machine-groups), andlist-machinesfor heartbeat. Do not declare Verify complete or start writing the final report until the applicableget-index, both binding reads,list-machines, and businessget-logs-v2calls have actually run. In particular, a newly created machine group requires a post-createget-machine-group; its Observe-phase 404 never counts as acceptance. Observe-phase reads describe the old state and never count as acceptance. Bounded polling (15s × up to 4) against U1-U6. After every config create/update/bind or onboarding flow, must execute a business-logstoreget-logs-v2acceptance query before the final answer — skipping it because “config already verified”, “no collector heartbeat”, or “data cannot arrive yet” is a task failure. Report observed count/progress; if no rows arrive (count=0/data=[]), you must say catalogdata_incompleteanddata_emptyplusreason. Forbidden on empty U5: every catalogforbidden_empty_successmarker. Never equate zero rows with successful delivery. The last user-facing answer (including after a later unbind) must contain catalogdata_arrivedordata_empty, plusreason. The user-facing final summary after onboarding/Verify must list the Project full name, Logstore, machine group, and config name (not only the three resource names). -
Rollback: only from the pre-execution snapshot or declared inverse; never rebuild config from memory.
Input contract and stop conditions: references/prerequisites.md.
Scope & adapter rules
- Cloud-only capabilities execute through
aliyun sls+ local validators. Existing cloud-only evals still forbid SSH/kubectl. install.deploymay usealiyun ecs run-command+aliyun ecs describe-invocation-results(ECS), user SSH (fingerprint confirmed, no private key in chat),aliyun cs(addon nameloongcollector), andkubectlfor self_k8s package install plus opt-in CR get/apply. Noworkbench exec/ OOS /kubectl exec/docker exec/ unbounded root shell. Never print kubeconfig or client certs; write a 0600 tempfile only.- After install, host path must call
onboarding.cloud. K8s path must reuse official Project/machine group, create the pipeline via SLS API by default, then U1–U6 andget-logs-v2. CRD apply only on explicit user request (references/crd-pipeline.md). - Fixed within one request: account/profile, region, project, cluster, group, config. Switching scope requires a fresh confirmation.
- One management plane per config: API or CRD, never both. If double-write is detected, STOP (
references/pipeline-config.md,references/crd-pipeline.md). A CR-owned config must not be updated via API. Defaulting to API does not authorize creating a second plane.
7. Risk, Approval & Rollback
| Level | Examples | Approval |
|---|---|---|
| R0 read | list/get, status, log query | auto |
| R1 local | schema validate, render, diff | auto |
| R2 reversible write | create resource, update config, apply/bind, create/update index, install collector, kubectl apply CR | show diff + one confirmation |
| R3 high impact | remove/unbind, bulk changes | confirm after stating impact |
| R4 destructive | delete project/logstore/config/group | second confirmation, restate resources |
R2 confirmation is a hard gate. Only an explicit affirmative answer authorizes execution; task wording is never implicit approval. Create-and-bind and unbind are two separate questions (see §6 SEPARATE UNBIND GATE). Rejection and timeout use the exact terminal statuses in §6, and rollback is a new approval workflow. The §6 TERMINAL-STATE HARD STOP covers
[BLOCKED: PERMISSION_REQUIRED]and every other terminal tag: once emitted, no further cloud write (including dry-run) until a fresh Plan + approval cycle completes.
Details, snapshot format, and rollback: references/risk-and-approval.md. Get before Update (Update is overwrite semantics — carry unchanged fields back). Get before Create (see §6 Execute — idempotent handling of AlreadyExist). New resources' rollback does not default to deletion.
8. Config / Index Coupling (hard rule)
When processors add, remove, or rename a field, you MUST: 1) generate the config diff and the corresponding index diff together; 2) present both diffs to the user in one combined approval and request a single confirmation; 3) after approval, run both direct dry-run calls first (config dry-run → index dry-run), then execute the two actual writes as two consecutive standalone invocations (config write, then immediately index write) with no other command, wait, status check, or pause between them. "Back-to-back" means two separate calls in a row — never join them with &&, ;, or any other shell chaining; §4 forbids that and a chained failure leaves the pair half-applied. Never "update config now, add index later", and never split the two changes into two confirmations. If a dry-run fails, issue neither actual write. If the index write fails after the config write succeeded, the batch is incomplete: immediately re-issue the identical index command through the §6 error-recovery branch in the same turn, keep the original errorCode/requestID, and report both outcomes together — do not redesign the payload by guessing, and do not end the turn with the config updated and the index pending. If the backend auto-syncs the index, say so in the plan and in the same confirmation.
status/status_code/http_statusdefault tolong; time/latency/bytes chosen by semantics.- SLS has no
floatfield-index type. Map a floating-point requirement (for examplerequest_time float) todoubleand show that mapping in the index diff. - JSON nested fields: declare parent
jsonindex or the flattened field index explicitly. - Prefix change must update index field names and the acceptance query together.
processor_renameis extended-only; JSON rename uses all-extendedprocessor_json→processor_renamewith pluralSourceKeys/DestKeys(neverSourceKey/DestKey/processor_rename_native). Index diff must drop old keys and add new ones;http_statusdefaults tolong. Full examples:references/index-coupling.md.- Use this exact Plan order: config/index snapshots →
scripts/validate_pipeline.pyon the full target config →scripts/normalize_diff.py --kind config→scripts/normalize_diff.py --kind index; rawdiffis not acceptable.
MANDATORY COMMAND MAPPING: use Pipeline names only (*-logtail-pipeline-config, get-applied-configs / get-applied-machine-groups, apply-config-to-machine-group / remove-config-from-machine-group). Create flags: --project-name / --logstore-name / --group-name; read/bind: --project / --logstore / --machine-group. update-log-store needs both --logstore and --logstore-name. Never classic get-config/get-logs/list-logstore aliases; classic inputDetail configs → [BLOCKED: CLASSIC_CONFIG_UNSUPPORTED]. Commands in this mapping are already verified, so do not spend a call on --help for them. Index: existing→update-index, missing→create-index; --line uses "chn": true; a text key's token is a JSON array of delimiter strings — a concatenated string returns IndexInfoInvalid: field token is of error format. On IndexInfoInvalid / token-format errors, repair only via aliyun sls (rewrite --keys as a JSON array of delimiter strings, or pass keys from a task-workspace JSON file — never ~/.aliyun/config.json). Do not install or call aliyun-log-python-sdk. If CLI-side repair still fails after the allowed retries, emit [BLOCKED: PARAMETER_UNRESOLVED] and stop. Full contracts: references/cli-contracts.yaml, references/index-coupling.md, references/field-conventions.md.
First-create flag invariant: the first dry-run and real invocation of create-project must use --project-name; never probe it with --project. Use --logstore-name for create-log-store and --group-name for create-machine-group. A failed exploratory write invocation can poison ordered evaluation even when a later retry succeeds.
9. SLS Lens (CloudLens for SLS) run logs
Lens is the primary data source for basic troubleshooting. Run-log query itself uses public aliyun sls get-logs-v2.
Entry discovery. Verify the business project with get-project first — a missing business project explains an empty query and must not be reported as a Lens failure. Then run get-logging once on it (even if missing) and python3 scripts/parse_lens_logging.py. State machine: (1) usable entry (loggingProject / parsed lens project+logstore) → verify+use and do not ask [AWAITING: LENS_ENTRY]; (2) else user-supplied lens_project/lens_logstore → verify+use; (3) else, after every independent non-Lens read has already been issued once (including get-applied-configs even on ProjectNotExist), ask once exactly catalog message lens_entry and emit [AWAITING: LENS_ENTRY] as the last line — that turn's user-facing content is only that question plus the tag: no conclusion, no evidence dump, and no catalog root_cause_located marker. End the turn and wait. Never reuse the business Project or invent internal-diagnostic_log on it; after the user answers, must run planned logtail_alarm + version-routed metric queries, then write the conclusion; (4) if user cannot provide → continue independent MG/config/business checks with resource_status: Resource not found, never fabricate “no alarms”. After the final Lens/troubleshoot report, stop — do not ask Lens location again and do not wait for a catalog end marker.
Forbidden: STAROps/JWT/console cookie; guessing log-service-{uid}-{region}; Lens topics on the business Project after discovery failed; skipping the ask-once HITL. [AWAITING: LENS_ENTRY] is only for troubleshoot.basic / lens.query. Missing collector version on onboarding/config uses [AWAITING: COLLECTOR_VERSION], never Lens.
Query hard constraints (full: references/sls-lens-contracts.md, references/monitoring-queries.yaml): no select *; version route >=3→loongcollector_metric, <3→logtail_profile/logtail_metric (must execute that get-logs-v2 even on failure); logtail_alarm filters by project (never config_name in where); Incomplete → [Query: Incomplete] attempt=<n>/4 + identical retry ≤4, then INCOMPLETE if still incomplete after 4; Complete → report Complete and never fabricate INCOMPLETE; report Lens project/logstore/topic/window/route/completeness.
Mandatory version-route call: for every known collector version >=3, including 3.2.6, execute a distinct aliyun sls get-logs-v2 query containing __topic__: loongcollector_metric. A logtail_status query is additional evidence and does not replace the metric query. Execute this call even when an earlier Lens query returns no rows or an error.
10. Troubleshooting (no-data, heartbeat)
Fixed chain: classify → heartbeat & version → config & binding → alarm/metric → business data → minimal fix → re-verify.
- HARD RULE — the diagnostic chain never stops early. If a read-only step returns
ResourceNotExist/ProjectNotExist/LogStoreNotExist/LoggingNotExist, execute every remaining independent diagnostic step.get-applied-configsis independent of Project existence — issue it once even afterget-project/get-machine-groupalready 404; do not substituteget-logtail-pipeline-configfor binding evidence. Businessget-logs-v2is also independent — issue it once against the named business logstore even when the Project is missing (expectProjectNotExist/LogStoreNotExist). The Lens-entry ask turn is only catalog messagelens_entryplus[AWAITING: LENS_ENTRY]; no report in that turn. After the final report, do not ask a follow-up. For Lens-dependent steps, follow §9 (ask once +[AWAITING: LENS_ENTRY]→ wait → then query) after any entry-discovery failure —LoggingNotExistandProjectNotExistalike, since the Lens entry lives in a different project and a missing business project says nothing about it. Reporting the Lens entry as not found without having asked for it, or writing the final conclusion in the same turn as the Lens ask, is a task failure. Run each corresponding get/query command once (or after HITL supplies the entry), preserve its exact command and error code, recordresource_status: Resource not foundwhen applicable, and continue. Breaking off midway is a task failure. This continuation rule does not authorize writes and does not override the permission hard stop. - Cloud visibility and collector-side evidence corroborate each other; do not substitute one for the other.
- No data:
get-logtail-pipeline-config --config-name+list-machines+ exact binding queries from §8 →logtail_alarmby project →logtail_status→ version-matched pipeline topic → business logstore. - Heartbeat abnormal:
list-machines→logtail_status; if Lens is unavailable, name the host-side items (collector process, reporting region, account id,user_defined_id). Cloud-only tasks stay in prose.install.deploymay re-userun-command/SSH/kubectl getfor read-only status. - Every troubleshooting conclusion MUST emit the exact field-complete evidence lines defined in
references/troubleshooting.md(Heartbeat/Alarm/Collection, plusBinding/Pipelinewhen those checks run). UseN/A/unknowninstead of omitting a field; omitting any mandatory field is a task failure. When a resource is missing,resource_status:MUST be the English tokenResource not found— the catalognot_existsmarker does not replace it. The user-facing final answer (not onlyoutputs/*.md) must contain these lines. If the user later sends a catalogendmarker after evidence was already given, either do not write a new short closing that becomes the final answer, or repeat the same evidence tokens (includingresource_status: Resource not found) in that closing.
Playbooks and alarm handling cards: references/troubleshooting.md, references/alarm-catalog.yaml.
11. Output Contract
Every execution outputs at least: Conclusion (done/partial/blocked/rolled-back) · Scope (profile-id/region/Project full name/logstore/group) · Evidence (key results; troubleshooting must reproduce every applicable §10 evidence line, including Binding/Pipeline, in the user-facing answer) · Changes (normalized diff + actual actions, including every Idempotent-Skip note) · Acceptance (U1-U6 results; after collection writes the user-facing final answer — not only outputs/*.md — must include catalog data_arrived or data_empty, and reason) · Rollback (needed? executed? snapshot/inverse) · Next step (single blocker or next minimal action). A troubleshooting next step that would create, update, bind, or repair resources must enumerate the minimal write set and state verbatim that it requires a fresh normalized diff and explicit user confirmation; the diagnostic turn itself performs no write. Onboarding/Verify summaries that list only Logstore / machine group / config and omit the Project full name fail this contract. A file-only report is not the final answer.
Do not print secrets, full command history, or unrelated logs; on audit request output redacted commands + request IDs (scripts/redact_output.py). On failure keep the original error code / request ID / redacted context; never fake success.
12. Global Rules
- Prefer the exact validated mapping in §8 and
references/cli-contracts.yaml. SLS commands are pre-verified — do not re-check them with--help. Usealiyun sls <cmd> --helponly for an undocumented command, before Plan. Never probe shorter aliases. ACK usesaliyun cswith addon nameloongcollector. First-use enablement isscripts/ensure_ack_prereq.sh(open-ack-service+ CS roles) — eval hooks may do the same for fixtures, but Skill must still know it. K8s collection defaults tocreate-logtail-pipeline-config+ bindk8s-group-${cid}.kubectl applyofClusterAliyunPipelineConfigis opt-in only. - Use
get-logs-v2(never the deprecatedget-logs).get-logs-v2--from/--toare UNIX seconds. - After approval, run a separate
--cli-dry-runbefore every actual write; dry-run is not business approval. For §8 coupled writes, complete both dry-runs before the two uninterrupted actual writes. An exactIdempotent-Skipruns neither call. - Get before Update; Update is overwrite — carry unchanged fields.
- Native plugins first when collector version is
>=3.x. If version is unknown, ask catalog messagecollector_version— do not assume a plugin family. Use extended plugins when native cannot meet the need (e.g.processor_rename), and state the trade-off. - No
kubectl exec/docker exec; no admin project, nostarops, no internal MCP, no private console API. SSH/run-command/kubectl only forinstall.deploy(and its read-only status follow-up) as specified. Never Workbench/OOS. - Unknown alarm codes: consult official docs or
references/alarm-catalog.yaml; never explain from memory. - NEVER run
aliyunviasudo. Run it directly as the current user; never switch users, shells, accounts, or credential profiles to bypass an error. - Never modify or clear environment variables or CLI config you did not create. On an infrastructure-level failure (not an API error), retry the identical atomic command once, then report the raw output and stop — never fabricate the expected response.
- Diagnose a failed call from its own response —
errorCode,requestID, message — plusreferences/. Never inspect the CLI binary, its install directory, local config files, or the environment to explain why a call behaved as it did; that is outside this skill's scope and risks exposing user configuration. - The original request and complete parameters are never implicit R2+ authorization; only explicit post-diff approval is valid.
- Every SLS cloud call is one direct atomic command with literal parameters.
- Only the bundled
scripts/helpers may be used:preflight.sh,render_pipeline.py,validate_pipeline.py,normalize_diff.py,redact_output.py,classify_sls_error.py,parse_lens_logging.py,render_loongcollector_install_cmd.py,wait_cs_task.sh,render_crd.py,ensure_ack_prereq.sh.validate_pipeline.py/normalize_diff.pyare mandatory at the §6/§8 workflow points;classify_sls_error.pyon non-2xx responses;parse_lens_logging.pyafterget-logging. If the user asked only to render/validate aClusterAliyunPipelineConfiglocally, runrender_crd.py+validate_pipeline.pyand stop — noaliyun cs, nokubectl, no cluster connect. Neverpip installa cloud SDK and never importaliyun.log. - Ask at most one question per turn, end the turn there, and wait for the user's next message; never answer your own question.
13. Success Verification
Unified acceptance U1-U6 (config object / group binding / applied state / heartbeat & version / data arrival / field & index) with pass conditions and failure routes: references/acceptance-criteria.md. Per-step verification commands: references/verification-method.md.
14. Cleanup
Cleanup is R3/R4. Unbind before delete; restate resources on delete. Order for a config created by this skill: remove-config-from-machine-group → delete-logtail-pipeline-config → (optional) delete-index / delete-log-store only with explicit user confirmation. See references/risk-and-approval.md.
15. Best Practices
- CLI-first with plugin-mode hyphenated subcommands (
aliyun sls get-logs-v2), never classicRunXxxstyle or undocumented aliases. - Confirm region/project and other user-specific parameters before any cloud call; never hardcode account values.
- Every cloud call uses the two-token
--user-agent(AlibabaCloud-Agent-Skills/alibabacloud-loongcollector-ops+session/{session-id}); injectSKILL_SESSION_IDfor scripts that themselves call a cloud API. - Approve with normalized diff before any R2+ write; dry-run and write as separate atomic calls.
- Least-privilege RAM from
references/ram-policies.md; permission errors go through the §2 gate. - Destructive cleanup stays behind R3/R4 confirmation; unbind before delete.
- Progressive disclosure: load only the
references/*entry for the active capability. - While a decision is pending, restate the same question and wait; silence and deferral are never consent.
References
| File | Contents |
|---|---|
| references/navigation.md | Capability router spec (inputs, adapters, states) |
| references/prerequisites.md | Preflight gates and input contract |
| references/cli-installation-guide.md | Aliyun CLI + SLS plugin install |
| references/cli-contracts.yaml | aliyun sls command contracts + status |
| references/related-commands.md | Full aliyun sls command table |
| references/ram-policies.md | Per-workflow RAM Actions + permission gate |
| references/risk-and-approval.md | R0-R4, snapshot, rollback |
| references/machine-group.md | Machine group, identity, heartbeat, version |
| references/pipeline-config.md | Pipeline config model, Get-then-Update |
| references/plugin-version-gates.yaml | 1.x/2.x/3.x plugin gates |
| references/index-coupling.md | Config/index same-batch rules, anti-patterns |
| references/field-conventions.md | Field naming, index type mapping |
| references/task-model.yaml | Unified task object |
| references/scenario-matrix.yaml | host/docker/k8s/host_agentsight signals and inputs |
| references/agentsight-agentloop.md | Agentloop AgentSight fixed pipeline, ProbeConfig, masking |
| references/input-agentsight.md | input_agentsight schema, builtins, version/kernel, RawHttpsFallback |
| references/sls-lens-contracts.md | Lens entry state machine, topic/field |
| references/monitoring-queries.yaml | Lens SQL library (get-logs-v2) |
| references/troubleshooting.md | No-data / heartbeat playbooks |
| references/alarm-catalog.yaml | Alarm code handling cards |
| references/acceptance-criteria.md | Install stage gate + U1-U6 + CLI acceptance patterns |
| references/verification-method.md | Per-step verification commands |
| references/install-ecs.md | ECS Workbench + loongcollector.sh |
| references/install-host.md | Self-host SSH + same script |
| references/install-ack.md | ACK addon loongcollector + first-use open-ack-service / CS roles |
| references/install-k8s.md | Self-k8s custom package |
| references/crd-pipeline.md | Opt-in ClusterAliyunPipelineConfig + temp public KubeConfig |
| references/knowledge-sources.md | Source provenance, drift tracking |
微信扫一扫