---
document_id: agent.collection.task.raw_html_v1
schema_version: 2
parent_document_id: agent.querying
section: querying
---

<!-- Generated by scripts.build_agent_guide; do not edit. -->

# Web collection / 网页采集

Fetch multiple public web pages by URL, including known pages the Agent cannot fetch locally, and return full-page Markdown or structured elements. Pages that need a login or interaction belong to the local browser MCP, not this task.

按 URL 批量采集多个公开网页,也适用于 Agent 本地 fetch 无法读取的已知页面;返回整页 Markdown 或结构化元素。需要登录或点击翻页的页面用本地浏览器 MCP,不是这个任务。

Task / 任务: `raw_html_v1`
Platform / 平台: Web / 网页
Billing / 计费: 3 Credits per successful URL / 每个成功 URL 3 Credits
Estimate / 估价: call `estimate_collection`; start directly when `approval_required=false`, otherwise obtain approval for `upper_bound_credits`.
估价：调用 `estimate_collection`；`approval_required=false` 时直接启动，否则请用户确认 `upper_bound_credits`。

## Limitations / 限制

- Markdown 单页 UTF-8 最多 2 MiB,序列化结果 data 总量最多 6 MiB;超限页面明确失败,不截断正文、不计费。
- 只抓公开页面,需要登录或反爬拦截的页面会以单条失败返回,不影响同批其他 URL。
- 目标域名必须解析到公网单播地址;请求固定连接已校验的 IP 并保留 Host/SNI,每次跳转最多 5 次且重新校验目标。
- 不返回原始 HTML:markdown 返回整页 Markdown,json 返回结构化元素。只剥掉脚本与样式,导航、侧栏、页脚都会保留,不做正文抽取。
- 不执行 JavaScript,靠脚本渲染的页面只能拿到首屏 HTML 的转换结果。
- 单次最多 10 个 URL,按成功条数结算。

## Inputs / 输入

- `urls` — Web URLs / 网页链接; kind=url_list; required=True; default=[]
- `output_format` — Output format / 输出格式; kind=select; required=False; default='markdown'; allowed=markdown, json

## Minimal call / 最小调用

`estimate_collection`

```json
{
  "task_query": "获取 https://example.com/ 的网页内容",
  "task_code": "raw_html_v1",
  "input": {
    "urls": [
      "https://example.com/"
    ],
    "output_format": "markdown"
  }
}
```


## Result and pagination / 结果与分页

- mode=none

Only the fields listed below are part of the stable result contract. /
只有下方列出的字段属于稳定结果契约。

## Result fields / 结果字段

- `meta.request_id` — type=string; Asklear request identifier for tracing. / 用于追踪的 Asklear 请求 ID。
- `meta.api_name` — type=string; Executed collection operation name. / 实际执行的采集操作名称。
- `meta.latency_ms` — type=integer; Collection execution latency in milliseconds. / 采集执行耗时（毫秒）。
- `meta.credits_charged` — type=integer; Asklear Credits settled for each successful URL in this task. / 本任务按每个成功 URL 结算的 Asklear Credits。
- `data.documents` — type=array<object>; One result document per submitted URL. / 每个提交 URL 对应一个结果文档。
- `data.documents[].url` — type=string; Normalized submitted URL. / 规范化后的提交 URL。
- `data.documents[].status` — type=enum; succeeded or failed. / succeeded 或 failed。
- `data.documents[].format` — type=enum; markdown or json. / markdown 或 json。
- `data.documents[].content` — type=string; Markdown for a successful document, up to 2 MiB in UTF-8; total serialized result data is limited to 6 MiB. Content is not truncated. / 成功文档的 Markdown，UTF-8 单页最多 2 MiB；序列化结果 data 总量最多 6 MiB。不截断正文。
- `data.documents[].elements` — type=array<object>; Structured JSON elements when format is json and status is succeeded. / format 为 json 且 status 为 succeeded 时的结构化元素。
- `data.documents[].elements[].type` — type=string; Element type, such as a visible HTML tag. / 元素类型，例如可见 HTML 标签。
- `data.documents[].elements[].text` — type=string; Visible text for the element. / 元素的可见文本。
- `data.documents[].elements[].metadata` — type=object; Bounded element metadata; currently includes the source tag. / 有界元素元数据；当前包含来源标签。
- `data.documents[].error_code` — type=string; collection_content_too_large for content exceeding the page or batch budget; collection_failed for other failures. Failed documents are not charged. / 单页或批次内容超限为 collection_content_too_large，其他失败为 collection_failed。失败文档不计费。

Use `list_collection_tasks` as the runtime authority, then `describe_collection_task` for one task's full input contract, examples and limits. Collection access is entitlement-gated. /
以 `list_collection_tasks` 返回为运行时准据，再用 `describe_collection_task` 取单个任务的完整入参契约、示例与限制；采集能力受 entitlement 控制。
