SEMANTIC DIRECTORY · plugin-directory-semantic-v3-18

Image Understanding & Analysis

让模型理解已有图片内容并输出语义结果:图像描述、视觉问答、截图或界面分析、物体与场景识别、图像比较及视觉模型桥接。不含以文字提取为核心的 OCR、确定性图像处理、屏幕采集与图像生成。

Use keyword search
Source catalog discovery.composite.2026-09-06.v1-0-5.b92b985e31dfCanonical structure · 823 nodes · 14829 plugins

Leaf membership

51 plugins

dsh-visionaryzhuiyueya/dsh-visionaryCatalog AnalyzedGive text-only DeepSeek models eyes — a DeepSeek Harness plugin that transparently converts chat images into OCR text + vision-model descriptions before they reach the LLM. Configure vision backends (GLM-4V, Qwen-VL, Gemini, Ollama…) right in the Models settings page; multi-backend fallback chain, double-layer caching, no config files.Vision & MultimodalGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsdsh-vlm-bridgeme9rez/dsh-vlm-bridgeCatalog AnalyzedDeepSeek Harness (dsh) bundle plugin: vision_analyze tool lets text-only LLM agents read images via SenseNova VLM, with Schemastery config and single-source credentialsVision & MultimodalGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsdsh-vqa-agentjypjypjypjyp/dsh-vqa-agentCatalog AnalyzedDSH Plugin:vqa_ask 双模型视觉问答 —— 主模型提问 → 视觉模型看图回答,UI 实时展示 QA 过程,支持多模态视觉模型选择UI EnhancementsGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals2 starseasy-visioneasy-visionCatalog AnalyzedA DeepSeek Harness tool plugin that lets a text-only agent 'see' local images: it reads an image file, auto-detects the real format, and returns (or writes to a Markdown file) a detailed text description produced by a separate OpenAI-compatible vision modVision & MultimodalnpmGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency audit0 reported in bounded auditArtifact0.1.1Evidence updated2026-09-08Public signals101 downloadsgewu-toolsgewu-toolsCatalog AnalyzedModel-agnostic visual-inspection pipeline for text-only agents: page-by-page HTML screenshots plus a ready-made vision-subagent briefing contract (gewu_prep), then source-code truth verification of every finding (gewu_locate); validated on mimo-v2.5 & qwen3.7-plus.Tools & CapabilitiesnpmGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency audit0 reported in bounded auditArtifact1.0.0Evidence updated2026-09-07Public signals1 stars · 184 downloadsmeow-visionmeimiaoji-creator/meow-visionCatalog Analyzedmeow-vision 是 DeepSeek Harness 的一款视觉Plugin,解决纯文本模型无视觉。另一方面vue页面开发视觉验证不闭环的问题。Vision & MultimodalGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals3 starsprismrelay-mcpArnoldkevin/prismrelay-mcpCatalog AnalyzedVision-first local MCP that gives text-only Agents image understanding through Agnes AI (BYOK).Vision & MultimodalGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals1 starstext-llm-visionshaoqiuyuavailable/text-llm-visionCatalog AnalyzedScene-aware vision routing layer for DeepSeek Harness (dsh): decides which engine/backend an image should go to (chat/UI/table/code) before other vision plugins route it. Switch-gated, front-loaded, never touches other plugins' tools. 场景级识图路由层。Vision & MultimodalGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals2 starsvision_kitSeom-ingit/vision_kitCatalog AnalyzedMake your AI agent a math tutor. Structured extraction of vectors, matrices & geometry from math figures, with dimension-consistency + geometric self-check. Vision plugins for DeepSeek Harness, opencode (MCP) & CLI. Verify, don't believe.Vision & MultimodalGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsVLM-of-deepseek-harnesswangziyi863/VLM-of-deepseek-harnessCatalog AnalyzedDSH plugin from wangziyi863/VLM-of-deepseek-harnessTools & CapabilitiesGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals1 starsxby-cellphone-detectionxby-skill/xby-cellphone-detectionCatalog Analyzed输入一张图像,对其中的手机进行检测,输出图片中所有目标的检测框、置信度和标签。Tools & CapabilitiesGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsxby-detect-vehiclexby-skill/xby-detect-vehicleCatalog Analyzed输入一张图像,检测图像中的车辆类型(car/truck/bus/motorbike/tricycle/carplate),输出所有目标的检测框、置信度和标签。Tools & CapabilitiesGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsxby-ebike-detectionxby-skill/xby-ebike-detectionCatalog Analyzed输入一张图像,对其中的电动自行车进行检测,输出图片中所有目标的检测框、置信度和标签。Tools & CapabilitiesGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsxby-reflective-vestxby-skill/xby-reflective-vestCatalog Analyzed输入一张图像,检测人员是否穿戴反光衣,输出图片中所有目标的检测框、置信度和标签(safe/unsafe)。Tools & CapabilitiesGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 starsxby-smoking-detectionxby-skill/xby-smoking-detectionCatalog Analyzed输入一张图像,对其中的香烟目标进行检测,输出图片中所有目标的检测框、置信度和标签。Tools & CapabilitiesGitHubCompatibilityDSH 0.1.2-rc.1 · webDependency auditNot testedArtifactUnresolvedEvidence updated2026-09-10Public signals0 stars

Categories are public semantic index facts, not a final recommendation. Compatibility, permissions, and source evidence remain separate.