YouMind
Sign in
Skill Evaluator

Skill Evaluator

A reference for evaluating AI Skills

Installed by
9
CategoryResearch
FromYouMind

Showcase

评测 · 技能图标生成器 · 2026-08-27

评级:★★★★☆(4.1 / 5)|综合分 82 / 100 → 达到「留存/上架」判定线

结论:值得留存/继续上架——核心能力(锚点图锁定一致性的三件套流水线)经实测成立,且来源完全可信(作者即本人);但「单张长图交付」依赖环境拼图工具、描述缺负面边界两处短板,建议打磨后长期保留。

五维评分表

维度得分关键证据(事实)归因(推断)
触发可靠性26/303/3 正面触发成功(2 条 showcase+1 条现场实测);实测用自然口语「给我做一套视觉物料:技能名……」即触发描述含明确的输入格式与用途(inputHint),触发面清晰
执行有效性20/25实测完整跑通锚点→封面→信息区三步,封面以锚点图为参考、风格一致;showcase 两套成品均三件套齐全一致性机制(锚点图+头图原图复用)设计成立;但本环境无拼图工具时只能降级交付两张图
成本效率11/15实测 3 次生图+2 次确认门控即完成一套;无无用重试摇骰子随机基调有「用户不满意→重生成」的潜在重复成本;确认门控增加交互次数
稳健与边界11/15实测中确认门控严格执行:未经锚点确认不生成封面、未经文案确认不生成长图防漂移设计是本技能最大优点;但描述无「何时不该用」负面边界,负面/边缘用例未实测
安全可信14/15作者=用户本人(own_skill);reviewStatus=approved;全文可查;仅调用 Create image 技能,无异常外发来源完全可信,无数据风险

综合分 = 26+20+11+11+14 = 82 ≥ 80 → 达到留存/上架判定线。

短板归因

  • 描述(触发边界不全):描述只讲「该用时怎么用」,没写「何时不该用」(如:只想做单张图、已有固定视觉品牌时不该用),会诱使 AI 在用户只想要一张图时也跑整套流程。

  • 步骤(阶段三交付形态依赖环境):最后一步「合成单张 9:16 长图」要求执行环境支持图片拼接;本环境不支持时只能交付「头图原图+信息区」两张图,「开箱即用单张长图」的承诺打折扣。

  • 步骤(showcase 与铁律轻微偏离):水果动画 showcase 实际用「单次生图+头图参照重绘」方式出整张长图,与「头图绝不重绘、必须原图复用」的铁律不完全一致——成功率依赖模型听话,存在潜在漂移。

改进建议(按优先级)

  1. 描述补一句负面边界(最易改):加「当只需单个图标、或已有既定视觉规范时,不要启动三件套流程」。

  2. 阶段三写明降级交付:在步骤里预置「环境无拼图工具时,明确告知用户需自行拼接头图与信息区,并给出拼接尺寸建议」,避免交付时突然降级。

  3. 收敛 showcase 执行路径:统一走「先拼图后交付」或「单图生图」其中一条,避免同一步骤两种做法并存导致行为不稳定。

迭代落地记录(2026-08-27)

评测 · 技能图标生成器 · 2026-08-27

Description

Enter the name, link, or instructions for the AI Skill to evaluate, then conduct an evidence-based initial quality review in real-world usage scenarios. Keep the six-question quick test (4 positive, 1 negative, and 1 ambiguous), and separately verify automatic selection and execution after explicit invocation; successful manual invocation does not count as successful automatic triggering. Use direct scoring across five dimensions: triggering 30%, execution 25%, cost 15%, robustness and boundaries 15%, and safety 15%. Distinguish between evidence from this test, historical run evidence, static checks, and unverified items. Do not provide an overall score or star rating when key evidence is incomplete. Assess safety based on data access, external transmission, and operational authorization—not on the author's identity or endorsement. Block recommendations in cases of serious unauthorized access. Produce a one-page, conclusion-first report showing the tested scope, unverified items, and prioritized improvement suggestions, and automatically create a record. Say “deep review” to add independent repeat tests and source checks; the budget will be confirmed before execution. This Skill only evaluates and provides recommendations; it does not automatically modify the Skill being evaluated. Suitable for quality checks of self-created Skills, rechecks of installed Skills, and pre-installation screening of marketplace Skills; the six-question results do not represent full certification.

Editor's Recommendation
S

Recommended by

Shuting@YouMind

Why we love this skill

An evidence-based framework that separates automatic selection from explicit execution, scores only traceable results, and includes safety risks and process efficiency in its conclusions.

Related Skills

View all

Skill Quality Audit (PDCA-QMS)

Evaluate Skills using a quality management system based on management science, not an off-the-cuff checklist. Based on Six Sigma CTQ/FMEA/DPMO × TQM × Deming PDSA cycle, through the Plan-Do-Check-Act four phases, it provides: a quality scorecard with Sigma level, a defect priority table ranked by RPN, root cause analysis down to the instruction design layer, and the full Skill text after Poka-Yoke correction. The fixer and reviewer are forcibly separated, and scores only increase, never decrease. v2.0 additions: ① Quantitative contract – closed enumeration of structural unit counting (only phase level), K-value gradient anti-scoring lock, Sigma table lookup with log axis interpolation, and full zero-padding to prevent division by zero, ensuring scores are reproducible and undistorted; ② Quality level × disposal path dual-axis gate, eliminating the contradiction of 'unqualified but recommended for release' labels; ③ Diagnostic mode – say 'only a report' or 'don't modify yet', then run through P-D-C and stop at Check, without modifying your script; ④ Output contract – pass/fail points fold, evidence limited to 50 characters, phase word count limits, values must include calculation process. When to use: evaluate whether a Skill is good, diagnose why its output is unstable, systematically optimize an existing Skill, benchmark multiple Skills, pre-release quality inspection of Skill drafts, or only want a diagnostic report without modifying the draft. Trigger words: detect skill quality, skill quality check, quality audit, rate a skill, PDCA optimize skill, skill health check, evaluate this skill, skill defect analysis, diagnose only without repair, pre-release skill quality check, skill benchmarking.

一
3500

Image Generation Grandmaster

If you want to extract production-ready image-generation Skills from reference images on YouMind in one streamlined process, this is the grandmaster-level option. Creating a new Skill costs around 300 credits, turning you into an image-generation Skill-making machine and helping you avoid Skills that cost thousands. With reference images, you can build what you need yourself. Personally tested and effective; see the task examples for details. It distills the visual mechanisms in reference images into reusable, testable image-generation Skills. The subject can change while the composition, spatial logic, material appearance, and visual hierarchy are preserved, rather than simply copying the original image. It is suited to creators and teams that need to establish a consistent visual style, reuse design principles, or evaluate existing generation Skills. Use it to compile visual rules, test existing Skills, calibrate based on feedback, run historical regressions, and publish versions. It clearly identifies the scope of application, fixed invariants, adjustable variables, adaptation rules, and common failure modes. It also distinguishes between an image that is high quality and one that truly preserves the visual mechanism, helping you identify why a generated result has deviated from the underlying rules. The final results may include a structured visual specification, an independent Skill, a test plan or audit report, along with calibration patches, regression results, and release records. When image generation or visual inspection is unavailable, complete test prompts are retained and the pending-verification status is clearly indicated. Once confirmed, rules, failure diagnoses, and version information can continue to be recorded in YouMind for future maintenance and reuse.

P
110k

Career Experience Analyzer

After years of work, the real challenge is often not a lack of experience, but knowing which parts are facts, which are your interpretations, and which experiences are still worth validating outside your original company. This Skill works with one anonymized, real-life experience: it organizes evidence cards, explains what the current material can and cannot prove, identifies the earliest evidence gap, and provides one low-cost validation path, one thing to hold off on, and a 7-day action plan. It is not a psychological test and does not assign an overall experience-asset score. It does not use your job title, years of experience, salary, certifications, praise from colleagues, or a single internal success as substitutes for external demand and payment evidence. When the material is insufficient, it will say directly: “Insufficient material; no judgment for now.” Please do not submit names, companies, clients, phone numbers, WeChat IDs, email addresses, government ID numbers, contracts, or unpublished business data. If identifiable information is detected, the Skill will stop the analysis and ask you to anonymize the material locally before resubmitting it. This is a preliminary self-guided assessment. It does not replace a professional evaluation, does not include manual review by Kevin, and does not promise income, traffic, career transitions, or business results.

K
01k

Information

Version
v4
Last updated
Runtime credits
Usage-based
Models
Auto

Ready to create something bolder?

Skill Evaluator