Audit Report
drhus/ai-auditor
8d681f7e9bee · ran in 34.8s · bundle fb69bd0f
Overall score
2.4 /4
Partial
Risk class
HIGH
1
Code passed
22 / 45
Attestation Yes
0
Outstanding ext.
1
CODE-CHECKED CLAUSES
- Strong15
- Adequate7
- Partial10
- Inadequate5
- Absent8
ATTESTATION QUESTIONS
- Yes0
- No0
- Not Applicable0
- Outstanding0
Harm
Don't hurt people
2.2/4
Partial
Truth
Don't deceive people
2.1/4
Partial
Responsibility
Don't abuse power
2.5/4
Partial
Order
Don't destabilize society
3.0/4
Adequate
⚠ 1 outstanding external confirmations — required for a complete Annex IV dossier. Complete now →
Safety
Don't harm people2.7/4Adequate6 clauses▸
Safety
Don't harm peoplePARTIALRisk management system established, implemented, documented EU AI Act, Art 9 LLM2/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.55 (2/3 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.62): Risk identification and analysis are evidenced (risk register and threat model), but the absence of documented CI/CD evaluation gates creates ambiguity about whether the 'continuous iterative process' and 'regular systematic review and updating' requirements are satisfied throughout the entire lifecycle. Human judgment needed to assess whether evaluation gates exist outside the repository or if alternative continuous monitoring mechanisms are documented elsewhere.
▸Evidence · 2 hits— click to view code
Suggested fix · we looked for these and found none
- ci_eval_gates
INADEQUATEResilience to errors, faults, inconsistencies EU AI Act, Art 15(4)0/3 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.16 (0/3 rules matched).
▸Evidence · 2 hits— click to view code
nippet": "import os from langchain.callbacks.streaming_stdout import StreamingStdOutCallbackHandler from langchain_openai import ChatOpenAI"
dOutCallbackHandler from langchain_openai import ChatOpenAI", "rule": "langchain_import" }, { "fSuggested fix · we looked for these and found none
- error_handling_at_tool_boundaries
- retry_logic
- fallback_behaviour
ADEQUATERisks and benefits to people identified NIST AI RMF, Art MAP 3.42/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.80 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
STRONGAI risk assessment process ISO/IEC 42001, Art 6.12/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
PARTIALOperational planning and control ISO/IEC 42001, Art 8.1 skip1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- presence_of_ci_workflows
STRONGData protection impact assessment (DPIA) GDPR, Art 351/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Privacy
Respect boundaries1.6/4Partial6 clauses▸
Privacy
Respect boundariesABSENTUntargeted facial image scraping for face databases EU AI Act, Art 5(1)(e)1/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched).
▸Evidence · 6 hits— click to view code
" description: "Biometric ID, categorisation, emotion recognition" - signal: critical_infra_signals category: "2" description:
otion recognition / biometric categorisation" - signal: agent_framework # any interactive AI is in scope paragraph: "50(1)"
calls (mediapipe, face_recognition, dlib, opencv haar cascade). score_mapping: { pass_default: 4, fail_on_match: 0 } remediation_h"^1.1.0", "@playwright/test": "^1.51.1", "babel-plugin-react-compiler": "*", "react": "^18.2.0 || 19.0.0-rc-de68d2f4
}, "@playwright/test": { "optional": true }, "babel-plugin-react-compiler": { "optional":…and 1 more.
ABSENTEmotion recognition in workplace and education EU AI Act, Art 5(1)(f)1/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched).
▸Evidence · 6 hits— click to view code
ms", pattern: /\b(?:emotion_detect|emotion_recognition|sentiment_score|affect_recognition|facial_emotion|micro_expression)\b/gi }, // ---
\b(?:emotion_detect|emotion_recognition|sentiment_score|affect_recognition|facial_emotion|micro_expression)\b/gi }, // ----- data_io ----
emotion_recognition|sentiment_score|affect_recognition|facial_emotion|micro_expression)\b/gi }, // ----- data_io ----- { signal: "data_ppet": "OMMANDS for candidate in tokens[index + 1:]: if candidate in {\"|\", \";\", \"&&\", \"||\"}: break f", "rule": "employmeens[index + 1:]: if candidate in {\"|\", \";\", \"&&\", \"||\"}: break f", "rule": "employment_terms" }, {…and 1 more.
EXTERNALReal-time remote biometric identification in public spaces EU AI Act, Art 5(1)(h)EXT▸
Why we flagged it
Deployment context (public space, real-time, law enforcement use, judicial authorisation) is operational, not knowable from code. Always external.
ABSENTData and data governance practices documented EU AI Act, Art 100/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.10 (0/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
"^1.1.0", "@playwright/test": "^1.51.1", "babel-plugin-react-compiler": "*", "react": "^18.2.0 || 19.0.0-rc-de68d2f4
}, "@playwright/test": { "optional": true }, "babel-plugin-react-compiler": { "optional":Suggested fix · we looked for these and found none
- presence_of_data_card
- data_loading_code_quality
- bias_evaluation_present
STRONGPrivacy risk of the AI system evaluated NIST AI RMF, Art MEASURE 2.82/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched).
▸Evidence · 7 hits— click to view code
description: "Code redacts or hashes PII before logging or sending to external models" - rule: privacy_documentation weig
p \"$SCRIPT_DIR/mcp/redaction.js\" \"$TARGET_ABS/mcp/\" cp \"$SCRIPT_DIR/mcp/lib/\"*.js \"$TARGET_ABS/mcp/lib/\" rm -rf \"$TARGET_ABS/mcp/li
t the proof needed, redact sensitive data, and report responsibly.</p></div> </div> </div> <div class=\"foot\"><span", "rule": "
h MCP, which writes redacted audit metadata and egress information.</p> </div> <div class=\"code-card\"> <div", "rule": "pii_red
n>Audited requests, redacted URLs, visible egress</span></div> </section> <section class=\"slide\" data-title=\"Egress\"> <", "r
…and 2 more.
STRONGData protection by design and by default GDPR, Art 252/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 6 hits— click to view code
crypto.createHash("sha256")crypto.createHash("sha256")redact
crypto.createHash("sha256")…and 1 more.
Transparency
Don't deceive people2.3/4Partial7 clauses▸
Transparency
Don't deceive peopleSTRONGSubliminal techniques distorting behaviour EU AI Act, Art 5(1)(a)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_manipulative_prompt_patterns
INADEQUATETransparent operation and instructions for use EU AI Act, Art 131/3 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.25 (1/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 3 hits— click to view code
2 required sections present
Suggested fix · we looked for these and found none
- output_interpretation_guidance
- limitations_section_present
INADEQUATEUsers informed they are interacting with an AI EU AI Act, Art 50(1)1/2 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.20 (1/2 rules matched).
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- ai_disclosure_in_user_facing_strings
ADEQUATEAI-generated content marked as such, machine-readable EU AI Act, Art 50(2)2/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.82 (2/2 rules matched).
▸Evidence · 8 hits— click to view code
ent provenance | no c2pa imports → ABSENT | "API response contains `aiGenerated:true` / C2PA header" | | GDPR Art 5(1)(c) — PII in logs | gr
aiGenerated:true` / C2PA header" | | GDPR Art 5(1)(c) — PII in logs | grep `user.email` near `logger.info` | "fake PII sent → grep all captu
` | 50(2) | **1** | C2PA / watermark library imports | | `art-50/p3-emotion-biometric-disclosure` | 50(3) | **1** | Deterministic disclosure
ovenance_hooks` | C2PA, watermarking, content labelling | Article 50(2) synthetic content disclosure
o/text) | C | C2PA / watermarking libraries; metadata writers; output post-processing. | | 50(3) | Emotion-recognition / biometri
…and 3 more.
ABSENTEmotion recognition / biometric categorisation disclosure EU AI Act, Art 50(3)0/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- emotion_or_biometric_disclosure_string
ADEQUATEDeepfake content labelled as artificially generated EU AI Act, Art 50(4) skip1/1 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.70 (1/1 rules matched).
▸Evidence · 2 hits— click to view code
ent provenance | no c2pa imports → ABSENT | "API response contains `aiGenerated:true` / C2PA header" | | GDPR Art 5(1)(c) — PII in logs | gr
aiGenerated:true` / C2PA header" | | GDPR Art 5(1)(c) — PII in logs | grep `user.email` near `logger.info` | "fake PII sent → grep all captu
STRONGPrinciples relating to processing of personal data GDPR, Art 52/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 3 hits— click to view code
Auditability
Actions must be traceable1.9/4Partial8 clauses▸
Auditability
Actions must be traceableINADEQUATETechnical documentation drawn up before placing on market EU AI Act, Art 111/3 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.15 (1/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
2 required sections present
Suggested fix · we looked for these and found none
- presence_of_model_card
- architecture_docs
ABSENTAutomatic recording of events over the lifetime EU AI Act, Art 12(1) LLM1/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.58 (1/3 rules matched). LLM judge (confidence 0.85): The evidence shows only ad-hoc console.log() statements at tool boundaries, which are runtime output rather than persistent, structured automatic logging. The clause requires technically enabled automatic recording over the system's lifetime, necessitating structured logging import and a persistent sink—neither of which matched the deterministic rules.
▸Evidence · 8 hits— click to view code
x.ts"), content); console.log(` rewrote index.ts (${entries.length} seeds)`); } async function main() { const args = process.argv.slic-1.0"], }; console.log(`\n=== ${owner}/${repo} (id=${input.auditId}) ===`); const t0 = Date.now(); try { const reportevt.kind === "log") console.log(` [${evt.stage}] ${evt.text}`); else if (evt.kind === "stage") console.log(` [${evt.stage}] phase=t.kind === "stage") console.log(` [${evt.stage}] phase=${evt.phase}${evt.durationMs ? ` (${evt.durationMs}ms)` : ""}`); else if (ev= "classification") console.log(` ✦ risk=${evt.classification} annex=${evt.annexIii.join("/")} art50=${evt.art50.join("/")}`); else…and 3 more.
Suggested fix · we looked for these and found none
- structured_logging_imported
- logging_persistent_sink
PARTIALLogging ensures traceability appropriate to risk EU AI Act, Art 12(2) LLM2/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.68 (2/3 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.62): The system demonstrates input/output pair logging and model identity tracking (0.75 weighted compliance), but lacks explicit request_id correlation in reviewed logs, preventing full traceability chain. Evidence shows audit-level logging rather than request-level granularity needed for deterministic request tracking.
▸Evidence · 8 hits— click to view code
x.ts"), content); console.log(` rewrote index.ts (${entries.length} seeds)`); } async function main() { const args = process.argv.slic-1.0"], }; console.log(`\n=== ${owner}/${repo} (id=${input.auditId}) ===`); const t0 = Date.now(); try { const reportevt.kind === "log") console.log(` [${evt.stage}] ${evt.text}`); else if (evt.kind === "stage") console.log(` [${evt.stage}] phase=t.kind === "stage") console.log(` [${evt.stage}] phase=${evt.phase}${evt.durationMs ? ` (${evt.durationMs}ms)` : ""}`); else if (ev= "classification") console.log(` ✦ risk=${evt.classification} annex=${evt.annexIii.join("/")} art50=${evt.art50.join("/")}`); else…and 3 more.
Suggested fix · we looked for these and found none
- logs_include_request_id
PARTIALContext of use established and understood NIST AI RMF, Art MAP 1.1 LLM1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.30 (1/2 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.65): Intended purposes are documented (README.md evidence), satisfying 60% of weighted requirements, but deployment context documentation is not confirmed in the repository scan. The clause requires both understanding of purposes AND prospective deployment settings; partial fulfillment warrants 'partial' pending external documentation review.
▸Evidence · 2 hits— click to view code
2 required sections present
Suggested fix · we looked for these and found none
- deployment_context_documented
PARTIALPost-deployment monitoring, appeal and override, change management NIST AI RMF, Art MANAGE 4.1 skip2/4 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (2/4 rules matched).
▸Evidence · 7 hits— click to view code
in-the-loop hooks: `human_input()`, interrupt nodes in LangGraph, approval-gate functions, manual- review flags, con
For LangGraph: use `interrupt()` nodes. For custom flows: build an approval-queue pattern. Document where humans can intervene i
ic: - rule: kill_switch_present weight: 0.7 description: | Code contains a documented kill-switch /
sm (function named `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin
med `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin endpoint that h
…and 2 more.
Suggested fix · we looked for these and found none
- feedback_capture_present
- structured_logging_imported
ADEQUATEDocumented information for the AI management system ISO/IEC 42001, Art 7.5 skip1/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.70 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
Suggested fix · we looked for these and found none
- presence_of_versioned_docs
INADEQUATEMonitoring, measurement, analysis and evaluation ISO/IEC 42001, Art 9.11/2 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.25 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
x.ts"), content); console.log(` rewrote index.ts (${entries.length} seeds)`); } async function main() { const args = process.argv.slic-1.0"], }; console.log(`\n=== ${owner}/${repo} (id=${input.auditId}) ===`); const t0 = Date.now(); try { const reportSuggested fix · we looked for these and found none
- presence_of_eval_suite
STRONGRecords of processing activities GDPR, Art 301/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Accountability
Don't abuse power2.5/4Adequate6 clauses▸
Accountability
Don't abuse powerABSENTDeployer log-retention capability supported EU AI Act, Art 26(6) LLM1/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.50 (1/1 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.85): Evidence shows only console.log statements at tool boundaries, which are runtime debugging outputs, not persistent audit logs meeting the 6-month retention requirement. No evidence of automatic log generation, storage mechanism, or retention policy controls required by Art. 26(6).
▸Evidence · 2 hits— click to view code
x.ts"), content); console.log(` rewrote index.ts (${entries.length} seeds)`); } async function main() { const args = process.argv.slic-1.0"], }; console.log(`\n=== ${owner}/${repo} (id=${input.auditId}) ===`); const t0 = Date.now(); try { const reportSTRONGRisk management process documented and accountable NIST AI RMF, Art GOVERN 1.42/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
ABSENTOngoing monitoring and periodic review of risk management NIST AI RMF, Art GOVERN 1.50/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- ci_eval_gates
- drift_monitoring_present
ADEQUATELeadership and commitment for AI management ISO/IEC 42001, Art 5.1 skip1/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.70 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- leadership_signoff_evidence
STRONGRoles, responsibilities and authorities ISO/IEC 42001, Art 5.31/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
STRONGInternal organization controls ISO/IEC 42001, Art A.51/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Human Oversight
Humans stay in control2.4/4Partial5 clauses▸
Human Oversight
Humans stay in controlPARTIALEffective human oversight designed and built-in EU AI Act, Art 14(1) LLM1/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/3 rules matched). LLM judge (confidence 0.62): Human-in-loop hooks are demonstrably present (LangGraph interrupt nodes, approval-gate functions, manual-review flags), satisfying the core intervention mechanism. However, the absence of a dedicated oversight UI and lack of dry-run/simulation capabilities for tool calls leaves the 'effectively overseen' requirement incomplete—operators cannot safely preview or test system actions before execution, limiting practical oversight capability.
▸Evidence · 6 hits— click to view code
in-the-loop hooks: `human_input()`, interrupt nodes in LangGraph, approval-gate functions, manual- review flags, con
For LangGraph: use `interrupt()` nodes. For custom flows: build an approval-queue pattern. Document where humans can intervene i
import ( AIMessage, HumanMessage,", "rule": "langchain_import" }, { "file": "gpt_engineer/core/aimport ( AIMessage, HumanMessage, SystemMessage, messages_from_dict, messages_", "rule": "langchain_import" },
chain.schema import HumanMessage, SystemMessage from termcolor import colored from gpt_engineer.core.ai import", "rule": "langch
…and 1 more.
Suggested fix · we looked for these and found none
- oversight_ui_present
- tool_calls_have_dry_run
PARTIALInterrupt / stop function reachable by overseer EU AI Act, Art 14(4)(d) LLM1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.70 (1/2 rules matched). LLM judge (confidence 0.72): A kill-switch mechanism is present and documented (rule matched at 0.7 weight), satisfying the core requirement for intervention capability. However, the absence of evidence for graceful_shutdown_handler (0.3 weight) creates ambiguity about whether the system reliably achieves a 'safe state' as required by the clause—implementation details on safe shutdown semantics must be verified.
▸Evidence · 6 hits— click to view code
ic: - rule: kill_switch_present weight: 0.7 description: | Code contains a documented kill-switch /
sm (function named `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin
med `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin endpoint that h
stop`, `disable_agent`, a feature flag with explicit disable, an admin endpoint that halts processing). - ru
ic: - rule: kill_switch_present weight: 0.6 - rule: feature_flag_for_disable weight: 0.4 descr
…and 1 more.
Suggested fix · we looked for these and found none
- graceful_shutdown_handler
PARTIALAbility to override / reverse the system's output EU AI Act, Art 14(4)(e) LLM1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.60 (1/2 rules matched). LLM judge (confidence 0.72): The system demonstrates override capability through documented human-in-loop mechanisms (interrupt nodes, approval gates, kill-switch functions), satisfying the override path requirement. However, evidence does not clearly establish that all AI decisions are addressable or reversible by humans, leaving the proportionality and completeness of override scope ambiguous.
▸Evidence · 6 hits— click to view code
in-the-loop hooks: `human_input()`, interrupt nodes in LangGraph, approval-gate functions, manual- review flags, con
For LangGraph: use `interrupt()` nodes. For custom flows: build an approval-queue pattern. Document where humans can intervene i
ic: - rule: kill_switch_present weight: 0.7 description: | Code contains a documented kill-switch /
sm (function named `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin
med `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin endpoint that h
…and 1 more.
Suggested fix · we looked for these and found none
- decisions_are_addressable
PARTIALMechanisms to supersede or deactivate AI systems NIST AI RMF, Art MANAGE 2.3 skip1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.60 (1/2 rules matched).
▸Evidence · 6 hits— click to view code
ic: - rule: kill_switch_present weight: 0.7 description: | Code contains a documented kill-switch /
sm (function named `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin
med `kill_switch`, `emergency_stop`, `disable_agent`, a feature flag with explicit disable, an admin endpoint that h
stop`, `disable_agent`, a feature flag with explicit disable, an admin endpoint that halts processing). - ru
ic: - rule: kill_switch_present weight: 0.6 - rule: feature_flag_for_disable weight: 0.4 descr
…and 1 more.
Suggested fix · we looked for these and found none
- feature_flag_for_disable
STRONGAutomated individual decision-making, including profiling GDPR, Art 222/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 12 hits— click to view code
in-the-loop hooks: `human_input()`, interrupt nodes in LangGraph, approval-gate functions, manual- review flags, con
For LangGraph: use `interrupt()` nodes. For custom flows: build an approval-queue pattern. Document where humans can intervene i
import ( AIMessage, HumanMessage,", "rule": "langchain_import" }, { "file": "gpt_engineer/core/aimport ( AIMessage, HumanMessage, SystemMessage, messages_from_dict, messages_", "rule": "langchain_import" },
chain.schema import HumanMessage, SystemMessage from termcolor import colored from gpt_engineer.core.ai import", "rule": "langch
…and 7 more.
Fairness
Treat people fairly3.0/4Adequate4 clauses▸
Fairness
Treat people fairlySTRONGExploiting vulnerabilities (age, disability, socio-economic) EU AI Act, Art 5(1)(b)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_protected_attribute_targeting
STRONGSocial scoring leading to detrimental treatment EU AI Act, Art 5(1)(c)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- detect_scoring_with_persistent_user_state
ABSENTPredictive policing solely from profiling EU AI Act, Art 5(1)(d)1/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.76 (1/1 rules matched).
▸Evidence · 6 hits— click to view code
crime-likelihood / recidivism / "risk to commit X" scores when the input contains only person profile data (no obje
names like `crime_risk`, `recidivism_score`, `offender_likelihood` combined with profile inputs. score_mapping: { parecidivism_score`, `offender_likelihood` combined with profile inputs. score_mapping: { pass_default: 4, fail_on_match: 0 }ms", pattern: /\b(?:crime_risk|recidivism|offender_likelihood|police_dispatch|criminal_record|sentencing_recommend)\b/gi }, // ----- migr
n: /\b(?:crime_risk|recidivism|offender_likelihood|police_dispatch|criminal_record|sentencing_recommend)\b/gi }, // ----- migration_signa
…and 1 more.
STRONGBiometric categorisation by protected attributes EU AI Act, Art 5(1)(g)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_biometric_categorisation_by_protected_attrs
Security & Governance
Don't destabilize society3.0/4Adequate4 clauses▸
Security & Governance
Don't destabilize societyADEQUATECybersecurity measures appropriate to circumstances EU AI Act, Art 15(5)3/4 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.80 (3/4 rules matched).
▸Evidence · 12 hits— click to view code
(hallucination, prompt injection, output bias, leakage, capability escalation), mitigation owner, status. Wire eval suite into CI
ic: - rule: prompt_injection_defences weight: 0.4 description: | Code includes prompt-injection miti
Code includes prompt-injection mitigations: output filters, input sanitisation, instruction-data segregation, system-promp
Eval suite includes prompt-injection / adversarial cases" score_mapping: ">=0.85": 4 ">=0.65": 3 ">=0.40": 2 ">=
hint: | Add a prompt-injection eval set (e.g. from `promptbench`, `garak`, or your own canonical injection prompts). Sanitise to
…and 7 more.
Suggested fix · we looked for these and found none
- adversarial_eval_present
PARTIALSecurity and resilience evaluated NIST AI RMF, Art MEASURE 2.7 LLM2/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.60 (2/3 rules matched). LLM judge (confidence 0.72): Prompt injection defences and rate limiting are documented and implemented (2 of 3 rules matched, 60% composite score), but adversarial evaluation is absent—a critical gap for demonstrating comprehensive security and resilience evaluation as required by MEASURE 2.7. The system has defensive controls but lacks the evaluation rigor needed for full compliance.
▸Evidence · 12 hits— click to view code
(hallucination, prompt injection, output bias, leakage, capability escalation), mitigation owner, status. Wire eval suite into CI
ic: - rule: prompt_injection_defences weight: 0.4 description: | Code includes prompt-injection miti
Code includes prompt-injection mitigations: output filters, input sanitisation, instruction-data segregation, system-promp
Eval suite includes prompt-injection / adversarial cases" score_mapping: ">=0.85": 4 ">=0.65": 3 ">=0.40": 2 ">=
hint: | Add a prompt-injection eval set (e.g. from `promptbench`, `garak`, or your own canonical injection prompts). Sanitise to
…and 7 more.
Suggested fix · we looked for these and found none
- adversarial_eval_present
STRONGResources for AI systems ISO/IEC 42001, Art A.72/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
ADEQUATESecurity of processing GDPR, Art 322/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.80 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 6 hits— click to view code
https://github.com/${owner}/${repo}`,https://github.com/owner/repo
https://github.com/${body.source.owner}/${body.source.repo}`,https://github.com/${body.source.owner}/${body.source.repo}`,…and 1 more.