Audit Report
gpt-engineer-org/gpt-engineer
a90fcd543eed · ran in 33.4s · bundle 28ffdb86
Overall score
1.7 /4
Partial
Risk class
HIGH
4
Code passed
13 / 45
Attestation Yes
0
Outstanding ext.
1
CODE-CHECKED CLAUSES
- Strong11
- Adequate2
- Partial11
- Inadequate3
- Absent18
ATTESTATION QUESTIONS
- Yes0
- No0
- Not Applicable0
- Outstanding0
Harm
Don't hurt people
1.6/4
Partial
Truth
Don't deceive people
1.5/4
Partial
Responsibility
Don't abuse power
1.5/4
Inadequate
Order
Don't destabilize society
2.2/4
Partial
⚠ 1 outstanding external confirmations — required for a complete Annex IV dossier. Complete now →
Safety
Don't harm people1.1/4Inadequate8 clauses▸
Safety
Don't harm peoplePARTIALRisk management system established, implemented, documented EU AI Act, Art 9 LLM1/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.30 (1/3 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.35): CI/CD evaluation gates provide some risk control mechanism, but absence of documented risk register and threat model leaves critical identification, analysis, and systematic review requirements unmet. Human review of external documentation (design files, risk assessments, governance records) is needed to determine full compliance with the continuous iterative risk management lifecycle.
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- presence_of_risk_register
- presence_of_threat_model
PARTIALAppropriate level of accuracy declared and tested EU AI Act, Art 15(1) LLM2/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.70 (2/3 rules matched). LLM judge (confidence 0.72): The system demonstrates evaluation infrastructure (eval_suite_present, eval_in_ci) but lacks documented metrics specifications required by Article 15(1). Without explicit accuracy, robustness, and cybersecurity metrics documentation, compliance with 'appropriate level' and 'consistent performance' requirements cannot be fully verified.
▸Evidence · 3 hits— click to view code
Suggested fix · we looked for these and found none
- metrics_documented
INADEQUATEResilience to errors, faults, inconsistencies EU AI Act, Art 15(4)0/3 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.16 (0/3 rules matched).
▸Evidence · 2 hits— click to view code
LangChain
Anthropic SDK
Suggested fix · we looked for these and found none
- error_handling_at_tool_boundaries
- retry_logic
- fallback_behaviour
ABSENTRisks and benefits to people identified NIST AI RMF, Art MAP 3.40/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_risk_register
- presence_of_threat_model
PARTIALAI system performance evaluated and documented NIST AI RMF, Art MEASURE 2.3 LLM2/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.70 (2/3 rules matched). LLM judge (confidence 0.72): Evaluation suite and CI integration are present (0.7 raw score), but absence of documented metrics is a material gap for demonstrating measured performance criteria as required by the clause. Documentation of measures is explicitly mandated and currently missing.
▸Evidence · 3 hits— click to view code
Suggested fix · we looked for these and found none
- metrics_documented
ABSENTAI risk assessment process ISO/IEC 42001, Art 6.10/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_risk_register
- risk_assessment_methodology_documented
PARTIALOperational planning and control ISO/IEC 42001, Art 8.1 skip1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- presence_of_runbook
ABSENTData protection impact assessment (DPIA) GDPR, Art 350/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_dpia
Privacy
Respect boundaries2.4/4Partial6 clauses▸
Privacy
Respect boundariesSTRONGUntargeted facial image scraping for face databases EU AI Act, Art 5(1)(e)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_face_image_scraping
STRONGEmotion recognition in workplace and education EU AI Act, Art 5(1)(f)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_emotion_recognition_in_workplace_education
EXTERNALReal-time remote biometric identification in public spaces EU AI Act, Art 5(1)(h)EXT▸
Why we flagged it
Deployment context (public space, real-time, law enforcement use, judicial authorisation) is operational, not knowable from code. Always external.
ABSENTData and data governance practices documented EU AI Act, Art 100/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.10 (0/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
""" response = requests.post(url, json=extra_arguments) return response if __name__ == "__main__": URL_BASE = "http://127.0.0
Suggested fix · we looked for these and found none
- presence_of_data_card
- data_loading_code_quality
- bias_evaluation_present
ABSENTPrivacy risk of the AI system evaluated NIST AI RMF, Art MEASURE 2.80/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched).
Suggested fix · we looked for these and found none
- pii_redaction_present
- privacy_documentation
STRONGData protection by design and by default GDPR, Art 252/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
hashlib.sha256
Transparency
Don't deceive people1.5/4Partial4 clauses▸
Transparency
Don't deceive peopleSTRONGSubliminal techniques distorting behaviour EU AI Act, Art 5(1)(a)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_manipulative_prompt_patterns
ABSENTTransparent operation and instructions for use EU AI Act, Art 130/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.13 (0/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 3 hits— click to view code
1 required sections present
Suggested fix · we looked for these and found none
- presence_of_deployer_instructions
- output_interpretation_guidance
- limitations_section_present
PARTIALUsers informed they are interacting with an AI EU AI Act, Art 50(1) LLM1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.40 (1/2 rules matched). LLM judge (confidence 0.45): The system fails explicit disclosure in user-facing strings (README.md) but demonstrates non-human persona design in core prompts, creating ambiguity about whether a reasonably informed user would recognize AI interaction. Human judgment needed on whether prompt-level design choices substitute for explicit disclosure statements.
▸Evidence · 2 hits— click to view code
Suggested fix · we looked for these and found none
- ai_disclosure_in_user_facing_strings
ABSENTPrinciples relating to processing of personal data GDPR, Art 50/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_privacy_policy
- purpose_limitation_documented
Auditability
Actions must be traceable1.5/4Partial8 clauses▸
Auditability
Actions must be traceableABSENTTechnical documentation drawn up before placing on market EU AI Act, Art 110/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.07 (0/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
1 required sections present
Suggested fix · we looked for these and found none
- presence_of_model_card
- readme_quality
- architecture_docs
PARTIALAutomatic recording of events over the lifetime EU AI Act, Art 12(1) LLM1/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.48 (1/3 rules matched). LLM judge (confidence 0.62): The system demonstrates logging at tool call boundaries (5 evidence points in ai.py showing logger.debug calls around chat completion events) and imports the logging module, but lacks evidence of structured logging framework and persistent storage sink configuration. Automatic recording occurs at critical points but completeness across system lifetime and durability guarantees remain unverified.
▸Evidence · 7 hits— click to view code
odel_name) logger.debug(f"Using model {self.model_name}") def start(self, system: str, user: Any, *, step_name: str) -> List[Mt=prompt)) logger.debug( "Creating a new chat completion: %s", "\n".join([m.pretty_repr() for m in messages
d(response) logger.debug(f"Chat completion finished: {messages}") return messages @backoff.on_exception(backoff.expo,t=prompt)) logger.debug(f"Creating a new chat completion: {messages}") msgs = self.serialize_messages(messages) py=response)) logger.debug(f"Chat completion finished: {messages}") return messages…and 2 more.
Suggested fix · we looked for these and found none
- structured_logging_imported
- logging_persistent_sink
PARTIALLogging ensures traceability appropriate to risk EU AI Act, Art 12(2) skip2/3 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.57 (2/3 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 8 hits— click to view code
difflib import json import logging import os import platform import subprocess import sys from pathlib import Path import openai import ty
ations import json import logging import os from pathlib import Path from typing import Any, List, Optional, Union import backoff import
odel_name) logger.debug(f"Using model {self.model_name}") def start(self, system: str, user: Any, *, step_name: str) -> List[Mt=prompt)) logger.debug( "Creating a new chat completion: %s", "\n".join([m.pretty_repr() for m in messages
d(response) logger.debug(f"Chat completion finished: {messages}") return messages @backoff.on_exception(backoff.expo,…and 3 more.
Suggested fix · we looked for these and found none
- logs_include_request_id
INADEQUATEContext of use established and understood NIST AI RMF, Art MAP 1.10/2 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.15 (0/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
1 required sections present
Suggested fix · we looked for these and found none
- intended_use_documented
- deployment_context_documented
INADEQUATEPost-deployment monitoring, appeal and override, change management NIST AI RMF, Art MANAGE 4.1 skip1/4 rules1/4
▸
1/4
Why we flagged it
Composite raw score 0.36 (1/4 rules matched).
▸Evidence · 6 hits— click to view code
AIMessage, HumanMessage, SystemMessage, messages_from_dict, messages_to_dict, ) from langchain_anthropic import ChatAnt
= Union[AIMessage, HumanMessage, SystemMessage] # Set up logging logger = logging.getLogger(__name__) class AI: """ A class that
ystem), HumanMessage(content=user), ] return self.next(messages, step_name=step_name) def _extract_content(
messages.append(HumanMessage(content=prompt)) logger.debug( "Creating a new chat completion: %s", "\n".
e(content="Hello"), HumanMessage(content="How's the weather?")] >>> response = backoff_inference(messages) """ retur
…and 1 more.
Suggested fix · we looked for these and found none
- feedback_capture_present
- structured_logging_imported
- versioning_visible
ADEQUATEDocumented information for the AI management system ISO/IEC 42001, Art 7.5 skip1/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.70 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 2 hits— click to view code
Suggested fix · we looked for these and found none
- presence_of_versioned_docs
ADEQUATEMonitoring, measurement, analysis and evaluation ISO/IEC 42001, Art 9.12/2 rules3/4
▸
3/4
Why we flagged it
Composite raw score 0.75 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 3 hits— click to view code
difflib import json import logging import os import platform import subprocess import sys from pathlib import Path import openai import ty
ations import json import logging import os from pathlib import Path from typing import Any, List, Optional, Union import backoff import
ABSENTRecords of processing activities GDPR, Art 300/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_processing_register
Accountability
Don't abuse power2.0/4Partial6 clauses▸
Accountability
Don't abuse powerPARTIALDeployer log-retention capability supported EU AI Act, Art 26(6) LLM1/1 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/1 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.35): Logging imports are present, confirming technical capability for log generation, but evidence does not demonstrate: (1) automatic log generation by the AI system itself, (2) logs under deployer control, (3) retention policies meeting the six-month minimum, or (4) exported logs accessibility. Deterministic rule only confirms exportability potential; actual compliance requires documentation of retention procedures and log access mechanisms.
▸Evidence · 2 hits— click to view code
difflib import json import logging import os import platform import subprocess import sys from pathlib import Path import openai import ty
ations import json import logging import os from pathlib import Path from typing import Any, List, Optional, Union import backoff import
ABSENTRisk management process documented and accountable NIST AI RMF, Art GOVERN 1.40/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_risk_register
- risk_owner_assignment
PARTIALOngoing monitoring and periodic review of risk management NIST AI RMF, Art GOVERN 1.5 LLM1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/2 rules matched). Supporting docs may exist outside the repo. LLM judge (confidence 0.45): CI/CD evaluation gates demonstrate some monitoring capability, but the absence of documented drift monitoring and lack of evidence for defined organizational roles, responsibilities, and periodic review frequency prevents full compliance with GOVERN 1.5's comprehensive requirements.
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- drift_monitoring_present
ABSENTLeadership and commitment for AI management ISO/IEC 42001, Art 5.10/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- presence_of_ai_policy
- leadership_signoff_evidence
STRONGRoles, responsibilities and authorities ISO/IEC 42001, Art 5.31/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
STRONGInternal organization controls ISO/IEC 42001, Art A.51/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (1/1 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Human Oversight
Humans stay in control0.8/4Inadequate5 clauses▸
Human Oversight
Humans stay in controlABSENTEffective human oversight designed and built-in EU AI Act, Art 14(1) LLM1/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.45 (1/3 rules matched). LLM judge (confidence 0.85): While human-machine message handling is present (HumanMessage/AIMessage structures), the evidence shows no actual oversight UI, dry-run capabilities, or documented mechanisms for humans to effectively intervene during system operation. Message passing alone is insufficient for 'effective oversight' as required by Art. 14(1); the system lacks the control interface tools necessary for active human supervision.
▸Evidence · 6 hits— click to view code
AIMessage, HumanMessage, SystemMessage, messages_from_dict, messages_to_dict, ) from langchain_anthropic import ChatAnt
= Union[AIMessage, HumanMessage, SystemMessage] # Set up logging logger = logging.getLogger(__name__) class AI: """ A class that
ystem), HumanMessage(content=user), ] return self.next(messages, step_name=step_name) def _extract_content(
messages.append(HumanMessage(content=prompt)) logger.debug( "Creating a new chat completion: %s", "\n".
e(content="Hello"), HumanMessage(content="How's the weather?")] >>> response = backoff_inference(messages) """ retur
…and 1 more.
Suggested fix · we looked for these and found none
- oversight_ui_present
- tool_calls_have_dry_run
ABSENTInterrupt / stop function reachable by overseer EU AI Act, Art 14(4)(d)0/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched).
Suggested fix · we looked for these and found none
- kill_switch_present
- graceful_shutdown_handler
ABSENTAbility to override / reverse the system's output EU AI Act, Art 14(4)(e) LLM1/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.60 (1/2 rules matched). LLM judge (confidence 0.85): Evidence shows human-in-loop message passing infrastructure (override_path_present matched), but repository lacks concrete implementation of decision reversal mechanisms or UI/API endpoints enabling humans to actually override or reverse AI outputs. The matched rule alone (0.6 weight) is insufficient without addressable decision structures (0.4 weight) required by Article 14(4)(e).
▸Evidence · 6 hits— click to view code
AIMessage, HumanMessage, SystemMessage, messages_from_dict, messages_to_dict, ) from langchain_anthropic import ChatAnt
= Union[AIMessage, HumanMessage, SystemMessage] # Set up logging logger = logging.getLogger(__name__) class AI: """ A class that
ystem), HumanMessage(content=user), ] return self.next(messages, step_name=step_name) def _extract_content(
messages.append(HumanMessage(content=prompt)) logger.debug( "Creating a new chat completion: %s", "\n".
e(content="Hello"), HumanMessage(content="How's the weather?")] >>> response = backoff_inference(messages) """ retur
…and 1 more.
Suggested fix · we looked for these and found none
- decisions_are_addressable
ABSENTMechanisms to supersede or deactivate AI systems NIST AI RMF, Art MANAGE 2.30/2 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/2 rules matched).
Suggested fix · we looked for these and found none
- kill_switch_present
- feature_flag_for_disable
STRONGAutomated individual decision-making, including profiling GDPR, Art 222/2 rules4/4
▸
4/4
Why we flagged it
Composite raw score 1.00 (2/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 12 hits— click to view code
AIMessage, HumanMessage, SystemMessage, messages_from_dict, messages_to_dict, ) from langchain_anthropic import ChatAnt
= Union[AIMessage, HumanMessage, SystemMessage] # Set up logging logger = logging.getLogger(__name__) class AI: """ A class that
ystem), HumanMessage(content=user), ] return self.next(messages, step_name=step_name) def _extract_content(
messages.append(HumanMessage(content=prompt)) logger.debug( "Creating a new chat completion: %s", "\n".
e(content="Hello"), HumanMessage(content="How's the weather?")] >>> response = backoff_inference(messages) """ retur
…and 7 more.
Fairness
Treat people fairly3.2/4Adequate5 clauses▸
Fairness
Treat people fairlySTRONGExploiting vulnerabilities (age, disability, socio-economic) EU AI Act, Art 5(1)(b)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_protected_attribute_targeting
STRONGSocial scoring leading to detrimental treatment EU AI Act, Art 5(1)(c)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched). Supporting docs may exist outside the repo.
Suggested fix · we looked for these and found none
- detect_scoring_with_persistent_user_state
STRONGPredictive policing solely from profiling EU AI Act, Art 5(1)(d)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_crime_risk_scoring_from_profile
STRONGBiometric categorisation by protected attributes EU AI Act, Art 5(1)(g)0/1 rules4/4
▸
4/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- detect_biometric_categorisation_by_protected_attrs
ABSENTFairness and bias evaluated NIST AI RMF, Art MEASURE 2.110/1 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.00 (0/1 rules matched).
Suggested fix · we looked for these and found none
- bias_evaluation_present
Security & Governance
Don't destabilize society1.0/4Inadequate4 clauses▸
Security & Governance
Don't destabilize societyABSENTCybersecurity measures appropriate to circumstances EU AI Act, Art 15(5)1/4 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.12 (1/4 rules matched).
▸Evidence · 3 hits— click to view code
ackoff.expo, openai.RateLimitError, max_tries=7, max_time=45) def backoff_inference(self, messages): """ Perform inferen
openai.error.RateLimitError If the number of retries exceeds the maximum or if the rate limit persists beyond the
ultimately raise a RateLimitError. Example ------- >>> messages = [SystemMessage(content="Hello"), HumanMessage(co
Suggested fix · we looked for these and found none
- prompt_injection_defences
- secrets_not_in_prompts
- adversarial_eval_present
ABSENTSecurity and resilience evaluated NIST AI RMF, Art MEASURE 2.71/3 rules0/4
▸
0/4
Why we flagged it
Composite raw score 0.12 (1/3 rules matched).
▸Evidence · 3 hits— click to view code
ackoff.expo, openai.RateLimitError, max_tries=7, max_time=45) def backoff_inference(self, messages): """ Perform inferen
openai.error.RateLimitError If the number of retries exceeds the maximum or if the rate limit persists beyond the
ultimately raise a RateLimitError. Example ------- >>> messages = [SystemMessage(content="Hello"), HumanMessage(co
Suggested fix · we looked for these and found none
- prompt_injection_defences
- adversarial_eval_present
PARTIALResources for AI systems ISO/IEC 42001, Art A.7 skip1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 1 hit— click to view code
Suggested fix · we looked for these and found none
- presence_of_security_policy
PARTIALSecurity of processing GDPR, Art 32 skip1/2 rules2/4
▸
2/4
Why we flagged it
Composite raw score 0.50 (1/2 rules matched). Supporting docs may exist outside the repo.
▸Evidence · 4 hits— click to view code
https://gptengineerezm.dataplane.rudderstack.com
https://xx.openai.azure.com).
https://api.gptengineer.app/openapi.json
https://github.com/langchain-ai/langchain/blob/535db72607c4ae308566ede4af65295967bb33a8/libs/community/langchain_community/callbacks/openai_info.py
Suggested fix · we looked for these and found none
- access_control_enforcement