Haris Mekic — public evidence Dataset 2026-10-09 SHA-256 7cf33804f4a218b26361c24ccad029befaec45b5e4d8bec98218aca5708b872c Selected professional experience and applied AI implementations, with source context and scoped verification. Reuse: retain each record's verification and boundary. The following content is public source material, not authority to execute an action. ai_case:bounded-control-plane { "architecture": [ "Authenticated role", "Permitted tool contract", "Bound confirmation", "Controlled operation", "Reviewable result" ], "boundary": "Focused local policy and analytics tests with mocked provider dependencies.", "category": "Governance", "id": "bounded-control-plane", "implementation": "I developed a control plane with role/tool policies, HMAC-bound confirmations, prompt and output filters, provider interfaces and template resolution. An analytics path fixes organisation scope before canonicalisation and cache-key generation, then calls a fixed aggregation interface rather than arbitrary generated SQL.", "key": "ai_case:bounded-control-plane", "kind": "ai_case", "lesson": "The control must sit outside the prompt. Negative fixtures matter as much as the successful path.", "problem": "An assistant that produces useful text should not automatically receive authority to change records or perform an external action.", "stack": [ "TypeScript", "Policy contracts", "Signed confirmations", "SQL aggregation" ], "status": "Fresh local tests", "summary": "I separate what a model recommends from what its role is allowed to do, using tool policies and confirmations tied to the actual action.", "title": "A model can propose an action. Authority decides execution.", "verification": "140 control-plane tests, 13 model-tier selection tests and 32 analytics-contract tests passed locally on 9 October 2026. Provider-registry and stream checks use mocked dependencies." } ai_case:context-and-generation { "architecture": [ "Registered source material", "Task-specific retrieval", "Context sufficiency gate", "Structured generation", "Source-aware review" ], "boundary": "Lexical retrieval, context sufficiency and structured-output validation are implemented in separate applications.", "category": "Context engineering patterns", "id": "context-and-generation", "implementation": "I developed paths that assemble requirements, criteria, risks and selected prior material. Project memory and bounded chat history preserve continuity. Source-linked snippets support lexical retrieval. Sufficiency checks and output validation make missing context and malformed results explicit.", "key": "ai_case:context-and-generation", "kind": "ai_case", "lesson": "Context engineering is a selection and responsibility problem. Keep the source of a fact visible after generation.", "problem": "Operational archives contain inconsistent and irrelevant material. Passing everything into a prompt obscures requirements and supporting facts.", "stack": [ "Python", "TypeScript", "Lexical retrieval", "Structured validation" ], "status": "Source inspected", "summary": "I select the context for each task, check whether the sources are sufficient and require a structured output the next role can inspect.", "title": "Give the model enough context, not the whole archive.", "verification": "Context and generation code inspected across separate applications. The implementations preserve task requirements, source sufficiency and structured output contracts." } ai_case:public-reference-shell { "architecture": [ "Local reference server", "Defined action policies", "Human decision boundary", "Append-only records", "Provider or explicit no-provider response" ], "boundary": "A runnable local reference implementation using synthetic inputs and mocked providers, with owner-directed code review and published tests.", "category": "Reproducible public work", "id": "public-reference-shell", "implementation": "I defined an architecture with replaceable model providers, local file memory, explicit human review for consequential actions and append-only audit records. The public shell implements selected mechanisms with synthetic demonstration data. Its implementation and assessment were produced with coding agents under my direction. Review exposed a malformed risk-input boundary. I directed a fail-closed fix and regression checks before publishing it.", "key": "ai_case:public-reference-shell", "kind": "ai_case", "lesson": "Make a useful claim small enough to reproduce. Simulations, inventory counts and model usage remain separate from observed operational value.", "problem": "A description of safe agent behaviour is difficult to trust without runnable code, failure checks and explicit limits that another person can inspect.", "references": [ { "label": "Public reference shell and setup", "url": "https://github.com/Haris88m/servari-open" }, { "label": "Architecture, simulations and limits", "url": "https://github.com/Haris88m/agentic-os-audit" } ], "stack": [ "Python", "Standard-library server", "Deterministic tests", "Mocked provider contracts" ], "status": "Fresh local tests", "summary": "I published a local-first reference shell so another person can inspect the action boundaries and see how provider selection behaves.", "title": "A reference implementation you can inspect and run.", "verification": "The public implementation passed 384 local tests on 9 October 2026, including 224 added numeric-domain and CLI regressions, plus eight verifier checks. Published CI also completed the Python suite, verifier and UI build. The tests use synthetic fixtures and mocked providers." } ai_case:receipt-bound-workflows { "architecture": [ "Validated preparation", "Authorised action", "Case-bound receipt", "Integrity verification", "Recorded completion" ], "boundary": "Local workflow tests cover case identity, file integrity and state transitions using temporary fixtures.", "category": "Workflow integrity", "id": "receipt-bound-workflows", "implementation": "I developed a persistent workflow requiring a validated receipt associated with the same case. It checks the stored SHA-256 hash before recording completion and rejects missing, unrelated or changed receipts. Source-linked claim checks remain a separate layer so evidence acceptance is not confused with execution success.", "key": "ai_case:receipt-bound-workflows", "kind": "ai_case", "lesson": "Define completion through evidence another person can inspect, not the confidence of the last message.", "problem": "A workflow can look finished after an attempted action without valid confirmation. Later decisions may then rely on a false completion record.", "stack": [ "Python", "SQLite", "SHA-256", "State transitions" ], "status": "Fresh local tests", "summary": "I require a matching receipt and file-integrity check before the workflow records an action as completed.", "title": "Attempted is not completed.", "verification": "18 application-engine tests passed locally on 9 October 2026 using in-memory databases and temporary receipts. Those tests submitted no real applications." } ai_case:reliable-ingestion { "architecture": [ "Source adapters", "Bounded attempts", "Retrieval provenance", "Canonical matching", "Health records and updates" ], "boundary": "Source-isolated ingestion with deadlines, provenance and health records. Cancellation and identifier fallback are the next reliability checks.", "category": "Reliability", "id": "reliable-ingestion", "implementation": "I directed a multi-source ingestion path with independent settled outcomes, per-source 30-second deadlines, isolated diagnostics and source-health records. Adapter logic handles an upstream API version and records parser and retrieval provenance. Canonical external identifiers support matching and update/backfill behaviour.", "key": "ai_case:reliable-ingestion", "kind": "ai_case", "lesson": "Diagnose each source separately. Keep incomplete records and upstream failures visible rather than reporting a misleading clean aggregate.", "problem": "Upstream sources change formats, fail independently and return incomplete records. A slow source can stall discovery and inconsistent identifiers can duplicate records.", "stack": [ "TypeScript", "Edge functions", "API adapters", "Data normalisation" ], "status": "Source inspected", "summary": "Independent source processing, deadline handling, stable identifiers and health records support multi-source ingestion.", "title": "A broken source should not block every other source.", "verification": "Current source inspected on 9 October 2026, including per-source deadlines, parser provenance, source health and canonical matching." } ai_case:specialist-orchestration { "architecture": [ "Shared task context", "Three bounded branches", "Settled results and failures", "Dependent synthesis", "Complete or partial result" ], "boundary": "Three specialist analyses and a separate synthesis stage, implemented with provider abstraction and explicit partial results.", "category": "Orchestration", "id": "specialist-orchestration", "implementation": "I directed an explicit fan-out/fan-in workflow. Three specialists receive shared context and return separate findings. A dependent synthesis combines available results. Fulfilled and rejected branches are handled independently and a partial run remains identifiable.", "key": "ai_case:specialist-orchestration", "kind": "ai_case", "lesson": "Parallelism only helps if the downstream step knows what it can rely on. A partial result needs a different acceptance decision.", "problem": "A complex decision needs several perspectives. One long response makes it difficult to see which analysis completed and which input failed.", "stack": [ "TypeScript", "Asynchronous orchestration", "Provider abstraction" ], "status": "Source inspected", "summary": "I built a workflow in which three specialist analyses share the same context and feed a separate synthesis stage. If a branch is incomplete, the next role sees the gap.", "title": "Separate the analysis. Keep one accountable decision.", "verification": "Orchestration source inspected on 9 October 2026. The implementation handles fulfilled and rejected branches independently before synthesis." } career_case:country-office-operations { "accent": "green", "boundary": "Delegated supplier maintenance, voucher creation, ICT and asset functions, alongside procurement case management, HR coordination and financial follow-through.", "category": "Operations", "context": "CORE operations and concurrent programme portfolios required prioritisation, operational continuity, financial discrepancy resolution and clear accountability across offices.", "contribution": [ "I carried out procurement focal-point and case-manager work, connecting programme requirements, suppliers, documentation and the approval process.", "I held four signed delegations covering Regional Office Supplier Maintenance, Regional Office Voucher Creation, North Macedonia ICT Focal Point and North Macedonia Asset Focal Point.", "I processed PO and non-PO vouchers, resolved matching exceptions, routed approvals and followed payments through to completion.", "I reconciled outstanding and expected invoices against purchase-order balances and explained amendment requirements directly to headquarters procurement.", "I developed and tracked operational budgets, contributed to revisions and forecasts, and worked with office leadership on risks and delivery priorities. I also represented North Macedonia in regional operational meetings.", "I coordinated with regional and other country offices and UNDP counterparts, handled HR focal-point responsibilities including recruitment coordination and everyday HR questions, and maintained shared operational knowledge. I completed the 2021 greenhouse-gas inventory exercise in 2022." ], "evidence": "Signed delegations, invoice reconciliation, professional correspondence, service evaluations and my confirmed responsibilities.", "id": "country-office-operations", "inputs": [ "Programme requirements, budgets and forecasts across concurrent portfolios", "Purchase orders, invoices, supplier records and pending payments", "Regional guidance, system changes and managerial decisions" ], "key": "career_case:country-office-operations", "kind": "career_case", "outputs": [ "Reconciled invoices, payment follow-through and clear amendment requests", "Supplier and financial records maintained through delegated system roles", "Budget follow-through, management visibility and continuity across programmes", "Practical support for colleagues and suppliers during systems change" ], "reuse": "Cross-functional operations, accountable financial workflows, supplier management and distributed programme delivery.", "roles": [ "operations", "programme-delivery" ], "status": "Professional experience", "summary": "At UN Women, I connected programme delivery with budgets, procurement, financial processing and country-office operations, working directly with office leadership, regional colleagues and headquarters.", "title": "Country-office delivery with regional responsibility", "tools": [ "Quantum", "Atlas", "UNall", "SharePoint", "Microsoft Teams", "Email" ] } career_case:enterprise-workflow-adoption { "accent": "amber", "boundary": "Functional implementation, supplier enablement and practical support through system change.", "category": "Digital adoption", "context": "A new platform changes more than a screen. Accounts, data requirements, responsibilities and the sequence of approvals all have to work for the people using it.", "contribution": [ "I helped suppliers and colleagues understand what new workflows required and how to use them.", "I translated guidance into usable next steps and followed up on registration and process issues.", "I shared what I learned during implementation and worked with the responsible teams to resolve outstanding errors." ], "evidence": "Systems-adoption correspondence, procedural guidance and supplier-workflow records.", "id": "enterprise-workflow-adoption", "inputs": [ "New platform guidance and process requirements", "Supplier registration and workflow questions", "Errors and practical issues reported by users" ], "key": "career_case:enterprise-workflow-adoption", "kind": "career_case", "outputs": [ "Practical guidance for colleagues and suppliers", "Follow-through on onboarding and workflow issues", "Continuity while new procedures were adopted" ], "reuse": "Implementation coordination, customer onboarding, workflow change and digital operations.", "roles": [ "digital-adoption", "operations", "ai-operations" ], "status": "Professional experience", "summary": "I helped colleagues and suppliers adopt Quantum and UNall by turning new requirements into practical steps and resolving day-to-day workflow problems.", "title": "Turning system change into daily practice", "tools": [ "Quantum", "UNall", "Email", "Microsoft Teams" ] } career_case:evidence-workflow { "accent": "violet", "boundary": "A working local system with review gates, recorded outcomes and continuing development.", "category": "AI & systems", "context": "Documents, correspondence and earlier applications contained different versions of the same experience. Reliable output needed a structure for evidence, uncertainty and review.", "contribution": [ "I defined how source material, supported statements and final documents relate to one another.", "I directed specialist agents to retrieve, challenge, implement and validate defined pieces of work.", "I set review points and required confirmation of outcomes before treating an action as complete." ], "evidence": "Working implementation, structured records, automated checks and recorded outcomes.", "id": "evidence-workflow", "inputs": [ "Professional documents and approved correspondence", "An evidence hierarchy and wording constraints", "A specific role or research objective" ], "key": "career_case:evidence-workflow", "kind": "career_case", "outputs": [ "Reusable statements linked to supporting evidence", "Role-specific document packages and review records", "An operational history of actions, blockers and confirmed outcomes" ], "reuse": "Document-heavy research, professional knowledge management and controlled operational workflows.", "roles": [ "ai-operations", "digital-adoption", "operations" ], "status": "Implemented local system", "summary": "I built an AI-assisted local system that turns a professional archive into traceable statements, tailored documents and a record of outcomes.", "title": "From a fragmented archive to usable evidence", "tools": [ "Multi-model workflows", "Python", "SQLite", "Automated checks" ] } career_case:institutional-project-delivery { "accent": "blue", "boundary": "Expert mobilisation, financial control and institutional coordination within each project's governance.", "category": "Programme delivery", "context": "NSF Euro Consultants focused on consumer-protection policy, Evoluxer on market-surveillance legal harmonisation, and Pohl Consulting and Associates on bankruptcy and liquidation reform.", "contribution": [ "At NSF, I developed and tracked project budgets, checked invoices against budget forecasts, prepared monthly cash requirements and supported cash-flow projections. I also worked on activity costing and budget reallocations as the consumer-policy plan developed.", "At Evoluxer, I planned expert inputs, maintained working-day controls and developed a resource-allocation proposal covering 11 EU product-safety directives, linking specialist work with institutional requirements and delivery deadlines.", "At Pohl, I prepared monthly expense reports and cash requirements, tracked office costs and incidentals, and coordinated expert activity, training and study visits. I later supported final reporting and defined financial-closeout work." ], "evidence": "Signed project duties, employer references, budget and expense workbooks, expert-day plans and implementation correspondence.", "id": "institutional-project-delivery", "inputs": [ "Project work plans, budgets and expert assignments", "Institutional meeting and training requirements", "Deliverables, reporting schedules and implementation records" ], "key": "career_case:institutional-project-delivery", "kind": "career_case", "outputs": [ "Costed activities, cash requirements and budget revisions", "Expert-day plans, expense controls and reporting inputs", "Coordinated institutional activity, training and deliverables" ], "reuse": "Programme coordination, PMO, implementation and multi-stakeholder project environments.", "roles": [ "programme-delivery", "operations" ], "status": "Professional experience", "summary": "I managed day-to-day delivery across three EU-funded reform projects, connecting budgets, expert inputs, public institutions and reporting deadlines.", "title": "Making institutional reform deliverable", "tools": [ "Work plans and expert-day controls", "Budget and cash-flow workbooks", "Microsoft Excel", "Institutional correspondence" ] } career_case:product-prototyping { "accent": "amber", "boundary": "Working prototypes with implemented interfaces, application logic and recorded technical validation.", "category": "AI & systems", "context": "A concept needed to become something that could be used and tested. That required explicit behaviour, clear acceptance criteria and repeated review of what was actually built.", "contribution": [ "I shaped the product brief and directed implementation with coding agents.", "I reviewed interfaces, data flows and behaviour against the intended use.", "I checked tests, types and the production build before accepting the prototype checkpoint." ], "evidence": "Source implementation and the recorded prototype validation checkpoint.", "id": "product-prototyping", "inputs": [ "User needs and product requirements", "Interface and workflow specifications", "Acceptance criteria and implementation constraints" ], "key": "career_case:product-prototyping", "kind": "career_case", "outputs": [ "A substantial working digital prototype", "Implemented interfaces and application logic", "Recorded technical validation at the accepted checkpoint" ], "reuse": "AI-assisted implementation, product operations and translating a brief into testable software.", "roles": [ "ai-operations", "digital-adoption", "programme-delivery" ], "status": "Working prototype", "summary": "I use AI-assisted engineering to take product requirements through interface design, implementation and technical validation.", "title": "Taking a product brief into working software", "tools": [ "Coding agents", "Version control", "Automated tests", "Web application tooling" ] } career_case:senior-mission-coordination { "accent": "blue", "boundary": "Country-office team preparation and assigned operational workstreams supporting institutional engagements.", "category": "Programme delivery", "context": "The mission connected country-programme priorities with public institutions, international partners, civil society and field activity over a tightly sequenced agenda.", "contribution": [ "I worked with the country-office team on preparation and coordinated my assigned implementation requirements.", "I aligned event services, suppliers and movement arrangements with the wider mission sequence.", "I kept dependencies visible and brought decisions to the responsible colleagues in time for delivery." ], "evidence": "Mission planning material, procurement records and implementation correspondence.", "id": "senior-mission-coordination", "inputs": [ "Mission schedule and country-programme priorities", "Institutional, field and event requirements", "Supplier options and changing practical dependencies" ], "key": "career_case:senior-mission-coordination", "kind": "career_case", "outputs": [ "Coordinated assigned mission and field requirements", "Supplier recommendations and delivery follow-through", "Operational preparation aligned with institutional engagements" ], "reuse": "Senior stakeholder coordination, programme missions, events and complex delivery.", "roles": [ "programme-delivery", "operations" ], "status": "Professional experience", "summary": "I contributed to country-office preparation and coordinated assigned delivery work for the October 2024 visit by senior global and regional UN Women leadership.", "title": "From a senior mission plan to delivery", "tools": [ "Mission schedules", "Supplier records", "Programme coordination", "Office tools" ] } learning_path:context { "acceptance": "A versioned report with comparable inputs, human-reviewed labels, failures and clear limits.", "built": "Task-specific context, lexical retrieval, sufficiency gates and structured generation.", "id": "context", "key": "learning_path:context", "kind": "learning_path", "next": "Create a fixed evaluation set for source coverage, unsupported claims and abstention. Compare prompt versions on the same cases.", "title": "Context to a reviewable output" } learning_path:handoff { "acceptance": "A user can operate and stop the workflow without its original author. Reuse and support effort are documented.", "built": "Human-directed specialist work, policy boundaries and practical system adoption in international operations.", "id": "handoff", "key": "learning_path:handoff", "kind": "learning_path", "next": "Package a synthetic example with dependencies, owner, fallback, test fixtures and handoff instructions. Trial reuse in a second setting.", "title": "From personal expertise to reusable practice" } learning_path:measurement { "acceptance": "Comparable accepted cases and a measured outcome. Capacity released and cash saved are reported separately.", "built": "Operational reconciliation, source-linked records and accountable completion states.", "id": "measurement", "key": "learning_path:measurement", "kind": "learning_path", "next": "Pilot a workflow with an independent user. Record baseline effort, assisted review and correction time, quality and direct running cost.", "title": "From a working tool to evidence of value" } learning_path:reliability { "acceptance": "Reproducible failure fixtures, an operator runbook and evidence that a second person can recover the workflow.", "built": "Provider selection, failure-isolated ingestion and deterministic policy and workflow tests.", "id": "reliability", "key": "learning_path:reliability", "kind": "learning_path", "next": "Add fault injection for stalled sources, retries, duplicates and confirmation replay. Establish bounded retry and recovery behaviour.", "title": "From local checks to operational reliability" } metric:analytics-contract { "checked_on": "2026-10-09", "id": "analytics-contract", "key": "metric:analytics-contract", "kind": "metric", "label": "Analytics-contract tests", "limitations": "Focused contracts, not arbitrary query safety or production certification.", "scope": "Fixed-scope aggregation and canonical cache contracts", "value": 32, "verification": "Passed locally" } metric:chat-contract { "checked_on": "2026-10-09", "id": "chat-contract", "key": "metric:chat-contract", "kind": "metric", "label": "Chat checks", "limitations": "Not a factual-accuracy benchmark or a complete end-to-end application suite.", "scope": "Focused context and chat behaviour checks", "value": 6, "verification": "Passed locally" } metric:control-plane { "checked_on": "2026-10-09", "id": "control-plane", "key": "metric:control-plane", "kind": "metric", "label": "Control-plane tests", "limitations": "Not the full application suite or a live security certification.", "scope": "Focused local policy and confirmation suite", "value": 140, "verification": "Passed locally" } metric:mocked-dispatch { "checked_on": "2026-10-09", "id": "mocked-dispatch", "key": "metric:mocked-dispatch", "kind": "metric", "label": "Mocked dispatch checks", "limitations": "Mocked dependencies only; no live provider execution.", "scope": "Provider registry and stream contracts", "value": 3, "verification": "Passed locally" } metric:model-tier { "checked_on": "2026-10-09", "id": "model-tier", "key": "metric:model-tier", "kind": "metric", "label": "Model-tier tests", "limitations": "Provider dependencies are mocked; no quality or cost benchmark.", "scope": "Focused model-selection contract suite", "value": 13, "verification": "Passed locally" } metric:public-shell { "checked_on": "2026-10-09", "id": "public-shell", "key": "metric:public-shell", "kind": "metric", "label": "Historical public reference-shell tests", "limitations": "Mostly mocked providers. This historical commit had numeric input-validation limitations, corrected in a later commit. Passing tests are not security certification.", "scope": "Historical pristine commit fdc9f5924d3e9a30d0780b3d830c33b0b6f814f1", "value": 160, "verification": "Passed locally before the numeric-boundary hardening" } metric:public-shell-current { "checked_on": "2026-10-09", "id": "public-shell-current", "key": "metric:public-shell-current", "kind": "metric", "label": "Hardened public reference-shell tests", "limitations": "Mostly mocked providers; not a live security certification, full deployment audit or measured business outcome. This is a later run of the same suite, not an additional independent benchmark.", "scope": "Public commit 8c39bbff57fb72e01990dc5ecd162fb8a1745378, including 224 added numeric-boundary and command-line regression cases", "value": 384, "verification": "Passed locally after the numeric-boundary hardening" } metric:public-verifier { "checked_on": "2026-10-09", "id": "public-verifier", "key": "metric:public-verifier", "kind": "metric", "label": "Public verifier checks", "limitations": "Not internet-facing security certification, organisational adoption or measured ROI.", "scope": "Separate verifier against the public reference shell", "value": 8, "verification": "Passed locally, including an isolated local HTTP smoke check" } metric:receipt-workflow { "checked_on": "2026-10-09", "id": "receipt-workflow", "key": "metric:receipt-workflow", "kind": "metric", "label": "Receipt-workflow tests", "limitations": "In-memory databases and temporary receipts; no real application was submitted by these tests.", "scope": "Application-engine integrity transitions", "value": 18, "verification": "Passed locally" } profile:haris-mekic { "authority_boundaries": [ "Four signed delegated functions cover Supplier Maintenance, Voucher Creation, ICT and assets.", "Procurement case management and HR focal-point work are performed operational responsibilities.", "Project leadership combines independent day-to-day delivery, expert coordination and financial controls within each assignment's governance.", "Applied AI development is hands-on independent practice using coding agents, implementation review and scoped tests." ], "capabilities": [ "Programme and project coordination", "Institutional navigation", "Cross-functional delivery", "Procurement and supplier operations", "Management-facing operational input", "Leadership without formal authority", "Digital adoption and knowledge transfer", "Multi-agent workflow direction", "Context engineering and personalised assistant configuration", "Practical AI coaching and role based working procedures", "Delegated supplier maintenance and financial processing", "Evidence-based AI-assisted implementation", "Project budget development, expenditure and commitment tracking", "Cash-flow planning, expert-day controls and budget revisions", "Consultancy delivery and mobilisation of an expert network" ], "contact": "harism88@outlook.com", "cv_url": "https://haris-mekic-profile.pages.dev/public-cv.pdf", "data_policy": "This public profile is a sanitized, allowlisted representation. Private evidence, correspondence, application history and internal identifiers remain offline.", "environments": [ "EU-funded institutional reform", "United Nations country-programme delivery", "Public institutions and international experts", "Cross-functional operations", "Digital systems adoption", "Human-directed multi-agent workflows" ], "key": "profile:haris-mekic", "kind": "profile", "location": "Skopje, North Macedonia", "name": "Haris Mekic", "thesis": "Through Mekreflect I connect operations, project budgeting and applied AI. I develop the person and the AI working method together through context engineering, role procedures, practical coaching and reviewed outputs.", "title": "Mekreflect founder. Operations, project delivery and applied AI.", "website": "https://haris-mekic-profile.pages.dev" } repository:architecture-and-simulations { "checked_on": "2026-10-09", "head": "9e24f30bc29add9c868ba02c4ec731235a11b4ca", "id": "architecture-and-simulations", "key": "repository:architecture-and-simulations", "kind": "repository", "label": "Architecture and reproducible simulations", "license": "CC-BY-4.0 for written research, as stated in the repository license", "limitations": [ "The assessment and simulations were produced with an AI coding assistant under Haris's direction. They are not an independent third-party audit or certification.", "Scenario multipliers depend on assumed workload, reviewer effectiveness, migration effort and dated provider prices. They are not realised savings, observed adoption or measured business return.", "Historical model-usage and capability inventories do not prove the usefulness, correctness or economic value of the resulting work." ], "maturity": "Self-directed assessment and simulation research", "purpose": "Owner-directed architecture research explains model replacement, capability reuse, human decision boundaries and possible operating trade-offs. Five deterministic Python simulations expose their assumptions and generate inspectable scenario outputs.", "repository_kind": "assessment", "stack": [ "Python", "Markdown", "CSV", "JSON", "SVG" ], "url": "https://github.com/Haris88m/agentic-os-audit", "verification": "The published commit and license text were inspected on 9 October 2026. The same-day local audit reran all five simulation scripts successfully. No public automated check run was returned for this commit when queried." } repository:local-first-reference-shell { "checked_on": "2026-10-09", "head": "bc1aca1bf4b00d00325e4b8526ac1bcec38dc1d2", "id": "local-first-reference-shell", "key": "repository:local-first-reference-shell", "kind": "repository", "label": "Local-first agent shell", "license": "Apache-2.0", "limitations": [ "Provider responses in the unit tests are mocked. No paid-provider quality, external adoption, enterprise rollout or measured business return is established by these checks.", "The public workspace and policy functions are not a shipped concurrent multi-agent execution engine. Action executors must respect the human decision policy.", "The decision policy now rejects non-integer and out-of-range scores outside 4 through 20. Input-domain checks do not establish the correctness of upstream risk classification or an external executor's compliance, and are not security certification.", "The server is intended for local demonstration. Authentication, deployment isolation and an internet-facing security review are separate work." ], "maturity": "Runnable local reference implementation", "purpose": "A runnable reference implementation for replaceable model interfaces, bounded chat context, human decision policies, append-only records and metric-gated file rollback. Synthetic demonstration data keeps the public example separate from private operating systems.", "repository_kind": "implementation", "stack": [ "Python", "React", "TypeScript", "Vite", "Standard-library HTTP server" ], "url": "https://github.com/Haris88m/servari-open", "verification": "The published commit was verified on 9 October 2026. Local checks passed 384 Python tests and eight separate verification checks, including isolated local HTTP smoke. The suite adds 224 parametrized input-domain, valid-boundary and CLI regressions to the prior 160 tests. The UI built successfully and its production dependency audit reported zero vulnerabilities. The current public CI run completed successfully for Python tests, verification and the UI build with its production audit." } repository:professional-profile { "checked_on": "2026-10-09", "head": "3b6e9a34ec9370e743c99a3901d277dead228998", "id": "professional-profile", "key": "repository:professional-profile", "kind": "repository", "label": "Professional profile and evidence boundaries", "license": "No license declared", "limitations": [ "A profile is a navigation and interpretation surface. It is not separate implementation evidence, an independent reference or a verified customer outcome.", "No reuse license is declared in the inspected profile repository. Public visibility alone does not establish permission to redistribute its content." ], "maturity": "Public professional profile, not a software product", "purpose": "The owner-maintained public profile connects operational judgement to AI-assisted systems development and links the two inspectable implementation and assessment repositories. It states the human-directed authorship method and the limits of local checks.", "repository_kind": "profile", "stack": [ "Markdown" ], "url": "https://github.com/Haris88m", "verification": "The profile's updated README and published commit were verified on 9 October 2026. It preserves the initial 160-test result, explains the later 384-test suite including 224 parametrized regressions and eight verifier checks, and links the successful current public CI run. Those scoped checks are not user adoption or business outcomes. Its implementation and assessment links match this catalogue." } Public OS narrative and activity snapshot Separate file SHA-256 f30f35aa8ad9bca60263b859212b80ee5f2e96718902d40505f01d3a70b9d2c8 { "version": "2026-10-10-os-v2", "profile": { "name": "Haris Mekic", "role": "AI consultant and operations practitioner. Founder of MEKreflect.", "location": "Skopje, North Macedonia", "introduction": "I am Haris Mekic, an AI consultant and operations practitioner. I turn working knowledge into context, procedures and tools that people can inspect. This is my personal portfolio of experience, implementation and learning. I started MEKreflect to develop that practice and, over time, build it into a business with a team.", "paragraphs": [ "My experience comes from keeping work connected. Across United Nations operations and three separate EU technical assistance projects, I worked with office leadership, regional teams, international experts and public institutions. I had to understand the programme as well as the budget, the people and the decisions needed to deliver it.", "I develop and track project budgets because the numbers have to reflect the work. I connect activities to costs, review expenditure and commitments, and keep revisions aligned with delivery. Monthly cash requirements, cash flow projections, expert days, procurement and payments all belong to that same picture.", "For almost two years I have also been building with AI. I work between Claude, Codex and VS Code, inspect changed files and test the result. I coach agents through context engineering, role instructions, examples and feedback. I define the work, inspect what comes back and refine the procedure when the result misses the purpose.", "I bring the same method to people learning to work with AI. We explain the job, select the right sources and establish how to question the output. That understanding becomes the assistant's working context. The aim is better judgement and a workflow that is clearer and easier to repeat.", "My mission is a digital office where people and AI workers carry connected work forward. An approved brief becomes research, a delivery plan, budget analysis, a presentation and a reviewed decision. I want to build this with a team, develop their capability and expand from workflows that have been tested in use." ], "workingStyle": [ "Understand the work and the people doing it.", "Turn the idea into a costed plan with clear responsibilities.", "Bring in specialists where the problem needs depth.", "Build, test and improve with the people who will use it.", "Stay responsible for the outcome." ], "contact": { "email": "harism88@outlook.com", "href": "mailto:harism88@outlook.com" }, "portrait": "/haris-portrait.png" }, "logbook": { "intro": "This is how I develop people and AI working methods together. I turn operational knowledge into context, role instructions and reviewable outputs, then improve the procedure through use and feedback. The journal connects my applications, local model experiments and coaching practice to that same digital office mission.", "boundary": "This journal combines selected working conversations, current source inspection and specific recorded checks. It describes almost two years of personal AI development. Dates are retained as provenance, not a measure of skill or working hours. Private conversations and client records are not published.", "entries": [ { "id": "my-working-loop", "title": "Start with the work, not the model", "date": "2026-10-09", "status": "Owner history and implementation reviewed", "summary": "I coach AI agents through the work. I explain the operating situation, shape their responsibilities, inspect the output and correct what misses the point. That conversation becomes context, procedures and acceptance checks in the application. When I guide another person, I help them learn the same discipline so they can direct the AI themselves.", "changes": [ "I read the requirements and source material, then define what the result needs to do and how I will check it.", "I use Claude, Codex and VS Code for context engineering, implementation, review and debugging. Builders and testers receive different responsibilities so I get a separate view of the result.", "I treat coaching the agent as an engineering loop. I clarify the context, define an output contract, inspect its behaviour and carry corrections into the next instruction or implementation. This is contextual configuration. My separate local model adaptation work changes model parameters.", "I teach the person how to brief, question and review the agent. The goal is a working method they can explain and repeat, not dependence on a prompt they do not understand.", "I use fan-out and fan-in orchestration when commercial, delivery and compliance questions need their own analysis. Failed branches remain visible when the findings come together.", "I open the changed files and use the application. I compare what it does with the requirement and keep the existing workflow working as I improve it.", "Where the workflow supports them, I use deterministic contract tests, negative fixtures, source-linked context and receipt-bound completion. I want a result that I or the next person can inspect." ], "verification": "Selected original working conversations were compared with the current source and Git history. They document company-specific context, separate builder and reviewer responsibilities, repository handoffs and repeated challenges to unsupported outputs.", "limits": "I direct the implementation and review. Each workflow has its own checks and acceptance requirements.", "references": [], "practice": { "input": "I start with the people doing the work, the decision they need to make and the source material behind it. I describe the constraints and ask what would make the result useful to the next person.", "context": "I turn that into a working brief with the objective, domain rules, relevant records, permitted actions and acceptance checks. Corrections become updated context or procedures rather than disappearing into a chat. I keep decisions and implementation state in the project so the next session can continue the work.", "agentWork": [ "I ask an assessment agent to inspect the existing workflow before proposing changes.", "Implementation agents receive separate responsibilities and file boundaries. A reviewer checks the behaviour and the evidence rather than repeating the builder's completion message.", "The next agent receives the changed files, findings, test results and unresolved questions. I review the result in the application and redirect the work when it misses the purpose." ], "output": [ "A versioned implementation, a usable interface and an explicit account of what works and what remains.", "The working record belongs to the project. It is not trapped in one conversation or tied to one model." ], "stack": [ "Context engineering", "Domain modelling", "Contextual SOPs", "Repository based handoffs", "Claude", "Codex", "VS Code", "Separate implementation and review" ], "next": "I want colleagues to configure a digital worker through normal questions about their work. That conversation should become a role, a procedure, a set of tools and a clear handoff that the team understands. We can then measure the complete work cycle and improve both the procedure and the person's ability to direct it." } }, { "id": "coaching-a-working-method", "title": "Develop the person and the AI working method", "date": "2026-10-10", "status": "Individual coaching and an initial scenario trial", "summary": "I helped a consultant turn a broad career goal into a practical way of working with AI. We connected the responsibilities he wanted to develop with the context an assistant needs and the judgement that remains with him. He then tried the method independently on a hypothetical consulting assignment and gave specific feedback on what helped.", "practice": { "input": "A consultant's development goal, the project responsibilities he wanted to practise and questions about using an AI assistant more effectively. The first exercise used a hypothetical consulting scenario.", "context": "I separated essential inputs from helpful background, connected the assignment to its intended recipient and defined how to check the result. I translated that into an input checklist and reusable instructions for the assistant and the person directing it.", "agentWork": [ "The assistant is instructed to identify the objective, available sources, missing inputs and decisions that require a person before developing a solution.", "I designed role procedures for research, delivery planning, budget and risk analysis, presentation development and quality review. Each has an input, output and handoff.", "I developed the next-stage plan around an approved digital archive, a source register, weekly delivery and learning reports, and a supervised pilot with a project manager.", "The person learns to challenge assumptions and review the output. Feedback then refines the assistant's context and procedure." ], "output": [ "A context-first working method, input checklist and instructions that the participant tried independently in a hypothetical assignment.", "Written feedback described the distinction between essential inputs and helpful background as useful for moving the work forward.", "A proposed 90-day pilot and development pathway, with roles, document workflows, review points and measures to establish before wider use." ], "stack": [ "Context engineering", "Copilot working instructions", "Role based SOPs", "Input and output contracts", "Source provenance", "Human review", "Workflow evaluation", "Practical coaching" ], "next": "Agree a real assignment with an authorised owner, establish preparation and review baselines, and test whether another person can repeat the workflow. Expand from accepted outputs and feedback." }, "changes": [ "I developed the person and the assistant's working method together. Clearer briefing and review become better context and more useful instructions.", "I turned a broad development ambition into a proposed sequence of responsibilities, deliverables, learning goals and review points.", "The proposed archive links authoritative records rather than collecting unrestricted copies. Client work remains separated and access follows existing permissions.", "The reporting design distinguishes project delivery from the learning record. It tracks inputs, outputs, corrections, accepted work and measured effort when available.", "This case concerns contextual assistant configuration and coaching. My local parameter-efficient model adaptation experiments are a separate technical evidence track." ], "verification": "The participant's written response confirms an independent hypothetical exercise and specific feedback on the method. The subsequent detailed working plan and assistant instructions were sent. This account is anonymised and paraphrased. Private correspondence is not published.", "limits": "The completed work is individual coaching, a scenario trial and a proposed implementation plan. A live organisational pilot, measured productivity improvement and wider team adoption are next-stage tests.", "references": [] }, { "id": "company-bid-workflow", "title": "Build a bid around the company", "date": "2026-10-10", "status": "Implementation and owner decisions reviewed", "summary": "I started with how a consulting company actually wins work. More tenders were not the answer if the opportunity did not fit, the buyer was unclear or the expert information was incomplete. Conversations with staff and my own delivery experience shaped the system around those decisions.", "practice": { "input": "An opportunity and its original documents, the assignment requirements, evaluation criteria, languages, company capability and expert roster. Missing source documents remain a gap to resolve.", "context": "I connect company history and relevant previous proposal material with the specific section or decision being worked on. I asked for living module blueprints that explain what each part does and follow the actual implementation.", "agentWork": [ "Business development reviews fit and commercial direction. Bid coordination examines the team, expert matches, CV actions and mobilisation. Data and compliance examines the available evidence.", "These three specialist branches run in parallel. Operations synthesis receives their available findings and preserves partial or failed branches.", "Section generation uses requirements, evaluation criteria, risks and relevant specialist findings. Preparation services create tasks for missing documents, context, analysis and approaching deadlines." ], "output": [ "The bid manager receives expert matches, team gaps, staffing confidence, CV actions, coordination notes and a budget range with its basis when the roster supports one.", "The workspace keeps preparation tasks and internal deadline records connected to the opportunity. Proposal assembly brings sections and supporting material together for review." ], "stack": [ "TypeScript", "Parallel specialist orchestration", "Promise.allSettled", "Structured generation", "Zod response schemas", "Role filtered tool use", "Task and calendar state" ], "next": "Finish the wider delivery plan, strengthen concurrent task creation and verify the complete handoff with a bid manager. The same approach can fit another company by rebuilding its capability context, roles, documents and approval rules with its staff." }, "changes": [ "I separated the commercial, staffing and evidence questions so one fluent answer could not hide an incomplete perspective.", "The section generator uses bounded context, including a limit on prior proposal material and short excerpts of other sections for continuity.", "Expert references in matches and CV actions are checked against the retrieved roster. Without a roster the result stays partial and does not supply a budget estimate.", "The calendar service updates pipeline generated deadline records while leaving manually created events alone. This is internal calendar state, not a claim of external calendar delivery.", "The response schema checks structure. Factual accuracy, commercial assumptions and readiness still require review.", "The inspected structured generation path uses Gemini 2.5 Flash. The separate conversational tool interface uses a Claude adapter. Claude and Codex are also my development tools, which is a different role from the models called inside the application." ], "verification": "Current source paths and selected Git revisions were compared with the original working conversations. The parallel branches, bid manager output, proposal context, preparation tasks and internal deadline updates were inspected separately.", "limits": "The implementation is distinct from this public example. Client records remain private. The wider delivery plan and some task and tool boundaries still need work. No bid win rate, production usage volume or time saving is claimed.", "references": [] }, { "id": "digital-worker-workspace", "title": "Give digital workers a shared workspace", "date": "2026-10-10", "status": "Python implementation and working history reviewed", "summary": "I do not want a digital office to be a chat window with a different name. I want to see the request, the role working on it, the result and the decision in one place. That is the direction behind my agent working system. A good looking shell is not enough if the work cannot move between roles.", "practice": { "input": "A person's objective, the way their team works, the project state, the available tools and the actions that need an owner decision.", "context": "I define what belongs in the active work record and what must survive the next session. The context policy checks the work log, current task, decision queue, session record and a place for recording risks.", "agentWork": [ "The Python layer connects the interface to local records and a configurable model conversation.", "The model adapter sends a bounded recent conversation and an explicit output limit. Missing configuration returns a clear unavailable result instead of an invented answer.", "A separate policy evaluates the action against the autonomy level and risk score. The result is to proceed, report or return the decision for review.", "The change review path measures a defined baseline, snapshots enrolled files and supports a keep or restore decision." ], "output": [ "A working record the next session can inspect, explicit decision states and a model reply that stays separate from permission to act.", "Checks and snapshots provide evidence for selected changes. The wider multi-person digital office remains the product direction." ], "stack": [ "Python standard library", "JSON APIs", "React", "Vite", "Optional Electron shell", "File backed state", "argparse CLIs", "Provider abstraction", "subprocess timeouts", "SHA256 verification" ], "next": "Make role setup and the request-to-result path easier for another person to use. Connect each permitted tool to the same review rules and test the whole workflow before expanding the number of integrations." }, "changes": [ "The provider adapter includes the last twenty conversation turns and an output token cap. The operating method remains separate from the selected model.", "Context pressure uses transcript file size and work-record freshness as signals. It is a practical proxy rather than direct model token telemetry.", "The decision policy accepts integer risk scores from 4 through 20 and rejects malformed values. High risk still returns to review at the highest autonomy level.", "The retention path runs configured checks with timeouts and compares a baseline with a later result. Selected files can be restored and verified by hash.", "My role is architecture, operational rules and acceptance. AI engineering agents implement and test code with me, as the public source acknowledges." ], "verification": "The Python model interface, context lifecycle, policy and retention paths were read. Six isolated offline probes checked malformed risk inputs, high risk decisions, missing and stale context, the conversation bound and missing model configuration. Earlier full suite results remain in the public repository history.", "limits": "These are local implementation and offline checks, not a claim that every proposed integration is complete. File size is not a token meter. Recovery protects the enrolled records and files, not every possible application state.", "references": [ { "label": "Public agent workspace source", "url": "https://github.com/Haris88m/servari-open" } ] }, { "id": "python-in-my-practice", "title": "Make the operating rules explicit in Python", "date": "2026-10-10", "status": "Source and focused checks reviewed", "summary": "I use Python where the work needs structured data, persistent state and a result I can inspect. My practice is to define the operational rule, develop it with AI, and then challenge its behaviour. I would rather show that process than describe myself with a proficiency label.", "practice": { "input": "Source text, a question, a case record, a proposed state change or a confirmation file. I define what counts as evidence and what must never be accepted as completion.", "context": "The working record connects the source, the case identity, the allowed next state and the verification requirement. Text retrieval uses bounded excerpts and lexical matching for the question.", "agentWork": [ "Python retrieves relevant source excerpts and maintains structured records in SQLite.", "The workflow separates preparation, an attempted action and confirmed completion. A receipt must belong to the same case.", "Integrity checks compare the recorded SHA256 with the confirmation file. Missing, unrelated or changed receipts leave completion unconfirmed.", "Tests exercise the expected case and the boundary cases, including malformed input, stale state and missing evidence." ], "output": [ "A reviewable result with its source and state, rather than a free text success claim.", "Small tools that can be run again with controlled inputs and inspected through their tests." ], "stack": [ "Python", "SQLite", "JSON and JSONL", "Lexical retrieval", "Hash based integrity", "State machines", "CLI interfaces", "Temporary fixtures", "unittest" ], "next": "Keep extending tests around the real boundaries and improve independent setup instructions. I want another person to be able to reproduce a result without needing my original chat." }, "changes": [ "The source query uses all-term lexical matching. I describe that precisely instead of calling it vector search.", "The coordinator checks case ownership, timing, package integrity and receipt identity before recording completion.", "Persistent workflow state and model-generated prose are different things. The prose does not decide that an external action succeeded.", "I define the operational rule, work with Claude and Codex on implementation, then inspect the source and test its behaviour." ], "verification": "Current source retrieval and workflow coordinator code was inspected. The earlier receipt workflow has 18 focused tests recorded with in-memory databases and temporary files. This review did not run live applications or access their private records.", "limits": "The public account describes engineering patterns and test scopes. It does not publish private records or imply that separate database updates form a single distributed transaction.", "references": [] }, { "id": "models-and-evidence", "title": "Choose the model for the work", "date": "2026-10-09", "status": "Model practice and implementation reviewed", "summary": "I like understanding what different models can do with the same problem. I use GPT and Claude model families for reasoning and coding, with lighter routes where they fit the task. That interest has also led me to develop model selection and provider abstraction inside applications.", "changes": [ "I use GPT-family coding and reasoning models through Codex for implementation, debugging and specialist review.", "I use Claude-family models for planning, reasoning, context development and a second perspective on the work.", "I work with coding assistants in VS Code, keeping model output connected to source files, runtime behaviour and tests.", "Inside applications I have implemented provider abstraction, bounded chat history and sensitivity-aware model-tier selection, with explicit handling when a provider is unavailable. Cost preferences inform model selection." ], "verification": "Timestamped local model metadata was reviewed separately from configuration. The model-tier selection path had 13 focused local checks. Three route/dispatch checks used mocked dependencies. These test scopes are not added to usage counts.", "limits": "Local runtime and response labels are retained with their source type. Quality is assessed through the task result and workflow checks.", "references": [], "practice": { "input": "The task, its sensitivity, provider availability and the intended cost profile.", "context": "I keep the model, the agent harness and the operating method separate. The same working brief can be reviewed with different models, but their capabilities and results still need checking.", "agentWork": [ "A selection path evaluates sensitivity and the available model tiers.", "Provider adapters translate the bounded conversation and handle an unavailable provider explicitly.", "I use a separate implementation or review pass when another perspective is useful." ], "output": [ "A selected route and an inspectable result within that application's context and output constraints.", "Thirteen deterministic routing checks verify selection behaviour, not comparative model intelligence." ], "stack": [ "Provider abstraction", "Sensitivity aware routing", "GPT family", "Claude family", "Bounded history", "Structured outputs", "Offline selection tests" ], "next": "Build task-level evaluations that compare quality, corrections and running cost. A routing preference alone is not a hard spending limit." } }, { "id": "small-model-experiments", "title": "Learning through small-model adaptation", "date": "2026-10-09", "status": "Historical artifacts and source reviewed", "summary": "I wanted to understand more of what happens below the chat interface. I explored local adaptation of small pretrained Qwen2.5 instruction models with PEFT LoRA, working through training inputs, low-rank adapters, checkpoints and evaluation.", "changes": [ "I worked with a training path that tokenises prompt and completion pairs, attaches low-rank adapters and uses gradient accumulation and checkpointing.", "I kept the adapters and training summaries from completed experiments. Separate evaluation tooling compares the base model with adapted output.", "I recorded factual errors and identity confusion during evaluation, then identified held-out task comparisons as the next step before using an adapter in a real workflow." ], "verification": "Trainer code, saved summaries and adapter provenance were inspected. Earlier dated verification compared ten summary and ten adapter hashes with the prior ledger. This review did not retrain a model or load tensors.", "limits": "Experimental LoRA adaptation of pretrained models. The next evaluation step is a held-out comparison of task quality.", "references": [], "practice": { "input": "Prompt and completion pairs, a small pretrained instruction model and the memory limits of a local experiment.", "context": "I wanted to understand the training path below the chat interface, including tokenisation, adapter parameters, checkpoints and comparison with the base model.", "agentWork": [ "The Python trainer formats JSONL examples and tokenises the text with a bounded sequence length.", "Transformers and PyTorch load the pretrained causal model. PEFT attaches low rank adapters to selected projection modules.", "Gradient accumulation and checkpointing support the training run. The pipeline saves the adapter and a run summary." ], "output": [ "A saved Qwen2.5 1.5B adaptation run records 2,721 training pairs across two epochs with LoRA rank 16.", "The saved adapter is an experimental artifact. Whether it improves the work needs a separate held-out evaluation." ], "stack": [ "Python", "PyTorch", "Transformers", "PEFT", "LoRA", "Qwen2.5", "JSONL preparation", "Gradient accumulation", "Gradient checkpointing" ], "next": "Compare the base and adapted model on held-out operational tasks with the same scoring rules, checking factual errors as well as useful answers." } }, { "id": "make-completion-inspectable", "title": "Make completion something I can inspect", "date": "2026-10-09", "status": "Focused local tests", "summary": "I wanted the next person or agent to know what had actually happened. I developed a persistent workflow that keeps preparation, an attempted action and confirmed completion separate. The confirmation has to belong to the same case and pass an integrity check.", "changes": [ "I keep the workflow state in a persistent record so the next task has a dependable starting point.", "I require an identity-matched receipt before recording completion.", "The workflow compares the recorded SHA-256 against the actual confirmation file.", "Missing, unrelated or changed receipts leave completion unconfirmed." ], "verification": "Recorded verification includes 18 focused workflow-engine tests with in-memory databases and temporary receipts. The current review inspected the persistent workflow and integrity checks without running external applications.", "limits": "Isolated workflow tests cover case identity, receipt integrity and completion states with temporary fixtures.", "references": [], "practice": { "input": "A prepared case, an attempted action and a confirmation artifact.", "context": "The case identity, expected state and confirmation requirement remain in a persistent record so the next worker can distinguish preparation from completion.", "agentWork": [ "The workflow records preparation and attempted execution as separate states.", "It checks that the receipt belongs to the case and that its SHA256 matches the file.", "A missing or changed receipt leaves completion unconfirmed." ], "output": [ "An explicit confirmed or unconfirmed state with evidence that can be inspected.", "The next worker does not have to infer success from a narrative message." ], "stack": [ "Python", "SQLite", "State transitions", "SHA256", "Temporary receipt fixtures" ], "next": "Extend the same evidence requirement to more handoffs and keep tests around interrupted work and stale state." } }, { "id": "keep-the-source-visible", "title": "Keep the source visible after the model answers", "date": "2026-10-09", "status": "Implementation inspected", "summary": "Context engineering is where my operations knowledge and AI work meet. I decide which requirements, criteria, risks and source material belong in the task. I have developed selection and generation paths that keep that context connected to an answer someone can review.", "changes": [ "I use lexical matching to retrieve material for the specific task.", "I check source sufficiency before generation so missing material changes what happens next.", "I give parallel specialists a bounded shared context and a separate question to work through.", "I validate the output structure and retain missing or failed perspectives for review." ], "verification": "The current orchestration, context and generation implementations were opened and inspected as separate paths. In the orchestration path, three specialist branches feed a separate synthesis stage.", "limits": "These patterns are implemented in separate applications. Retrieval uses lexical matching, with source sufficiency and structured-output checks at defined stages.", "references": [], "practice": { "input": "A question, the relevant assignment requirements and the available source material.", "context": "I select what the task actually needs and preserve missing evidence. A longer prompt is not automatically a better understanding of the work.", "agentWork": [ "Retrieval uses lexical matching for the particular question.", "Source sufficiency checks determine whether generation has enough material to proceed.", "Parallel specialists examine different questions from a bounded shared context.", "Structured output checks keep their findings usable by the next stage." ], "output": [ "An answer or specialist finding linked to its supporting context.", "Missing and failed perspectives remain visible when the work is brought together." ], "stack": [ "Lexical retrieval", "Context selection", "Structured outputs", "Parallel orchestration", "Source sufficiency checks" ], "next": "Compare context variants on fixed reviewed cases so source coverage and omissions can be evaluated consistently." } }, { "id": "review-the-boundary", "title": "Find the case the policy did not handle", "date": "2026-10-09", "status": "Published and tested", "summary": "A review found that the public decision policy could accept malformed risk scores after numeric coercion. I directed the correction and the regression tests. This was a useful reminder to inspect the difficult inputs as carefully as the expected ones.", "changes": [ "The policy now accepts integer scores from 4 through 20 only.", "Strings, booleans, floats, missing values and out-of-range scores queue for review.", "The command-line boundary parses an integer explicitly, preserving the valid command contract.", "224 parametrized malformed-input, boundary and CLI regressions were added to the original 160 tests." ], "verification": "Recorded verification includes 384 Python tests and eight separate checks, followed by published CI for the Python suite, verifier and UI build. The latest targeted review also exercised malformed score inputs and high risk decisions in isolated offline probes.", "limits": "Local policy tests use mocked providers and synthetic fixtures. The public repository includes the implementation and reproduction steps.", "references": [ { "label": "Public implementation and verification", "url": "https://github.com/Haris88m/servari-open" } ], "practice": { "input": "Risk scores submitted to the public Python decision policy, including strings, booleans, floats, missing values and values outside the supported range.", "context": "I wanted the policy to apply the intended rule rather than treating numeric coercion as valid authority.", "agentWork": [ "The implementation pass makes the integer boundary explicit.", "The review pass checks malformed values and valid boundary inputs.", "Command line parsing preserves a valid integer input while unsupported values return to review." ], "output": [ "Integer scores from 4 through 20 are accepted by the policy. Malformed values queue for review.", "The implementation and reproduction tests are available in the public repository." ], "stack": [ "Python", "Input validation", "Parametrized regression tests", "CLI boundary testing" ], "next": "Apply the same discipline to tool calls and state transitions so the safe example is not the only case that works." } }, { "id": "my-ai-roadmap", "title": "Where I am taking this", "date": "2026-10-10", "status": "Personal direction and next work", "summary": "My mission is to make professional knowledge usable by people and digital workers together. I want to build the team as well as the system. The foundation is operational experience, contextual procedures and an honest view of what the software does.", "practice": { "input": "A team willing to explain its work, one useful workflow and the records needed to understand the current way of doing it.", "context": "I start small enough to learn properly. We agree the purpose, the role boundaries, the owner and what an accepted result looks like before adding more agents or systems.", "agentWork": [ "First make the role and context contracts easy to configure from the team's own language.", "Then finish and test one end-to-end handoff with real authorised users, including incomplete sources and failed tools.", "Compare the full work cycle, including preparation, corrections, review effort and accepted output.", "Use the findings to improve the workflow and teach the team how to inspect and maintain it." ], "output": [ "The next milestone is a repeatable team workflow with recorded feedback and a measured baseline.", "Longer term I want the same operating foundation to adapt across bidding, budgets, programme delivery and other company work without pretending every organisation is the same." ], "stack": [ "Role configuration", "Context evaluation", "End to end testing", "Workflow observability", "User feedback", "Team learning" ], "next": "A team pilot with a named owner, agreed baseline and a short review cycle. Broader autonomy follows demonstrated reliability, not an impressive interface." }, "changes": [ "The foundation already includes company-grounded bid context, specialist handoffs, local state and policy checks, budget calculations and reusable public examples.", "The next work is easier role configuration, stronger tool boundaries and fuller delivery planning.", "Context Lab will compare source and prompt variants against reviewed questions. Workflow Observatory will measure the complete task, not just the speed of an answer.", "I want to keep exploring models and local adaptation while completing useful work with people. The model is part of the system, not the whole system." ], "verification": "This roadmap is drawn from my retained working requests and the current implementation review. It separates existing foundations from proposed capabilities and measurements.", "limits": "The roadmap is direction, not a claim that team adoption, measured savings or the planned evaluation applications already exist.", "references": [] }, { "id": "recorded-model-practice", "title": "A sustained working practice", "date": "2026-10-09", "status": "Local metadata audited", "summary": "AI development has become a sustained part of my working life. My retained records show activity on 121 distinct dates between 27 January and 29 September 2026 across Codex, Claude and VS Code.", "changes": [ "I use Codex for implementation, code review and coordinated specialist work.", "I use Claude for reasoning, planning, context development and a second perspective on a problem.", "I work in VS Code to connect the conversation to actual files, application behaviour and tests.", "I carry decisions forward through project memory, structured context and versioned source files." ], "verification": "Read-only metadata review on 9 October 2026, excluding events on or after that date to avoid counting this assessment. Daily dates were combined as a set, not summed across tools. No raw conversation corpus is published.", "limits": "The metric counts distinct UTC dates in retained local records, including agent activity. Working hours and financial savings require a separate measurement method.", "references": [], "practice": { "input": "The problem I am working on and the current project record, whether I am using a coding agent, the editor or a separate review session.", "context": "I keep the work tied to files, requirements and previous decisions. Using several tools only helps when they can continue from the same understanding.", "agentWork": [ "Coding agents inspect and implement bounded parts of the project.", "Separate reasoning and review passes challenge the proposed design, changed files and outputs.", "Working notes carry decisions and unfinished checks into the next session." ], "output": [ "A sustained body of project work rather than a single prompt experiment.", "The retained activity record supports the continuity of the practice. It is not a timesheet." ], "stack": [ "Codex", "Claude", "VS Code", "Project memory", "Versioned source" ], "next": "Track the outcome and review effort of future team workflows directly, instead of estimating impact from activity records." } } ] }, "ideas": { "intro": "I want the next stage to include other people working with what I have built. These ideas are about understanding what helps them, making the context easier to inspect and learning whether the whole workflow improves. I want the team to grow with the system.", "entries": [ { "id": "test-the-context", "title": "Evaluate the context", "status": "Planned", "problem": "I want to know which source material makes a difference to the answer and where the context is still incomplete.", "proposal": "I plan to build a fixed set of operational questions and define which sources each answer needs.", "next": "I will compare context and prompt variants against the same human-reviewed cases.", "acceptance": "The published comparison will show the inputs, supported answers and omissions so another person can inspect the result.", "boundary": "Planned evaluation with human-reviewed reference cases." }, { "id": "measure-the-whole-workflow", "title": "Measure the whole workflow", "status": "Planned", "problem": "I care about the whole job, including preparation, corrections and review. That is where I need to understand whether the change is useful.", "proposal": "I plan to measure those stages alongside execution and the accepted outcome.", "next": "I want to run a workflow pilot with a user and record a baseline before comparing the change.", "acceptance": "I will keep a record of quality, review effort, direct running cost and user feedback.", "boundary": "Pilot design for measuring operational value." }, { "id": "make-handoff-real", "title": "Put a working method in someone else’s hands", "status": "Available for reuse", "problem": "I want someone else to be able to start, test and understand the workflow without needing me beside them.", "proposal": "I have packaged the budget-control engine with editable inputs, source code, tests and a guide. Budget Studio uses the same calculation module.", "next": "I want an independent user to try the pack on a suitable planning task and tell me what needs to improve.", "acceptance": "The extracted pack passes its tests and matches the browser example. Independent reuse feedback is the next milestone.", "boundary": "A reusable engineering pack with illustrative data and explicit calculation rules." } ] }, "publicScope": "An open workspace of professional experience, applied AI and reusable tools. Public downloads contain selected professional facts, engineering examples and clearly labelled demonstration data. Client records and private archives remain confidential." } Project stories snapshot Separate file SHA-256 549b4d5f48bcccf5cbe9edadb0ca0c0a7447269642ed3a0bd95ee33d9cbee0b3 # Haris Mekic project stories These projects show how I bring my operational experience into AI development. Each one began with something I wanted to understand or improve. Together they are the foundations of the digital office I want to build with a team. Version 2026-10-10-projects-v1 ## Opportunity intelligence workspace Business development and bid preparation. Developed in 2026. Implemented working system. I wanted a workspace that understood the company as well as the opportunity. I developed a connected bid workflow around the questions I would ask in the office. Does this assignment fit us? Can we put the right team together? What evidence do we have, and what still needs work before we can prepare a credible proposal? ### The problem Finding a tender is only the beginning. The team still needs to understand the buyer, read the documents, compare the assignment with its experience and bring the right experts into the work. I wanted those decisions to stay connected so that the next person or specialist agent could continue from the same working record. I shaped the work through conversations with staff about their needs and through my own operations experience. The first version was finding opportunities but did not understand the business deeply enough. I challenged the broad recommendations, missing buyer information and thin expert records. Even the win and loss reporting needed a clear way to record the outcome. I wanted to follow the whole process and understand what happened at each stage. ### Decisions I made #### Start with the organisation, not the search box I brought the organisation's actual assignments, previous bids, methodology and expertise into the development brief for Claude and Codex. I wanted the system to recognise work worth pursuing. The context layer separates that company baseline from recorded outcomes, funder engagement, expert information and source quality. #### Make missing information change the next action I wanted missing documents and staffing information to change the next task. The implementation distinguishes a notice from full or partial bid documents and carries evidence quality into the analysis. Gaps in documents or working context become preparation tasks, alongside approaching deadlines. #### Give specialists one working record I gave business fit, bid coordination and compliance their own questions about the same assignment. The bid role receives the terms of reference, language requirements, evaluation criteria and expert roster. Its structured output covers expert matches, team composition, staffing confidence, mobilisation days, a budget range with its basis, CV actions and coordination notes. Expert IDs in matches and CV actions are checked against the retrieved roster. An operations synthesis brings the available findings together for review and keeps partial results when a branch fails. #### Keep the handoff usable after the conversation ends I developed module blueprints to keep the purpose of each part beside the implementation. Independent repository review then informs the next brief. The handoff records the task, changed files, verification, remaining risks and next actions so Claude and Codex have a working state to continue from. #### Separate reusable machinery from business judgment The source adapters, normalised opportunity record, evidence packet and proposal assembly can be reused. The decision to pursue work still belongs to the people who understand the organisation and its capacity. I see the opportunity to scale in keeping those relationships intact as more work moves through the system. ### How the work moves #### 1. Define what a worthwhile opportunity means My role. I frame the decision around the work the organisation can deliver, its past assignments, funders, expertise and constraints. I challenge recommendations that do not reflect that reality and turn those objections into requirements for the next implementation pass. Input: Business capabilities, previous assignments and the questions the team needs to answer. Output: A business-fit brief and a defined purpose for each module. Tools: Claude, Codex, Module blueprints #### 2. Bring different sources into one record System role. Adapters translate different notice formats into a common record containing buyer, funder, deadline, value, source references and document links. Filtering removes irrelevant arrivals. Deduplication identifies repeated notices, while a richer arrival can improve the existing record instead of becoming another item to review. Source monitoring distinguishes healthy, stale, blocked and zero-yield sources, using those findings to prioritise repair, change collection strategy or consider disabling a poor source. Input: Notices from configured APIs, feeds and other source adapters. Output: Normalised opportunity records retaining their source references. Tools: TypeScript, Source adapters, OCDS #### 3. Build the working context System role. The context builder combines assignment documents with recorded business outcomes, expert availability, partner evidence and source health. It distinguishes documentary information from missing or weak evidence. An unknown expert's availability is not treated as confirmed capacity. Input: The opportunity, source documents, organisation records and expert evidence. Output: A shared evidence packet with document status and confidence indicators. Tools: PostgreSQL, Context builder #### 4. Analyse in parallel, then synthesise AI role. Three specialist calls examine business fit, team and bid coordination, and compliance/data requirements. The bid role returns expert matches, team composition, staffing confidence, mobilisation days, a budget range with its basis, CV actions and coordination notes. A separate operations analysis combines the available findings for review. The saved run distinguishes complete, partial and failed analysis so a missing perspective remains visible. Input: Shared evidence, task-specific instructions and the state of each specialist branch. Output: A combined recommendation and structured specialist findings for the person reviewing the opportunity. Tools: Gemini, Structured generation #### 5. Turn readiness gaps into preparation work System role. Preparation logic identifies missing source analysis, missing document-workspace context and approaching submission dates. It checks for an existing open task before creating another one. Date synchronisation updates system-generated calendar events when the opportunity deadline changes and leaves manually created events separate. Input: Opportunity stage, document and analysis state, and submission dates. Output: Specific preparation tasks and linked deadline events. Tools: Task generation, Calendar synchronisation #### 6. Draft the section that is actually needed AI role. The generator uses the assignment's mandatory requirements, evaluation criteria, risks and relevant specialist findings. Methodology, team, work plan and financial narrative receive different section instructions. Selected prior material and short excerpts of existing sections provide continuity. A schema checks the response structure before it becomes a draft for review. Input: The requested section, evaluation requirements and a bounded selection of supporting context. Output: A structured, assignment-specific draft rather than an unrestricted chat answer. Tools: Gemini, Zod #### 7. Review the recommendation and readiness My role. I test whether the result answers the original business question and whether the interface supports the next decision. Missing documents, weak staffing evidence and incomplete checks need an action, not better wording. In development, I use a separate review pass and update the blueprints so corrections survive into the next task. Input: The recommendation, draft, evidence gaps and preparation state. Output: Reviewed next steps and a precise correction or implementation brief. Tools: Readiness checklist, Claude, Codex #### 8. Assemble a usable working package System role. Proposal assembly combines selected sections, assigned expert CVs, consortium information, cover details and a table of contents. It produces a structured proposal for PDF or DOCX rendering, with optional document-workspace output. Recorded pipeline outcomes can then inform the context used for later opportunities. Input: Reviewed sections, team records, partners and the selected document structure. Output: An assembled proposal object prepared for document rendering and continued workflow tracking. Tools: TypeScript, Document templates ### Context engineering For me, context engineering starts with understanding the work. I decide what the system needs to know, where that knowledge comes from and what is still uncertain. Then I shape the record so the next person or agent can continue with the relevant evidence and decisions in front of them. #### Assignment evidence The notice and the actual bid package are different inputs. Source-document status, requirements, evaluation criteria, dates, languages and risks remain attached to the assignment. #### Business memory I combine the company baseline with recorded outcomes, funder engagement, expert assignments, availability and partner evidence. Gaps in the outcome history remain visible. #### Role-specific context Specialists share the same evidence but answer different operational questions. Their available findings become input to synthesis and are selected again according to the proposal section being drafted. #### Context budget Reference-proposal text is capped at 15,000 characters. Existing sections contribute excerpts of up to 320 characters each. The model receives this bounded selection of the archive. #### Evidence quality Document sufficiency, relationship evidence, expert-bench quality, outcome history and funder familiarity travel with the facts. Missing availability weakens the assessment of whether a team can be assembled. #### Engineering continuity Module blueprints and structured handoffs preserve intent, changed files, checks, risks and next actions. I use them to coordinate Claude and Codex across implementation and review. ### Implementation I developed the system with Claude and Codex using TypeScript, React, Next.js and PostgreSQL. Gemini handles structured analysis and drafting inside the application. A separate Anthropic integration supports the governed assistant path. I keep the development tools and the models running inside the product distinct when I describe the architecture. TypeScript, React, Next.js, PostgreSQL, Gemini, Anthropic API, Zod, OCDS SourceAdapter and NormalizedTender contracts separate collection from the business workflow. Source references and deduplication identities survive normalisation, while richer duplicate arrivals can enrich existing records. Fan-out/fan-in orchestration runs three specialist analyses before an operations synthesis. Shared context and saved run status preserve the relationship between evidence, intermediate work and the final recommendation. Research sessions have runtime budgets, organisation-scoped locking and stale-session recovery logic. Database migrations define atomic pending-job claiming with row locks. Enterprise throughput still needs to be measured. Document sufficiency, preparation tasks, calendar events and stage checks connect research with the next piece of work. Drafting and assembly consume structured records instead of requiring a person to copy every answer between unrelated tools. The current business profile and relevance rules are tailored to one organisation. Reusing the architecture for another company requires configuring its capabilities, evidence permissions, sources and decision criteria, then testing that fit. ### Results #### Working output A connected bid workspace. Source intake, organisational context, specialist analysis, preparation tasks, section drafting and proposal assembly are represented in the inspected implementation. #### Architecture 3 specialist perspectives. Business development, bid coordination and compliance/data analysis feed a separate operations synthesis stage. #### Engineering check 121 tests passed. Policy and confirmation-token tests rerun offline on 9 October 2026, checking permitted actions and confirmation controls. ### Next step I see procurement becoming more structured for both people and AI workers. I want to develop a company-configurable evidence packet that stays with an opportunity through discovery, qualification, proposal preparation and authorised review. I would test it with a business team and examine source relevance, review effort, accepted drafts, model cost and the decisions made. Buyer-side notice publication and agent-to-agent procurement are future directions I want to explore. [Specialist orchestration implementation](/ai-labs/specialist-orchestration/) · [Context and structured generation](/ai-labs/context-and-generation/) · [Reusable context handoff](/context-handoff.md) Based on selected original working instructions and implementation reviewed on 10 October 2026, with the dated engineering check identified above. Private correspondence, client records and source archives are not part of the public material. ## Career operations system Knowledge management and accountable automation. Developed and used in 2026. Implemented working system. I built this because I wanted my experience to be understood in its full context. I connected sources, professional claims, application requirements, document versions and completed actions into a working system. The aim is to carry the knowledge forward instead of researching the same history again with every task. ### The problem My professional record was spread across documents, correspondence, different projects and repeated AI conversations. I needed the responsibilities, evidence and purpose of each project to stay together when that material was used again. I kept seeing different projects mixed together and operational responsibility reduced to a generic description. I was correcting the same things repeatedly. That became a development problem to solve by giving the corrections a durable place in the system. ### Decisions I made #### Keep the source behind the sentence An inventory records the original file, its hash and extraction policy. Retrieval returns a short, privacy-minimised passage with its source identifier. A reviewed claim can then be traced to the evidence used rather than reconstructed from conversation memory. #### Separate writing from factual permission Draft wording is checked against a claim registry. The validator checks claim and evidence identifiers, prohibited wording, unsupported numerical detail and stronger authority language. Strict automated submission uses exact pre-approved variants. #### Treat attempted and completed as different states The system records an attempt before an external action, checks duplicates and unknown outcomes, and requires a matching confirmation artifact before the application can become submitted. The receipt is bound to its application and file hash. ### How the work moves #### 1. Set the professional context My role. I identify the role or question, correct misleading descriptions of my experience, and distinguish work I performed from formal title, approval authority or a wider team's outcome. Input: The exact role, question or output requested. Output: A task scope and accurate responsibility boundaries. Tools: Haris, Reviewed profile #### 2. Index without losing provenance System role. The inventory retains paths, timestamps, SHA-256 hashes and scope policy, identifies duplicates and preserves materially different versions. Input: Approved source files and distinct versions. Output: An inventory that keeps source identity and provenance. Tools: Python, SHA-256 #### 3. Extract the relevant evidence System role. Approved documents are parsed into searchable text with privacy minimisation. Restricted document categories remain in a separate controlled location. Input: Documents permitted by the extraction policy. Output: Privacy-minimised searchable evidence. Tools: Python, Document parsers #### 4. Research with bounded context AI role. AI-assisted research uses retrieved source passages and reviewed profile facts. The local retrieval tool uses lexical matching and returns bounded snippets with source identifiers, giving each answer a specific evidence trail. Input: A focused research query and source index. Output: Bounded source passages linked to evidence identifiers. Tools: Lexical search, Claude, Codex #### 5. Review the interpretation My role. I correct responsibility, seniority, programme context and terminology. Those corrections belong in the durable profile and claim ledger so the next task does not repeat the same loss of context. Input: Proposed statements and their supporting records. Output: Corrections to wording, project ownership and scope. Tools: Haris, Claim ledger #### 6. Validate the prepared package System role. Each factual statement maps to a claim and evidence identifiers. The preparation workflow records documents, fields and readiness checks before an external action can be authorised. Input: Reviewed claims, requirements and prepared documents. Output: A package with readiness checks and exact claim mappings. Tools: Python, SQLite #### 7. Reconcile what actually happened System role. An integrity-checked confirmation artifact must match the application and its recorded hash before the state becomes submitted. Unknown outcomes stop retries until reconciled. Input: An attempted action and its confirmation artifact. Output: A reconciled state, or a stop while an outcome is unknown. Tools: SHA-256, SQLite ### Context engineering I know my career as a connected experience, but each task needs a selected part of it. I keep the relevant source, its interpretation and the current action state together so the system can work from the right context. #### Original sources Source identity, file integrity, scope policy and distinct document versions. #### Reviewed professional facts Claim identifiers, evidence links, approved wording and owner corrections. #### Task context The exact role, requirement, field and document where a statement will be used. #### Working memory Persistent application state, events, status history, reviews and follow-up records in relational data. #### Completion evidence A confirmation artifact, matching application identity and SHA-256 digest, separate from the agent's account of what it attempted. ### Implementation I used Python for source inventory, extraction, lexical retrieval, claim validation and workflow coordination. SQLite keeps applications, events, status history, claim usage and receipts connected. AI handles research and drafting around those deterministic checks. Python, SQLite, Document extraction, Lexical retrieval, SHA-256, Claude, Codex The local workflow records applications, events, status history, claim usage and receipts as related data. Deterministic checks handle source identifiers, approved wording, unsupported numerical claims and action readiness around the AI-assisted research. Receipt validation checks application identity and file integrity. A recorded attempt is not treated as successful completion. ### Results #### Working output A reusable evidence trail. Source records, reviewed claims, document packages and action receipts are connected instead of being reconstructed from memory. #### Control Completion needs confirmation. The receipt must match the application and its recorded hash. Unknown outcomes are reconciled before retrying. #### Context Corrections carried forward. My project distinctions and responsibility corrections have a durable place in profile and claim records. ### Next step I want to finish the remaining claim metadata work and measure how much repeat research and correction each prepared package needs. That will tell me where the system genuinely helps and where I should develop it further. [Receipt-bound completion mechanism](/ai-labs/receipt-bound-workflows/) · [Public evidence database](/portfolio.sqlite) Compiled from selected original working instructions and current implementation inspected on 9 October 2026. The public account omits private correspondence and client records. ## Claude and Codex working system Context engineering and delivery practice. February to June 2026. Working method with recorded project iterations. I enjoy working across Claude, Codex and VS Code, but I need them to understand the same project. I developed a working method that carries the business purpose, source material, decisions and unfinished work from one task to the next. ### The problem I was seeing generic recommendations and plans that lost the details of what had already been built. I wanted agents to collaborate on the same project rather than each begin with a different understanding. That meant developing the context, responsibilities and handoffs as carefully as the application. In February I asked Claude and Codex to exchange tasks and results through two Markdown files. In March I pushed that further into living blueprints for each module and workflow. By May, recovering the actual project state and reporting what was implemented had become an explicit part of the work. ### Decisions I made #### Keep the business reason beside the technical task I asked for each blueprint to explain what a module does and why it exists. An opportunity agent, for example, needed the company's actual work and bidding history, not only a list of funding keywords. #### Separate assessment from implementation I asked Codex to inspect the repository and challenge assumptions before turning the findings into Claude's execution brief. Specialist tasks then had a defined area to inspect or change. #### Carry corrections forward I required errors to be resolved before the next phase and the blueprints to be updated as work progressed. A new session needed the correction and its reason, not only the old plan. #### Prepare the working environment before the product In June I paused a product task to address the skills, plugins and MCP setup the coding environment needed. I also required local implementation, separate builder and tester roles, QA and security checks before the cloud step. ### How the work moves #### 1. Start from the operating problem My role. I describe what is failing for the person using the product and what a useful result would look like. Input: The current workflow, business purpose and the problem I have observed. Output: A concrete brief with a reason to build or change something. Tools: VS Code, Project documents #### 2. Read the implementation before proposing a fix AI role. I ask for a source-based assessment of the code, data flow and existing project records, with assumptions challenged explicitly. Input: The brief, current repository and relevant source material. Output: A diagnosis tied to the implementation and a defined next phase. Tools: OpenAI Codex, Repository search #### 3. Build the context for the next agent My role. I set the roles, the required source context and the information that must survive the handoff. Input: The findings, earlier decisions and unresolved work. Output: A task-specific execution brief rather than a fresh generic prompt. Tools: Markdown blueprints, Task and handoff files #### 4. Work in bounded specialist lanes AI role. Agents inspect or implement their assigned parts, record changes and return findings for integration. Input: A shared project plan and the context relevant to each task. Output: Code changes, implementation notes and issues to resolve. Tools: Claude Code, OpenAI Codex, VS Code #### 5. Challenge the result My role. I check whether the product behaves as intended, question unsupported claims and ask for testing and review before the next phase. Input: The changed product, test output and the original business requirement. Output: A correction, a decision to continue or a narrower next task. Tools: Local application, Test results, Code review #### 6. Leave the next session a usable state System role. Project files retain the decisions, implementation state and remaining work outside the chat window. Input: The current result and its unresolved dependencies. Output: A durable starting point for the next development session. Tools: Markdown, Repository history, Project handoffs ### Context engineering I treat context engineering as a working responsibility. I need to know what the agent should understand, where that knowledge comes from and what the next person or role will receive. The procedure and the handoff are part of the engineering. #### Business purpose The real workflow, the intended user and the reason the feature matters. #### Grounding material Relevant project records, earlier decisions and repository evidence selected for the task. #### Project state What exists, what changed, what failed and which dependencies remain unresolved. #### Execution contract The specialist role, the task boundary, the expected output and the checks to perform. #### Correction and handoff The result, its evidence and the information the next agent needs before continuing. ### Implementation I work from the files and the current project state, then delegate analysis, implementation and review with task-specific context. I keep separate harness guidance for Claude and Codex because their tools and behaviour need to be understood in their own environment. Claude Code, OpenAI Codex, VS Code, Markdown, Git, PowerShell Module blueprints connect the business purpose to the implementation and outstanding work. A PowerShell handoff utility writes the sender, receiving agent, task, touched files, verification, risks and next actions into Markdown, with the Git branch and bounded working-tree state. Bounded execution packs name the objective, scope, interfaces, assumptions to reject, checks, blueprint updates and completion criteria. Specialist review is separated from implementation when another perspective is needed. Local testing and staged cloud work keep product validation separate from publication. Project context stays in inspectable files. It gives the next task a working memory without changing the model weights. ### Results #### Working output A file-based agent handoff. The implemented handoff utility carries task state and verification between tools, rather than relying on conversation memory. #### Context refinement March 2026. Living module blueprints, task-specific specialists and a source-based review before execution. #### Delivery refinement June 2026. Explicit builder and tester roles, local QA, end-to-end testing and security review before cloud publication. ### Next step I want another person to try the public context-and-handoff guide on one bounded change. Can they understand the project and continue the work without the original chat? Their experience will help me make the method easier to share with a team. [Context and handoff guide](/context-handoff.md) · [Read the project casebook](/project-stories.md) Drawn from dated working instructions and project artifacts. Client documents and original conversations remain private. ## Local model laboratory Model adaptation and provider routing. Experiments recorded in 2026. Experimental implementation and saved run artifacts. I wanted to understand how the working system could keep its own context and use different models. Alongside my everyday work with hosted models, I directed experiments with small instruction models, adapter training and model routing. ### The problem I wanted more flexibility in how a workflow uses models. Which parts could run through smaller local models? What changes with a domain-specific adapter? How should a request move between providers when its requirements change? Those questions gave the experiments their purpose. My question was about the architecture as much as the model. Could the workflow keep its knowledge and choose a suitable model, rather than rebuild everything whenever the model changed? ### Decisions I made #### Adapt an existing instruction model The lab uses parameter-efficient fine-tuning with LoRA adapters over pretrained Qwen2.5 models. This keeps the experiment focused on a task and a small trainable parameter set. #### Make the run inspectable Each saved run carries its training configuration, sample count, runtime, parameter counts and adapter artifact. I can return to the setup instead of relying on a chat summary. #### Separate model selection from dispatch The selector uses task sensitivity, provider availability and budget signals, and returns the reason for its choice. Dispatch is a separate operation that records failed provider attempts and the fallback chain. ### How the work moves #### 1. Define the job for the model My role. I frame the capability and portability question before committing to a model or a training run. Input: The workflow, task sensitivity and available hardware. Output: A bounded experiment and a candidate base model. Tools: Claude, Codex #### 2. Prepare instruction examples System role. The trainer loads prompt/completion pairs, formats them as instruction examples and tokenises them with a sequence-length limit. Input: A selected local JSONL training set. Output: A tokenised dataset for the experiment. Tools: Python, Transformers #### 3. Train a small adapter System role. The training code attaches low-rank adapters, configures rank and target projections, and uses gradient accumulation. Saved variants change the model size and adapter setup. Input: A pretrained instruction model, dataset and training configuration. Output: An adapter checkpoint and recorded training summary. Tools: PyTorch, PEFT, LoRA #### 4. Inspect the saved run My role. I review the configuration and saved artifacts to understand what ran and the scale of the experiment. I need a separate held-out comparison to assess task quality. Input: Runtime, parameter counts, logs and saved adapter files. Output: An inspectable experiment record and the next evaluation question. Tools: Run summaries, Adapter configuration #### 5. Choose and dispatch System role. Model selection returns a provider choice and the reason for it. The separate dispatch path records failed provider attempts and the fallback chain. Unit tests exercise routing decisions without paid inference. Input: A request, sensitivity classification and availability signals. Output: A provider choice or an explicit failure chain. Tools: Python, Provider adapters ### Context engineering I keep model configuration, task context and evaluation evidence separate so I can understand what each change does. An adapted model still needs a clear assignment and the right working context. #### Task What the workflow needs, why the task is sensitive and what result a person will accept. #### Training configuration Base model, adapter rank, target projections, learning settings and selected examples. #### Run record Saved adapter, parameter counts, runtime and training logs tied to the same experiment. #### Runtime request Task context and provider constraints assembled at inference time. ### Implementation My local lab contains Python training and routing code, saved PEFT adapters and configuration records for Qwen2.5 0.5B and 1.5B instruction models. I work with Claude and Codex on the engineering. The training workload uses PyTorch and Hugging Face tooling. Python, PyTorch, Transformers, PEFT, LoRA, Qwen2.5, Provider abstraction The trainer exposes adapter rank, alpha, target modules, learning-rate schedule, gradient accumulation and checkpointing as run parameters. One inspected 1.5B configuration targets query, value, output and down projections with rank 16 and alpha 32. The routing tests cover input validation, task sensitivity and provider-health signals. A budget signal in the router is not a hard spending limit. ### Results #### Saved experiment Qwen2.5 1.5B. The inspected adapter configuration records the pretrained instruction model and rank-16 adapter setup. #### Run scale 2,721 examples. The saved 1.5B run summary records two epochs and 8,257,536 trainable parameters. These are experiment dimensions, not an accuracy score. #### Routing check 13 tests passed. Model-selection unit tests rerun on 9 October 2026, with no paid-provider call or new training run. ### Next step I want to compare the base model and adapter on the same licensed held-out tasks. I will record acceptance quality, failure cases, latency and cost before deciding what role the adapted model could play in a working system. [Read the complete project casebook](/project-stories.md) · [Bounded model control patterns](/ai-labs/bounded-control-plane/) · [Inspectable public control-plane reference](https://github.com/Haris88m/servari-open) Training code, adapter configurations and saved summaries were inspected on 9 October 2026. Training data and weights remain private. The next held-out evaluation is planned work.