{
  "version": "2026-10-10-os-v2",
  "profile": {
    "name": "Haris Mekic",
    "role": "AI consultant and operations practitioner. Founder of MEKreflect.",
    "location": "Skopje, North Macedonia",
    "introduction": "I am Haris Mekic, an AI consultant and operations practitioner. I turn working knowledge into context, procedures and tools that people can inspect. This is my personal portfolio of experience, implementation and learning. I started MEKreflect to develop that practice and, over time, build it into a business with a team.",
    "paragraphs": [
      "My experience comes from keeping work connected. Across United Nations operations and three separate EU technical assistance projects, I worked with office leadership, regional teams, international experts and public institutions. I had to understand the programme as well as the budget, the people and the decisions needed to deliver it.",
      "I develop and track project budgets because the numbers have to reflect the work. I connect activities to costs, review expenditure and commitments, and keep revisions aligned with delivery. Monthly cash requirements, cash flow projections, expert days, procurement and payments all belong to that same picture.",
      "For almost two years I have also been building with AI. I work between Claude, Codex and VS Code, inspect changed files and test the result. I coach agents through context engineering, role instructions, examples and feedback. I define the work, inspect what comes back and refine the procedure when the result misses the purpose.",
      "I bring the same method to people learning to work with AI. We explain the job, select the right sources and establish how to question the output. That understanding becomes the assistant's working context. The aim is better judgement and a workflow that is clearer and easier to repeat.",
      "My mission is a digital office where people and AI workers carry connected work forward. An approved brief becomes research, a delivery plan, budget analysis, a presentation and a reviewed decision. I want to build this with a team, develop their capability and expand from workflows that have been tested in use."
    ],
    "workingStyle": [
      "Understand the work and the people doing it.",
      "Turn the idea into a costed plan with clear responsibilities.",
      "Bring in specialists where the problem needs depth.",
      "Build, test and improve with the people who will use it.",
      "Stay responsible for the outcome."
    ],
    "contact": {
      "email": "harism88@outlook.com",
      "href": "mailto:harism88@outlook.com"
    },
    "portrait": "/haris-portrait.png"
  },
  "logbook": {
    "intro": "This is how I develop people and AI working methods together. I turn operational knowledge into context, role instructions and reviewable outputs, then improve the procedure through use and feedback. The journal connects my applications, local model experiments and coaching practice to that same digital office mission.",
    "boundary": "This journal combines selected working conversations, current source inspection and specific recorded checks. It describes almost two years of personal AI development. Dates are retained as provenance, not a measure of skill or working hours. Private conversations and client records are not published.",
    "entries": [
      {
        "id": "my-working-loop",
        "title": "Start with the work, not the model",
        "date": "2026-10-09",
        "status": "Owner history and implementation reviewed",
        "summary": "I coach AI agents through the work. I explain the operating situation, shape their responsibilities, inspect the output and correct what misses the point. That conversation becomes context, procedures and acceptance checks in the application. When I guide another person, I help them learn the same discipline so they can direct the AI themselves.",
        "changes": [
          "I read the requirements and source material, then define what the result needs to do and how I will check it.",
          "I use Claude, Codex and VS Code for context engineering, implementation, review and debugging. Builders and testers receive different responsibilities so I get a separate view of the result.",
          "I treat coaching the agent as an engineering loop. I clarify the context, define an output contract, inspect its behaviour and carry corrections into the next instruction or implementation. This is contextual configuration. My separate local model adaptation work changes model parameters.",
          "I teach the person how to brief, question and review the agent. The goal is a working method they can explain and repeat, not dependence on a prompt they do not understand.",
          "I use fan-out and fan-in orchestration when commercial, delivery and compliance questions need their own analysis. Failed branches remain visible when the findings come together.",
          "I open the changed files and use the application. I compare what it does with the requirement and keep the existing workflow working as I improve it.",
          "Where the workflow supports them, I use deterministic contract tests, negative fixtures, source-linked context and receipt-bound completion. I want a result that I or the next person can inspect."
        ],
        "verification": "Selected original working conversations were compared with the current source and Git history. They document company-specific context, separate builder and reviewer responsibilities, repository handoffs and repeated challenges to unsupported outputs.",
        "limits": "I direct the implementation and review. Each workflow has its own checks and acceptance requirements.",
        "references": [],
        "practice": {
          "input": "I start with the people doing the work, the decision they need to make and the source material behind it. I describe the constraints and ask what would make the result useful to the next person.",
          "context": "I turn that into a working brief with the objective, domain rules, relevant records, permitted actions and acceptance checks. Corrections become updated context or procedures rather than disappearing into a chat. I keep decisions and implementation state in the project so the next session can continue the work.",
          "agentWork": [
            "I ask an assessment agent to inspect the existing workflow before proposing changes.",
            "Implementation agents receive separate responsibilities and file boundaries. A reviewer checks the behaviour and the evidence rather than repeating the builder's completion message.",
            "The next agent receives the changed files, findings, test results and unresolved questions. I review the result in the application and redirect the work when it misses the purpose."
          ],
          "output": [
            "A versioned implementation, a usable interface and an explicit account of what works and what remains.",
            "The working record belongs to the project. It is not trapped in one conversation or tied to one model."
          ],
          "stack": [
            "Context engineering",
            "Domain modelling",
            "Contextual SOPs",
            "Repository based handoffs",
            "Claude",
            "Codex",
            "VS Code",
            "Separate implementation and review"
          ],
          "next": "I want colleagues to configure a digital worker through normal questions about their work. That conversation should become a role, a procedure, a set of tools and a clear handoff that the team understands. We can then measure the complete work cycle and improve both the procedure and the person's ability to direct it."
        }
      },
      {
        "id": "coaching-a-working-method",
        "title": "Develop the person and the AI working method",
        "date": "2026-10-10",
        "status": "Individual coaching and an initial scenario trial",
        "summary": "I helped a consultant turn a broad career goal into a practical way of working with AI. We connected the responsibilities he wanted to develop with the context an assistant needs and the judgement that remains with him. He then tried the method independently on a hypothetical consulting assignment and gave specific feedback on what helped.",
        "practice": {
          "input": "A consultant's development goal, the project responsibilities he wanted to practise and questions about using an AI assistant more effectively. The first exercise used a hypothetical consulting scenario.",
          "context": "I separated essential inputs from helpful background, connected the assignment to its intended recipient and defined how to check the result. I translated that into an input checklist and reusable instructions for the assistant and the person directing it.",
          "agentWork": [
            "The assistant is instructed to identify the objective, available sources, missing inputs and decisions that require a person before developing a solution.",
            "I designed role procedures for research, delivery planning, budget and risk analysis, presentation development and quality review. Each has an input, output and handoff.",
            "I developed the next-stage plan around an approved digital archive, a source register, weekly delivery and learning reports, and a supervised pilot with a project manager.",
            "The person learns to challenge assumptions and review the output. Feedback then refines the assistant's context and procedure."
          ],
          "output": [
            "A context-first working method, input checklist and instructions that the participant tried independently in a hypothetical assignment.",
            "Written feedback described the distinction between essential inputs and helpful background as useful for moving the work forward.",
            "A proposed 90-day pilot and development pathway, with roles, document workflows, review points and measures to establish before wider use."
          ],
          "stack": [
            "Context engineering",
            "Copilot working instructions",
            "Role based SOPs",
            "Input and output contracts",
            "Source provenance",
            "Human review",
            "Workflow evaluation",
            "Practical coaching"
          ],
          "next": "Agree a real assignment with an authorised owner, establish preparation and review baselines, and test whether another person can repeat the workflow. Expand from accepted outputs and feedback."
        },
        "changes": [
          "I developed the person and the assistant's working method together. Clearer briefing and review become better context and more useful instructions.",
          "I turned a broad development ambition into a proposed sequence of responsibilities, deliverables, learning goals and review points.",
          "The proposed archive links authoritative records rather than collecting unrestricted copies. Client work remains separated and access follows existing permissions.",
          "The reporting design distinguishes project delivery from the learning record. It tracks inputs, outputs, corrections, accepted work and measured effort when available.",
          "This case concerns contextual assistant configuration and coaching. My local parameter-efficient model adaptation experiments are a separate technical evidence track."
        ],
        "verification": "The participant's written response confirms an independent hypothetical exercise and specific feedback on the method. The subsequent detailed working plan and assistant instructions were sent. This account is anonymised and paraphrased. Private correspondence is not published.",
        "limits": "The completed work is individual coaching, a scenario trial and a proposed implementation plan. A live organisational pilot, measured productivity improvement and wider team adoption are next-stage tests.",
        "references": []
      },
      {
        "id": "company-bid-workflow",
        "title": "Build a bid around the company",
        "date": "2026-10-10",
        "status": "Implementation and owner decisions reviewed",
        "summary": "I started with how a consulting company actually wins work. More tenders were not the answer if the opportunity did not fit, the buyer was unclear or the expert information was incomplete. Conversations with staff and my own delivery experience shaped the system around those decisions.",
        "practice": {
          "input": "An opportunity and its original documents, the assignment requirements, evaluation criteria, languages, company capability and expert roster. Missing source documents remain a gap to resolve.",
          "context": "I connect company history and relevant previous proposal material with the specific section or decision being worked on. I asked for living module blueprints that explain what each part does and follow the actual implementation.",
          "agentWork": [
            "Business development reviews fit and commercial direction. Bid coordination examines the team, expert matches, CV actions and mobilisation. Data and compliance examines the available evidence.",
            "These three specialist branches run in parallel. Operations synthesis receives their available findings and preserves partial or failed branches.",
            "Section generation uses requirements, evaluation criteria, risks and relevant specialist findings. Preparation services create tasks for missing documents, context, analysis and approaching deadlines."
          ],
          "output": [
            "The bid manager receives expert matches, team gaps, staffing confidence, CV actions, coordination notes and a budget range with its basis when the roster supports one.",
            "The workspace keeps preparation tasks and internal deadline records connected to the opportunity. Proposal assembly brings sections and supporting material together for review."
          ],
          "stack": [
            "TypeScript",
            "Parallel specialist orchestration",
            "Promise.allSettled",
            "Structured generation",
            "Zod response schemas",
            "Role filtered tool use",
            "Task and calendar state"
          ],
          "next": "Finish the wider delivery plan, strengthen concurrent task creation and verify the complete handoff with a bid manager. The same approach can fit another company by rebuilding its capability context, roles, documents and approval rules with its staff."
        },
        "changes": [
          "I separated the commercial, staffing and evidence questions so one fluent answer could not hide an incomplete perspective.",
          "The section generator uses bounded context, including a limit on prior proposal material and short excerpts of other sections for continuity.",
          "Expert references in matches and CV actions are checked against the retrieved roster. Without a roster the result stays partial and does not supply a budget estimate.",
          "The calendar service updates pipeline generated deadline records while leaving manually created events alone. This is internal calendar state, not a claim of external calendar delivery.",
          "The response schema checks structure. Factual accuracy, commercial assumptions and readiness still require review.",
          "The inspected structured generation path uses Gemini 2.5 Flash. The separate conversational tool interface uses a Claude adapter. Claude and Codex are also my development tools, which is a different role from the models called inside the application."
        ],
        "verification": "Current source paths and selected Git revisions were compared with the original working conversations. The parallel branches, bid manager output, proposal context, preparation tasks and internal deadline updates were inspected separately.",
        "limits": "The implementation is distinct from this public example. Client records remain private. The wider delivery plan and some task and tool boundaries still need work. No bid win rate, production usage volume or time saving is claimed.",
        "references": []
      },
      {
        "id": "digital-worker-workspace",
        "title": "Give digital workers a shared workspace",
        "date": "2026-10-10",
        "status": "Python implementation and working history reviewed",
        "summary": "I do not want a digital office to be a chat window with a different name. I want to see the request, the role working on it, the result and the decision in one place. That is the direction behind my agent working system. A good looking shell is not enough if the work cannot move between roles.",
        "practice": {
          "input": "A person's objective, the way their team works, the project state, the available tools and the actions that need an owner decision.",
          "context": "I define what belongs in the active work record and what must survive the next session. The context policy checks the work log, current task, decision queue, session record and a place for recording risks.",
          "agentWork": [
            "The Python layer connects the interface to local records and a configurable model conversation.",
            "The model adapter sends a bounded recent conversation and an explicit output limit. Missing configuration returns a clear unavailable result instead of an invented answer.",
            "A separate policy evaluates the action against the autonomy level and risk score. The result is to proceed, report or return the decision for review.",
            "The change review path measures a defined baseline, snapshots enrolled files and supports a keep or restore decision."
          ],
          "output": [
            "A working record the next session can inspect, explicit decision states and a model reply that stays separate from permission to act.",
            "Checks and snapshots provide evidence for selected changes. The wider multi-person digital office remains the product direction."
          ],
          "stack": [
            "Python standard library",
            "JSON APIs",
            "React",
            "Vite",
            "Optional Electron shell",
            "File backed state",
            "argparse CLIs",
            "Provider abstraction",
            "subprocess timeouts",
            "SHA256 verification"
          ],
          "next": "Make role setup and the request-to-result path easier for another person to use. Connect each permitted tool to the same review rules and test the whole workflow before expanding the number of integrations."
        },
        "changes": [
          "The provider adapter includes the last twenty conversation turns and an output token cap. The operating method remains separate from the selected model.",
          "Context pressure uses transcript file size and work-record freshness as signals. It is a practical proxy rather than direct model token telemetry.",
          "The decision policy accepts integer risk scores from 4 through 20 and rejects malformed values. High risk still returns to review at the highest autonomy level.",
          "The retention path runs configured checks with timeouts and compares a baseline with a later result. Selected files can be restored and verified by hash.",
          "My role is architecture, operational rules and acceptance. AI engineering agents implement and test code with me, as the public source acknowledges."
        ],
        "verification": "The Python model interface, context lifecycle, policy and retention paths were read. Six isolated offline probes checked malformed risk inputs, high risk decisions, missing and stale context, the conversation bound and missing model configuration. Earlier full suite results remain in the public repository history.",
        "limits": "These are local implementation and offline checks, not a claim that every proposed integration is complete. File size is not a token meter. Recovery protects the enrolled records and files, not every possible application state.",
        "references": [
          {
            "label": "Public agent workspace source",
            "url": "https://github.com/Haris88m/servari-open"
          }
        ]
      },
      {
        "id": "python-in-my-practice",
        "title": "Make the operating rules explicit in Python",
        "date": "2026-10-10",
        "status": "Source and focused checks reviewed",
        "summary": "I use Python where the work needs structured data, persistent state and a result I can inspect. My practice is to define the operational rule, develop it with AI, and then challenge its behaviour. I would rather show that process than describe myself with a proficiency label.",
        "practice": {
          "input": "Source text, a question, a case record, a proposed state change or a confirmation file. I define what counts as evidence and what must never be accepted as completion.",
          "context": "The working record connects the source, the case identity, the allowed next state and the verification requirement. Text retrieval uses bounded excerpts and lexical matching for the question.",
          "agentWork": [
            "Python retrieves relevant source excerpts and maintains structured records in SQLite.",
            "The workflow separates preparation, an attempted action and confirmed completion. A receipt must belong to the same case.",
            "Integrity checks compare the recorded SHA256 with the confirmation file. Missing, unrelated or changed receipts leave completion unconfirmed.",
            "Tests exercise the expected case and the boundary cases, including malformed input, stale state and missing evidence."
          ],
          "output": [
            "A reviewable result with its source and state, rather than a free text success claim.",
            "Small tools that can be run again with controlled inputs and inspected through their tests."
          ],
          "stack": [
            "Python",
            "SQLite",
            "JSON and JSONL",
            "Lexical retrieval",
            "Hash based integrity",
            "State machines",
            "CLI interfaces",
            "Temporary fixtures",
            "unittest"
          ],
          "next": "Keep extending tests around the real boundaries and improve independent setup instructions. I want another person to be able to reproduce a result without needing my original chat."
        },
        "changes": [
          "The source query uses all-term lexical matching. I describe that precisely instead of calling it vector search.",
          "The coordinator checks case ownership, timing, package integrity and receipt identity before recording completion.",
          "Persistent workflow state and model-generated prose are different things. The prose does not decide that an external action succeeded.",
          "I define the operational rule, work with Claude and Codex on implementation, then inspect the source and test its behaviour."
        ],
        "verification": "Current source retrieval and workflow coordinator code was inspected. The earlier receipt workflow has 18 focused tests recorded with in-memory databases and temporary files. This review did not run live applications or access their private records.",
        "limits": "The public account describes engineering patterns and test scopes. It does not publish private records or imply that separate database updates form a single distributed transaction.",
        "references": []
      },
      {
        "id": "models-and-evidence",
        "title": "Choose the model for the work",
        "date": "2026-10-09",
        "status": "Model practice and implementation reviewed",
        "summary": "I like understanding what different models can do with the same problem. I use GPT and Claude model families for reasoning and coding, with lighter routes where they fit the task. That interest has also led me to develop model selection and provider abstraction inside applications.",
        "changes": [
          "I use GPT-family coding and reasoning models through Codex for implementation, debugging and specialist review.",
          "I use Claude-family models for planning, reasoning, context development and a second perspective on the work.",
          "I work with coding assistants in VS Code, keeping model output connected to source files, runtime behaviour and tests.",
          "Inside applications I have implemented provider abstraction, bounded chat history and sensitivity-aware model-tier selection, with explicit handling when a provider is unavailable. Cost preferences inform model selection."
        ],
        "verification": "Timestamped local model metadata was reviewed separately from configuration. The model-tier selection path had 13 focused local checks. Three route/dispatch checks used mocked dependencies. These test scopes are not added to usage counts.",
        "limits": "Local runtime and response labels are retained with their source type. Quality is assessed through the task result and workflow checks.",
        "references": [],
        "practice": {
          "input": "The task, its sensitivity, provider availability and the intended cost profile.",
          "context": "I keep the model, the agent harness and the operating method separate. The same working brief can be reviewed with different models, but their capabilities and results still need checking.",
          "agentWork": [
            "A selection path evaluates sensitivity and the available model tiers.",
            "Provider adapters translate the bounded conversation and handle an unavailable provider explicitly.",
            "I use a separate implementation or review pass when another perspective is useful."
          ],
          "output": [
            "A selected route and an inspectable result within that application's context and output constraints.",
            "Thirteen deterministic routing checks verify selection behaviour, not comparative model intelligence."
          ],
          "stack": [
            "Provider abstraction",
            "Sensitivity aware routing",
            "GPT family",
            "Claude family",
            "Bounded history",
            "Structured outputs",
            "Offline selection tests"
          ],
          "next": "Build task-level evaluations that compare quality, corrections and running cost. A routing preference alone is not a hard spending limit."
        }
      },
      {
        "id": "small-model-experiments",
        "title": "Learning through small-model adaptation",
        "date": "2026-10-09",
        "status": "Historical artifacts and source reviewed",
        "summary": "I wanted to understand more of what happens below the chat interface. I explored local adaptation of small pretrained Qwen2.5 instruction models with PEFT LoRA, working through training inputs, low-rank adapters, checkpoints and evaluation.",
        "changes": [
          "I worked with a training path that tokenises prompt and completion pairs, attaches low-rank adapters and uses gradient accumulation and checkpointing.",
          "I kept the adapters and training summaries from completed experiments. Separate evaluation tooling compares the base model with adapted output.",
          "I recorded factual errors and identity confusion during evaluation, then identified held-out task comparisons as the next step before using an adapter in a real workflow."
        ],
        "verification": "Trainer code, saved summaries and adapter provenance were inspected. Earlier dated verification compared ten summary and ten adapter hashes with the prior ledger. This review did not retrain a model or load tensors.",
        "limits": "Experimental LoRA adaptation of pretrained models. The next evaluation step is a held-out comparison of task quality.",
        "references": [],
        "practice": {
          "input": "Prompt and completion pairs, a small pretrained instruction model and the memory limits of a local experiment.",
          "context": "I wanted to understand the training path below the chat interface, including tokenisation, adapter parameters, checkpoints and comparison with the base model.",
          "agentWork": [
            "The Python trainer formats JSONL examples and tokenises the text with a bounded sequence length.",
            "Transformers and PyTorch load the pretrained causal model. PEFT attaches low rank adapters to selected projection modules.",
            "Gradient accumulation and checkpointing support the training run. The pipeline saves the adapter and a run summary."
          ],
          "output": [
            "A saved Qwen2.5 1.5B adaptation run records 2,721 training pairs across two epochs with LoRA rank 16.",
            "The saved adapter is an experimental artifact. Whether it improves the work needs a separate held-out evaluation."
          ],
          "stack": [
            "Python",
            "PyTorch",
            "Transformers",
            "PEFT",
            "LoRA",
            "Qwen2.5",
            "JSONL preparation",
            "Gradient accumulation",
            "Gradient checkpointing"
          ],
          "next": "Compare the base and adapted model on held-out operational tasks with the same scoring rules, checking factual errors as well as useful answers."
        }
      },
      {
        "id": "make-completion-inspectable",
        "title": "Make completion something I can inspect",
        "date": "2026-10-09",
        "status": "Focused local tests",
        "summary": "I wanted the next person or agent to know what had actually happened. I developed a persistent workflow that keeps preparation, an attempted action and confirmed completion separate. The confirmation has to belong to the same case and pass an integrity check.",
        "changes": [
          "I keep the workflow state in a persistent record so the next task has a dependable starting point.",
          "I require an identity-matched receipt before recording completion.",
          "The workflow compares the recorded SHA-256 against the actual confirmation file.",
          "Missing, unrelated or changed receipts leave completion unconfirmed."
        ],
        "verification": "Recorded verification includes 18 focused workflow-engine tests with in-memory databases and temporary receipts. The current review inspected the persistent workflow and integrity checks without running external applications.",
        "limits": "Isolated workflow tests cover case identity, receipt integrity and completion states with temporary fixtures.",
        "references": [],
        "practice": {
          "input": "A prepared case, an attempted action and a confirmation artifact.",
          "context": "The case identity, expected state and confirmation requirement remain in a persistent record so the next worker can distinguish preparation from completion.",
          "agentWork": [
            "The workflow records preparation and attempted execution as separate states.",
            "It checks that the receipt belongs to the case and that its SHA256 matches the file.",
            "A missing or changed receipt leaves completion unconfirmed."
          ],
          "output": [
            "An explicit confirmed or unconfirmed state with evidence that can be inspected.",
            "The next worker does not have to infer success from a narrative message."
          ],
          "stack": [
            "Python",
            "SQLite",
            "State transitions",
            "SHA256",
            "Temporary receipt fixtures"
          ],
          "next": "Extend the same evidence requirement to more handoffs and keep tests around interrupted work and stale state."
        }
      },
      {
        "id": "keep-the-source-visible",
        "title": "Keep the source visible after the model answers",
        "date": "2026-10-09",
        "status": "Implementation inspected",
        "summary": "Context engineering is where my operations knowledge and AI work meet. I decide which requirements, criteria, risks and source material belong in the task. I have developed selection and generation paths that keep that context connected to an answer someone can review.",
        "changes": [
          "I use lexical matching to retrieve material for the specific task.",
          "I check source sufficiency before generation so missing material changes what happens next.",
          "I give parallel specialists a bounded shared context and a separate question to work through.",
          "I validate the output structure and retain missing or failed perspectives for review."
        ],
        "verification": "The current orchestration, context and generation implementations were opened and inspected as separate paths. In the orchestration path, three specialist branches feed a separate synthesis stage.",
        "limits": "These patterns are implemented in separate applications. Retrieval uses lexical matching, with source sufficiency and structured-output checks at defined stages.",
        "references": [],
        "practice": {
          "input": "A question, the relevant assignment requirements and the available source material.",
          "context": "I select what the task actually needs and preserve missing evidence. A longer prompt is not automatically a better understanding of the work.",
          "agentWork": [
            "Retrieval uses lexical matching for the particular question.",
            "Source sufficiency checks determine whether generation has enough material to proceed.",
            "Parallel specialists examine different questions from a bounded shared context.",
            "Structured output checks keep their findings usable by the next stage."
          ],
          "output": [
            "An answer or specialist finding linked to its supporting context.",
            "Missing and failed perspectives remain visible when the work is brought together."
          ],
          "stack": [
            "Lexical retrieval",
            "Context selection",
            "Structured outputs",
            "Parallel orchestration",
            "Source sufficiency checks"
          ],
          "next": "Compare context variants on fixed reviewed cases so source coverage and omissions can be evaluated consistently."
        }
      },
      {
        "id": "review-the-boundary",
        "title": "Find the case the policy did not handle",
        "date": "2026-10-09",
        "status": "Published and tested",
        "summary": "A review found that the public decision policy could accept malformed risk scores after numeric coercion. I directed the correction and the regression tests. This was a useful reminder to inspect the difficult inputs as carefully as the expected ones.",
        "changes": [
          "The policy now accepts integer scores from 4 through 20 only.",
          "Strings, booleans, floats, missing values and out-of-range scores queue for review.",
          "The command-line boundary parses an integer explicitly, preserving the valid command contract.",
          "224 parametrized malformed-input, boundary and CLI regressions were added to the original 160 tests."
        ],
        "verification": "Recorded verification includes 384 Python tests and eight separate checks, followed by published CI for the Python suite, verifier and UI build. The latest targeted review also exercised malformed score inputs and high risk decisions in isolated offline probes.",
        "limits": "Local policy tests use mocked providers and synthetic fixtures. The public repository includes the implementation and reproduction steps.",
        "references": [
          {
            "label": "Public implementation and verification",
            "url": "https://github.com/Haris88m/servari-open"
          }
        ],
        "practice": {
          "input": "Risk scores submitted to the public Python decision policy, including strings, booleans, floats, missing values and values outside the supported range.",
          "context": "I wanted the policy to apply the intended rule rather than treating numeric coercion as valid authority.",
          "agentWork": [
            "The implementation pass makes the integer boundary explicit.",
            "The review pass checks malformed values and valid boundary inputs.",
            "Command line parsing preserves a valid integer input while unsupported values return to review."
          ],
          "output": [
            "Integer scores from 4 through 20 are accepted by the policy. Malformed values queue for review.",
            "The implementation and reproduction tests are available in the public repository."
          ],
          "stack": [
            "Python",
            "Input validation",
            "Parametrized regression tests",
            "CLI boundary testing"
          ],
          "next": "Apply the same discipline to tool calls and state transitions so the safe example is not the only case that works."
        }
      },
      {
        "id": "my-ai-roadmap",
        "title": "Where I am taking this",
        "date": "2026-10-10",
        "status": "Personal direction and next work",
        "summary": "My mission is to make professional knowledge usable by people and digital workers together. I want to build the team as well as the system. The foundation is operational experience, contextual procedures and an honest view of what the software does.",
        "practice": {
          "input": "A team willing to explain its work, one useful workflow and the records needed to understand the current way of doing it.",
          "context": "I start small enough to learn properly. We agree the purpose, the role boundaries, the owner and what an accepted result looks like before adding more agents or systems.",
          "agentWork": [
            "First make the role and context contracts easy to configure from the team's own language.",
            "Then finish and test one end-to-end handoff with real authorised users, including incomplete sources and failed tools.",
            "Compare the full work cycle, including preparation, corrections, review effort and accepted output.",
            "Use the findings to improve the workflow and teach the team how to inspect and maintain it."
          ],
          "output": [
            "The next milestone is a repeatable team workflow with recorded feedback and a measured baseline.",
            "Longer term I want the same operating foundation to adapt across bidding, budgets, programme delivery and other company work without pretending every organisation is the same."
          ],
          "stack": [
            "Role configuration",
            "Context evaluation",
            "End to end testing",
            "Workflow observability",
            "User feedback",
            "Team learning"
          ],
          "next": "A team pilot with a named owner, agreed baseline and a short review cycle. Broader autonomy follows demonstrated reliability, not an impressive interface."
        },
        "changes": [
          "The foundation already includes company-grounded bid context, specialist handoffs, local state and policy checks, budget calculations and reusable public examples.",
          "The next work is easier role configuration, stronger tool boundaries and fuller delivery planning.",
          "Context Lab will compare source and prompt variants against reviewed questions. Workflow Observatory will measure the complete task, not just the speed of an answer.",
          "I want to keep exploring models and local adaptation while completing useful work with people. The model is part of the system, not the whole system."
        ],
        "verification": "This roadmap is drawn from my retained working requests and the current implementation review. It separates existing foundations from proposed capabilities and measurements.",
        "limits": "The roadmap is direction, not a claim that team adoption, measured savings or the planned evaluation applications already exist.",
        "references": []
      },
      {
        "id": "recorded-model-practice",
        "title": "A sustained working practice",
        "date": "2026-10-09",
        "status": "Local metadata audited",
        "summary": "AI development has become a sustained part of my working life. My retained records show activity on 121 distinct dates between 27 January and 29 September 2026 across Codex, Claude and VS Code.",
        "changes": [
          "I use Codex for implementation, code review and coordinated specialist work.",
          "I use Claude for reasoning, planning, context development and a second perspective on a problem.",
          "I work in VS Code to connect the conversation to actual files, application behaviour and tests.",
          "I carry decisions forward through project memory, structured context and versioned source files."
        ],
        "verification": "Read-only metadata review on 9 October 2026, excluding events on or after that date to avoid counting this assessment. Daily dates were combined as a set, not summed across tools. No raw conversation corpus is published.",
        "limits": "The metric counts distinct UTC dates in retained local records, including agent activity. Working hours and financial savings require a separate measurement method.",
        "references": [],
        "practice": {
          "input": "The problem I am working on and the current project record, whether I am using a coding agent, the editor or a separate review session.",
          "context": "I keep the work tied to files, requirements and previous decisions. Using several tools only helps when they can continue from the same understanding.",
          "agentWork": [
            "Coding agents inspect and implement bounded parts of the project.",
            "Separate reasoning and review passes challenge the proposed design, changed files and outputs.",
            "Working notes carry decisions and unfinished checks into the next session."
          ],
          "output": [
            "A sustained body of project work rather than a single prompt experiment.",
            "The retained activity record supports the continuity of the practice. It is not a timesheet."
          ],
          "stack": [
            "Codex",
            "Claude",
            "VS Code",
            "Project memory",
            "Versioned source"
          ],
          "next": "Track the outcome and review effort of future team workflows directly, instead of estimating impact from activity records."
        }
      }
    ]
  },
  "ideas": {
    "intro": "I want the next stage to include other people working with what I have built. These ideas are about understanding what helps them, making the context easier to inspect and learning whether the whole workflow improves. I want the team to grow with the system.",
    "entries": [
      {
        "id": "test-the-context",
        "title": "Evaluate the context",
        "status": "Planned",
        "problem": "I want to know which source material makes a difference to the answer and where the context is still incomplete.",
        "proposal": "I plan to build a fixed set of operational questions and define which sources each answer needs.",
        "next": "I will compare context and prompt variants against the same human-reviewed cases.",
        "acceptance": "The published comparison will show the inputs, supported answers and omissions so another person can inspect the result.",
        "boundary": "Planned evaluation with human-reviewed reference cases."
      },
      {
        "id": "measure-the-whole-workflow",
        "title": "Measure the whole workflow",
        "status": "Planned",
        "problem": "I care about the whole job, including preparation, corrections and review. That is where I need to understand whether the change is useful.",
        "proposal": "I plan to measure those stages alongside execution and the accepted outcome.",
        "next": "I want to run a workflow pilot with a user and record a baseline before comparing the change.",
        "acceptance": "I will keep a record of quality, review effort, direct running cost and user feedback.",
        "boundary": "Pilot design for measuring operational value."
      },
      {
        "id": "make-handoff-real",
        "title": "Put a working method in someone else’s hands",
        "status": "Available for reuse",
        "problem": "I want someone else to be able to start, test and understand the workflow without needing me beside them.",
        "proposal": "I have packaged the budget-control engine with editable inputs, source code, tests and a guide. Budget Studio uses the same calculation module.",
        "next": "I want an independent user to try the pack on a suitable planning task and tell me what needs to improve.",
        "acceptance": "The extracted pack passes its tests and matches the browser example. Independent reuse feedback is the next milestone.",
        "boundary": "A reusable engineering pack with illustrative data and explicit calculation rules."
      }
    ]
  },
  "publicScope": "An open workspace of professional experience, applied AI and reusable tools. Public downloads contain selected professional facts, engineering examples and clearly labelled demonstration data. Client records and private archives remain confidential."
}
