跳到正文
原文
BasicPharma搬砖工· 西门君·· 5 小时前精选AI 评分84

EMA 公布新附录 22《人工智能》研讨会纪要及录屏,为起草收集行业意见

EMA公布 新附录 22《人工智能》最新研讨报告和讨论录屏

AI 导读

EMA 公布 2026 年 6 月 30 日研讨会纪要(EMA/156789/2026)及录屏,为起草 EU GMP 新附录 22《人工智能》收集行业意见。附录 22 是附录 11 的补充,规范嵌入 AI 模型的系统;草案原本禁止生成式 AI/LLM、概率模型和自适应模型用于关键 GMP 应用,公众咨询多数意见反对。

推荐理由

材料梳理 EMA 附录 22 研讨会六个专题的行业共识与分歧,可帮助读者了解该附录在范围和控制策略上的讨论方向。

正文

2026.10.02欧盟公布了EMA 2026年6月30日研讨会纪要(EMA/156789/2026)文字报告和长达六小时的录屏:为起草 EU GMP 新附录 22《人工智能》收集行业意见。附录 22 是附录 11(计算机化系统)的补充,规范嵌入 AI 模型的系统在 GMP 中的使用。

草案原本禁止生成式 AI/LLM、概率模型、自适应模型用于关键 GMP;公开咨询多数意见反对,要求放开。研讨会即为此收集行业做法。

六个专题的核心共识

专题一(监管路径):不应因技术标签排除 AI;判据是能否证明适合预期用途,且风险可在现有质量体系内受控。不确定性不等于禁止,应驱动更强的控制、监测和再验证。

专题五(战略风险):反对一刀切禁止;应逐案评估,综合考虑决策后果、模型影响力、不确定性、复杂度、可检测性。当不确定性无法受控时,有效的质量风险管理结论可以是"不使用",并记录在案。

专题三(人工监督):监督不等于人在回路——前者是贯穿生命周期的治理目标,后者只是一种实现方式。监督按风险分级设计,要求人员能理解、质疑、干预、否决;监督不能弥补不足的验证。

专题二(技术可靠性):采用分层控制——输入预防、模型内检测、输出遏制三层叠加;每个控制针对特定失效模式,不能假设覆盖未知失效。主附录用"控制机制"而非易过时的"护栏",技术细节放附录或可更新指南。

专题四(验证与生命周期):验证是生命周期而非单次事件;测试数据必须独立于训练数据;监测、变更控制、事件处理、CAPA、供应商监督均属验证范畴。

专题六(网络安全与外包):AI 系统本质上仍是计算机化系统,附录 11 第 7 章(外包)完全适用,无需另设制度;即便使用第三方模型或云服务,受监管用户仍负全部责任,靠协议、证据、认证与变更可见性管控。

第一,技术中立、基于结果——附录 22 应是结果导向框架,不是技术处方清单。第二,质量风险管理原则(ICH Q9R1)适用于 AI,严重度乘概率并考虑可检测性。第三,控制状态须可证明——风险场景、控制映射、验证监控证据构成检查证据包。第四,需新增一组定义:预期用途、模型影响力、复杂度、不确定性、控制策略、回退、安全状态等。

以下正文

引言

Annex 22 "Artificial Intelligence" will be an Annex of the EU Good Manufacturing Practices (GMP) Guide intended to provide additional guidance to Annex 11 for computerised systems in which AI models are embedded.

附录22《人工智能》将成为欧盟《良好生产规范(GMP)指南》的一个附录,旨在为嵌入AI模型的计算机化系统提供针对附录11的补充指南。

The draft version of Annex 22 currently states "the document does not apply to Generative AI and Large Language Models (LLM), models with a probabilistic output, dynamic models and models which adapt performance, such models should not be used in critical GMP applications."

附录22草案版本目前声明:"本文件不适用于生成式AI和大语言模型(LLM)、具有概率性输出的模型、动态模型以及性能自适应的模型,此类模型不应用于关键GMP应用。"

Comments received during the public consultation on Annex 22 indicated support for allowing manufacturers to use Large Language Models and generative AI systems in GMP applications, whether critical or non-critical. This position reflects the need to support innovation and encourage investment by pharmaceutical companies.

在附录22公众咨询期间收到的意见表明,

各方支持允许生产商在GMP应用中使用大语言模型和生成式AI系统

,无论其属于关键还是非关键应用。这一立场反映了支持创新和鼓励制药公司投资的必要性。

The drafting group has therefore discussed the scope of Annex 22 to allow dynamic or adaptive models, probabilistic models, and Generative AI/LLMs in GMP applications, provided they comply with the requirements of the annex and are supported by a fully documented, robust, risk-based control strategy.

因此,起草组已讨论附录22的范围,允许在GMP应用中使用动态或自适应模型、概率性模型以及生成式AI/LLM,

前提是这些模型符合本附录的要求,并有充分记录的、稳健的风险控制策略作为支撑。

However, concerns remain about the elements needed for such a control strategy. The drafting group therefore requested a meeting with AI experts to assess whether a risk-based approach can be applied to Generative AI/LLMs, including the use control mechanisms as mitigations for their use in GMP applications.

然而,对于此类控制策略所需的要素仍存在关切。因此,起草组请求与AI专家举行会议,以评估基于风险的方法是否可应用于生成式AI/LLM,包括使用控制机制作为其在GMP应用中使用的缓解措施。

The objective of this workshop was therefore to gather structured, experience-based input from industry on the control, governance, and mitigation measures applied throughout the AI system lifecycle, with particular emphasis on the application of ICH Q9 quality risk management principles. To facilitate this discussion, industry stakeholder associations registered with European Medicines Agency (EMA) collaborated to provide a consolidated response to the List of Questions previously shared with industry and delivered during the workshop a series of presentations reflecting a collective industry perspective across the six predefined topics.

因此,本次研讨会的目的是收集来自行业的、基于经验的结构化输入,涵盖AI系统整个生命周期中应用的控制、治理和缓解措施,特别强调ICH Q9质量风险管理原则的应用。为促进此次讨论,已在欧洲药品管理局(EMA)注册的行业利益相关方协会合作,对先前与行业分享的问题清单提供了统一回复,并在研讨会期间作了一系列演讲,反映了在六个预定义专题上的行业集体观点。

These 6 topics addressed key areas including:

这6个专题涵盖的关键领域包括:

Topic 1: Regulatory Pathways for Adaptive/probabilistic AI models in GMP

专题1:GMP中适应性/概率性AI模型的监管路径

Topic 2: Technical Reliability & Incident Response

专题2:技术可靠性与事件响应

Topic 3: Human Oversight & Accountability

专题3:人类监督与问责制

Topic 4: Validation & Lifecycle Management

专题4:验证与生命周期管理

Topic 5: Strategic Risk & Compliance Limits

专题5:战略风险与合规限度

Topic 6: Cybersecurity and Outsourced activities

专题6:网络安全与外包活动

Through the presentations and subsequent discussions, the workshop was intended to enhance the drafting group's understanding of current practices, challenges, and emerging approaches for managing risks associated with AI systems in GMP environments. The insights obtained were intended to inform the revision of the scope of Annex 22 and support the development of a risk-proportionate and implementable regulatory framework that both safeguards product quality and patient safety while enabling responsible innovation.

通过演讲和随后的讨论,本次研讨会旨在增强起草组对当前实践、挑战以及GMP环境中AI系统相关风险管理新兴方法的理解。所获得的见解旨在为附录22范围的修订提供参考,并支持制定一个风险相称且可实施的监管框架,在保障产品质量和患者安全的同时,实现负责任的创新。

1. 专题讨论

1.1 专题1:GMP中适应性/概率性AI模型的监管路径

起草组提出的问题

How could we accommodate adaptive and probabilistic models in Annex 22?

我们如何在附录22中容纳适应性模型和概率性模型?

What would be the validation paradigm for adaptive models that should be in annex 22?

附录22中应包含的适应性模型验证范式是什么?

Please comment on the application of a quality risk management approach to the entire lifecycle of such AI system models that may be used in GMP high risk areas where there may be a low detectability of deficiencies but a high impact on the patient and taking into account ICH Q9(R1) principles. In applying quality risk management, what should the regulated user take into account when selecting and using an adaptive or probabilistic model?

请评论将质量风险管理方法应用于此类AI系统模型整个生命周期的适用性,这些模型可能用于GMP高风险领域,在这些领域缺陷的可检测性可能较低但对患者的影响较大,并需考虑ICH Q9(R1)原则。在应用质量风险管理时,受监管用户在选用适应性或概率性模型时应考虑哪些因素?

行业观点

The industry position, presented as a consolidated interested parties view, was that adaptive, probabilistic and generative AI models should not be excluded from GMP applications on the basis of a technology label alone. Instead, the central criterion should be whether the model can be demonstrated to be fit for its intended use and whether the associated risks can be controlled within an existing pharmaceutical quality system. The concepts of importance, uncertainty, complexity and level of formality should be taken into consideration when a risk assessment is conducted.

以利益相关方统一观点呈现的行业立场认为,不应仅基于技术标签将适应性、概率性和生成式AI模型排除在GMP应用之外。相反,核心标准应是模型能否被证明适合其预期用途,以及相关风险能否在现有药品质量体系内得到控制。在进行风险评估时,应考虑重要性、不确定性、复杂性和正式程度等概念。

The speaker emphasised that AI is a new and rapidly evolving technology, but that uncertainty and model drift are not conceptually new to GMP. Measurement systems drift, calibration intervals are adjusted on a risk basis, and lifecycle review is already used to confirm that a system remains fit for use. This analogy was not intended to equate all forms of AI adaptation with measurement drift, but to illustrate that GMP already contains mechanisms for reassessing continued suitability when system behaviour changes over time.

演讲者强调,AI是一项新兴且快速发展的技术,但不确定性和模型漂移在GMP中并非新概念。测量系统会漂移,校准间隔基于风险进行调整,生命周期审查已用于确认系统持续适用。这一类比并非旨在将所有形式的AI适应等同于测量漂移,而是说明当系统行为随时间变化时,GMP已包含重新评估持续适用性的机制。

A key theme was the distinction between uncertainty and unacceptable risk. The presentation argued that uncertainty should drive increased knowledge generation, stronger lifecycle controls, more formal validation, monitoring and re-verification, rather than an automatic prohibition. The proposed Annex 22 wording presented by the speaker would broaden the scope to cover deterministic and probabilistic outputs, static and dynamic models, and models whose performance may adapt during use, provided controls are implemented according to intended use and risk to patient safety and product quality.

一个关键主题是不确定性与不可接受风险之间的区别。演讲认为,不确定性应驱动更多的知识生成、更强的生命周期控制、更正式的验证、监测和重新验证,而非自动禁止。演讲者提出的附录22措辞建议将范围扩大至涵盖确定性输出和概率性输出、静态模型和动态模型,以及在使用过程中性能可能自适应的模型,前提是按照预期用途和对患者安全及产品质量的风险实施控制。

问答环节

The Q&A focused on whether the framework sufficiently addresses specific features of generative AI and LLMs, including hallucination, fabrication, sycophancy or epistemic overconfidence. Regulators questioned whether these phenomena are merely ordinary statistical uncertainty or a fundamentally different failure mode requiring specific attention. The response was that such behaviours must be identified and evaluated during development and validation; if the manufacturer cannot demonstrate fitness for intended use in view of those failure modes, the use cannot be justified. The discussion did not set a specific technical method for estimating confidence intervals or hallucination rates; rather, it deferred detailed technical controls to later.

问答环节聚焦于该框架是否充分解决了生成式AI和LLM的特定特征,包括幻觉、虚构、谄媚或认知过度自信。监管方质疑这些现象是否仅仅是普通的统计不确定性,还是需要特别关注的根本性不同的失效模式。回应是,此类行为必须在开发和验证过程中被识别和评估;如果生产商无法在这些失效模式存在的情况下证明其适合预期用途,则该使用无法被证明合理。讨论未设定估计置信区间或幻觉率的具体技术方法,而是将详细技术控制推迟到后续阶段。

Detectability was another point for discussion. The moderator noted that in adaptive or generative AI, inability to detect failure may itself be a central risk factor. The speaker acknowledged that, although the general definition of risk in ICH Q9 refers to severity and probability, methods such as Failure Mode and Effects Analysis (FMEA) may use detectability because it is useful in certain contexts. For AI, detectability may be highly appropriate as a dimension of the assessment, particularly where errors are difficult to observe. The discussion therefore supported consideration of detectability, without mandating a single risk scoring method.

可检测性是另一个讨论要点。主持人指出,在适应性或生成式AI中,无法检测到失效本身可能就是一个核心风险因素。演讲者承认,尽管ICH Q9中风险的一般定义涉及严重性和可能性,但失效模式与影响分析(FMEA)等方法可使用可检测性,因为它在某些情境下很有用。对于AI,可检测性可能作为评估维度非常适用,特别是在错误难以观察的情况下。因此,讨论支持考虑可检测性,但不强制使用单一的风险评分方法。

The session also touched on complexity. One regulator asked whether inclusion of complexity might be controversial for industry. The speaker responded that the interested parties involved in preparing the workshop had not objected; on the contrary, complexity is explicitly part of the ICH Q9(R1) formality concept and should be included when determining the necessary rigour of controls. Overall, the session created broad consensus that Annex 22 should remain technology-neutral and principle-based, but should make clear that AI models require documented intended use, documented understanding of uncertainties, lifecycle monitoring, risk review, and proportionate control before use in GMP Applicationas.

会议还涉及复杂性。一位监管方询问纳入复杂性是否会引起行业争议。演讲者回应称,参与筹备研讨会的利益相关方并未反对;相反,复杂性明确属于ICH Q9(R1)正式程度概念的一部分,应在确定控制所需严格程度时予以纳入。总体而言,会议形成了广泛共识,即附录22应保持技术中立和基于原则,但应明确AI模型在GMP应用中使用前需要有记录的预期用途、对不确定性的记录理解、生命周期监测、风险审查以及相称的控制。

会议关键信息

Existing QRM and lifecycle principles are applicable to AI, including adaptive and probabilistic models.

现有的QRM和生命周期原则适用于AI,包括适应性模型和概率性模型。

Technology labels should not determine acceptability; intended use, impact, uncertainty, complexity and control strategy should.

技术标签不应决定可接受性;预期用途、影响、不确定性、复杂性和控制策略才应决定。

Higher risk should drive increased formality and stronger controls.

更高的风险应驱动更高的正式程度和更强的控制。

Detectability is a relevant AI risk consideration, especially where deficiencies are difficult to identify.

可检测性是AI风险的相关考虑因素,特别是在缺陷难以识别的情况下。

Annex 22 should avoid detailed technical prescriptions and remain coherent with Annex 11, Chapter 4 and ICH Q9.

附录22应避免详细的技术规定,并与附录11、第4章和ICH Q9保持一致。

对附录22起草的潜在影响

The discussion explored the considerations relevant to the possible use of probabilistic or adaptive models, including defined intended use, risk assessment, documented knowledge of the model and data, lifecycle validation, monitoring and risk-based change control. Definitions may need to clarify intended use, model influence, complexity, uncertainty, control strategy, detectability and continued fitness for use.

讨论探讨了与可能使用概率性或适应性模型相关的考虑因素,包括明确的预期用途、风险评估、对模型和数据的记录知识、生命周期验证、监测和基于风险的变更控制。定义可能需要澄清

预期用途、模型影响、复杂性、不确定性、控制策略、可检测性和持续适用性

。

Industry would be expected to justify use through documented evidence and lifecycle controls.

行业应通过记录的证据和生命周期控制来证明使用的合理性。

1.2 专题5:战略风险与合规限度

起草组提出的问题

Do current guardrail approaches sufficiently address GMP high risk areas such as data integrity, process control, auditability, and traceability—especially where GenAI systems may influence or automate decision-making?

当前的

护栏方法

是否充分解决了GMP高风险领域的问题,如数据完整性、过程控制、可审计性和可追溯性——特别是当生成式AI系统可能影响或自动化决策时?

Where do experts believe the limits of risk-based mitigation lie? Are there classes of critical decisions where no combination of guardrails and oversight would be sufficient?

专家认为基于风险的缓解措施的限度在哪里?是否存在任何护栏和监督的组合都不足以应对的关键决策类别?

Are there scenarios in which the level of guardrail effort required to mitigate GenAI risks becomes disproportionate or operationally impractical, thereby indicating that such systems should not be used in certain critical GMP functions?

是否存在缓解生成式AI风险所需的护栏工作量变得不相称或在操作上不切实际的情形,从而表明此类系统不应在某些关键GMP功能中使用?

What other elements of an overall control strategy should be considered for inclusion in the Annex 22 to allow the use of dynamic or adaptive models, probabilistic models, and Generative AI/LLMs in GMP applications?

为允许在GMP应用中使用动态或自适应模型、概率性模型以及生成式AI/LLM,整体控制策略中还应考虑纳入附录22的哪些其他要素?

行业观点

The second session addressed whether there are strategic limits to the use of generative AI or other probabilistic/adaptive AI models in critical GMP decision-making. The interested parties' position was that guardrails and other controls, when designed under QRM principles, can contribute to a state of control for data integrity, traceability, auditability and process control. The presentation argued that these attributes are not new GMP expectations; industry has long managed them through systems, records, audit trails, verification and monitoring. However, AI requires that established controls and newer AI-specific controls operate together as an integrated framework.

第二场会议讨论了在关键GMP决策中使用生成式AI或其他概率性/适应性AI模型是否存在战略性限度。利益相关方的立场是,当护栏和其他控制按照QRM原则设计时,可以为数据完整性、可追溯性、可审计性和过程控制的状态受控做出贡献。演讲认为,这些属性并非新的GMP期望;行业长期以来已通过系统、记录、审计追踪、验证和监测来管理它们。然而,AI要求既定控制与新型AI特定控制作为一个集成框架协同运作。

The industry position was strongly against categorical prohibitions. The speaker stated that there should be no class of decisions considered unacceptable solely by technology category or criticality label. Instead, acceptability should be assessed case by case by considering decision consequence, model influence, uncertainty, complexity and the ability to reduce residual risk to an acceptable level. The presentation acknowledged that in some use cases the mitigation effort may be heavy, impractical, or insufficient. In such cases the conclusion would be that the use is not justified at that time and in that context, and the conclusion should be documented.

行业立场强烈反对分类禁止。演讲者声明,不应仅因技术类别或关键性标签而将任何决策类别视为不可接受。相反,应通过考虑决策后果、模型影响、不确定性、复杂性以及将残余风险降至可接受水平的能力,逐案评估可接受性。演讲承认,在某些使用案例中,缓解工作量可能很大、不切实际或不足。在这种情况下,结论应是在当时和该情境下使用不合理,且该结论应予以记录。

问答环节

A significant regulatory challenge raised during Q&A was how a manufacturer could demonstrate to inspectors that the combined layered controls do provide reliable output for a critical decision. The response identified the risk assessment as the starting point: it should describe the risk scenarios, the controls mapped to those scenarios, and the evidence generated through verification and monitoring. Monitoring history, responsiveness to signals, and documented state of control would then be inspection evidence. Additional discussion introduced the concept of controls by design, continuous monitoring of expected residual risk, and automated controls where appropriate.

问答环节提出的一个重大监管挑战是,生产商如何向检查员证明组合的分层控制确实为关键决策提供了可靠的输出。回应将风险评估确定为起点:它应描述风险情景、映射到这些情景的控制,以及通过验证和监测生成的证据。监测历史、对信号的响应性以及记录的状态受控将成为检查证据。额外讨论引入了设计控制、对预期残余风险的持续监测以及适当情况下的自动化控制等概念。

Responsibility and accountability were also explored. Regulators asked how a responsible person could take responsibility for a model that is not fully transparent. Responses emphasised that the use of AI should not remove predicate-rule responsibilities and that accountable persons may rely on qualified experts, as they already do for complex manufacturing technologies. The idea of an "AI qualification dossier" was introduced by another industry speaker as a package containing validation evidence, training data information, guardrails around prompting and output, and evidence that the system works as expected.

还探讨了责任和问责制。监管方询问负责人如何对不完全透明的模型承担责任。回应强调,AI的使用不应免除前置规则责任,责任人可以依赖合格的专家,就像他们已经在复杂制造技术中所做的那样。另一位行业演讲者提出了"AI资质档案"的概念,作为一个包含验证证据、训练数据信息、围绕提示和输出的护栏,以及系统按预期工作的证据的成套资料。

Overall, the strategic session reinforced the preference for Annex 22 as an outcome-based framework, not a technology-prescriptive one.

总体而言,战略会议强化了将附录22作为基于结果的框架而非技术规定性框架的偏好。

讨论中涌现的关键信息

No categorical prohibition was supported by industry; use and risk should be assessed in context.

行业不支持分类禁止;应在情境中评估使用和风险。

A valid QRM outcome may still be 'do not use' where uncertainty cannot be controlled.

当不确定性无法控制时,有效的QRM结果仍可能是"不使用"。

State of control should be demonstrated through mapped risk scenarios, controls, verification and monitoring evidence.

应通过映射的风险情景、控制、验证和监测证据来证明状态受控。

Criticality alone is insufficient; model influence, uncertainty, complexity and detectability must be considered.

仅有关键性是不够的;必须考虑模型影响、不确定性、复杂性和可检测性。

Accountability remains with qualified/responsible persons and cannot be delegated to AI.

问责制仍由合格/负责人承担,不能委托给AI。

对附录22起草的潜在影响

Annex 22 may be influenced towards an explicit performance-based acceptability framework: identify intended use; evaluate decision consequence, model influence, uncertainty, complexity and detectability; define controls; verify them; monitor residual risk; and document the go/no-go conclusion. Definitions of direct impact, model influence, control strategy and residual-risk acceptance may need to be developed.

附录22可能会朝着明确的基于绩效的可接受性框架发展:

识别预期用途;评估决策后果、模型影响、不确定性、复杂性和可检测性;定义控制;验证控制;监测残余风险;并记录进行/不进行的结论。可能需要制定直接影响、模型影响、控制策略和残余风险可接受性的定义。

Industry expectations would include an inspectable evidence trail.

行业期望将包括可检查的证据链。

1.3 专题3:人类监督与问责制

起草组提出的问题

What level and form of human-in-the-loop oversight is still required when guardrails are implemented, and is this oversight sufficient to ensure accuracy, traceability, and accountability?

当实施护栏时,仍需要何种程度和形式的人机回环监督,这种监督是否足以确保准确性、可追溯性和问责制?

行业观点

The third session examined the role of human oversight when technical guardrails and AI-supported workflows are used. The interested parties emphasised that human oversight should be designed through a documented QRM process and should be proportionate to the AI model's risk profile and context of use. A key distinction was made between "human oversight" as a broad lifecycle governance objective and "human in the loop" as one specific implementation pattern. Human oversight asks whether a qualified, accountable human can understand, challenge, intervene, override, reject or escalate. Human-in-the-loop, by contrast, is where the process waits for a human action before a GMP-relevant workflow continues.

第三场会议探讨了在使用技术护栏和AI支持的工作流程时人类监督的作用。利益相关方强调,人类监督应通过记录的QRM过程设计,并应与AI模型的风险特征和使用情境相称。一个重要的区分是"人类监督"作为广泛的生命周期治理目标与"人机回环"作为一种特定实施模式之间的区别。人类监督询问的是合格、负责任的人类是否能够理解、质疑、干预、覆盖、拒绝或升级。相比之下,人机回环是指流程在GMP相关工作流继续之前等待人类行动。

The presentation placed human oversight within an integrated control framework, not as a standalone safeguard. Guardrails may reduce risk but do not eliminate the need for risk-proportionate human oversight or accountability. EU GMP Part I Chapter 2 on personnel qualification and training was cited as applicable, and the EU AI Act was referenced in relation to human oversight commensurate with risk, autonomy and context of use. The oversight model proposed criteria such as impact of incorrect decisions, model influence, uncertainty, complexity and importance. These criteria determine the degree of human involvement and the permissible level of autonomy.

演讲将人类监督置于集成控制框架内,而非独立保障。护栏可能降低风险,但不能消除风险相称的人类监督或问责制的需要。引用了欧盟GMP第一部分第2章关于人员资质和培训的适用性,并参考了欧盟AI法案中关于与风险、自主性和使用情境相称的人类监督。监督模型提出了诸如错误决策的影响、模型影响、不确定性、复杂性和重要性等标准。这些标准决定人类参与的程度和允许的自主性水平。

A deviation management support use case illustrated how human involvement might be high when a model has high impact, variable uncertainty and complexity, even if the AI influence is initially more informational. Over time, as knowledge and experience increase, the risk profile may shift, potentially allowing more autonomy and less human involvement. Conversely, if new risks emerge, oversight may increase. This reinforced the lifecycle nature of oversight.

一个偏差管理支持用例说明了当模型具有高影响、可变不确定性和复杂性时,人类参与程度可能很高,即使AI影响最初更多是信息性的。随着时间的推移,随着知识和经验的增加,风险特征可能发生变化,可能允许更多的自主性和更少的人类参与。相反,如果出现新风险,监督可能会增加。这强化了监督的生命周期性质。

问答环节

The Q&A focused on whether human oversight can be meaningful in practice. Regulators asked how to monitor the performance of the human performing oversight; how to manage automation bias; what information and competence a human reviewer needs; whether human-in-the-loop can become a bottleneck; how fallback/override should be validated; and how oversight can be evidenced during inspection. Responses identified training, qualification, knowledge testing, audit trails, recording of AI outputs and human modifications, and monitoring of human-AI team performance as potential evidence. Automation bias was acknowledged as requiring meaningful oversight, appropriate training and technical measures. A human reviewer does not need to understand every internal detail of the model, but must understand the relevant model principles, the process, the output and the decision context sufficiently to challenge the AI.

问答环节聚焦于人类监督在实践中是否有意义。监管方询问如何监督执行监督的人的表现;如何管理自动化偏见;人类审查者需要什么信息和能力;人机回环是否可能成为瓶颈;回退/覆盖应如何验证;以及如何在检查期间证明监督。回应将培训、资质、知识测试、审计追踪、AI输出和人类修改的记录,以及人机团队绩效的监测确定为潜在证据。自动化偏见被认为需要有意义的监督、适当的培训和技术措施。人类审查者不需要理解模型的每个内部细节,但必须充分理解相关的模型原理、流程、输出和决策情境,以便对AI输出提出质疑。

A regulator asked if human-in-the-loop could ever justify reduced validation. The question was deferred to the validation topic; an industry speaker noted that the answer is risk-based rather than a categorical yes or no. Separately, one speaker noted that oversight alone does not directly correct risk; it provides knowledge, detection and potential triggers for action. Therefore, human oversight should be judged by whether it produces relevant knowledge and supports effective risk reduction through follow-up actions. The discussion also recognised that human reviewers may face operational pressure or fatigue. The response was that such risks are not AI-specific and should be managed through existing GMP provisions. Overall, the session supports Annex 22 language requiring meaningful, risk-based and documented oversight rather than default, symbolic or one-size-fits-all

human-in-the-loop

.

一位监管方询问人机回环是否可以为减少验证提供理由。该问题被推迟到验证专题;一位行业演讲者指出,答案是基于风险的,而不是简单的"是"或"否"。另外,一位演讲者指出,监督本身并不直接纠正风险;它提供知识、检测和行动的潜在触发因素。因此,人类监督应根据其是否产生相关知识并通过后续行动支持有效的风险降低来评判。讨论还承认,人类审查者可能面临操作压力或疲劳。回应是,此类风险并非AI特有,应通过现有GMP条款进行管理。总体而言,会议支持附录22要求有意义的、基于风险的、有记录的监督,而非默认的、象征性的或一刀切的

人机回环

监督。

讨论中涌现的关键信息

Human oversight is broader than human-in-the-loop and should be lifecycle-wide.

人类监督比人机回环更广泛,应基于全面的生命周期。

Human-in-the-loop is one possible control, not a default requirement for all AI uses.

人机回环是一种可能的控制,而非所有AI系统的默认要求。

Oversight must be proportionate to risk, autonomy, uncertainty, complexity and decision impact.

监督必须与风险、自主性、不确定性、复杂性和决策影响相称。

Training, qualification, authority and ability to challenge are essential.

培训、资质、权限和质疑能力至关重要。

Oversight should be evidenced and periodically reviewed; it cannot compensate for inadequate validation.

监督应有证据支持并定期审查;它不能弥补验证不足。

对附录22起草的潜在影响

The discussion supports Annex 22 text requiring human oversight to be defined in the control strategy according to intended use and QRM. Definitions should distinguish oversight, Human in the Loop (HITL), Human on the Loop (HOTL) monitoring, fallback/override and accountability. Requirements may include training, qualification, documentation, auditability and periodic review of oversight effectiveness.

讨论支持附录22要求在控制策略中根据预期用途和QRM定义人类监督。定义应区分监督、人机回环(HITL)、人机在环(HOTL)监测、回退/覆盖和问责制。要求可能包括培训、资质、文档、可审计性和监督有效性的定期审查。

Industry expectations should include evidence that reviewers can understand, challenge and intervene.

行业期望应包括审查者能够理解、质疑和干预的证据。

1.4 专题2:技术可靠性与事件响应

起草组提出的问题

To what extent can available guardrail mechanisms reliably prevent, detect, or contain hallucinations, incorrect recommendations, or fabricated data when dynamic or adaptive models, probabilistic models, and Generative AI/LLMs are used within GMP workflows?

当在GMP工作流程中使用动态或自适应模型、概率性模型以及生成式AI/LLM时,现有护栏机制能在多大程度上可靠地预防、检测或遏制幻觉、错误建议或虚构数据?

When guardrails fail or detect uncertainty, what mechanisms are required to ensure timely escalation and prevention of GMP impact?

当护栏失效或检测到不确定性时,需要哪些机制来确保及时升级并防止GMP影响?

行业观点

The fourth session provided the most technical discussion of the day. It examined whether available guardrail mechanisms can prevent, detect or contain hallucinations, incorrect recommendations and fabricated data, and how failures should be escalated. The presentation used five use cases:

第四场会议提供了当天最技术性的讨论。它探讨了现有护栏机制能否预防、检测或遏制幻觉、错误建议和虚构数据,以及失效应如何升级。演讲使用了五个用例:

adaptive Heating, Ventilation and Air Conditioning (HVAC) control within validated environmental boundaries;

在已验证环境边界内的自适应供暖、通风和空调(HVAC)控制;

an agentic AI system supporting deviation management;

支持偏差管理的智能体AI系统;

an LLM supporting GMP risk assessment for an electronic batch record system;

支持电子批记录系统GMP风险评估的LLM;

text-to-SQL generation for annual product review data extraction;

用于年度产品质量回顾数据提取的文本转SQL生成;

and model predictive control for bioreactor feeding.

以及生物反应器培养的模型预测控制。

Some examples were stated to be pilots or theoretical frameworks rather than production implementations.

某些示例被说明为试点或理论框架,而非生产实施。

A central model was layered control:

input/prevention controls,

in-model/detection controls,

and output/containment controls.

一个核心模型是分层控制:

输入/预防控制,

模型内/检测控制,

以及输出/遏制控制。

This was linked to the "Swiss cheese" concept: each control addresses a particular failure mode and may be imperfect alone, but layers can combine to reduce risk.

这与"瑞士奶酪"概念相关联:每项控制针对特定的失效模式,单独可能不完美,但各层可以组合以降低风险。

Examples of input controls included context-of-use verification, out-of-distribution detection, semantic validation and rule-based limits. In-model controls included retrieval augmented generation with source/citation binding, conformal prediction, uncertainty quantification, self-consistency scoring and access to model logs or rationale where available. Output controls included hard constraints, confidence gates, cumulative sum monitoring, exponentially weighted moving averages, ensemble disagreement, human gates and fallback to validated states.

输入控制的示例包括使用情境验证、分布外检测、语义验证和基于规则的限值。模型内控制包括具有来源/引用绑定的检索增强生成、共形预测、不确定性量化、自一致性评分以及在可用时访问模型日志或推理依据。输出控制包括硬约束、置信度门控、累积和监测、指数加权移动平均、集成分歧、人类门控和回退到已验证状态。

The presenter repeatedly emphasised that guardrails are specific controls for specific risks and that unanticipated failure modes cannot be assumed to be covered by controls designed for known failure modes.

演讲者反复强调,护栏是针对特定风险的特定控制,不能假设针对已知失效模式设计的控制能够覆盖未预见的失效模式。

问答环节

This led to important Q&A. One regulator asked whether all controls must be derived from risk assessment, given that GMP often uses tacit knowledge and experience to create controls beyond formally identified risks. The response, supplemented by another speaker, was that risk assessments are limited by available knowledge and must be updated through lifecycle learning, deviations and defects. New knowledge should feed back into risk assessments and therefore into new or revised controls.

这引发了重要的问答。一位监管方询问,

鉴于GMP经常使用隐性知识和经验来创建超出正式识别风险的控制,是否所有控制都必须源自风险评估。

由另一位演讲者补充的回应是,风险评估受可用知识的限制,必须通过生命周期学习、偏差和缺陷来更新。新知识应反馈到风险评估中,从而反馈到新的或修订的控制中。

Terminology was a major issue. Several participants questioned whether the term "guardrail" should appear in Annex 22. Industry responses suggested that "control mechanisms" may be preferable in the main annex because guardrails can be technology-specific and transient. The discussion reflects agreement that the main annex should contain principles and objectives — such as prevention, detection and containment — while technical methods and use cases may be better placed in appendices, Q&A, or another updateable format.

术语是一个主要问题。几位参与者质疑"护栏"一词是否应出现在附录22中。行业回应建议,在主附录中"控制机制"可能更可取,因为护栏可能是技术特定且短暂的。讨论反映了共识,即主附录应包含原则和目标——如预防、检测和遏制——而技术方法和用例可能更适合放在附录、问答或另一种可更新的格式中。

The Q&A also explored whether these methods imply a preference for local or self-developed models. The presenter stated that the key is not ownership of the model but access to sufficient evidence to measure how it works. If a third-party model or Application Programming Interface (API) cannot provide necessary outputs, rationale, logs or evidence needed for control, the manufacturer cannot measure and justify the system adequately.

问答还探讨了这些方法是否意味着偏好本地或自研模型。演讲者声明,关键不在于模型的所有权,而在于能否获得足够的证据来衡量其工作方式。如果第三方模型或应用程序编程接口(API)无法提供控制所需的必要输出、推理依据、日志或证据,生产商就无法充分测量和证明系统的合理性。

The session also addressed fallback. For HVAC and bioreactor examples, failure or out-of-bound behaviour led to reversion to a pre-validated classical control state. Regulators noted that a fully model-operated process might not have such a classical fallback. The response was that this would need to be considered in the contingency strategy; sometimes accepting stoppage may be the contingency.

会议还讨论了回退。对于HVAC和生物反应器示例,失效或超出边界的行为导致恢复到预先验证的经典控制状态。

监管方指出,完全由模型操作的过程可能没有这样的经典回退。回应是,这需要在应急策略中考虑;有时接受停机可能是应急措施。

Overall, this session strongly supports Annex 22 requiring layered, risk-based controls and documented escalation pathways, while keeping technical examples outside the main body or at least non-prescriptive.

总体而言,本次会议强烈支持附录22要求分层的、基于风险的控制和记录的升级路径,同时将技术示例保留在正文之外或至少以

非规定性表述。

。

讨论中涌现的关键信息

Controls should be layered across prevention, detection and containment.

控制应在预防、检测和遏制方面分层设置。

Each control should be linked to intended use, failure modes and measurable acceptance criteria.

每项控制应与预期用途、失效模式和可测量的验收标准相关联。

Unanticipated failures require lifecycle learning and risk assessment updates.

未预见的失效需要生命周期学习和风险评估更新。

The main annex should avoid transient technical details; examples may be placed in updateable guidance.

主附录应避免短暂的技术细节;示例可放在可更新的指南中。

Supplier or API models must provide sufficient evidence and transparency for control.

供应商或API模型必须提供足够的证据和透明度以进行控制。

对附录22起草的潜在影响

This discussion may significantly shape Annex 22's control-strategy sections. The annex could require layered controls and escalation pathways while avoiding detailed references to any one technical method. Definitions may need to cover control mechanisms, unacceptable output, failure mode, escalation, fallback, safe state and monitoring thresholds.

本次讨论可能对附录22的控制策略部分产生重大影响。附录可能要求分层控制和升级路径,同时避免详细引用任何单一技术方法。定义可能需要涵盖控制机制、不可接受输出、失效模式、升级、回退、安全状态和监测阈值。

Industry expectations include documented qualification of controls, evidence of thresholds and lifecycle monitoring.

行业期望包括控制的记录确认、阈值证据和生命周期监测。

1.5 Topic 4: Validation & Lifecycle Management专题4:验证与生命周期管理

起草组提出的问题

How can guardrails remain effective when the underlying AI model evolves (e.g., updates, retraining, drift), and what controls are necessary to ensure guardrail performance remains validated over time?

当底层AI模型演进时(如更新、再训练、漂移),护栏如何保持有效,以及需要哪些控制来确保护栏性能随时间保持已验证状态?

What type and level of evidence (e.g., validation data, stress testing results, failure analyses) would be needed to justify that guardrails are a reliable risk mitigation measure enabling GenAI use in GMP functions?

需要什么类型和级别的证据(如验证数据、压力测试结果、失效分析)来证明护栏是可靠的风险缓解措施,从而能够在GMP功能中使用生成式AI?

Can you demonstrate the effectiveness of such guardrails or similar measures?

您能否证明此类护栏或类似措施的有效性?

行业观点

The fifth session addressed validation and lifecycle management as the integrative framework connecting the other topics. The presenters framed validation lifecycle management as central to the six-pillar architecture because it interfaces with regulatory pathways, technical reliability, human oversight, change control, monitoring and incident response. The session answered questions on how controls remain effective when models evolve, what evidence is needed to justify control mechanisms as reliable risk mitigations, and how the effectiveness of such controls can be demonstrated.

第五场会议将验证和生命周期管理作为连接其他专题的集成框架。演讲者将验证生命周期管理定位为六支柱架构的核心,因为它与监管路径、技术可靠性、人类监督、变更控制、监测和事件响应相连接。会议回答了当模型演进时控制如何保持有效、需要什么证据来证明控制机制是可靠的风险缓解措施,以及如何证明此类控制的有效性等问题。

The presentation emphasised a lifecycle decision-making framework rather than a single validation event. A model's intended use, design assumptions, context, data, training and verification strategy should define the depth and type of validation. Risk factors discussed included: uncertainty, importance, complexity, novelty of the AI approach, supplier experience, model limitations, decision consequence, interpretability/explainability, degree of autonomy and adaptiveness.

演讲强调生命周期决策框架,而非单一验证事件。模型的预期用途、设计假设、情境、数据、训练和验证策略应决定验证的深度和类型。讨论的风险因素包括:不确定性、重要性、复杂性、AI方法的新颖性、供应商经验、模型局限性、决策后果、可解释性/可说明性、自主性程度和适应性。

These factors should determine validation activities, selection of control mechanisms, evidence requirements, monitoring and revalidation/change triggers.

这些因素应决定验证活动、控制机制的选择、证据要求、监测和再验证/变更的触发。

The presenters highlighted that data used for validation/testing should remain independent from development and training data.

演讲者强调,

用于验证/测试的数据应保持独立于开发和训练数据。

问答环节

During Q&A, regulators asked how independence can be maintained where a model continuously retrains on operational data. The answer was to use lifecycle mechanisms, defined samples, test sets, procedures and controls to preserve a means of independent verification; where the model evolves, monitoring and change-control triggers should determine when additional verification or re-optimisation is required. The response was not fully prescriptive and did not define one mandatory technical approach.

在问答环节中,监管方询问在模型基于运行数据持续再训练的情况下如何保持独立性。答案是使用生命周期机制、定义的样本、测试集、程序和控制来保留独立验证的手段;当模型演进时,监测和变更控制触发因素应确定何时需要额外验证或再优化。回应并非完全规定性的,也未定义一个强制性的技术解决方案。

The session also dealt with human-in-the-loop, fallback and oversight in the validation context. Discussion from earlier sessions was carried forward: human oversight cannot be a substitute for inadequate validation, and control effectiveness must be evidenced. There was recognition that automation bias, human bias, skill level, AI literacy, explainability, confidence scores and challenge/testing of operators may be relevant. Lifecycle monitoring should include incidents and problems, Corrective and Preventive Action (CAPA), rollback or intervention criteria, supplier oversight and cloud/provider change visibility. The presenters linked supplier oversight to validation because the regulated user's responsibility cannot be transferred to the supplier.

会议还讨论了验证背景下的人机回环、回退和监督。先前会议的讨论被延续:人类监督不能替代不足的验证,控制有效性必须有证据支持。与会者认识到,自动化偏见、人类偏见、技能水平、AI素养、可解释性、置信度分数以及对操作员的挑战/测试可能相关。生命周期监测应包括事件和问题、纠正和预防措施(CAPA)、回滚或干预标准、供应商监督以及云/提供商变更可见性。演讲者将供应商监督与验证联系起来,因为受监管用户的责任不能转移给供应商。

A recurring topic was use-case patterns. Presenters suggested that risk patterns may be developed for similar AI use-case types, including risks such as cybersecurity, model evasion or data poisoning, and that applicable control strategies could be mapped accordingly. However, not all controls fit every use case. The workshop did not conclude a definitive template, but the concept of reusable patterns emerged as an idea for further discussion.

一个反复出现的话题是用例模式。演讲者建议可以为类似的AI用例类型开发风险模式,包括网络安全、模型规避或数据投毒等风险,并相应映射适用的控制策略。然而,并非所有控制都适用于每个用例。研讨会未得出明确的模板,但可重用模式的概念作为进一步讨论的想法浮现出来。

Overall, this topic reinforced that Annex 22 should require a documented AI lifecycle: intended use, design/qualification, independent verification, acceptance criteria, monitoring, change control, revalidation/re-verification, incident and problem management, CAPA, supplier oversight, and periodic risk review. It also reinforced the need to keep technical examples flexible and updateable.

总体而言,本专题强化了附录22应要求记录的AI生命周期:预期用途、设计/确认、独立验证、验收标准、监测、变更控制、再验证/重新验证、事件和问题管理、CAPA、供应商监督以及定期风险审查。它还强化了保持技术示例灵活和可更新的必要性。

讨论中涌现的关键信息

Validation should be a lifecycle, not a one-time event.

验证应是一个生命周期,而非一次性事件。

Validation depth should be driven by intended use, uncertainty, importance, complexity, autonomy and adaptiveness.

验证深度应由预期用途、不确定性、重要性、复杂性、自主性和适应性驱动。

Independent verification evidence remains a critical expectation, but implementation for continuously learning systems needs clarification.

独立验证证据仍是一个关键期望,但对于持续学习系统的实施需要澄清。

Monitoring, change control, incident/problem management and CAPA are part of validation lifecycle management.

监测、变更控制、事件/问题管理和CAPA是验证生命周期管理的一部分。

Supplier and cloud-provider oversight are integral because accountability remains with the regulated user.

供应商和云提供商监督是不可或缺的,因为问责制仍由受监管用户承担。

对附录22起草的潜在影响

The validation session may influence the guideline text on lifecycle control, continued model performance, revalidation/re-verification and change management. Definitions may need to include lifecycle validation, independent test set, operational monitoring, model drift, re-optimisation, change trigger and control effectiveness.

验证会议可能影响关于生命周期控制、持续模型性能、再验证/重新验证和变更管理的指南文本。定义可能需要包括生命周期验证、独立测试集、运行监测、模型漂移、再优化、变更触发和控制有效性。

Industry expectations would include an AI validation/qualification package aligned to risk and intended use.

行业期望将包括与风险和预期用途一致的AI验证/确认包。

1.6 Topic 6: Cybersecurity and Outsourced activities专题6:网络安全与外包活动

起草组提出的问题

Is there a need to address prevention and protection of AI systems from tampering or unauthorized access in Annex 22 or is this already sufficiently addressed in Annex 11? What would you propose in terms of appropriate wording?

是否有必要在附录22中解决AI系统防篡改或未经授权访问的预防和保护问题,还是附录11已充分解决?您会建议什么样的适当措辞?

What risks arise when guardrail infrastructure is outside the manufacturer's direct control, and how can supplier qualification, change-control visibility, and independent audit capabilities be preserved in such cloud-based AI supply chains?

当护栏基础设施超出生产商的直接控制时,会产生哪些风险,以及如何在基于云的AI供应链中保持供应商资质、变更控制可见性和独立审计能力?

行业观点

The final topic addressed cybersecurity and outsourced activities. The presenter began from the premise that AI systems consist of data, model/software and hardware; therefore, from a GMP perspective they remain computerised systems. On this basis, the interested parties' position was that existing Annex 11 principles, particularly Chapter 7 on outsourced activities, already provide an appropriate framework for cybersecurity and supplier management, including for AI systems. The presentation also referenced broader cybersecurity frameworks and standards such as NIS2, SOC 2 and ISO 27000 series, and noted that the EU AI Act contains extensive cybersecurity references.

最后一个专题涉及网络安全和外包活动。演讲者从AI系统由数据、模型/软件和硬件组成的前提出发;因此,从GMP角度看,它们仍然是计算机化系统。在此基础上,利益相关方的立场是,现有的附录11原则,特别是关于外包活动的第7章,已经为网络安全和供应商管理(包括AI系统)提供了适当的框架。演讲还参考了更广泛的网络安全框架和标准,如NIS2、SOC 2和ISO 27000系列,并指出欧盟AI法案包含广泛的网络安全条款。

The first core message was that AI does not require a separate cybersecurity regime in Annex 22 if Annex 11 and applicable cybersecurity standards are properly applied. Access controls, audit trails, supplier qualification, continuous monitoring, agreements and legal arrangements were described as already covered by Annex 11. The second message was that the regulated pharmaceutical user remains accountable for the computerised/AI system even when a third party provides the model, platform or cloud infrastructure. Responsibilities may be shared operationally, but accountability for GMP use remains with the regulated user.

第一个核心信息是,如果附录11和适用的网络安全标准得到正确应用,AI不需要在附录22中单独的网络安全制度。访问控制、审计追踪、供应商资质、持续监测、协议和法律安排被描述为已被附录11涵盖。第二个核心信息是,即使第三方提供模型、平台或云基础设施,受监管的药品使用者仍对计算机化/AI系统负责。责任可能在运营上共享,但GMP使用的问责制仍由受监管用户承担。

Supplier and cloud oversight were discussed in detail. The presenter argued that agreements must clarify responsibilities, configuration expectations, change control participation, evidence access, audit/certification arrangements and transparency of controls. The supplier must understand pharmaceutical requirements and language. For cloud and AI services, the pharmaceutical user should require evidence for relevant controls, data integrity, anonymisation/minimisation, zero data retention where applicable, confidential computing or trusted execution environments, regional requirements, federated learning and data loss prevention where relevant. The presentation framed this as a shared responsibility model in which the provider and customer responsibilities vary depending on infrastructure-as-a-service, platform-as-a-service or software-as-a-service models.

详细讨论了供应商和云监督。演讲者认为,协议必须明确责任、配置期望、变更控制参与、证据访问、审计/认证安排和控制的透明度。供应商必须理解药品要求和语言。对于云和AI服务,

药品使用者应要求相关控制、数据完整性、匿名化/最小化、适用情况下的零数据保留、机密计算或可信执行环境、区域要求、联邦学习以及适用情况下的数据丢失防护的证据

。演讲将此框架为共享责任模型,其中提供商和客户责任因基础设施即服务、平台即服务或软件即服务模型而异。

问答环节

The Q&A focused on what objective assurance evidence is needed where direct physical audit of major cloud suppliers is difficult. The answer was to use documented evidence such as certifications and compliance information, including NIS2, SOC 2, ISO 27000 series and European framework suitability, and then assess alignment with the company's own requirements. If the provider cannot meet requirements, another option should be chosen. A chat question asked whether current GMP requirements from Annex 11 Chapter 7 are sufficient for AI suppliers/service providers. The response was that Annex 11 Chapter 7 is fully applicable because AI systems are computerised systems; however, agreement types may vary depending on the service and intended use.

问答环节聚焦于在难以对主要云供应商进行直接物理审计的情况下需要什么客观保证证据。答案是使用记录的证据,如认证和合规信息,包括NIS2、SOC 2、ISO 27000系列和欧洲框架适用性,然后评估与公司自身要求的一致性。如果提供商无法满足要求,应选择另一选项。一个聊天问题询问附录11第7章的当前GMP要求是否足以适用于AI供应商/服务提供商。回应是,附录11第7章完全适用,因为AI系统是计算机化系统;但是,协议类型可能因服务和预期用途而异。

The discussion also linked back to previous topics: where there is limited visibility of supplier changes, regulated users need to be part of change control and have visibility over control logic. Control of outsourced AI therefore links to validation lifecycle and monitoring. The final session supports an Annex 22 approach that references Annex 11 rather than duplicating it, while potentially adding AI-specific considerations around evidence, transparency, change visibility, data retention, confidential computing and shared responsibility.

讨论还回溯到先前专题:在供应商变更可见性有限的情况下,受监管用户需要参与变更控制并对控制逻辑具有可见性。因此,外包AI的控制与验证生命周期和监测相关联。最后一场会议支持附录22采用引用附录11而非重复它的方法,同时可能添加围绕证据、透明度、变更可见性、数据保留、机密计算和共享责任的AI特定考虑因素。

讨论中涌现的关键信息

AI systems should be treated as computerised systems for cybersecurity and outsourcing governance.

出于网络安全和外包治理目的,AI系统应被视为计算机化系统。

Annex 11 Chapter 7 was considered applicable and sufficient at a principle level.

附录11第7章被认为在原则层面适用且充分。

The regulated user remains accountable even when using third-party AI or cloud services.

即使使用第三方AI或云服务,受监管用户仍承担问责制。

Supplier agreements must address responsibilities, evidence, configuration and change control.

供应商协议必须解决责任、证据、配置和变更控制问题。

AI-specific examples may be useful, but duplication of Annex 11 should be avoided.

AI特定示例可能有用,但应避免重复附录11。

对附录22起草的潜在影响

Annex 22 may be influenced to cross-reference Annex 11 for cybersecurity and outsourced activities while adding targeted AI-specific considerations. Definitions may include shared responsibility, third-party AI provider, regulated user accountability, supplier evidence and change visibility.

附录22可能会受到影响,对网络安全和外包活动交叉引用附录11,同时添加针对性的AI特定考虑因素。定义可能包括共享责任、第三方AI提供商、受监管用户问责制、供应商证据和变更可见性。

Industry expectations would include supplier qualification, agreements, certifications/evidence, documented responsibilities and continuous monitoring.

行业期望将包括供应商资质、协议、认证/证据、记录的责任和持续监测。

2. 结论

The workshop provided valuable expert input to support the ongoing development of Annex 22 and demonstrated that a broad range of industry stakeholders were able to articulate a largely aligned position across the six discussion topics. While recognising the challenges associated with adaptive, probabilistic and generative AI systems, participants consistently advocated for a risk-based, technology-neutral approach in which the acceptability of an AI application is determined by its intended use, the risks it presents, and the effectiveness of the controls applied throughout its lifecycle rather than by the underlying technology type itself.

研讨会为附录22的持续制定提供了宝贵的专家意见,并表明广泛的行业利益相关方能够在六个讨论专题上表达基本一致的立场。尽管在认识到适应性、概率性和生成式AI系统带来的挑战的同时,与会者始终倡导基于风险的、技术中立的方法,其中AI应用的可接受性由其预期用途、所呈现的风险以及在其整个生命周期中应用的控制的有效性决定,而非由底层技术类型决定。

Across all sessions, several common themes emerged. These included the importance of documented intended use, robust quality risk management, lifecycle validation and monitoring, risk-proportionate human oversight, effective control mechanisms, clear accountability, supplier oversight, and the maintenance of a demonstrable state of control. Participants also emphasised that uncertainty, complexity, model influence and detectability are important considerations when evaluating AI systems used in GMP-regulated activities.

在所有会议中,出现了几个共同主题。这些包括记录的预期用途的重要性、稳健的质量风险管理、生命周期验证和监测、风险相称的人类监督、有效的控制机制、明确的问责制、供应商监督以及保持可证明的状态受控。与会者还强调,在评估用于GMP监管活动的AI系统时,不确定性、复杂性、模型影响和可检测性是重要的考虑因素。

The discussions further demonstrated that existing GMP principles, particularly those relating to quality risk management, validation, lifecycle management, oversight, and outsourced activities, remain relevant and applicable to AI systems. At the same time, the workshop helped identify areas where additional clarification, definitions, or AI-specific expectations may be needed to ensure that Annex 22 provides a practical and scientifically sound framework for the use of increasingly sophisticated AI technologies in GMP environments.

讨论进一步表明,现有的GMP原则,特别是与质量风险管理、验证、生命周期管理、监督和外包活动相关的原则,仍然与AI系统相关且适用。同时,研讨会帮助识别了可能需要额外澄清、定义或AI特定期望的领域,以确保附录22为在GMP环境中使用日益复杂的AI技术提供实用且科学合理的框架。

Overall, the workshop achieved its objective of gathering structured, experience-based input from industry on the governance, control, validation and oversight of AI systems and provided the drafting group with valuable insights to support the revision of Annex 22. The discussions suggest that subject to appropriate controls and a robust risk-based framework, there may be viable pathways for the use of adaptive, probabilistic and generative AI technologies in GMP-regulated applications while continuing to protect product quality and patient safety.

总体而言,研讨会实现了其收集行业关于AI系统治理、控制、验证和监督的结构化、基于经验的输入的目标,并为起草组提供了宝贵的见解以支持附录22的修订。讨论表明,在适当控制和稳健的风险框架的前提下,适应性、概率性和生成式AI技术在GMP监管应用中的使用可能存在可行路径,同时继续保护产品质量和患者安全。

3. 后续步骤

The information presented during the workshop, including the consolidated responses provided by the participating industry associations, will be carefully reviewed and considered by the Annex 22 drafting group.

研讨会期间提交的信息,包括参与行业协会提供的统一回复,将由附录22起草组仔细审查和考虑。

Following the workshop, the drafting group will assess the key themes, common approaches, areas of divergence, and identified challenges discussed across the six topics. The outcomes of these discussions will be used to inform the ongoing revision of Annex 22, with the aim of developing a science-based, risk-proportionate, and implementable framework for the use of AI in GMP-regulated environments. Where necessary, additional clarification and internal discussion may be undertaken to ensure that the revised guidance appropriately balances the protection of product quality and patient safety with the potential benefits of technological innovation.

研讨会结束后,起草组将评估在六个专题中讨论的关键主题、共同方法、分歧领域和已识别的挑战。这些讨论的结果将用于指导附录22的持续修订,目的是制定一个基于科学的、风险相称的且可实施的框架,用于在GMP监管环境中使用AI。必要时,可能会进行额外的澄清和内部讨论,以确保修订后的指南在保护产品质量和患者安全与技术创新潜在益处之间取得适当平衡。

来源:BasicPharma搬砖工 · mp.weixin.qq.com