This invention discloses a
reinforcement learning-based standard alignment method for automatically determining and optimizing terminology definitions. The method first constructs a structured instruction fine-tuning dataset integrating expert instructions and standard terminology, and performs supervised fine-tuning on a basic large
language model to obtain a
concept extraction model and a multi-dimensional judgment model. Second, the
concept extraction model is used to analyze the superordinate concepts and distinguishing features of the target term, and the multi-dimensional judgment model automatically evaluates the definition to be optimized across six dimensions: accuracy, conciseness, appropriateness, substitution principle, circular definition principle, and
negation form principle. Finally, a composite reward model is constructed, combining the multi-dimensional evaluation results with the
semantic similarity to the
standard definition, and the GRPO
reinforcement learning algorithm is used to iteratively optimize the definition generation strategy until the
semantic similarity between the generated revised definition and the
standard definition exceeds a preset threshold. This invention achieves automated, standardized, and standard-aligned terminology definitions, significantly improving the quality and consistency of terminology definitions in standardized documents, and is applicable to intelligent standard-setting scenarios in professional fields such as
medicine, law, and finance.