Text generation processing method and device based on term alignment and related equipment

By constructing a business standard terminology graph for terminology candidate matching and contextual disambiguation, the problem of inaccurate terminology recognition in the financial field by large language models is solved, and highly accurate and standardized text generation is achieved.

CN121145852APending Publication Date: 2025-12-16PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511060335.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Large language models have low accuracy in specific business scenarios, especially in the financial field, where there are problems such as inaccurate terminology recognition and inconsistent generation results, which affect the professionalism and credibility of the output.

Method used

By constructing a business standard terminology graph, performing term candidate matching and context disambiguation, identifying business terms in the input text, generating target prompt words, and using a large language model to generate accurate output text.

Benefits of technology

It improves the accuracy and controllability of the output of large language models in specific domain business scenarios, and ensures the standardization and credibility of terminology usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145852A_ABST
    Figure CN121145852A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, discloses a text generation processing method and device based on term alignment, equipment and a medium, and aims to solve the problem of relatively low output accuracy of a large language model in a specific field business scene in the traditional technology. Term candidate matching and context disambiguation are carried out on service terms contained in a user input text to obtain target service standard terms with the highest text context integrating degree with the input text, and term-level semantic alignment of the user text is achieved; generating a target cue word according to the input text and a target business standard term, and generating an output text based on a large language model to realize text generation control and semantic enhancement based on term perception, so as to combine term-level semantic alignment with term perception semantic control by means of a preset business standard term atlas; the method can improve the accuracy of large model language text generation, and can be applied to but not limited to the financial field, the insurance field and the medical health field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of natural language processing and financial technology, and particularly relates to a text generation processing method and device based on term alignment, equipment and a medium. BACKGROUND

[0002] A large language model (LLM) is a natural language processing (NLP) model based on artificial intelligence (AI) and deep learning, which can understand, generate and manipulate human language. With the development of large language models, large language models have been widely applied in more and more fields and scenarios. For example, with the wide application of large language models (LLM) in the financial field, it has shown good performance in financial customer service, intelligent investment, report analysis and other tasks.

[0003] In the traditional technology, the application of large language models in specific fields and business scenarios is generally based on fine-tuning of general large language models to improve the performance of the corresponding models to make the large language models suitable for specific fields and business scenarios. For example, in the financial field, a general large language model is trained using corresponding training samples in the financial field to make the large language model suitable for business scenarios in the banking, securities and other financial fields, or in the insurance field, a general large language model is trained using corresponding training samples in the insurance field to make the large language model suitable for business scenarios in the insurance field.

[0004] However, the inventors realized that due to the limitations of fine-tuning large language models in the above-mentioned traditional technology and the complexity of the model usage environment corresponding to specific fields and business scenarios, the large language model has inaccurate recognition of related terms, and further generates incorrect or inconsistent terms, which makes it difficult for the fine-tuned large language model to have high output accuracy, thereby affecting the professionalism and credibility of the output.

[0005] For example, in the financial field, financial terms are often ambiguous and have industry-specific meanings. For instance, the "discount rate" has different meanings in accounting, macroeconomics, and banking, leading to strong ambiguity and making it difficult for traditional large language models to distinguish them correctly. Furthermore, there are many non-standardized expressions in users' natural language, such as "return on equity" and "ROE", "total revenue" and "total operating income". As a result, the expression of these terms is diverse, leading to a lack of a unified understanding framework for large language models. Therefore, due to the lack of sufficient domain knowledge injection, general-purpose large language models struggle to accurately map user expressions to corresponding standard terms. Mainstream pre-trained large language models (such as ChatGPT, BERT, and their financial fine-tuned versions) often fail to accurately identify terms in financial scenarios when dealing with professional terms and concepts, resulting in incorrect or inconsistent terminology usage in the generated results. Significant issues such as terminology ambiguity, semantic drift, and difficulties in knowledge alignment persist, especially in complex financial product scenarios (such as structured deposits and equity swaps). This severely impacts the accuracy and controllability of the model's output, making it difficult to achieve high output accuracy and ultimately affecting the professionalism and credibility of the output.

[0006] Therefore, improving the output accuracy of large language models has become an urgent technical problem to be solved in natural language processing business scenarios, including but not limited to the financial and insurance sectors. Summary of the Invention

[0007] This invention provides a text generation and processing method, apparatus, computer device, and medium based on term alignment, to solve the technical problem of low output accuracy of large language models in specific domain business scenarios in traditional technologies.

[0008] Firstly, a text generation processing method based on terminology alignment is provided, comprising: responding to a text generation instruction and determining the input text corresponding to user input; identifying business terms contained in the input text; performing terminology candidate matching on the business terms based on a preset business standard terminology graph to obtain a number of candidate business standard terms corresponding to the business terms; performing context disambiguation on the candidate business standard terms according to the input text and the candidate business standard terms to obtain the target business standard term with the highest text context fit corresponding to the input text; generating target prompt words according to the input text and the target business standard terms; and generating output text corresponding to the input text based on the target prompt words and a preset text generation large language model.

[0009] Secondly, a text generation processing device based on terminology alignment is provided, comprising: a first determining module, configured to determine the input text corresponding to user input in response to a text generation instruction; a first recognizing module, configured to recognize the business terms contained in the input text; a first matching module, configured to perform terminology candidate matching on the business terms based on a preset business standard terminology map, to obtain a plurality of candidate business standard terms corresponding to the business terms; a context disambiguation module, configured to perform context disambiguation on the candidate business standard terms according to the input text and the candidate business standard terms, to obtain the target business standard term with the highest text context fit corresponding to the input text; a first generating module, configured to generate target prompt words according to the input text and the target business standard terms; and a second generating module, configured to generate output text corresponding to the input text according to the target prompt words and based on a preset text generation large language model.

[0010] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0011] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described method.

[0012] In the aforementioned text generation processing method, apparatus, computer equipment, and storage medium based on terminology alignment, the method performs terminology candidate matching and context disambiguation on the business terms contained in the user's input text to obtain the target business standard terms with the highest text context fit corresponding to the input text, thus achieving terminology-level semantic alignment of the user text; and generates target prompt words based on the input text and target business standard terms, and then generates the output text corresponding to the input text based on a preset text generation large language model, thereby achieving text generation control and semantic enhancement based on accurate terminology awareness, and thus leveraging a preset business standard terminology map and through terminology candidate matching and context disambiguation. This approach achieves terminology-level semantic alignment and terminology disambiguation in the input text. Combined with the text generation control of a large language model, it enables corresponding semantic control of text generation based on accurate terminology awareness. By leveraging a pre-defined business standard terminology graph and combining terminology-level semantic alignment with semantic control of text generation based on accurate terminology awareness, a "alignment-guidance" text generation terminology usage control mechanism is implemented. This improves the controllability, accuracy, and standardization of terminology usage, thereby enhancing the controllability, accuracy, and standardization of the output and response of the large model's language text generation in specific business scenarios. It can be applied to fields including but not limited to finance, insurance, and healthcare. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating the text generation and processing method based on term alignment provided in an embodiment of the present invention;

[0015] Figure 2 This is a schematic diagram of the first sub-process of the text generation and processing method based on term alignment provided in an embodiment of the present invention;

[0016] Figure 3 A schematic diagram of the second sub-process of the text generation and processing method based on term alignment provided in an embodiment of the present invention;

[0017] Figure 4 A schematic block diagram of a text generation and processing apparatus based on term alignment provided in an embodiment of the present invention;

[0018] Figure 5 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0019] Figure 6 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0022] This invention provides a text generation and processing method based on term alignment. The method can be applied to computer devices including but not limited to smartphones, tablets, desktop computers, servers, and cloud platforms. It is used in text generation based on large language models in fields including but not limited to finance, insurance, and healthcare. For example, in the financial field, it is used in business scenarios such as intelligent customer service, investment advisory systems, insurance Q&A, and financial statement analysis in banks and insurance companies.

[0023] To address the technical problem of low output accuracy of large language models in specific domain business scenarios in traditional technologies, the inventors propose a text generation processing method based on terminology alignment in this invention. The core idea of ​​this invention is as follows: for specific domain business scenarios, a corresponding business standard terminology graph is constructed. The business terms in the user-input text are matched with the business standard terms in the business standard terminology graph, and candidate term matching and contextual disambiguation are performed to determine the target business standard terms that best fit the input text, thereby achieving terminology-level semantic alignment. Based on the target business standard terms, the large language model is guided to generate and output corresponding text, achieving terminology-aware text generation control. This improves the accuracy and standardization of output and response of large model language text generation in specific domain business scenarios and can be applied to fields including but not limited to finance, insurance, and healthcare.

[0024] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0025] Please see Figure 1 , Figure 1 This is a flowchart illustrating the text generation and processing method based on term alignment provided in an embodiment of the present invention. Figure 1 As shown, the method includes, but is not limited to, the following steps S11-S16:

[0026] S11. Respond to the text generation instruction and determine the input text corresponding to the user input.

[0027] Explained, in specific business scenarios, text generation typically involves user input of relevant text. The user inputs the corresponding text, triggering a text generation command. This command then retrieves and determines the user's input text. For example, in the financial sector, during business processes related to intelligent customer service, investment advisory systems, and financial statement analysis in banking and insurance, users input relevant text, including but not limited to information related to banking, insurance, finance, and investment. This triggers a text generation command, thus determining the relevant user-input text. Similarly, in the healthcare sector, in digital healthcare intelligent customer service scenarios, when a patient inputs a question requiring an answer, a text generation command is triggered, thus determining the patient's question text.

[0028] S12. Identify the business terms contained in the input text.

[0029] Explained, business terms refer to terms related to relevant business operations. Therefore, the input text can be divided into two parts: business terms and non-business terms. Non-business terms are everyday expressions or terms. A preliminary distinction is made between business and non-business terms. In specific business scenarios, non-business terms are generally more universal, frequently used, and numerous. Therefore, a non-business term dictionary can be pre-set, and the input text can be segmented. Combining this with the non-business term dictionary, the business terms contained in the input text can be identified. This step can borrow techniques commonly used in natural language processing. For example, in the financial field, if a user asks, "How is our gross profit this quarter?", based on word segmentation and the non-business term dictionary, phrases like "How is our gross profit this quarter?", "situation", and "how is it?" corresponding to non-financial business terms are first removed, leaving only a limited set of terms related to "gross profit" as the identified business terms in the user's input text. Further targeted processing of these business terms can greatly improve data processing efficiency.

[0030] S13. Based on a preset business standard terminology map, perform terminology candidate matching on the business terms to obtain several candidate business standard terms corresponding to the business terms.

[0031] Explained, a pre-defined business standard terminology graph is established, i.e., a preset business standard terminology graph. This graph represents a graph constructed based on business standard terms. It includes preset business standard terminology nodes and edges. Nodes represent preset business standard terms, and edges represent the relationships between these nodes. For example, in the financial field, a preset financial business standard terminology graph is constructed based on industry standards and policy documents. Each node represents a financial business standard term, which includes, but is not limited to, net profit, gross profit margin, net assets, and liabilities. Edges include, but are not limited to, the following semantic relationships: synonym / alias relationships (e.g., "net profit" – "Net Income" – "Profit after tax"); hierarchical relationships (i.e., domains, e.g., "revenue" – "main business revenue"); attribute relationships (e.g., "interest rate" associated with "interest calculation method"); and regulatory relationships (e.g., "reserve requirement ratio" ←→ "People's Bank of China policy").

[0032] Based on the above concept and setup, the similarity between the business term and the preset business standard terms contained in the preset business standard terminology graph is calculated and compared based on vector distance. Vector retrieval based on, but not limited to, FAISS vector index and Top-K retrieval can be used to perform term candidate matching between the business term and the preset business standard terms. This allows for the selection of several candidate business standard terms that are relatively similar to the business term from the preset business standard terminology graph, resulting in several candidate business standard terms corresponding to the business term. Candidate business standard terms represent preset business standard terms that are in the candidate state and have not yet been determined to be the most similar to the business term. The preset business standard terms correspond to the preset business standard term nodes contained in the preset business standard terminology graph.

[0033] Since a pre-defined business standard terminology map typically contains a large number of pre-defined business standard terms, and may also include variations and abbreviations of the corresponding business standard terms, it is not suitable to use model-based contextual disambiguation to make a contextual semantic judgment for each pre-defined business standard term. Therefore, by first using vector distance, we can quickly and initially retrieve and screen out the candidate business standard terms that are "possibly correct" from the huge number of pre-defined business standard terms contained in the pre-defined business standard terminology map. This allows for fast and efficient retrieval and screening of candidate business standard terms. Thus, candidate retrieval relies on vector distance to achieve efficient retrieval, which greatly narrows the scope of further screening of target business standard terms, so as to further refine the screening of candidate business standard terms.

[0034] S14. Based on the input text and the candidate business standard terms, perform context disambiguation on the candidate business standard terms to obtain the target business standard terms with the highest text context matching degree corresponding to the input text.

[0035] Explained, based on the input text and candidate business standard terms, the candidate business standard terms are combined with the contextual semantics corresponding to the input text to perform semantic adaptation. This allows for the precise retrieval of the business standard terms that best match the contextual semantics of the input text from the candidate business standard terms, while filtering out candidate business standard terms whose contextual semantics do not match the input text's context. This process of context disambiguation yields the target business standard terms with the highest textual contextual fit to the input text. Here, context disambiguation refers to using contextual semantics to eliminate ambiguous candidate business standard terms, i.e., using the input text... This process uses the contextual semantic context to identify and filter out candidate business standard terms that cause ambiguity in the input text, in order to select and obtain the accurate target business standard terms that best fit the contextual semantic context of the input text. This addresses the problem of inaccurate recognition of related terms by large language models, leading to incorrect or inconsistent terminology usage in the generated results, when there are ambiguous terms in the input text. Contextual disambiguation is generally carried out through modeling, for example, by constructing a context-aware discriminator based on, but not limited to, Siamese networks and Siamese neural networks, to determine which candidate business standard term best matches the contextual semantics of the input text.

[0036] For example, in the financial field, in a financial business scenario, if a user asks, "How is our gross profit this quarter?", the financial term "gross profit" is used. After the above-mentioned term candidate matching and context disambiguation, "gross profit margin" is finally obtained as the target business standard term with the highest text context matching degree to the input text. Subsequently, the corresponding text is generated based on "gross profit margin".

[0037] Therefore, based on the above-mentioned term candidate matching, efficient retrieval is achieved, which greatly narrows the scope of further screening of target business standard terms. Furthermore, contextual disambiguation judgment relies on context modeling to achieve further precise screening of candidate business standard terms. That is, candidate retrieval relies on vector distance (efficient), and disambiguation judgment relies on context modeling (precise). Thus, by combining the two stages corresponding to term candidate matching and contextual disambiguation, term-level semantic alignment that balances efficiency and accuracy is achieved, so that the corresponding text can be generated based on more accurate target business standard terms in the future.

[0038] S15. Generate target prompt words based on the input text and the target business standard terms.

[0039] Explained, a prompt in a large language model is a text determined based on user input, used to guide the model to generate a specific response or perform a specific task. Thus, a target prompt is generated based on the input text and target business standard terms. For example, the input text and target business standard terms are embedded into a preset prompt template to generate the target prompt, or a generative rewriting model is used to generate the target prompt.

[0040] Further, based on the input text and the target business standard terms, target prompt words are generated, including:

[0041] Based on the input text and the target business standard terms, the target business standard terms are inserted and displayed to generate target prompt words.

[0042] Specifically, "explicit terminology insertion" refers to embedding key terms into prompts in some form to guide the generative model to focus on and use these terms. Target prompts can be generated in, but are not limited to, the following two ways:

[0043] Method 1: Match templates from a template library, select and rewrite prompt word templates, and prepare multiple rewritten templates for different question structures or intent types, such as:

[0044] "Please analyze in conjunction with '[Terminology]'...";

[0045] "Consider the indicator '[terminology]'..."

[0046] Identify the question structure or intent type corresponding to the user's input text, such as inquiring about profitability or cost analysis, and obtain the corresponding term tags, such as "gross profit margin." For example, use a pre-set intent recognition module (or simple rule matching) to determine the current intent category; then match the corresponding template and insert the corresponding terms into it.

[0047] Method 2: Generative rewriting. Use a rewriting model (such as T5) to take "input text + term labels" as input, generate a rewritten Prompt, and generate target prompt words.

[0048] By inserting target business standard terms into the display and generating target prompt words, the corresponding generation model is guided to pay attention to and use these target business standard terms. This addresses the problem of inaccurate recognition of relevant terms by large language models when there are ambiguous terms in the input text, which leads to incorrect or inconsistent terminology usage in the generated results. This further improves the controllability, accuracy, and standardization of the use of relevant terms, and thus further improves the controllability, accuracy, and standardization of the output and response of the large model's language text generation in specific business scenarios.

[0049] S16. Based on the target prompt word and a preset text generation large language model, generate the output text corresponding to the input text.

[0050] Explainedly, based on target prompts and a pre-defined text generation language model, the output text corresponding to the input text is generated. Since the target business standard terms are used accurately, the terms used in the pre-defined text generation language model are also accurate, resulting in correspondingly accurate generated text. This improves the controllability, accuracy, and standardization of the use of corresponding terms, thereby enhancing the controllability, accuracy, and standardization of the output and response of the large model language text generation in specific domain business scenarios.

[0051] Furthermore, based on a pre-defined text generation language model, the output text corresponding to the input text can be generated. Logit bias can also be employed to guide the pre-defined text generation language model to accurately apply the corresponding terms. Logit bias includes the following processes: fine-tuning the word probability distribution of the output layer softmax; setting a positive bias for the corresponding terms or their synonyms to increase their generation probability; and controlling the appearance of the corresponding terms in appropriate positions. For example, in the financial field, when a user inputs "How is our gross profit this quarter?", the generated target prompt could be: "Please answer in conjunction with the 'gross profit margin' indicator: How is our profitability this quarter?" The output could be a structured and terminologically compliant natural language response text, such as: "The company's gross profit margin this quarter was 43.5%, an increase of 2.4% year-on-year, demonstrating good profitability."

[0052] This invention, in its embodiments, achieves terminology-level semantic alignment of user text by performing terminology candidate matching and context disambiguation on the business terms contained in the user's input text to obtain the target business standard terms with the highest text context fit corresponding to the input text. Based on the input text and the target business standard terms, target prompt words are generated. Then, based on a preset text generation large language model, the output text corresponding to the input text is generated, achieving text generation control based on accurate terminology awareness. This leverages a preset business standard terminology map and achieves terminology-level semantic alignment through terminology candidate matching and context disambiguation, thus achieving terminology disambiguation of the input text—that is, aligning the input text with business standard terms. Combined with text generation control using a large language model, this achieves semantic control and semantic enhancement based on accurate terminology usage in the corresponding text generation. By leveraging a pre-defined business standard terminology map and combining terminology-level semantic alignment with text generation semantic control based on accurate terminology awareness, a "alignment-guidance" text generation terminology usage control mechanism is achieved. This effectively alleviates the problems of weak professionalism and insufficient terminology understanding in specific fields such as finance, insurance, and healthcare that general-purpose large language models lack. It has strong generalization capabilities for various user expressions, and even if non-professional users describe vaguely, it can accurately map to the corresponding business standard terms, significantly reducing errors in terminology recognition and generation confusion. This improves the controllability, accuracy, and standardization of terminology usage, thereby enhancing the controllability, accuracy, and standardization of the output and response of large-scale model language text generation in specific business scenarios. It can be applied to fields including but not limited to finance, insurance, and healthcare.

[0053] In one embodiment, please refer to Figure 2 , Figure 2 This is a schematic diagram of the first sub-process of the text generation and processing method based on term alignment provided in an embodiment of the present invention. For example... Figure 2 As shown, the preset business standard terminology map includes preset business standard terminology nodes, and each preset business standard terminology node corresponds to a preset business standard term. Based on the preset business standard terminology map, terminology candidate matching is performed on the business terms to obtain several candidate business standard terms corresponding to the business terms, including:

[0054] S21. Determine the business term vector corresponding to the business term, and determine the business standard term vector corresponding to the preset business standard term;

[0055] S22. Calculate the vector similarity between the business term vector and the business standard term vector to obtain the first vector similarity.

[0056] S23. Sort the similarity of several first vectors in descending order, and determine the top several preset business standard terms that are most similar to the business term vector as candidate business standard terms, thereby obtaining several candidate business standard terms corresponding to the business term.

[0057] Explained, when the preset business standard terminology map contains preset business standard terminology nodes and the preset business standard terminology nodes correspond to preset business standard terms, the business terminology vector corresponding to the business term is determined, and several business standard terminology vectors corresponding to preset business standard terms are determined. The vector similarity between the business terminology vector and each business standard terminology vector is calculated to obtain several first vector similarities. Thus, the similarity between the business terminology and each preset business standard term is calculated and compared based on the vector distance to determine the degree of similarity between the two. The vector distance represents the distance or similarity between two vectors in space calculated by mathematical measurement methods, thereby measuring the correlation between them.

[0058] Then, the similarity of several first vectors is sorted in descending order, and the top-ranked similarity of several first vectors is selected. The corresponding preset business standard terms are the top few preset business standard terms that are most similar to the business term vector, and are determined as candidate business standard terms, thus obtaining several candidate business standard terms corresponding to the business term.

[0059] Further, determining the business standard term vector corresponding to the preset business standard terminology includes:

[0060] Determine the preset structured standard term tag corresponding to the preset business standard term node, wherein the preset structured standard term tag includes the preset business standard term and its corresponding preset standard term embedding vector;

[0061] The preset standard term embedding vector is used as the business standard term vector corresponding to the preset business standard term.

[0062] Specifically, pre-set structured standard term tags, i.e., preset structured standard term tags, represent the structured tags corresponding to preset business standard terms. Pre-set structured standard term tags correspond to preset business standard term nodes. Pre-set structured standard term tags can be in the form of a set, used to represent the relevant information data of preset business standard terms. Pre-set structured standard term tags contain preset business standard terms and their corresponding preset standard term embedding vectors. Among them, the preset standard term embedding vector represents the embedding vector corresponding to the preset business standard terms. The embedding vector represents the vector based on embedding technology corresponding to the preset business standard terms obtained based on, but not limited to, BERT or other encoders. Embeddings are a technology that maps high-dimensional data to a low-dimensional space. The core idea of ​​embedding is to transform data points into vectors with semantic understanding through mathematical models.

[0063] For example, in the financial field, a preset financial business standard terminology map includes preset financial business standard terminology nodes, preset financial business standard terminology nodes correspond to preset structured financial standard terminology tags, and preset structured financial standard terminology tags include preset financial business standard terms and their corresponding preset financial standard terminology embedding vectors. The preset structured financial standard term tags corresponding to the preset financial business standard term nodes can be represented as: T_i={name,aliases,domain,definition,relations,embedding}, where name represents the preset financial business standard term, aliases represent the synonyms corresponding to the preset financial business standard term, domain represents the hierarchical relationship (i.e., domain) corresponding to the preset financial business standard term, definition represents the definition corresponding to the preset financial business standard term, relations represent the association relationship corresponding to the preset financial business standard term, and embedding represents the embedding vector corresponding to the preset financial business standard term. As mentioned above, the relationships are as follows: synonym / alias relationship (e.g., “net profit” – “Net Income” – “Profit after tax”); hierarchical relationship (i.e., domain) (e.g., “revenue” – “main business revenue”); attribute relationship (e.g., “interest rate” associated with “interest calculation method”); regulatory relationship (e.g., “reserve requirement ratio” ←→ “People’s Bank of China policy”).

[0064] Based on the above concept and setup, the preset structured standard term tags corresponding to the preset business standard term nodes are determined. The preset structured standard term tags contain preset business standard terms and their corresponding preset standard term embedding vectors, as shown in the example above. Then, the preset standard term embedding vectors are used as the business standard term vectors corresponding to the preset business standard terms, thereby determining the business standard term vectors corresponding to the preset business standard terms. Since the preset business standard terms and their corresponding preset standard term embedding vectors are preset, they can be used directly when performing term candidate matching, which can improve the computational efficiency of term candidate matching and thus improve the processing efficiency of text generation. Especially in real-time data analysis business application scenarios, it can achieve real-time and high-efficiency user input response.

[0065] This invention employs a method of terminology candidate matching by performing vector distance calculations between business terms and preset business standard terms contained in a preset business standard terminology graph. This maps business terms to preset business standard terms within the graph, identifying several candidate business standard terms corresponding to each term. This process first uses vector distance to quickly retrieve and filter potentially correct candidate business standard terms from the vast number of preset business standard terms in the graph. Based on this vector distance, rapid and efficient retrieval and filtering of candidate business standard terms is achieved, significantly narrowing the scope for further filtering of target business standard terms. Furthermore, contextual disambiguation judgment relies on context modeling to achieve further precise filtering of candidate business standard terms. By combining terminology candidate matching and contextual disambiguation in a two-stage process, efficient and accurate term-level semantic alignment is achieved. This provides semantic anchors corresponding to business standard terms for subsequent text generation, ensuring that the generated text is accurately mapped to the corresponding business standard terms. This improves the accuracy and standardization of large-model language text generation output and response in specific domain business scenarios.

[0066] In one embodiment, before sorting the similarity scores of the plurality of first vectors in descending order, the method further includes:

[0067] Determine whether the similarity of the first vector is equal to or less than a preset first vector similarity threshold;

[0068] If the above judgment is true, the business term is determined to be a fuzzy term, and the step of "sorting the similarity of several first vectors in descending order" is executed;

[0069] If the above judgment is not true, the business term is determined to be a vague term, and based on the input text, a large language model is generated based on a preset text to generate the output text corresponding to the input text.

[0070] Explained, it is determined whether the first vector similarity is equal to or less than a preset first vector similarity threshold. The preset first vector similarity threshold represents the threshold value corresponding to whether a business term is a vague term. When the corresponding vector similarity threshold is equal to or less than the preset first vector similarity threshold, it is determined to be a vague term. Here, a vague term refers to a term that is unclear or ambiguous and whose corresponding meaning cannot be clearly and accurately understood. Vague terms include, but are not limited to, variants, abbreviations, polysemous terms, and unclear or inaccurate terms corresponding to non-standard business terms.

[0071] If the above judgment is true, that is, the first vector similarity is equal to or less than the preset first vector similarity threshold, the business term is determined to be a vague term. In this case, the step of "sorting several first vector similarities in descending order" is executed to perform subsequent term candidate matching and context disambiguation to achieve term-level semantic alignment. If the above judgment is false, that is, the first vector similarity is greater than the preset first vector similarity threshold, the business term is determined to be not a vague term, that is, the business term is a clear and accurate business standard term. The large language model is generated directly based on the input text and the preset text to generate the output text corresponding to the input text. Subsequent term candidate matching and context disambiguation are no longer performed to achieve term-level semantic alignment. For example, when the user inputs financial standard terms such as "net profit" and "debt-to-equity ratio," the system can directly enter the generation module to generate the corresponding text without needing to perform term candidate matching and context-based semantic alignment. However, if the user inputs terms such as "insurance policy," "interest rate," and "service," which have multiple meanings and are therefore ambiguous, further contextual judgment is required. In this case, term candidate matching and context-based semantic alignment are then performed to achieve term disambiguation. By introducing a "fuzzy term recognition" mechanism as a preprocessing step—that is, by introducing a pre-judgment of "whether the business term is a fuzzy term"—the system can balance text accuracy and text generation efficiency, thereby further improving the timeliness of text generation.

[0072] In this embodiment of the invention, term-level semantic alignment is achieved by leveraging the vector distance during term candidate matching and triggering term candidate matching and context disambiguation only when the business term is determined to be a fuzzy term. Otherwise, the corresponding text is generated directly. By determining whether the business term is a fuzzy term a step earlier, the accuracy and standardization of the output and response of the large model language text generation in specific domain business scenarios can be improved, while taking into account both text accuracy and text generation efficiency, thereby further improving the timeliness of text generation.

[0073] In one embodiment, based on the input text and the candidate business standard terms, context disambiguation is performed on the candidate business standard terms to obtain the target business standard terms with the highest text context fit corresponding to the input text, including:

[0074] Determine the semantic encoding of the input text corresponding to the input text, and determine the semantic encoding of the candidate standard terms corresponding to the candidate business standard terms;

[0075] Based on the semantic encoding of the input text and the semantic encoding of the candidate standard terms, and using a preset context-aware discriminator, calculate the semantic matching score between each pair of the semantic encoding of the input text and the semantic encoding of the candidate standard terms;

[0076] The candidate business standard term with the highest semantic matching score is taken as the target business standard term, thus obtaining the target business standard term with the highest text context fit to the input text.

[0077] Explained, a context-aware discriminator is pre-set, i.e., a pre-defined context-aware discriminator. A pre-defined context-aware discriminator is a model that judges which corresponding term best fits the contextual semantics of the current context based on contextual semantic perception. Pre-defined context-aware discriminators include, but are not limited to, discriminators corresponding to Siamese networks and Siamese neural networks.

[0078] Based on the above concept and setup, the semantic code of the input text corresponding to the input text is determined, and the semantic code of the candidate standard terms corresponding to the candidate business standard terms is determined. Based on the semantic codes of the input text and the candidate standard terms, and using a preset context-aware discriminator, the semantic matching score between each pair of input text semantic codes and candidate standard term semantic codes is calculated. The candidate business standard term with the highest semantic matching score is taken as the target business standard term, thus obtaining the target business standard term with the highest text context fit to the input text. In other words, the target business standard term is the business standard term that best matches the contextual semantics of the current input text. This achieves context disambiguation of the candidate business standard terms, enabling the generation of corresponding text using accurate and standardized target business standard terms even when the user inputs ambiguous terms. This allows for text generation control based on accurate term awareness, achieving semantic control and semantic enhancement based on the use of accurate terms in the corresponding text generation. By improving the controllability, accuracy, and standardization of the use of corresponding terms, the controllability, accuracy, and standardization of the output and response of the large-scale model language text generation in specific domain business scenarios can be improved.

[0079] For example, in the financial field, within a financial business scenario, based on a pre-set context-aware discriminator, candidate business standard terms are disambiguated to obtain the target business standard term with the highest textual context fit to the input text. This can be achieved through the following judgment process:

[0080] 1) Input components:

[0081] C_input: Represents the semantic encoding of contextual statements (such as "gross profit has improved this quarter");

[0082] T_i: Top-K candidate terms (such as "gross profit", "gross profit margin", "gross profit") and their semantic representations retrieved from the aforementioned preset financial business standard terminology map;

[0083] Typically, C_input and each T_i are fed together into a pre-defined context-aware discriminator (such as a Siamese network or a cross-attention network).

[0084] 2) Judgment Logic:

[0085] For each pair (C_input, T_i), output a matching score P_i;

[0086] The score represents the degree of semantic fit of the candidate term T_i in the context;

[0087] Choose T_i corresponding to argmax(P_i) as the most suitable term, that is, the target business standard term with the highest text context fit corresponding to the input text.

[0088] Among them, the "pairing scoring" mechanism corresponding to the preset context-aware discriminator can use a Siamese network (dual encoder to calculate cosine similarity) or a Cross-Encoder (such as BERT cross-attention encoding) to achieve more refined context modeling.

[0089] In this embodiment of the invention, a pre-set context-aware discriminator performs context disambiguation on several candidate business standard terms to obtain the target business standard term with the highest text context fit corresponding to the input text. Based on the above term candidate matching, context disambiguation judgment is performed based on context modeling to accurately select the target business standard term with the highest text context fit corresponding to the input text from a limited number of candidate business standard terms. Thus, by combining term candidate matching and context disambiguation in a two-stage process, term-level semantic alignment that balances efficiency and accuracy is achieved. This provides semantic anchors corresponding to business standard terms for subsequent corresponding text generation, thereby enabling the generated corresponding text to be accurately mapped to the corresponding business standard terms. This improves the accuracy and standardization of output and response of large model language text generation in specific domain business scenarios.

[0090] In one embodiment, before performing term candidate matching on the business terms based on a preset business standard terminology map to obtain a plurality of candidate business standard terms corresponding to the business terms, the method further includes:

[0091] Based on a preset fuzzy terminology judgment method, determine whether the business term is a fuzzy term;

[0092] If the above judgment is true, execute the step of "based on the preset business standard terminology map, perform terminology candidate matching on the business terms to obtain several candidate business standard terms corresponding to the business terms";

[0093] If the above judgment is not true, based on the input text, a large language model is generated based on a preset text generation method to generate the output text corresponding to the input text.

[0094] Explained, a pre-set fuzzy term judgment method is defined as a pre-set method for judging whether the business terms contained in the input text are fuzzy terms. The pre-set fuzzy term judgment method includes, but is not limited to, the following methods: comparing the business terms with the pre-set fuzzy terms to determine whether they are fuzzy terms, or comparing the business terms with the pre-set business standard terms to determine whether they are fuzzy terms.

[0095] Based on the above concept and setup, a pre-defined fuzzy term judgment method is used to determine whether a business term is a fuzzy term. If the judgment is yes, i.e., the business term is a fuzzy term, the step of "matching the business term with candidate terms based on a pre-defined business standard term map to obtain several candidate business standard terms corresponding to the business term" is executed to perform subsequent candidate term matching and context disambiguation to achieve term-level semantic alignment. If the judgment is no, i.e., the business term is not a fuzzy term, the output text corresponding to the input text is generated directly based on the input text and a pre-defined text generation large language model, without performing subsequent candidate term matching and context disambiguation to achieve term-level semantic alignment. As mentioned above, the introduction of a pre-judgment of "whether the business term is a fuzzy term" achieves a balance between text accuracy and text generation efficiency in text generation, which can further improve the timeliness of text generation.

[0096] In this embodiment of the invention, before performing terminology-level semantic alignment on business terms, it first determines whether a business term is an ambiguous term, and only if the above determination is yes, executes the step of "based on a preset business standard terminology map, performing terminology candidate matching on the business term to obtain several candidate business standard terms corresponding to the business term". This can improve the accuracy and standardization of the output and response of large model language text generation in specific domain business scenarios, while taking into account both text accuracy and text generation efficiency, thereby further improving the timeliness of text generation.

[0097] Please see Figure 3 , Figure 3 This is a schematic diagram of the second sub-process of the text generation and processing method based on term alignment provided in an embodiment of the present invention. For example... Figure 3 As shown, in this embodiment, determining whether the business term is a fuzzy term based on a preset fuzzy term determination method includes:

[0098] S31. Determine the preset business fuzzy term map and the preset business fuzzy term nodes contained therein, wherein the preset business fuzzy term nodes correspond to preset business fuzzy terms.

[0099] S32. Determine the business term vector corresponding to the business term, and determine the business fuzzy term vector corresponding to the preset business fuzzy term;

[0100] S33. Calculate the vector similarity between the business term vector and the business fuzzy term vector to obtain the second vector similarity.

[0101] S34. Determine whether the similarity of the second vector is greater than or equal to the preset similarity threshold of the second vector;

[0102] S35. If the above judgment is correct, the business term is determined to be a vague term.

[0103] S36. If the above judgment is negative, the business term is determined to be a non-vague term.

[0104] Explained, similar to the aforementioned preset business standard terminology map and its corresponding preset business standard terminology nodes, a preset business fuzzy terminology map and its contained preset business fuzzy terminology nodes are also set up. That is, a preset business fuzzy terminology map and its corresponding preset business fuzzy terminology nodes are set up, with preset business fuzzy terminology nodes corresponding to preset business fuzzy terms. The preset business fuzzy terminology map represents a map constructed based on preset business fuzzy terms, and the preset business fuzzy terminology nodes represent nodes in the aforementioned map corresponding to the preset business fuzzy terms. Business fuzzy terms refer to terms related to business that are unclear or ambiguous and whose meanings cannot be clearly and accurately understood. For example, in the financial field, words such as "insurance policy," "interest rate," and "service" have multiple meanings and are ambiguous, thus constituting financial business fuzzy terms. Among them, the nodes corresponding to the preset business fuzzy terms may include preset business fuzzy term polysemy tags. Polysemy tags are a labeling mechanism designed to solve the problem of polysemy. The design of polysemy tags includes, but is not limited to, at least one of the following: creating different tag IDs for each polysemous word; defining clear semantic boundaries for each tag; establishing hierarchical or relational relationships between tags; and recording typical contexts in which different meanings appear, thereby realizing the construction of a preset business fuzzy term graph by combining graph lexicon matching with polysemy tags.

[0105] Based on the above concept and setup, a preset business fuzzy terminology map and its contained preset business fuzzy terminology nodes are determined, with each preset business fuzzy terminology node corresponding to a preset business fuzzy term. Then, the business terminology vector corresponding to each business term is determined, and the business fuzzy terminology vector corresponding to each preset business fuzzy term is also determined. The business terminology vector and the business fuzzy terminology vector are then compared to calculate a second vector similarity. It is then determined whether the second vector similarity is greater than or equal to a preset second vector similarity threshold, where the preset second vector similarity threshold represents the critical value for whether a business term is a fuzzy term. If the corresponding vector similarity threshold is greater than or equal to the preset second vector similarity threshold, it is determined to be a fuzzy term. Therefore, if the above determination is yes, the business term is determined to be a fuzzy term; if the above determination is no, the business term is determined not to be a fuzzy term.

[0106] In this embodiment of the invention, by using graph-based vocabulary matching, it is possible to quickly determine whether business terms are fuzzy terms. This allows for the improvement of the accuracy and standardization of output and response in specific domain business scenarios for large-scale model language text generation, while also taking into account text accuracy and text generation efficiency, thereby further improving the timeliness of text generation.

[0107] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0108] In one embodiment, a term-alignment-based text generation processing apparatus is provided, which corresponds one-to-one with the term-alignment-based text generation processing method described in the above embodiments. Please refer to... Figure 4 , Figure 4 This is a schematic block diagram of a text generation and processing apparatus based on term alignment provided in an embodiment of the present invention. Figure 4 As shown, the terminology alignment-based text generation processing device 40 includes a first determining module 41, a first recognizing module 42, a first matching module 43, a context disambiguation module 44, a first generating module 45, and a second generating module 46. The detailed descriptions of each functional module are as follows: The first determining module 41 is used to respond to a text generation command and determine the input text corresponding to the user input; the first recognizing module 42 is used to recognize the business terms contained in the input text; the first matching module 43 is used to perform terminology candidate matching on the business terms based on a preset business standard terminology map to obtain several candidate business standard terms corresponding to the business terms; the context disambiguation module 44 is used to perform context disambiguation on the candidate business standard terms based on the input text and the candidate business standard terms to obtain the target business standard term with the highest text context fit corresponding to the input text; the first generating module 45 is used to generate target prompt words based on the input text and the target business standard terms; the second generating module 46 is used to generate the output text corresponding to the input text based on the target prompt words and a preset text generation large language model.

[0109] In one embodiment, the preset business standard terminology map includes preset business standard terminology nodes, and the preset business standard terminology nodes correspond to preset business standard terms; the first matching module 43 includes: a first determining submodule, used to determine the business terminology vector corresponding to the business term, and to determine the business standard terminology vector corresponding to the preset business standard term; a first similarity calculation submodule, used to perform vector similarity calculation between the business terminology vector and the business standard terminology vector to obtain a first vector similarity; and a sorting submodule, used to sort several first vector similarities in descending order, and to determine the top several preset business standard terms most similar to the business terminology vector as candidate business standard terms, thereby obtaining several candidate business standard terms corresponding to the business term.

[0110] In one embodiment, the first determining submodule includes: a second determining submodule, configured to determine a preset structured standard term tag corresponding to the preset business standard term node, wherein the preset structured standard term tag includes the preset business standard term and its corresponding preset standard term embedding vector; and a third determining submodule, configured to use the preset standard term embedding vector as the business standard term vector corresponding to the preset business standard term.

[0111] In one embodiment, the first matching module 43 further includes: a first judgment submodule, configured to judge whether the first vector similarity is equal to or less than a preset first vector similarity threshold; and a first execution submodule, configured to, if the above judgment is yes, determine that the business term is a fuzzy term and execute the step of "sorting the similarities of several first vectors in descending order".

[0112] In one embodiment, the context disambiguation module 44 includes: a fourth determining submodule, used to determine the semantic code of the input text corresponding to the input text, and to determine the semantic code of the candidate standard term corresponding to the candidate business standard term; a scoring calculation submodule, used to calculate the semantic matching score between each pair of the input text semantic code and the candidate standard term semantic code based on the input text semantic code and the candidate standard term semantic code, and based on a preset context-aware discriminator; and a fifth determining submodule, used to take the candidate business standard term with the highest semantic matching score as the target business standard term, thereby obtaining the target business standard term with the highest text context fit corresponding to the input text.

[0113] In one embodiment, the text generation processing device 40 further includes: a first judgment module, configured to determine whether the business term is a fuzzy term based on a preset fuzzy term judgment method; and a first execution module, configured to, if the above judgment is yes, execute the step of "matching the business term with term candidates based on a preset business standard term map to obtain a number of candidate business standard terms corresponding to the business term".

[0114] In one embodiment, the first determination module includes: a sixth determination submodule, configured to determine a preset business fuzzy term map and the preset business fuzzy term nodes contained therein, wherein the preset business fuzzy term nodes correspond to preset business fuzzy terms; a seventh determination submodule, configured to determine the business term vector corresponding to the business term, and determine the business fuzzy term vector corresponding to the preset business fuzzy term; a second similarity calculation submodule, configured to perform vector similarity calculation between the business term vector and the business fuzzy term vector to obtain a second vector similarity; a second determination submodule, configured to determine whether the second vector similarity is greater than or equal to a preset second vector similarity threshold; and a first determination submodule, configured to determine that the business term is a fuzzy term if the above determination is yes.

[0115] This invention provides a text generation processing device based on terminology alignment. By performing terminology candidate matching and context disambiguation on the business terms contained in the user's input text, the device obtains the target business standard terms with the highest text context fit corresponding to the input text, thereby achieving terminology-level semantic alignment of the user text. Based on the input text and the target business standard terms, the device generates target prompt words, and then generates the output text corresponding to the input text based on a preset text generation large language model. This achieves semantic control and semantic enhancement of text generation based on terminology awareness, which can improve the controllability, accuracy, and standardization of the output and response of the large model language text generation in specific domain business scenarios. It can be applied to fields including but not limited to finance, insurance, and healthcare.

[0116] Specific limitations regarding the term-alignment-based text generation processing apparatus can be found in the limitations of the term-alignment-based text generation processing method described above, and will not be repeated here. Each module in the aforementioned term-alignment-based text generation processing apparatus can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0117] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a term-alignment-based text generation processing method on the server side.

[0118] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a term-alignment-based text generation processing method.

[0119] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the text generation processing method described in the above embodiments.

[0120] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the text generation processing method described in the above embodiments.

[0121] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0123] The software tools or components not belonging to our company that appear in the embodiments of this invention are merely illustrative examples and do not represent actual use.

[0124] The data collection in this embodiment of the invention complies with the requirements of relevant laws and regulations, such as China's Personal Information Protection Law, GDPR (General Data Protection Regulation of the European Union), or information security standards of other countries and regions.

[0125] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A text generation and processing method based on term alignment, characterized in that, include: Respond to text generation instructions to determine the input text corresponding to the user input; Identify the business terms contained in the input text; Based on a preset business standard terminology map, the business terms are matched with candidate terms to obtain several candidate business standard terms corresponding to the business terms. Based on the input text and the candidate business standard terms, the candidate business standard terms are disambiguated in context to obtain the target business standard terms that have the highest text context fit with the input text. Based on the input text and the target business standard terms, generate target prompt words; Based on the target prompt words and a preset text generation large language model, the output text corresponding to the input text is generated.

2. The text generation processing method based on term alignment as described in claim 1, characterized in that, The preset business standard terminology map includes preset business standard terminology nodes, and the preset business standard terminology nodes correspond to preset business standard terms. Based on a preset business standard terminology map, the business terms are matched with candidate terms to obtain several candidate business standard terms corresponding to the business terms, including: Determine the business term vector corresponding to the business term, and determine the business standard term vector corresponding to the preset business standard term; The first vector similarity is obtained by calculating the vector similarity between the business term vector and the business standard term vector. The similarity scores of several first vectors are sorted in descending order, and the top several preset business standard terms that are most similar to the business term vector are determined as candidate business standard terms, thus obtaining several candidate business standard terms corresponding to the business term.

3. The text generation processing method based on term alignment as described in claim 2, characterized in that, Determining the business standard term vector corresponding to the preset business standard terminology includes: Determine the preset structured standard term tag corresponding to the preset business standard term node, wherein the preset structured standard term tag includes the preset business standard term and its corresponding preset standard term embedding vector; The preset standard term embedding vector is used as the business standard term vector corresponding to the preset business standard term.

4. The text generation processing method based on term alignment as described in claim 2 or 3, characterized in that, Before sorting the similarity scores of several of the first vectors in descending order, the process also includes: Determine whether the similarity of the first vector is equal to or less than a preset first vector similarity threshold; If the above judgment is true, the business term is determined to be a fuzzy term, and the step of "sorting the similarity of several first vectors in descending order" is executed.

5. The text generation and processing method based on term alignment as described in any one of claims 1-3, characterized in that, Based on the input text and the candidate business standard terms, the candidate business standard terms are disambiguated to obtain the target business standard terms with the highest text context fit corresponding to the input text, including: Determine the semantic encoding of the input text corresponding to the input text, and determine the semantic encoding of the candidate standard terms corresponding to the candidate business standard terms; Based on the semantic encoding of the input text and the semantic encoding of the candidate standard terms, and using a preset context-aware discriminator, calculate the semantic matching score between each pair of the semantic encoding of the input text and the semantic encoding of the candidate standard terms; The candidate business standard term with the highest semantic matching score is taken as the target business standard term, thus obtaining the target business standard term with the highest text context fit to the input text.

6. The text generation processing method based on term alignment as described in any one of claims 1-3, characterized in that, Before performing terminology candidate matching on the business terms based on a preset business standard terminology map to obtain several candidate business standard terms corresponding to the business terms, the method further includes: Based on a preset fuzzy terminology judgment method, determine whether the business term is a fuzzy term; If the above judgment is true, execute the step of "based on the preset business standard terminology map, perform terminology candidate matching on the business terms to obtain several candidate business standard terms corresponding to the business terms".

7. The text generation processing method based on term alignment as described in claim 6, characterized in that, Based on a preset fuzzy term determination method, the determination of whether the business term is a fuzzy term includes: Determine the preset business fuzzy term map and the preset business fuzzy term nodes contained therein, wherein the preset business fuzzy term nodes correspond to preset business fuzzy terms; Determine the business term vector corresponding to the business term, and determine the business fuzzy term vector corresponding to the preset business fuzzy term; The second vector similarity is obtained by calculating the vector similarity between the business term vector and the business fuzzy term vector. Determine whether the similarity of the second vector is greater than or equal to a preset second vector similarity threshold; If the above judgment is true, the business term is determined to be a vague term.

8. A text generation and processing apparatus based on term alignment, characterized in that, include: The first determining module is used to respond to the text generation instruction and determine the input text corresponding to the user input; The first recognition module is used to recognize the business terms contained in the input text; The first matching module is used to perform term candidate matching on the business terms based on a preset business standard terminology map to obtain a number of candidate business standard terms corresponding to the business terms. The context disambiguation module is used to perform context disambiguation on the candidate business standard terms based on the input text and the candidate business standard terms, so as to obtain the target business standard terms with the highest text context matching degree corresponding to the input text. The first generation module is used to generate target prompt words based on the input text and the target business standard terms; The second generation module is used to generate the output text corresponding to the input text based on the target prompt word and a large language model generated based on the preset text.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the term alignment-based text generation processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the term alignment-based text generation processing method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Health term ambiguity resolution and standardization method and system based on multi-modal comparative learning and context perception

    CN121766304A

  • Health term disambiguation and standardization method and system based on multi-modal contrast learning and context perception

    CN121766304B