Training data generation and model training method, product and device, electronic equipment and storage medium

By extracting the inventive concepts of the claims from the patent publication text as training data and adjusting the language model parameters, the problem of inconsistency in the patent text generated by the language model was solved, thus improving the quality and accuracy of the generated patent text.

CN121920568APending Publication Date: 2026-04-24BEIJING MENGZHIWANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING MENGZHIWANG TECH CO LTD
Filing Date
2025-11-20
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing language models struggle to strictly adhere to the technical solutions outlined in the technical disclosure documents when generating patent texts, resulting in inconsistencies between the output patent text content and actual requirements, a large amount of fabricated content, and poor quality.

Method used

By identifying inventive concepts from the claims text in patent publications and using them as training data for a language model, and combining the claims text as output labels, the model parameters are adjusted to improve the quality of patent text generation.

Benefits of technology

It improves the accuracy and quality of patent texts generated by the language model, ensures that the output patent texts conform to the technical solutions in the technical disclosure, reduces redundant content, and enhances the model's ability to draft patent texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920568A_ABST
    Figure CN121920568A_ABST
Patent Text Reader

Abstract

The invention provides a training data generation and model training method, a product, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. Based on the training data generation method provided by the invention, the invention conception for writing each claim can be identified from the patent disclosure text, so that the patent disclosure text can be used as alternative training data of a patent disclosure book during language model training of patent document writing; the technical disclosure book-patent text mapping relation missing caused by technical disclosure data confidentiality is avoided, and then the language model has the capacity of writing patent texts (especially claims) based on patent disclosure books. Besides, the specification text of the patent text is analyzed according to the technical problem and the technical means corresponding to the technical problem for all the claims, so that the core invention concept of all the claims is extracted, the influence of redundant texts in the specification text on understanding of the technical scheme is reduced, and the quality of the model generation claims is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, specifically to a training data generation and model training method, product, electronic device, and storage medium. Background Technology

[0002] With the development of artificial intelligence, large language models (LMs, also known as generative large models, hereinafter referred to as language models), represented by OpenAI, DeepSeek, and Qianwen, have demonstrated powerful capabilities in the field of text processing, thus assisting users in completing text processing tasks such as article writing and plot creation under their control.

[0003] Generating patent application texts is essentially a textual processing task. However, due to the inherent complexity of the patent application text logic, when using a language model to directly generate patent texts (i.e., application texts) based on the technical disclosure of the patent document, a text that is similar in form to the requirements of the patent text may be obtained, but whose content cannot strictly follow the technical solutions recorded in the relevant laws and technical disclosures, and often includes fabricated content.

[0004] Therefore, how to improve the quality of model training data and generate high-quality patents is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, embodiments of this application provide a training data generation and model training method, product, electronic device and storage medium, aiming to solve the problem of poor quality of current model training data, thereby improving the quality of patent text generated by the model.

[0006] In a first aspect, this application provides a training data generation method, comprising: acquiring a patent publication text, wherein the patent publication text includes a specification text and at least one claim text, the specification text being used to explain each claim text recorded in the at least one claim text; identifying the inventive concept of each claim text in the at least one claim text from the specification text, wherein the inventive concept corresponds one-to-one with the claim text, the inventive concept including a technical means text and a technical problem text of the corresponding claim text, the technical problem text recording the technical problem solved and / or the technical effect achieved by the corresponding claim text; determining at least one set of claim concept pairs in the training data of the patent publication text based on the at least one claim text and its corresponding inventive concept, wherein the claim concept pair includes one claim text and its corresponding inventive concept in the at least one claim text; when the training data of the patent publication text is used to train the claim generation capability of a language model, the claim concept pair is configured as the claim training data of the language model, the inventive concept in the claim concept pair is configured as the sample data of the language model, and the claim text in the claim concept pair is configured as the model output label of the language model.

[0007] Secondly, this application provides a training data generation apparatus, comprising: an acquisition module for acquiring a patent disclosure text, wherein the patent disclosure text includes a specification text and at least one claim text, the specification text being used to interpret each claim text recorded in the at least one claim text; and an identification module for identifying the inventive concept of each claim text in the at least one claim text from the specification text, wherein the inventive concept corresponds one-to-one with the claim text, and the inventive concept includes a technical means text and a technical problem text corresponding to the claim text, the technical problem text recording the technical solution provided by the corresponding claim text. The processing module is configured to determine at least one set of claim concept pairs in the training data of the patent disclosure text based on the at least one claim text and its corresponding inventive concept, wherein the claim concept pair includes one claim text and its corresponding inventive concept in the at least one claim text; when the training data of the patent disclosure text is used to train the claim generation capability of the language model, the claim concept pair is configured as the claim training data of the language model, the inventive concept in the claim concept pair is configured as the sample data of the language model, and the claim text in the claim concept pair is configured as the model output label of the language model.

[0008] Thirdly, this application provides a training method for a patent text generation model, comprising: determining training data of patent disclosure text determined based on the training data generation method described in the first aspect, and configuring at least one set of claim idea pairs in the training data as claim training data for the language model; for a target claim idea pair in the at least one set of claim idea pairs, inputting the inventive concept in the target claim idea pair into the language model to be trained, and instructing the language model to be trained to generate claims, thereby determining the claim text output of the language model to be trained; and adjusting the model parameters of the language model to be trained based on the difference between the claim text output and the claim text in the target claim idea pair.

[0009] Fourthly, this application provides a training apparatus for a patent text generation model, comprising: a determining module, configured to determine training data of patent disclosure text determined based on the training data generation method described in the first aspect, and configure at least one set of claim idea pairs in the training data as claim training data of the language model; an input module, configured to input the inventive concept of the target claim idea pair into the language model to be trained for a target claim idea pair among the at least one set of claim idea pairs, and instruct the language model to be trained to generate claims, thereby determining the claim text output of the language model to be trained; and an adjusting module, configured to adjust the model parameters of the language model to be trained based on the differences between the claim text output and the claim text in the target claim idea pair.

[0010] Fifthly, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the training data generation method of the first aspect or the training method of the patented text generation model of the third aspect.

[0011] In a sixth aspect, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, it implements the training data generation method of the first aspect or the training method of the patented text generation model of the third aspect.

[0012] In a seventh aspect, this application provides a computer program product on which a computer program is stored, and when the computer program is executed by a processor, it implements the training data generation method of the first aspect or the training method of the patented text generation model of the third aspect.

[0013] This application provides a method, product, electronic device, and storage medium for generating training data and training a model. First, it obtains patent publication text and identifies inventive concepts corresponding to the claims. These inventive concepts include the technical means of the corresponding claims, the technical problem to be solved, and / or the technical effect achieved. Then, based on the claim text and inventive concepts, it determines training data for training a language model, thus broadening the channels for obtaining training data and ensuring its quality. The training data includes pairs of claim concepts, each consisting of claim text and the corresponding inventive concept. These are grouped into a pair of training data and configured as the output labels and sample data for training the model. This allows the model to generate corresponding claims based on the inventive concepts when generating claims, reducing redundant content in complex technical solutions and improving the quality of claims generated by the model.

[0014] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a training data generation method provided in an exemplary embodiment of this application.

[0018] Figure 2 This is a data diagram illustrating the model training process of a language model provided in an exemplary embodiment of this application.

[0019] Figure 3 This is an exemplary flowchart of a model training method for shaping model thinking ability provided by an exemplary embodiment of this application.

[0020] Figure 4 This is an exemplary flowchart of a method for determining the interpretable text of the claims provided in an exemplary embodiment of this application.

[0021] Figure 5 This is an exemplary flowchart of an inventive concept extraction method based on a directed graph of technology provided in an exemplary embodiment of this application.

[0022] Figure 6 This is an exemplary flowchart of a construction technique for splitting data pairs provided in an exemplary embodiment of this application.

[0023] Figure 7 This is an exemplary flowchart of a decoupling method in training data generation provided in an exemplary embodiment of this application.

[0024] Figure 8 This is an exemplary flowchart of a decoupling method during training provided in an exemplary embodiment of this application.

[0025] Figure 9 This is a block diagram of a patent text generation electronic device provided in an exemplary embodiment of this application. Detailed Implementation

[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0027] Application Overview: As mentioned in the background, the core of LLM training lies in training with massive amounts of text data. That is, the model can learn the syntax, semantics, and contextual relationships of a language to acquire the ability to understand, generate, and predict natural language. Specifically, LLM training often relies on publicly available materials, such as manually annotating relevant technical content in complex texts, which significantly increases the cost of obtaining training data of a certain quality for LLM models.

[0028] On the other hand, assuming that the technical disclosure document and its processing and analysis are not publicly disclosed (generally kept as a trade secret), the mapping relationship between the technical disclosure document and the patent text cannot be directly used as training data for LLM parsing. Consequently, after inputting the technical disclosure document, the LLM can only superficially understand it, and its output text content cannot strictly correspond to the technical disclosure document, containing a large amount of model illusion. In particular, when the technical disclosure document is directly input into the LLM model, the technical solution reflected in the patent text directly output by the LLM model may not be completely consistent with the technical disclosure document. For example, when a technical solution is input into the LLM, its output of the sole right may correspond to the technical disclosure document, while subsequent rights may be fabricated content by the LLM model.

[0029] To address the aforementioned issues, embodiments of this application provide a training data generation and model training method, product, electronic device, and storage medium. The training data generation method provided in this application can identify the inventive concepts for drafting each claim from the patent disclosure text, thus serving as alternative training data for the patent disclosure document during language model training for patent document drafting. This avoids the loss of the "technology disclosure document - patent text" mapping relationship due to the confidentiality of technology disclosure materials, thereby enabling the language model to draft patent text (especially claims) based on the patent disclosure document. Furthermore, this application analyzes the specification text of the patent text for each claim using the "technical problem and its corresponding technical means" approach to extract the core inventive concepts of each claim, reducing the impact of redundant text in the specification text on the understanding of the technical solution and improving the quality of claims generated by the model.

[0030] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0031] Exemplary training data generation method: Figure 1 This is a flowchart illustrating a training data generation method provided in an exemplary embodiment of this application. That is, based on... Figure 1 The process P100 shown can generate training data for training a language model. In practical applications, the computing devices used to generate the training data and the computing devices used for training can be executed by different computing devices, or by the same device, depending on their actual computing power.

[0032] As an example only, the aforementioned P100 can be executed by a computing device configured with a high-performance language model.

[0033] like Figure 1 As shown, P100 may include the following steps: S110. Obtain the patent publication text.

[0034] S120. Identify, from the text of the specification, the inventive concept of each claim text in at least one claim text.

[0035] S130. Determine at least one set of claim concept pairs in the training data of the patent disclosure text based on at least one claim text and its corresponding inventive concept.

[0036] In S110, patent publication text refers to publicly available patent documents that can be obtained by the public through the internet and other means. That is, according to patent-related regulations (such as Article 29 of the Agreement on Trade-Related Aspects of Intellectual Property Rights (TRIPS Agreement), the text of a patent application is fully disclosed. In practice, the State Intellectual Property Office of China generally publishes patent documents periodically.

[0037] Patent publications can be further configured into readable text data of the published patent documents. That is, patent publication text can be transformed from a published patent document into a readable document. The readable document records the text data of various parts of the patent document using text data.

[0038] According to the formal requirements of patent documents, the patent publication may contain a specification and at least one claim. Considering the readable format of the aforementioned patent publication, the specification text and at least one claim text contained in the patent publication can be stored separately so that step S110 can determine the specification text and at least one claim text of the patent publication when reading the patent publication.

[0039] In some embodiments, the specification text and at least one claim text can also be determined from the patent publication text based on relevant algorithms (such as regular expression algorithms, keyword segmentation, text segmentation algorithms, etc.). In this case, each claim in the at least one claim text can be read independently to subsequently determine the inventive concept of each claim.

[0040] It should be noted that, considering the structural requirements of the specification itself and the subsequent processing of the specification, the aforementioned specification text can also be understood as the text content of the specific implementation section of the specification, in order to determine the interpretation of each claim in the specification.

[0041] In the aforementioned S120, the inventive concept is a text constructed by this application to reflect the core content (i.e., the inventive point) when the corresponding claim was drafted.

[0042] Specifically, considering the actual compliance process of patent document drafting, staff need to draft patent application documents based on the technical disclosure, reflecting the innovative technical means in the technical data through the claims, and then supporting and interpreting the claims through the specification. However, after the patent documents are published, the technical disclosure is not disclosed at the same time, so the specific data on which the inventors based their drafting is unknown.

[0043] Furthermore, considering that patent documents require a thorough description of the claims in the specification to support their implementation, especially for independent claims which should describe the technical problem and its effects, it can be understood that the technical information and claim-related parts of the patent document are included in the specification to support those skilled in the art in implementing the corresponding claims.

[0044] Based on the foregoing, this application creatively extracts the inventive concepts of each claim from the specification text as alternative textual content to the technical disclosure or other relevant materials, and as semantic data for model training.

[0045] Furthermore, considering that a compliant claim should be "a collection of technical means that utilize natural laws to solve a technical problem," the aforementioned inventive concept can be characterized as "technical problem + technical means." That is, the inventive concept includes a technical means text corresponding to the claim text and a technical problem text, wherein the technical problem text describes the technical problem solved and / or the technical effect achieved by the corresponding claim text, and the technical means text describes the technical means used to solve / improve the technical problem / achieve the technical effect.

[0046] In some embodiments, the aforementioned S120 can be implemented through a language model, that is, the specification text corresponding to the claim text can be input into the language model, and the language model can be instructed to extract the technical problem and technical means from the input content as the inventive concept of the claim.

[0047] In some embodiments, to avoid mutual interference among claims, the inventive concept of each claim is determined individually. That is, for the target claim text selected for inventive concept identification from at least one claim text, the corresponding specification text can be input into a language model, instructing the language model to extract the technical problem and technical means from the input content as the inventive concept of that target claim text. The inventive concept of each claim is then determined by continuously changing the target claim text.

[0048] In some embodiments, to improve recognition accuracy, the process can be further broken down into multiple model processing steps. More details regarding the high-precision extraction inventive concept can be found at [link to relevant documentation]. Figure 4 , Figure 5 And its related descriptions.

[0049] In some embodiments, considering the strong correlation between the technical problem extracted from the inventive concept and the technical means, the inventive concept can be verified after the inventive concept of each claim is extracted, so as to ensure that the extracted technical means can solve the corresponding technical problem.

[0050] That is, for the target claim text in the at least one claim text, the technical means text and technical problem text of the target claim text are logically verified to determine the logical verification result of the target claim text, so as to determine the logical verification result of each claim text in the at least one claim text, wherein the logical verification result reflects the consistency between the technical means text and the claim text and the effectiveness of the technical means text in solving the technical problem text.

[0051] For example, this logical verification process can be implemented through a language model. Generally, the technical problem and technical means can be input into the language model, which instructs the language model to judge whether the technical means can solve the technical problem and provide the solution logic of the technical means to solve the technical problem as the aforementioned logical verification result.

[0052] Based on the aforementioned logical verification results, generally only technical solutions that can solve the technical problem are considered valid inventive concepts; other inventive concepts may have extraction anomalies or original text anomalies. That is, in response to the logical verification results of the target claim text satisfying the consistency and validity requirements, a set of claim training data can be generated based on the target claim text and its inventive concept to generate at least one set of claim training data for the patent disclosure text.

[0053] In the aforementioned S130, the claim-concept pair can refer to the associated stored claim text and its inventive concept. That is, after confirming the inventive concept of the claim text, the inventive concept and its corresponding claim text can be associated and stored as a claim-concept pair.

[0054] It should be noted that each claim in a patent publication represents a set of technical means employed to address a specific technical problem. However, in actual patent publications, considering the varying quality of drafting, some claims may be supplementary, potentially failing to extract accurate inventive concepts. Therefore, while the inventive concepts obtained in the aforementioned S120 process should theoretically correspond one-to-one with the claim texts, in practice, each inventive concept corresponds to a different claim text, and some claim texts fail to extract inventive concepts. Consequently, the essential concepts constructed in S130 only pertain to the claims from which the inventive concepts can be identified.

[0055] Based on the aforementioned identified pairs of claims, these can be used as training data for the patent publication text. This training data can be used to train the claim generation capability of the language model. Specifically, when the training data of the patent publication text is used to train the claim generation capability of the language model, the pairs of claims are configured as the claim training data of the language model, the inventive concepts within the pairs of claims are configured as the sample data of the language model, and the claim texts within the pairs of claims are configured as the model output labels of the language model. The training logic of the language model is similar to that of conventional model training; for details, please refer to [link to relevant documentation]. Figure 2 And its related descriptions.

[0056] In some embodiments, the training data determined based on the aforementioned S130 can be stored based on the patent publication text, that is, the training data of the same patent publication text (such as each claim concept pair) is stored based on the identification information (such as text number) of the patent publication text, so that the training data of the entire patent publication text can be called during subsequent training.

[0057] It should be noted that the language model trained in S130 is not necessarily related to the language models processed in other steps; they can be the same model or different models. In practical applications, the language models involved in other steps (such as S120) generally have a larger number of parameters for accurate data labeling, while the language model with smaller parameters can be used in S130 to meet the processing needs of specific domains.

[0058] Furthermore, considering that the patent text also includes other data that can be further processed, when constructing the training data for the patent disclosure text, other data from the patent disclosure text can be written into the training data to facilitate the training of the language model. For details regarding other data in the training data, please refer to the following description (e.g., Figure 2 (Related description).

[0059] Therefore, the training data generation method provided in this application can identify the inventive concepts for drafting each claim from the patent disclosure text. This data can then serve as alternative training data for the patent disclosure document when training the language model for patent document drafting. This avoids the loss of the "technology disclosure document-patent text" mapping relationship due to the confidentiality of technology disclosure materials, thereby enabling the language model to draft patent texts (especially claims) based on the patent disclosure document. Furthermore, this application analyzes the specification text of the patent document for each claim using the principle of "technical problem and its corresponding technical means" to extract the core inventive concepts of each claim. This reduces the impact of redundant text in the specification text on the understanding of the technical solution and improves the quality of the claims generated by the model.

[0060] In some embodiments, the training data obtained from the aforementioned data processing of patent publication texts are further organized to construct specialized models for patents in specific fields. That is, when determining the training data for patent publication texts, type information of the patent publication texts is generated based on other public information of the patent document texts (such as applicant, subject, and receiving office). The type information indicates the technical field, legal basis (such as which country's patent law it is based on), text language, and subject type of the patent publication text.

[0061] Type information can be used as classification labels for training data, allowing for separate training of different types of patents to provide specialized models when needed.

[0062] Exemplary model training process: To further illustrate the training process of the language model, this application also provides a data diagram illustrating the training process of the language model. Figure 2 ).

[0063] like Figure 2 As shown, to distinguish between the language models before and after training, the language model during training can be denoted as the language model to be trained 221, and the language model after training can be denoted as the trained language model 222. The trained language model 222 is based on the trained language model to be trained 221. Considering the current basic training methods for language models, then... Figure 2 In China, the training of language models 221 is often not a zero-based training, but a further training (also known as fine-tuning) of a language model with certain processing capabilities, so that it can adapt to the processing work of specific scenarios.

[0064] The training process of the language model 221 to be trained is similar to that of general artificial intelligence training. It requires collecting samples and labeling the samples, and adjusting the parameters of the language model 221 to be trained based on the difference between the model output based on the samples and the sample labels.

[0065] In some embodiments, the fine-tuning process of the language model (such as difference calculation and parameter iteration algorithms) can be implemented based on existing fine-tuning algorithms or platforms (such as LLMTune, LLaMA Factory, etc.). Further details will not be provided here.

[0066] Specifically, during the model training process, the training data for the language model 221 to be trained can be determined first, and then the language model 221 to be trained can be trained based on the training data. That is, the training data 210 of multiple patent disclosure texts determined by the aforementioned training data generation method can be determined first, and at least one set of claim concept pairs in the training data can be configured as the claim training data of the language model. In the specific training process, for the target claim concept pair in the at least one set of claim concept pairs, the inventive concept 211 in the target claim concept pair is input into the language model 221 to be trained, and the language model to be trained is instructed to generate claims, thereby determining the claim text output 230 of the language model to be trained. Then, the model parameters of the language model 221 to be trained are adjusted based on the difference between the claim text output 230 and the claim text 212 in the target claim concept pair. Then, the above specific training process is repeated, and after the training is completed, the language model 221 to be trained is configured as a trained language model 222, so that the trained language model 222 can generate claim output 242 based on the inventive concept input 241. The inventive concept input 241 and the claim output 242 are generally data involved in the actual application of the trained language model 222, and not the data in the aforementioned training data.

[0067] In some fine-tuning platforms, the difference between the model output and the sample labels can be automatically determined by the algorithm provided by the platform, or it can be calibrated by the user. To further improve the accuracy of model training, the differences in the aforementioned fine-tuning process can be addressed through claim text scoring. Specifically, the first weight score of claim text 212 and the second weight score of weight text output 230 can be determined first. Then, the model parameters of the language model to be trained are adjusted based on the difference between the first and second weight scores and the difference between the weight text output and the claim text.

[0068] In some embodiments, the aforementioned weighted score can be evaluated by someone skilled in the art. In some embodiments, the aforementioned weighted score can also be evaluated by a language model. For example, the weighted score rules are input into the language model as injected information, and the language model is instructed to score based on the scoring rules. In some embodiments, the aforementioned weighted score can also be determined based on a machine learning model, that is, the score can be labeled by someone skilled in the art, and the mapping relationship between the score and the weighted text features can be recognized through machine learning, thereby outputting the score.

[0069] In some embodiments, similar to the aforementioned training process, the language model 221 to be trained can also be trained based on other training data from the patent disclosure text to further improve its processing capabilities in other stages of the patent generation process. For example, the language model 221 to be trained can be fine-tuned based on the "claim-specification pair" to enhance its ability to generate specification text based on the claim text. As another example, a correspondence can be established between the overall technical solution summarizing the patent disclosure text and the inventive concept corresponding to the claim to train the language model 221's ability to decompose technical solutions. More information on technical solution decomposition can be found in [link to relevant documentation]. Figure 6 And its related descriptions.

[0070] In some embodiments, considering that some claims in the patent publication (generally dependent claims) are often based on the claims they reference, training / generation directly based on their inventive concept may result in significant differences between the model's output claims and the labels. To overcome this, the claims in the patent publication can be decoupled, allowing the prior claims of the referenced claims to be input into the language model as injection information during training / generation, thus configuring each claim as an independent training process. Further description of inventive concept decoupling (or claim decoupling) can be found in [link to relevant documentation]. Figure 7 , 8 And its related descriptions.

[0071] With the development of language models, current text-based language models generally possess a thought chain function. This means that in the model's output, part of the content represents thought content (also known as model inference content), while the other part is the actual final output. In principle, these two parts are actually presented using tag information (generally...) <think>(This refers to two pieces of text information that are distinguished.)

[0072] To enable the aforementioned process to shape the language model's thinking ability, the model's thinking ability can also be trained during the training process. To further illustrate this process, this application also provides an exemplary flowchart of a model training method for shaping the model's thinking ability (…). Figure 3 ).

[0073] like Figure 3 As shown, the model training process P300, which involves model thinking ability, may include the following steps: S310. Determine the training data of the patent disclosure text, and configure at least one set of authoritative ideas in the training data as authoritative training data for the language model.

[0074] S320. Input the inventive concept of the target claim into the language model to be trained, and instruct the language model to be trained to generate claims, thereby determining the thinking field output and claim text output of the language model to be trained.

[0075] S330. Adjust the model parameters of the language model to be trained based on the differences between the output of the claim text and the claim text in the target claim concept pair.

[0076] S340. Adjust the model parameters of the language model to be trained based on the differences between the output of the thinking field and the writing of thinking text.

[0077] In the aforementioned S310, with Figure 2 The training process is similar. However, considering that P300 involves training the model's thinking process, the training data provided in S310 should also include drafting thought texts for training the model's thinking ability. These drafting thought texts can reflect the use of claim drafting rules in the process of generating corresponding claim texts from corresponding inventive concepts. Claim drafting rules can be pre-configured normative documents reflecting the claim drafting logic (claim type classification and corresponding drafting sentence structures) or drafting requirements (such as drafting grammar).

[0078] In some embodiments, the aforementioned writing reflection text may be generated during the generation of training data. That is, for a target hierarchy of ideas in at least one set of hierarchy of ideas, writing reflection texts for the target hierarchy of ideas can be generated based on hierarchy writing rules to determine the writing reflection texts for each hierarchy of ideas in at least one set of hierarchy of ideas.

[0079] In some embodiments, the aforementioned drafting thought text can be manually annotated by relevant personnel. That is, the "Rules for Drafting Official Documents" can be provided to staff so that they can apply and manually annotate the text based on the "Rules for Drafting Official Documents".

[0080] In some embodiments, the aforementioned writing thought text can also be automatically generated by a language model. That is, the claim writing rules and the target claim concept can be used as input information to the target language model, instructing the target language model to generate the thought process of generating the corresponding claim text from the corresponding inventive concept based on the claim writing rules and configure it as the writing thought text.

[0081] It should be noted that the thought process output by the aforementioned language model may deviate to some extent from the final claim text. This part can be modified by inputting more information so that the language model can be actively modified or by staff to modify it manually.

[0082] As mentioned above Figure 2 Unlike the previous S320, the model output was further segmented, specifically represented as the thinking field output and the weighted text output. The weighted text output differs from the aforementioned... Figure 2 The content shown is consistent, and the relevant discussion objects of the thinking ability of the aforementioned model can be represented as the text content within the thinking label (generally presented as...). <think> and< / think> The text content between the text fields, i.e., the output of the thinking field is based on thinking labels.

[0083] The subsequent model parameter adjustment process for S330 and S340 can be performed based on existing algorithms (e.g., by existing fine-tuning platforms). Traditional language models without thinking ability often output empty thinking fields when executing S340, but can output the content of that field after adjustment based on feedback from S340.

[0084] Therefore, based on the aforementioned training method, the language model can be equipped with the ability to generate claims in the patent field based on the determined inventive concept. Furthermore, the aforementioned training method can also shape the language model's thinking ability when generating claims, enabling the language model to fully apply the "Consider the Drafting Rules," thereby further improving the language model's claim generation capability.

[0085] Exemplary method for extracting inventive concepts: In the process of identifying the aforementioned inventive concept, this application found that if simple identification is performed based solely on a language model, it may be impossible to extract the accurate inventive concept from the specification text in some scenarios.

[0086] Specifically, this application conducted multiple tests on the actual identification process and found that the extraction of the aforementioned inventive concepts mainly suffers from two problems: the accuracy of the correspondence between the claims and specifications, and the accuracy of the identification of the concept content. These two phenomena and their solutions will be discussed in detail below.

[0087] The accuracy of the correspondence between the claims and the specification refers to the fact that when extracting inventive concepts, the extracted specification text should be the text that interprets the corresponding claims. If the correspondence between the two is inaccurate, the language model may extract incorrect inventive concepts from an incorrectly or incompletely matched specification text.

[0088] In practice, to improve the accuracy of the correspondence between authority items in the specification, staff can manually annotate them. However, this would require a significant amount of manpower.

[0089] To achieve automated and highly accurate matching of claims, a text similarity algorithm is generally used to match the description text with the claims text. Specifically, for at least one target claim in the claim text, the text similarity between each text segment in the description text and the target claim can be calculated. Then, based on the text similarity, the explanatory text of the target claim is determined from each text segment of the description text, thus determining the explanatory text of each claim. Subsequently, in subsequent steps, the inventive concept corresponding to the claim can be determined based on the explanatory text. The text segments in the description can be segmented according to actual needs; they can be a single paragraph or sentences separated by punctuation marks such as "." and ":".

[0090] Further testing in this application revealed that technical terms are shared in the claims text, while text similarity algorithms often rely on keyword recall for calculation. This can lead to at least one claim with a citation relationship having the same / similar keywords, making it impossible to distinguish the description of each claim in the specification.

[0091] To address this issue, this application further optimizes the keyword extraction logic and the target claim text selection logic to extract more accurate inventive concepts from the specification text. Specifically, during keyword extraction, keywords can be extracted from each claim text first, and then specific keywords or keyword combinations in each claim can be identified based on citation relationships, thereby optimizing the text matching results. In further matching, claims without subsequent claims can be prioritized to avoid overlapping interpretations.

[0092] To further illustrate this situation, this application also provides an exemplary flowchart of a method for determining the interpretation text of the claims ( Figure 4 ).like Figure 4 As shown, process P400 may include the following: S410. Determine the set of keywords for each claim text in at least one claim text.

[0093] S420. For the first claim text in at least one claim text, a first feature set and a second feature set are determined from the keyword set of the first claim text based on the prior claims of the first claim text, so as to determine the first feature set and the second feature set of each claim text.

[0094] S430. For the second claim text in at least one claim text, determine the text similarity between each text segment in the specification text and the second claim text based on the first feature set and the second feature set of the second claim text.

[0095] S440. Determine the explanatory text of the second claim text based on text similarity, so as to determine the explanatory text of each claim text.

[0096] The aforementioned P400 can be an iterative process. That is, after the aforementioned S440 is executed, the second claim text that has been determined as an explanatory text can be removed and the aforementioned S430 can be re-executed to iterate until the explanatory text of each claim text is determined.

[0097] In the aforementioned S410, the keyword set can refer to a text set constructed from keywords, which may include multiple keywords appearing in the corresponding claim. Keywords are generally presented as word text or phrase text containing semantic content. For example, keywords in the claim text are generally mainly noun technical terms, and also include action, form, relationship, and other modifying elements.

[0098] The keyword recognition described in S410 above generally involves extracting feature words from sentences. Considering that claim texts are often highly condensed textual expressions, traditional text-meaning-based keyword extraction methods (such as TF-IDF (combining term frequency and inverse document frequency), CSI (based on co-occurrence sparsity), ECC (based on graph model eccentricity), and TextRank (iterative propagation weights)) can be avoided when performing keyword recognition. Instead, sentence analysis can be used for keyword recognition.

[0099] This involves splitting the statements in the claims (e.g., by using semicolons ";" and periods "."), determining the sentence structure of each statement (e.g., subject-verb-object, attributive, adverbial, complement, predicate, appositive, etc.), and then extracting keywords from the sentence structure based on the expression method of the claims (as mentioned above, keywords are determined by the position of noun technical terms (e.g., subject, object) and the modifying components of the technical terms). Considering the relative simplicity of this process, it can generally be implemented using one or more models and / or algorithms with a low number of parameters (e.g., specific sentence splitting can be implemented by regularization algorithms, and sentence structure extraction and keyword recognition can be implemented by language models with a low number of parameters) to improve processing speed.

[0100] In the aforementioned S420, the text of the first claim can refer to any claim text in at least one claim text. A prior claim can refer to a claim that is cited or indirectly cited by a particular claim. For example, if claim 2 of this application cites claim 1, and claim 3 cites claim 2, then the corresponding claims 1 (indirect citation) and 2 (direct citation) are prior claims of claim 3.

[0101] Considering that in drafting claims, dependent claims often further define certain technical features of the claims they reference, and the keywords in dependent claims are often related to the prior claims and introduce new descriptions based on them, a first feature set and a second feature set can be determined in S420 above to reflect this. The first feature set can be keywords or combinations of keywords in the keyword set of the corresponding claim text that do not appear in the prior claims, and the second feature set can be keywords or combinations of keywords in the keyword set of the corresponding claim text that appear in the prior claims.

[0102] The aforementioned first and second feature sets are generally established by comparing keywords with the text of prior claims to determine newly appearing wording in the claim. Furthermore, considering that some further limitations may manifest as further limitations on existing keywords, combinations of related words can be incorporated into the first feature set. Specifically, for potentially related keyword combinations, their relationships can be defined by identifying regular or positional relationships (such as character distance between two keywords, whether they are in the same sentence, their order, etc.) (this can also be similar to the relevant grammatical rules of patent searches) to construct a phrase.

[0103] It should be noted that, considering that the determination of the aforementioned first feature set and second feature set can also be based directly on the text content of the prior claims without relying on the prior claim keyword set, the aforementioned step of identifying the keyword set can also be performed after selecting the text of the second claim, without first extracting the keyword set of each claim text.

[0104] In the aforementioned S430, the text of the second claim can be a claim that does not have a subsequent claim. A subsequent claim can refer to other claims that apply to that claim; that is, in the aforementioned example, claims 2 and 3 are subsequent claims to claim 1.

[0105] Considering that when performing text similarity recognition, the explanatory text of the later claim often affects the recognition of the explanatory text of the earlier claim, in the aforementioned S430 and S440, the explanatory text of the claim that does not have a later claim (i.e., the text of the second claim) can be identified first to avoid its influence on the earlier claim.

[0106] In the aforementioned S430, the selection of the text of the second claim can be determined based on the citation relationship between at least one claim text. When there are multiple claim texts that can serve as the text of the second claim, arbitrary selection or non-repeated selection (such as selection by serial number) can be adopted.

[0107] Text similarity refers to the degree of similarity / relevance in content between the text of the claims and the text of the specification. In conventional methods for calculating text similarity, it is necessary to compare the keywords and other relevant parameters of the two objects to be calculated (i.e., the claims and the text segments) (such as calculating parameters like cosine similarity, Jaccard similarity, and extended similarity fusion) to obtain the text similarity score.

[0108] Considering that the text segment is generally characterized as an explanatory text of the claims, its content is generally not highly condensed, and the aforementioned traditional keyword extraction method can be used for identification.

[0109] In some embodiments, the aforementioned text similarity calculation can also be based on keyword retrieval technology, that is, by recalling / detecting whether the corresponding keywords appear in the text segment, it can be determined whether the corresponding text segment is a description of the corresponding keyword / keyword combination.

[0110] The text similarity calculation based on the aforementioned first and second feature sets can output either one similarity score or two similarity scores. Regardless of whether one or two similarity scores are output, the first feature set has greater representational power than the second feature set. That is, if only one similarity score is output, the weight corresponding to the first feature set is greater than the weight corresponding to the second feature set when calculating the similarity. If two similarity scores are output, the similarity weight corresponding to the first feature set is greater than the similarity weight corresponding to the second feature set when determining relevance.

[0111] In some embodiments, it is further considered that the text segments of the specification text may contain exemplary text content. That is, a certain text segment may explain the content of the preceding text from the perspective of examples. Considering that it is a discussion at the level of examples, it may not involve the keywords of the preceding text. In order to avoid such exemplary text not being recalled, the corresponding text segment can be associated with the preceding text based on wording such as "for example" or "example" and treated as a text segment.

[0112] In the aforementioned S440, the explanatory text may refer to at least a portion of the specification text used to interpret the corresponding claim text.

[0113] It should be noted that in compliant patent publications, the specification should interpret the text of each claim to support its legal scope of protection. That is, in compliant patent publications, the interpretations of each claim should be decoupled.

[0114] For example, claim 3 is a further limitation of claim 2. In addition to the implementation shown in claim 3, the specification of this application also discusses other implementations of the limited step (such as manual annotation) and general implementation requirements (i.e., thinking about the actual representational meaning and formation requirements of the text) to avoid claim 3 being regarded as the only implementation of claim 2, thereby supporting the scope of protection of claim 2.

[0115] Therefore, the explanatory texts of each claim determined in S440 can be primarily based on the text matching degree corresponding to the first feature set, so as to ensure that the determined explanatory texts can interpret the corresponding claims as much as possible. Furthermore, considering that the explanatory texts of subsequent claims often fall within the scope of the explanatory texts of earlier claims, to further decouple the process, after determining the explanatory texts of subsequent claims, the explanatory texts of the second claim, which already has a determined explanatory text, can be removed simultaneously to avoid interfering with the identification results of the earlier claims during the determination of their explanatory texts.

[0116] The aforementioned process of determining the explanatory text is generally based on a text similarity threshold. That is, when the text similarity is greater than a preset threshold, the corresponding text segment can be determined as the explanatory text of the second claim. Furthermore, considering that the explanatory text needs to be removed after determination, when determining the explanatory text of the second claim, it is necessary to ensure that the determined explanatory text corresponds only to the second claim. Therefore, in the similarity calculation of S430, the similarity between the first feature set and the second feature set is calculated separately, thereby determining the explanatory text based on the similarity threshold of the first feature set. In addition, to ensure the comprehensiveness of the explanatory text, the explanatory text may include text segments whose similarity to the second feature set is higher than the threshold and which contain keywords from the first feature set. However, these segments do not need to be deleted from the specification text and can still be considered as candidate text segments in the process of determining the explanatory text of the prior claims.

[0117] Furthermore, considering that the instruction manual text is relatively coherent in terms of paragraph organization, that is, the explanatory text that follows is generally presented as a continuous text segment, the similarity between the preceding and following text segments can also be determined when performing the aforementioned S440, thereby determining the accurate explanatory text.

[0118] Therefore, based on the aforementioned process, the explanatory text corresponding to each claim text can be determined, so that the inventive concept can be automatically identified through the language model based on the corresponding explanatory text.

[0119] Furthermore, considering the referencing relationships between the claims, the following two situations may occur when identifying the inventive concept from the explanatory text of the claims: ① When identifying inventive concepts solely based on the explanatory text of the target claim, the technical solution may be incomplete. This is particularly evident when the target claim is a dependent claim. Specifically, if the explanatory texts of the claims referenced by the target claim (i.e., claims directly or indirectly referenced in the target claim's reference section, denoted as prior claims) are not identified together, technical concepts described in the explanatory text of the prior claims may not appear in the explanatory text of the target claim. This could prevent the language model from acquiring relevant data, thus hindering its understanding of the complete technical solution and key technical terms, resulting in the inventive concept identified by the language model deviating from the corresponding claim.

[0120] ② At the same time, when identifying inventive concepts based on the target claims and the explanatory text of the prior claims, the explanatory text of the prior claims often serves as interference information to interfere with the language model's identification of the content of the explanatory text of the target claims, which may result in the identified inventive concept actually being the inventive concept of the prior claims.

[0121] Based on the above two situations, before identifying the inventive concept after S440, this application can further improve the input content of the model so as to provide a complete technical solution while distinguishing it from existing solutions. Specifically, when identifying the inventive concept of the claims, different labels can be configured for the explanatory text of the prior concept and the explanatory text of the target claim, thereby instructing the language model to perform different processing based on the text labels.

[0122] Therefore, in the process shown on P400 above, after S440 determines the interpretive text of each claim, the following steps may also be included: S450. For a target claim text in at least one claim text, determine the explanatory text of the prior claim of the target claim text and configure it as the prior art text.

[0123] S460. Input the prior art text and the explanatory text of the target claims into the language model, and instruct the language model to identify the inventive concept from the explanatory text of the target claims to determine the inventive concept of each claim.

[0124] In the aforementioned S450, prior art text can refer to the basic technology required to implement the target claim text. Considering the claim organization of patent publications, dependent claims are generally further clarifications of prior claims, and therefore, at the implementation level, they often rely on the prior claims they reference. Thus, the technical solutions (interpretive text set) of the prior claims in the target claim text can be regarded as the prior art text for implementing the target claim text.

[0125] In some embodiments, considering that some independent claims may omit certain technical features of the prior art and only describe the differences from the prior art (also known as background art), the technical solutions that need to be supplemented in this part are often reflected in the discussion of the prior art / technical problem. Therefore, the aforementioned prior art text may also include an explanatory text of the prior art. The method for determining the explanatory text of the prior art is similar to that for the aforementioned claim text. Specifically, the discussion of the background art section in the specification can be drafted as claim text (e.g., denoted as claim 0), and this claim text is considered to be referenced by each independent claim. Thus, the explanatory text of the prior art is determined from the specification using the aforementioned method and incorporated into the aforementioned prior art text.

[0126] In some embodiments, the aforementioned explanatory text of prior claims (and explanatory text of existing technology) is configured as prior art text, which can generally be understood as storing the data using the same variable (or storing multiple variables in association), and attaching additional identification information when loading the language model. Examples include "[Prior Art Text]" and "<Prior Art Text>". Furthermore, "prior art text" here is merely a name for the tag; other names may be used in actual operation, such as prior content, known technology, existing technology, etc.

[0127] Similar to the aforementioned prior art text, the explanatory text of the aforementioned target claim is often configured with identification information when input into the language model. This allows the language model's prompts to easily distinguish between two different text contents, enabling the language model to identify the inventive concept based on the explanatory text of the target claim. For example, the identification information for the explanatory text of the aforementioned target claim could be "[Target Text]". Then, the system prompts of the language model could be: "Please identify the technical problem to be solved by the technical solution and the core technical means to solve the technical problem from [Target Text]. The relevant content of [Prior Art Text] can assist in understanding [Target Text], and [Prior Art Text] can be regarded as basic technology."

[0128] Therefore, based on the aforementioned P400, the accurate interpretation of the claim text can be determined, thereby enabling the identification of the inventive concept.

[0129] Besides the aforementioned issue of the accuracy of the correspondence between the claims and the specification, there may also be accuracy issues in the process of identifying inventive concepts. Specifically, this application has found that the expression of technical problems and technical means has a certain degree of arbitrariness, and overly broad or general extraction results may lead to the extracted inventive concepts failing to actually reflect the technical solutions.

[0130] For example, taking the sole claim of this application as an example, the actual technical problem it solves is "how to construct model training data that reflects the mapping relationship between the disclosure data and the claims when the disclosure data is not publicly available, so that the language model can learn this nonlinear relationship." Besides this specific problem, this technical problem can also be characterized as broader technical problems such as "how to improve the language model's ability to generate patents" and "how to train the language model." If the technical problem solved by the aforementioned claims is positioned as a broader technical problem, it may lead to the technical problem failing to actually reflect the technical elements that need improvement or the relationships between technical elements, thereby causing the generated claims to fail to reflect the defects of the prior art and reducing the quality of the generated claims.

[0131] To overcome this situation, this application further defines the technical problem and technical means to accurately reflect the actual defects existing in the prior art. Specifically, this application introduces relevant concepts from the core theory of modern systems engineering, "Systems Analysis," to systematically decompose the technical solution into multiple technical elements and the interaction relationships between them. The aforementioned technical problem can be characterized as defects in the technical elements themselves, defects in the interaction relationships between technical elements, or a combination thereof. The corresponding technical means should then be characterized as new technical elements or new interaction relationships within the technical solution. General methods for decomposing systems can be found in relevant textbooks, papers, and theories of "Systems Analysis," and will not be elaborated upon here.

[0132] Taking the technical problem of claim 1 of this application as an example, based on the aforementioned system analysis, its core technical problem can be characterized as "the disclosure information is a non-public resource (the defect of the technical element of disclosure information), which makes it impossible for the disclosure information to be used as training data input into the language model (the defect of the interaction between the technical element of disclosure information and the technical element of language model)".

[0133] To meet patent drafting requirements, system analysis can be performed from various analytical dimensions, including objective existence breakdown, functional module breakdown, and workflow breakdown, to construct a directed graph containing technical elements and their relationships. Objective existence can refer to real, existing objects, including but not limited to physical structures (such as various mechanical structures) and objectively existing data (such as the aforementioned text data). Functional modules often refer to a functional module built upon one or more objective existences. Workflow can be characterized as the flow of energy, information, and matter between different stages (or understood as the orderly combination of steps in a work process).

[0134] The aforementioned different splitting angles can be selected in the same patent document to construct at least one corresponding directed graph based on actual needs (wherein, a mapping relationship generally needs to exist between different types of directed graphs). Based on this directed graph, the defects of the prior art and the technical means of the claim can be determined.

[0135] In the process described above, the language model can be executed based on the relevant materials from the "System Analysis," enabling it to identify, in the aforementioned form, the technical problems reflecting the actual defects and the technical means to resolve those defects from the instruction manual. The aforementioned relevant materials from the "System Analysis" can be used as injected information, or training pairs can be constructed based on the "System Analysis" to train the model (e.g., fine-tuning).

[0136] To further illustrate this point, this application also provides an exemplary flowchart of a method for extracting inventive concepts based on directed technical graphs (PQP). Figure 5 ).

[0137] like Figure 5 As shown, when extracting the inventive concept (i.e., performing the aforementioned S120), process P500 may include the following steps: S510. For the target claim text in at least one claim text, input the specification text corresponding to the target claim text into the language model, and instruct the language model to determine at least one technical directed graph based on the specification text and the system analysis method.

[0138] S520, The indicator language model determines the technical problems and technical means in the inventive concept of the target claim text based on the directed technical graph.

[0139] In the aforementioned S510, the directed technical graph can be graph data that reflects technical solutions, meaning that technical solutions can be analyzed using graph data. Following the relevant discussion in "System Analysis," when using graph data to analyze technical solutions, the graph data generally includes nodes and edges, where edges connect nodes. To analyze technical solutions, the nodes in the directed technical graph reflect technical elements, and the edges between nodes reflect the interaction relationships between technical elements.

[0140] Following the preceding systematic analysis, when constructing a directed graph of technology, the technical elements can be divided into multiple graph data based on their type, and mappings can be established between these graph data. For example, objective existence, functional modules, workflows, etc., can be divided into multiple graph data, and mapping relationships can be constructed between these graph data (such as edges connecting nodes in the graph data to indicate the mapping relationship). In practical applications, patent publications may only involve some types of graph data; in such cases, the corresponding graph data can be constructed.

[0141] Similar to the keyword identification in the aforementioned claims, when constructing the graph data, it can be based on noun-based technical data (the steps can be regarded as a noun term as a whole). The modifying components of the noun-based technical term (such as attributives, subject-predicate structures, etc.) can be used as node attributes, the relationships between noun-based technical terms (generally predicates) can be used as edges between nodes, and the modifications / limitations of the relationship (such as adverbs, etc.) can be used as edge attributes.

[0142] Furthermore, considering that the technical solution defined in the claims may be dynamic, the aforementioned construction of the map data also constructs map data for different states and describes the state transition through the mapping relationship between states, so as to reflect the dynamic technical solution.

[0143] Therefore, based on the aforementioned S510, a directed technical graph of the corresponding technical solution can be constructed based on the specification text corresponding to the claims (such as the aforementioned explanatory text). As discussed above, the technical problem should manifest as defects in the technical elements themselves (node ​​attribute defects) and / or defects in the interaction relationships between technical elements (edge ​​defects). Therefore, when executing the aforementioned S520, the technical problem can be determined based on the aforementioned node defects and edge defects, and then the technical means (such as presented as a local directed technical graph) can be determined by improving the technical problem within the directed technical graph.

[0144] In some embodiments, considering the legal impact on the scope of protection of patent texts, some wording in patent publications (mainly reflected in the claims text) may contradict actual technical wording (e.g., using more accurate or broader terminology) and is often further defined / explained in the specification text, while the aforementioned wording is often directly used as technical elements. Directly using some wording from the patent publications as technical elements may not accurately convey the meaning. To overcome this, the wording in the claims text can be paraphrased when forming the directed technical graph.

[0145] In some embodiments, when paraphrasing target wording in the patent disclosure text, the wording interpretation of the target wording in the specification text can be determined (generally presented in a subject-verb-complement structure, i.e., A is...). Then, based on the wording interpretation of the target wording in the specification text, the general terminology of the target wording is determined. As an alternative embodiment, a technical dictionary method can also be used to determine this (i.e., technical terminology matching). Further, the corresponding technical dictionary can be invoked based on the IPC classification number or other technical classification information, and then the general terminology can be matched from the technical dictionary based on the target wording or its wording interpretation.

[0146] In some embodiments, considering that the technical problem and the technical means can be determined by comparing the technical directed graphs, when performing the aforementioned S510, the technical directed graph corresponding to the target claim text and the technical directed graph of the current technical solution that does not adopt the target claim text can be determined, so that in S520, the technical problem and its technical means can be determined based on the difference between the two.

[0147] Wherein, the current technical solution not adopted in the target claim text can be the technical solution corresponding to a claim text directly referenced by the target claim text. If the target claim text is an independent claim, it can be the technical solution corresponding to the background art.

[0148] Furthermore, the logical verification of the aforementioned technical means text and technical problem text can also be based on the aforementioned directed technical graph. That is, inferences can be made based on the interaction relationships or technical node defects corresponding to the technical problem text (such as causal chain inferences, see related Triz theories for details), thereby verifying the validity of the technical problem. Similar inferences can be made based on the interaction relationships corresponding to the technical means text, thereby determining the validity of the technical means text.

[0149] Therefore, based on the aforementioned processes P400 and P500, the inventive concept of the claim text can be accurately identified, thereby improving the accuracy of the training samples constructed.

[0150] Exemplary inventive concept decomposition method: Given that the aforementioned training data actually constitutes a mapping relationship between "inventive concept - claims", the language model has a certain processing capability for this process. In patent practice, it is often necessary to determine the text of the claims from technical documents such as technical disclosure documents, and then form the text of the specification to construct the patent disclosure text.

[0151] To further improve the language model's processing capabilities in the above process, corresponding training data can also be constructed for the mapping process from technical disclosure materials to various inventive concepts, so that the language model has the corresponding mapping capabilities.

[0152] Specifically, in this mapping process, the citation relationships of the claims in the patent publication can reflect how each inventive concept breaks down the overall technical solution.

[0153] To further describe this situation, this application also provides an exemplary flowchart for constructing technology-split data pairs ( Figure 6 ).

[0154] like Figure 6 As shown, P600 may include the following steps: S610. Determine the overall technical solution based on the specification text and / or at least one set of authoritative concepts and the various technical means texts.

[0155] S620. Based on the overall technical solution and each inventive concept, determine the technical breakdown data pairs in the training data of the patent disclosure text.

[0156] In the aforementioned S610, the overall technical solution records all the technical features of each technical means text. In some embodiments, the overall technical solution may be presented as the aforementioned directed technical graph or other forms of data (such as plain text data).

[0157] When performing the aforementioned S610, the content of the patent disclosure text can be summarized using a language model to determine the solution. For example, a technical summary can be made directly from the specification text, or the technical solutions protected by each claim can be summarized. For example, the directed technical graphs corresponding to each claim can be pieced together to construct a directed technical graph of a complete technical solution.

[0158] In particular, when there are multiple alternatives, matching can be achieved by expanding the dimensions of the directed technology graph. That is, different directed technology graphs can be constructed based on different alternatives.

[0159] The aforementioned S620 can directly call the determined inventive concepts so that each inventive concept is associated with and stored as a technical breakdown data pair with the overall technical solution. That is, the technical breakdown data pair can include the overall technical solution and each inventive concept.

[0160] Based on the technology segmentation data pairs, the training data of the patent publication text can be used to train the technology solution segmentation capability of the language model. In this case, the technology segmentation data pairs are configured as the solution segmentation training data of the language model, the overall technical solution in the technology segmentation data pairs is configured as the sample data of the language model, and the individual inventive ideas in the technology segmentation data pairs are configured as the model output labels of the language model.

[0161] Furthermore, considering that reference relationships can exist when a technical solution is broken down, these reference relationships can be supplemented when executing the aforementioned S620, thereby achieving the following sub-steps: S621. Determine the reference relationship of at least one claim text.

[0162] S622. Determine the dependency relationship of the inventive concept corresponding to each claim text based on the reference relationship of at least one claim text.

[0163] S623. Based on the overall technical solution, the various inventive concepts and their dependencies, determine the technical breakdown data pairs in the training data of the patent disclosure text.

[0164] Therefore, when the training data of the patent disclosure text is used to train the language model's ability to decompose technical solutions, the dependencies between the various inventive concepts in the technical decomposition data pairs are also configured as the model's output labels. In this case, the technical solutions of the overall solution can be hierarchically decomposed based on the reference relationships.

[0165] The aforementioned training process based on splitting data pairs using technology is similar to the conventional training process, and can be followed as follows: First, identify the technical breakdown data pairs in the training data of the patent disclosure text. The technical breakdown data pairs include the overall technical solution of the patent disclosure text and each inventive concept. The overall technical solution records all the technical features of the patent disclosure text.

[0166] Next, the overall technical solution of the technical decomposition data pair is input into the language model to be trained, and the language model to be trained is instructed to perform technical decomposition processing to determine at least one concept output.

[0167] Finally, the model parameters of the language model to be trained are adjusted based on the differences between at least one idea output and the various inventive ideas of the technology split data pairs.

[0168] If the aforementioned reference relationships are involved, the reference output of the concept output can be output simultaneously when outputting the concept output, and the correction is based on the reference relationships in the training data.

[0169] Exemplary inventive concept decoupling method: This application further discovers that in patent publications, dependent claims often rely on the claims they reference, and considering that their technical solutions are built upon prior claims, training / using a separate model based on dependent claims often leads to output results that differ significantly from the dependent claims themselves (e.g., the output may be an independent claim, lacking necessary technical features, or the model may arbitrarily add technical features).

[0170] To address the aforementioned issues, this application decouples the claim concepts constructed from the dependent claims, enabling the model to be generated and trained based on the decoupled dependent claims without encountering the problems described above. Specifically, the decoupling logic can be understood as writing the prior claims, on which the dependent claims depend, as injection information into the model input. This process can be performed during training data generation or during model training.

[0171] To further illustrate this decoupling process, this application provides an exemplary flowchart of the decoupling method in training data generation ( Figure 7 ) and an exemplary flowchart of the decoupling method during training ( Figure 8 ).

[0172] like Figure 7 As shown, process P700 may include the following steps: S710. For the target claim text in at least one claim text, determine the prior claim text of the target claim text.

[0173] S720. Update the inventive concept of the target claim text based on the prior claim text and the inventive concept corresponding to the prior claim text, so that the inventive concept of the target claim text records the prior claim text and the inventive concept corresponding to the prior claim text.

[0174] S730. Based on the text of the target claim and the inventive concept corresponding to the text of the target claim, determine the claim concept pair of the target claim, so as to determine at least one set of claim concept pairs corresponding to at least one claim text.

[0175] In S710 above, similarly to the preceding statement, the target claim text refers to an optional version of the claim text. The prior claim text may refer to a claim directly or indirectly referenced by the target claim text. Optionally, an independent claim may be considered as not referencing any claim or referencing the background art (i.e., the background art can be considered as claim 0 with reference to the aforementioned scheme).

[0176] In the aforementioned S720, the update of the inventive concept of the target claim text can be understood as incorporating the inventive concept of the prior claim text. This process can be implemented by constructing additional stored data or by segmenting the original text data with special identifiers. For details, please refer to the relevant technology in this application.

[0177] In the aforementioned S730, each claim can be processed based on the processing of the target claim to update the claim pair. In this updated claim pair, the dependent claims are decoupled and can be used directly.

[0178] like Figure 8 As shown, P800 may include the following steps: S810. Determine the prior claim text of the target claim text and the inventive concept corresponding to the prior claim text.

[0179] S820. Determine the injection information of the target claim concept pair based on the text of the prior claims and the inventive concept corresponding to the text of the prior claims.

[0180] S830, The inventive concept aligned with the target claim concept and the injected information are input into the language model to be trained, and the language model to be trained is instructed to generate claims, thereby determining the claim text output of the language model to be trained.

[0181] Unlike the aforementioned P700, P800 takes into account that the decoupling process is performed during training, so the relevant data of the prior claims can be directly written into the language model as injected information without needing to be associated with the target claims for storage.

[0182] Exemplary virtual devices and other equipment: Figure 9 This is a block diagram of a patent text generation electronic device provided in an exemplary embodiment of this application.

[0183] Reference Figure 9 The electronic device 900 includes a processing component 910, which further includes one or more processors, and memory resources represented by a memory 920 for storing instructions, such as application programs, that can be executed by the processing component 910. The application programs stored in the memory 920 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 910 is configured to execute instructions to perform the training data generation method described above.

[0184] Electronic device 900 may also include a power supply component configured to manage the power supply of electronic device 900, a wired or wireless network interface configured to connect electronic device 900 to a network, and an input / output (I / O) interface. Electronic device 900 can operate based on an operating system stored in memory 920, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0185] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device 900, enables the electronic device 900 to execute a training data generation method, comprising: determining at least one inventive concept, wherein the inventive concept includes a technical problem text and a technical means text, the technical problem recorded in the technical problem text being determined based on technical documents in a patent text, and the technical means text recording a set of technical means in the technical documents for solving the technical problem recorded in the technical problem text; processing the at least one inventive concept through a language model to determine a set of claim texts for the at least one inventive concept, wherein each claim text in the claim text set corresponds to a different inventive concept, and the claim text includes technical feature text, the technical feature text corresponding to the technical means text in the corresponding inventive concept.

[0186] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.

[0187] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0188] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0189] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0190] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0191] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0192] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0193] It should be noted that in the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0194] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications or equivalent substitutions made within the spirit and principles of this application should be included within the protection scope of this application.< / think>

Claims

1. A method for generating training data, characterized in that, include: Obtain a patent publication text, wherein the patent publication text includes a specification text and at least one claim text, and the specification text is used to interpret each claim text recorded in the at least one claim text; Identify the inventive concept of each claim text in the at least one claim text from the specification text, wherein the inventive concept corresponds to the claim text, and the inventive concept includes a technical means text and a technical problem text of the corresponding claim text, wherein the technical problem text records the technical problem solved by the corresponding claim text and / or the technical effect achieved; Based on the at least one claim text and its corresponding inventive concept, at least one set of claim concept pairs in the training data of the patent disclosure text is determined, wherein the claim concept pair includes one claim text and its corresponding inventive concept from the at least one claim text; When the training data of the patent disclosure text is used to train the claim generation capability of the language model, the claim concept pair is configured as the claim training data of the language model, the inventive concept in the claim concept pair is configured as the sample data of the language model, and the claim text in the claim concept pair is configured as the model output label of the language model.

2. The training data generation method according to claim 1, characterized in that, The training data generation method further includes: For a target pair of claims in the at least one set of claims, a writing thought text for the target pair of claims is generated based on the claims writing rules, so as to determine the writing thought text for each pair of claims in the at least one set of claims, wherein the writing thought text reflects the use of the claims writing rules in the process of generating the corresponding claim text from the corresponding inventive concept. Wherein, when the training data is used to train the claim generation capability of the language model, the writing thought text of the claim idea is configured as the distillation label of the language model, and the distillation label is used to shape the thought field text in the model output of the language model.

3. The training data generation method according to claim 2, characterized in that, The generation of writing thought texts for the target authority idea pairs based on authority writing rules, to determine the writing thought texts for each authority idea pair in the at least one set of authority idea pairs, includes: The supremacy drafting rules and the target supremacy concept are used as input information to the target language model, instructing the target language model to generate the thought process of generating the corresponding claim text from the corresponding inventive concept based on the supremacy drafting rules, and configuring it as the drafting thought text.

4. The training data generation method according to claim 1, characterized in that, The training data generation method further includes: The overall technical solution is determined based on the specification text and / or the texts of various technical means in the at least one set of claims, wherein the overall technical solution records all the technical features of each text of technical means. Based on the overall technical solution and each inventive concept, determine the technical breakdown data pairs in the training data of the patent disclosure text, wherein the technical breakdown data pairs include the overall technical solution and each inventive concept; When the training data of the patent disclosure text is used to train the technical solution segmentation capability of the language model, the technical segmentation data is configured to segment the solutions of the language model. The overall technical solution in the technical split data pair is configured as the sample data of the language model, and each inventive idea in the technical split data pair is configured as the model output label of the language model.

5. The training data generation method according to claim 4, characterized in that, The technical segmentation data pairs in the training data for determining the patent disclosure text based on the overall technical solution and each inventive concept include: Determine the reference relationships of the text of at least one claim; The dependency relationship of the inventive concept corresponding to each claim text is determined based on the citation relationship of the at least one claim text; Based on the overall technical solution, each inventive concept and their dependencies, the technical segmentation data pairs in the training data of the patent disclosure text are determined. When the training data of the patent disclosure text is used to train the technical solution segmentation capability of the language model, the dependencies of each inventive concept in the technical segmentation data pairs are also configured as the model output labels of the language model.

6. The training data generation method according to claim 1, characterized in that, The determination of at least one set of claim pairings of the patent disclosure text based on the at least one claim text and its corresponding inventive concept includes: For the target claim text in the at least one claim text, a logical verification is performed on the technical means text and the technical problem text of the target claim text to determine the logical verification result of the target claim text, so as to determine the logical verification result of each claim text in the at least one claim text, wherein the logical verification result reflects the consistency between the technical means text and the claim text and the effectiveness of the technical means text in solving the technical problem text; In response to the logical verification results of the target claim text satisfying the consistency and validity requirements, a set of claim concept pairs is generated based on the target claim text and its inventive concept to generate at least one set of claim concept pairs of the patent disclosure text.

7. The training data generation method according to claim 1, characterized in that, The process of determining at least one set of claim concept pairs in the training data of the patent disclosure text based on the at least one claim text and its corresponding inventive concept includes: For the target claim text in the at least one claim text, determine the prior claim text of the target claim text; The inventive concept of the target claim text is updated based on the prior claim text and the inventive concept corresponding to the prior claim text, so that the inventive concept of the target claim text records the prior claim text and the inventive concept corresponding to the prior claim text; Based on the text of the target claim and the inventive concept corresponding to the text of the target claim, the claim concept pair of the target claim is determined, so as to determine at least one set of claim concept pairs corresponding to the at least one claim text.

8. The training data generation method according to claim 1, characterized in that, The inventive concept of identifying each claim text in the at least one claim text from the specification text includes: Determine the set of keywords for each claim text in the at least one claim text; For the first claim text in the at least one claim text, a first feature set and a second feature set are determined from the keyword set of the first claim text based on the prior claims of the first claim text, so as to determine the first feature set and the second feature set of each claim text, wherein the first feature set is the keywords or keyword combinations in the keyword set of the corresponding claim text that do not appear in the prior claims, and the second feature set is the keywords or keyword combinations in the keyword set of the corresponding claim text that appear in the prior claims; For the second claim text in the text of the at least one claim, the text similarity between each text segment in the specification text and the second claim text is determined based on the first feature set and the second feature set of the second claim text; Based on the text similarity, the explanatory text of the second claim text is determined, so as to determine the explanatory text of each claim text in the at least one claim text; Identify the inventive concept of each claim text from the explanatory text of each claim text.

9. The training data generation method according to claim 1, characterized in that, The process of identifying the inventive concept of each claim text from the explanatory text of each claim text includes: For the target claim text in the at least one claim text, determine the interpretation text of the prior claim of the target claim text and configure it as the prior art text; The prior art text and the explanatory text of the target claim are input into the language model, and the language model is instructed to identify the inventive concept from the explanatory text of the target claim to determine the inventive concept of each claim.

10. The training data generation method according to claim 1, characterized in that, The inventive concept of identifying each claim text in the at least one claim text from the specification text includes: For the target claim text in the at least one claim text, the specification text corresponding to the target claim text is input into the language model, and the language model is instructed to determine at least one technical directed graph based on the specification text and the system analysis method. The language model is instructed to determine the technical problems and technical means in the inventive concept of the target claim text based on the directed graph of the technology.

11. A training method for a patent text generation model, characterized in that, The training method includes: Determine training data for the patent disclosure text based on the training data generation method according to any one of claims 1 to 10, and configure at least one set of authoritative ideas in the training data as authoritative training data for the language model; For a target claim idea pair in the at least one set of claim idea pairs, the inventive idea in the target claim idea pair is input into the language model to be trained, and the language model to be trained is instructed to generate claims, thereby determining the claim text output of the language model to be trained. The model parameters of the language model to be trained are adjusted based on the differences between the output of the claim text and the claim text in the target claim concept pair.

12. The training method for the patent text generation model according to claim 11, characterized in that, The step of inputting the inventive concept from the target claim concept pair into the language model to be trained, instructing the language model to be trained to generate claims, and determining the claim text output of the language model to be trained includes: Determine the prior claim text of the target claim concept and the inventive concept corresponding to the prior claim text; The injection information of the target claim concept pair is determined based on the text of the prior claims and the inventive concept corresponding to the text of the prior claims; The inventive concept and injected information of the target claim concept are input into the language model to be trained, and the language model to be trained is instructed to generate claims, thereby determining the claim text output of the language model to be trained.

13. The training method for the patent text generation model according to claim 11, characterized in that, The claim concept pair also includes drafting thought text, wherein the drafting thought text reflects the use of the claim drafting rules in the process of generating the corresponding claim text from the corresponding inventive concept; The step of inputting the inventive concept from the target claim concept pair into the language model to be trained, instructing the language model to be trained to generate claims, and determining the claim text output of the language model to be trained includes: The inventive concept in the target claim concept pair is input into the language model to be trained, and the language model to be trained is instructed to generate claims, thereby determining the thinking field output and claim text output of the language model to be trained, wherein the thinking field output is based on thinking tag separation; The training method further includes adjusting the model parameters of the language model to be trained based on the differences between the output of the thinking field and the written thinking text.

14. The training method for the patent text generation model according to claim 11, characterized in that, The training method also includes: The technical segmentation data pairs in the training data of the patent disclosure text are determined, wherein the technical segmentation data pairs include the overall technical solution of the patent disclosure text and each inventive concept, and the overall technical solution records all the technical features of the patent disclosure text. The overall technical solution of the technical decomposition data pair is input into the language model to be trained, and the language model to be trained is instructed to perform technical decomposition processing to determine at least one concept output; The model parameters of the language model to be trained are adjusted based on the differences between the output of at least one concept and the various inventive concepts of the technology split data pair.

15. The training method for the patent text generation model according to claim 11, characterized in that, The step of adjusting the model parameters of the language model to be trained based on the difference between the output of the authoritative text and the claim text in the target authoritative concept pair includes: Determine a first assertion score for the claim text and a second assertion score for the output of the claim text; The model parameters of the language model to be trained are adjusted based on the difference between the first and second weighted scores and the difference between the weighted text output and the claim text.

16. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for performing the method described in any one of claims 1 to 15.

17. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions. The processor is used to execute the method according to any one of claims 1 to 15.

18. A computer program product having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 15.