Text processing method, text generation method, equipment, storage medium and program product

By segmenting and building text libraries, and using large models to adjust the legal and regulatory text units, the problems of low generation efficiency and low accuracy in the existing technology are solved, and efficient and high-quality legal and regulatory text generation are achieved.

CN120493872APending Publication Date: 2025-08-15SHIXIN KEMING (BEIJING) INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510571211.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing legal and regulatory text generation is inefficient and has low accuracy, making it difficult to meet the needs of quickly generating high-quality legal and regulatory texts.

Method used

By obtaining multiple legal and regulatory texts of the target legal and regulatory type, dividing them into multiple text units and extracting text features, building a target text library, using a large model to determine text units that meet the similarity requirements for adjustments, and generating the second legal and regulatory text.

Benefits of technology

It improves the accuracy and efficiency of legal and regulatory text generation, and realizes the rapid generation of high-quality legal and regulatory texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493872A_ABST
    Figure CN120493872A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text processing method, a text generation method, equipment, a storage medium and a program product, which are applied to the field of data processing, and comprise the following steps: obtaining a plurality of law and regulation texts of a target law and regulation type; for any law and regulation text, segmenting the law and regulation text into a plurality of text units; extracting text features respectively corresponding to the plurality of text units; constructing a target text library corresponding to the target law and regulation type based on the text features corresponding to the plurality of law and regulation texts; wherein the target text library is used for determining at least one text unit which meets the similarity requirement with the text features of the first law and regulation text of the target law and regulation type by utilizing the first large model, and the at least one text unit is used for adjusting the first law and regulation text by utilizing the second large model so as to generate a second law and regulation text. According to the scheme of the embodiment of the invention, the generation efficiency and accuracy of the law and regulation text are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing, and in particular to a text processing method, a text generation method, a device, a storage medium, and a program product. Background Art

[0002] In the current scenario of generating legal and regulatory texts, relevant personnel usually manually generate new legal and regulatory texts by searching for relevant legal and regulatory texts as references, which is inefficient and not very accurate.

[0003] Therefore, how to improve the efficiency and accuracy of generating legal and regulatory texts has become a technical problem that needs to be solved urgently. Summary of the Invention

[0004] The embodiments of the present application provide a text processing method, a text generation method, a device, a storage medium, and a program product to improve the efficiency and accuracy of generating legal and regulatory texts.

[0005] In a first aspect, an embodiment of the present application provides a text processing method, comprising:

[0006] Obtain multiple legal and regulatory texts of the target legal and regulatory type;

[0007] For any legal or regulatory text, the legal or regulatory text is divided into multiple text units;

[0008] Extracting text features corresponding to the plurality of text units respectively;

[0009] Building a target text library corresponding to the target legal and regulatory type based on the text features corresponding to the plurality of legal and regulatory texts;

[0010] Among them, the target text library is used to use the first large model to determine at least one text unit whose text features meet the similarity requirements with the first legal and regulatory text of the target legal and regulatory type, and the at least one text unit is used to use the second large model to adjust the first legal and regulatory text to generate a second legal and regulatory text.

[0011] In a second aspect, an embodiment of the present application provides a text generation method, comprising:

[0012] Responding to the text generation request, obtaining a first legal and regulatory text;

[0013] Determining a text library corresponding to the legal and regulatory type of the first legal and regulatory text; wherein the text library is constructed by text features corresponding to a plurality of legal and regulatory texts corresponding to the legal and regulatory type, the text features being extracted from text units obtained by segmenting the legal and regulatory text;

[0014] Determining, from the text library using the first large model, at least one text unit that meets similarity requirements with text features of the first legal and regulatory text;

[0015] The first legal and regulatory text is adjusted according to the at least one text unit using the second large model to generate a second legal and regulatory text.

[0016] In a third aspect, an embodiment of the present application provides a computing device comprising a storage component and a processing component; the storage component stores one or more computer program instructions, the computer program instructions being called and executed by the processing component, and the processing component executing the one or more computer program instructions to implement the text processing method described in the first aspect, or the text generation method described in the second aspect.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, wherein the computer program is executed by a computer to implement the text processing method described in the first aspect or the text generation method described in the second aspect.

[0018] In a fifth aspect, an embodiment of the present application provides a computer program product storing a computer program, which, when executed by a computer, implements the text processing method described in the first aspect or the text generation method described in the second aspect.

[0019] In an embodiment of the present application, by constructing a target text library corresponding to the target legal and regulatory type, a large model can be used to determine, from the corresponding target text library, at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text based on the first legal and regulatory text. Furthermore, a second large model can be used to adjust the first legal and regulatory text based on the at least one text unit to generate a second legal and regulatory text, thereby improving the accuracy of legal and regulatory text generation. Furthermore, by obtaining multiple legal and regulatory texts of the target legal and regulatory type, dividing the obtained legal and regulatory texts into multiple text units and extracting text features, and constructing a target text library based on the text features, similarity calculations can be performed based on the text features to determine at least one text unit, thereby improving computational efficiency and, in turn, improving the efficiency of legal and regulatory text generation.

[0020] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0022] Figure 1 A flowchart of an embodiment of a text processing method provided by the present application is shown;

[0023] Figure 2 A flowchart of an embodiment of a text generation method provided by the present application is shown;

[0024] Figure 3 A schematic diagram showing a system architecture in a practical application is shown;

[0025] Figure 4 A schematic structural diagram of an embodiment of a text processing device provided by the present application is shown;

[0026] Figure 5 A schematic structural diagram of an embodiment of a text generation device provided by the present application is shown;

[0027] Figure 6 A schematic structural diagram of an embodiment of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0028] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0029] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.

[0030] As described in the background technology section, current legal and regulatory texts are usually manually generated by relevant personnel by searching for relevant legal and regulatory texts as references, which is inefficient and not very accurate.

[0031] In order to improve the efficiency and accuracy of legal and regulatory text generation, the inventors proposed the technical solution of the present application, including obtaining multiple legal and regulatory texts of a target legal and regulatory type; for any legal and regulatory text, dividing the legal and regulatory text into multiple text units; extracting text features corresponding to the multiple text units; and constructing a target text library corresponding to the target legal and regulatory type based on the text features corresponding to the multiple legal and regulatory texts; wherein the target text library is used to use a first large model to determine at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text of the target legal and regulatory type, and the at least one text unit is used to use a second large model to adjust the first legal and regulatory text to generate a second legal and regulatory text.

[0032] By constructing a target text library corresponding to the target legal and regulatory type, a large model can be used to identify at least one text unit from the corresponding target text library based on a first legal and regulatory text of the target legal and regulatory type, whose text features meet the similarity requirements with the first legal and regulatory text. A second large model can then be used to adjust the first legal and regulatory text based on the at least one text unit to generate a second legal and regulatory text, thereby improving the accuracy of legal and regulatory text generation. Furthermore, by obtaining multiple legal and regulatory texts of the target legal and regulatory type, segmenting the obtained legal and regulatory texts into multiple text units, and extracting text features, the target text library is constructed based on the text features, enabling similarity calculation based on the text features to determine the at least one text unit. This improves computational efficiency, thereby improving the efficiency of legal and regulatory text generation.

[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0034] The technical solutions of the embodiments of the present application can be applied to a system architecture including a user terminal and a server terminal, wherein the user terminal and the server terminal are connected via a network. The network provides a medium for the communication link between the user terminal and the server terminal. The network can include various connection types, such as wired or wireless communication links or fiber optic cables.

[0035] The user end can interact with the server end through the network to send a first legal and regulatory text or receive a second legal and regulatory text, etc.

[0036] Among them, the user end can be a browser, APP (Application), or web application such as H5 (Hyper Text Markup Language 5, Hypertext Markup Language 5th edition) application, or light application (also known as applet, a lightweight application) or cloud application, etc. The user end can be deployed in an electronic device and needs to rely on the device to run or certain apps in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, tablet computer, personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0037] The server may include servers that provide various services, such as a server that builds a text library, a server that adjusts a first legal and regulatory text sent by a user to generate a second legal and regulatory text, etc.

[0038] It should be noted that the server can be implemented as a distributed server cluster consisting of multiple servers or a single server. The server can also be a server in a distributed system or a server integrated with blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology.

[0039] It should be noted that the text processing method and text generation method provided in the embodiments of the present application are generally executed by the server, and the corresponding text processing device and text generation model are generally deployed in the server. However, in other embodiments of the present application, the user terminal may also have similar functions to the server, thereby executing the text processing method and text generation method provided in the embodiments of the present application. In other embodiments, the text processing method and text generation method provided in the embodiments of the present application may also be jointly executed by the user terminal and the server.

[0040] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).

[0041] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0042] like Figure 1 FIG. 1 is a flowchart of an embodiment of a text processing method provided by the present application, which can be executed by a server. The method may include the following steps:

[0043] 101: Get multiple legal and regulatory texts of the target legal and regulatory type.

[0044] In the embodiments of the present application, the types of laws and regulations can be divided according to different classification criteria. As an optional implementation method, the types of laws and regulations can be divided according to different regions. As another optional implementation method, the types of laws and regulations can be divided according to different fields, for example, they can include civil law, commercial law, and other types. As another optional implementation method, the types of laws and regulations can be divided according to different effectiveness, for example, they can include higher-level laws, equivalent laws, and other types.

[0045] The target legal and regulatory type can be any legal and regulatory type. Acquiring multiple legal and regulatory texts of the target legal and regulatory type can, for example, involve acquiring multiple legal and regulatory texts uploaded by target users, where the target users can refer to users associated with the generation of legal and regulatory texts, such as legislators. Alternatively, multiple legal and regulatory texts can be obtained from a general legal and regulatory text library and classified into different types to determine the multiple legal and regulatory texts corresponding to the target legal and regulatory type. This application does not impose any limitations on this.

[0046] 102: For any legal or regulatory text, divide the legal or regulatory text into multiple text units.

[0047] To facilitate subsequent text processing, optionally, before text segmentation, the obtained legal and regulatory texts may be subjected to pre-processing operations such as text cleaning and text verification.

[0048] There are many ways to segment legal and regulatory texts, such as segmenting according to text paragraphs, text sentence patterns, etc., which will be explained in subsequent embodiments and will not be repeated here.

[0049] 103: Extract text features corresponding to multiple text units.

[0050] Specifically, a feature extraction model can be used to extract corresponding text features from the multiple text units obtained through segmentation. In a practical application, a vector feature extraction model (an embedding model used to map high-dimensional data (such as text, images, and audio) to a low-dimensional space while preserving semantic relationships, making the high-dimensional raw data separable after being mapped to a low-dimensional space, essentially mapping from a semantic space to a vector space) can be used to extract text feature vectors corresponding to the multiple text units.

[0051] 104: Based on the text features corresponding to the multiple legal and regulatory texts, a target text library corresponding to the target legal and regulatory type is constructed.

[0052] Among them, the target text library can be used to use the first large model to determine at least one text unit whose text features meet the similarity requirements with the first legal and regulatory text of the target legal and regulatory type, and the at least one text unit can be used to use the second large model to adjust the first legal and regulatory text to generate a second legal and regulatory text.

[0053] That is to say, the constructed target text library can support the use of a large model to determine at least one text unit that has similar text features to the first legal and regulatory text of the target legal and regulatory type, thereby supporting the second large model to adjust the first legal and regulatory text according to the at least one text unit to generate a second legal and regulatory text, thereby improving the accuracy of legal and regulatory text generation.

[0054] The large model (also called the foundation model, i.e., FoundationModel) involved in the embodiments of the present application includes the first large model, the second large model, and the third large model, the fourth large model, and the fifth large model involved below. It refers to a machine learning model with a large number of parameters and a complex structure. It can process massive amounts of data and complete various complex tasks, such as natural language processing, computer vision, speech recognition, etc. It is an AI (Artificial Intelligence) model. The large model can be implemented using LLM (Large Language Model) or MLM (Multimodal Large Model), for example, a generative pre-training model can be used, and this application does not limit this.

[0055] Processing using a large model can be inputting prompt information into the large model. Prompt information is a form of input used to prompt or guide the large model to give expected output, indicating what actions the large model should take or what output it should generate when performing a specific task. Prompt information is a natural language input, similar to a command or instruction, to let the large model know what it needs to do. Specifically, the implementation process of determining at least one text unit using the first large model and generating a second legal text using the second large model will be described in subsequent embodiments.

[0056] It should be noted that the first large model and the second large model can be the same large model or different large models, and this application does not impose any restrictions on this.

[0057] In this embodiment, by constructing a target text library corresponding to the target legal and regulatory type, a large model can be used to determine, from the corresponding target text library, at least one text unit whose text features meet similarity requirements with the first legal and regulatory text based on a first legal and regulatory text of the target legal and regulatory type. The second large model can then be used to adjust the first legal and regulatory text based on the at least one text unit to generate a second legal and regulatory text, thereby improving the accuracy of legal and regulatory text generation. Furthermore, by obtaining multiple legal and regulatory texts of the target legal and regulatory type, segmenting the obtained legal and regulatory texts into multiple text units, extracting text features, and constructing a target text library based on the text features, similarity calculations based on the text features are performed to determine the at least one text unit, thereby improving computational efficiency and, in turn, improving the efficiency of legal and regulatory text generation.

[0058] The following explains the process of segmenting legal and regulatory texts.

[0059] As an optional implementation method, the legal and regulatory text can be segmented into multiple first text segments corresponding to the text structure according to the document format of the legal and regulatory text, and any first text segment can be used as a text unit. The document format can include, for example, word, txt (text format), pdf (Portable Document Format), md (Markdown, a lightweight markup language), csv (Comma-Separated Values, whose files store tabular data (numbers and text) in plain text form) and other formats. The text structure can include, for example, chapters, sections, articles, etc. For example, the legal and regulatory text can be segmented according to chapters to obtain multiple first text segments corresponding to each chapter, and the first text segment corresponding to each chapter can be used as a text unit. For another example, the legal and regulatory text can be segmented according to articles to obtain multiple first text segments corresponding to each article, and the first text segment corresponding to each article can be used as a text unit, and so on.

[0060] As another optional implementation, the legal text can be segmented into multiple text units according to preset symbols. Preset symbols may include punctuation marks such as periods, semicolons, exclamation marks, and question marks. For example, the legal text can be segmented into multiple sentences according to periods, with each sentence being considered a text unit.

[0061] As another optional implementation, the legal and regulatory text can be segmented into a plurality of first text segments corresponding to respective text structures according to the document format of the legal and regulatory text, and the first text segments corresponding to the same text structure can be segmented into a plurality of second text segments according to preset symbols, with any second text segment being regarded as a text unit. The first text segments corresponding to the same text structure can, for example, include the first text segments corresponding to the same chapter, the same section, or the same article, and the first text segments corresponding to the same text structure can be segmented into a plurality of second text segments according to punctuation marks such as periods, semicolons, exclamation marks, and question marks, with each second text segment being regarded as a text unit.

[0062] By segmenting a larger legal and regulatory text into multiple smaller text units, the efficiency of subsequent text feature extraction and text feature similarity calculation based on text units is improved.

[0063] In order to further improve the efficiency of subsequent text feature extraction and text feature similarity calculation, the text segments can be further segmented. Therefore, in some embodiments, the above method can also include:

[0064] When the number of bytes in the second text segment is greater than the first threshold, the second text segment is divided into a plurality of third text segments, where the number of bytes in the third text segments is less than the first threshold;

[0065] For two adjacent third text segments, preset bytes in the previous third text segment are added to the latter third text segment to obtain multiple fourth text segments; wherein the number of the preset bytes is less than the second threshold.

[0066] In this embodiment, the first threshold and the second threshold can be set according to actual needs. For example, the first threshold can be set to 140 bytes, 150 bytes, etc., and the second threshold can be set to 20 bytes, 30 bytes, etc.

[0067] Specifically, when the number of bytes in the second text segment is greater than the first threshold, the second text segment can be further segmented to obtain a plurality of third text segments with a number of bytes less than the first threshold. Moreover, in order to improve the text continuity between the third text segments, an overlapping window with a number of bytes less than the second threshold can be set, that is, at a preset position of the latter third text segment, such as the head position, add the preset bytes of another preset position in the previous third text segment, such as the tail position. For example, if the content of the previous third text segment includes: xxx for processing, then the preset bytes of "the processing" can be added to the head position of the latter third text segment, thereby obtaining a plurality of fourth text segments, and the content of the fourth text segment can include: the processing xxx, so as to improve text continuity and readability.

[0068] At this time, any second text segment as a text unit may include:

[0069] Any fourth text segment is regarded as a text unit.

[0070] By setting a first threshold and a second threshold, if the number of bytes in the second text segment exceeds the first threshold, the second text segment is further segmented to obtain multiple smaller third text segments, further improving the efficiency of subsequent text feature extraction and text similarity calculation. In addition, by adding preset bytes from the previous third text segment to the next third text segment between two adjacent text segments, multiple fourth text segments are obtained, which improves the text continuity and readability between adjacent text segments, improves the rationality of text unit division, and further improves the accuracy of subsequent text feature extraction and text feature similarity calculation.

[0071] To further improve the accuracy of text processing, in some embodiments, the above method may further include:

[0072] Use the third model to perform semantic analysis on legal and regulatory texts to obtain semantic data corresponding to legal and regulatory texts;

[0073] Extracting semantic features of semantic data;

[0074] The semantic features and text features are stored in the target text library in correspondence.

[0075] The semantic feature can be used to determine, using the first large model, at least one text unit that meets similarity requirements with the text feature of the first legal and regulatory text of the target legal and regulatory type.

[0076] In this embodiment, the corresponding prompt information may be, for example: Please perform semantic analysis on the given legal text and explain the professional terms and concepts included in the legal text. By inputting this prompt information into the third model, the semantic data corresponding to the legal text can be obtained.

[0077] It should be noted that the third large model can be the same large model as the aforementioned first large model and second large model, or can be different, and this application does not impose any restrictions on this.

[0078] For the obtained semantic data, a feature extraction model, such as an embedding model, can be used to extract the corresponding semantic features, and the semantic features and multiple text features corresponding to the legal and regulatory text can be stored in the target text library so that the semantic features and text features can be combined to perform similarity calculations in the future to improve the accuracy of text unit searches.

[0079] By using the third model to perform semantic analysis on the legal and regulatory text, the semantic data corresponding to the legal and regulatory text is obtained, and the corresponding semantic features are stored together with the text features in the target text library, thereby improving the semantic clarity of the text unit, helping the target users to better understand and interpret the legal and regulatory text, and facilitating the search and determination of at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text, thereby improving the accuracy of text unit determination and thereby improving the accuracy of legal and regulatory text generation.

[0080] To further improve the accuracy of text processing, in some embodiments, the above method may further include:

[0081] Get multiple case data;

[0082] Determine at least one target case data corresponding to the legal and regulatory text from the plurality of case data using the fourth model;

[0083] At least one target case data is stored in the target text library.

[0084] The at least one target case data may be used to determine, using the first large model, at least one text unit that meets similarity requirements with the text features of the first legal and regulatory text of the target legal and regulatory type.

[0085] In the embodiment of the present application, case data may refer to data related to cases judged in accordance with legal texts. The case data may be provided by the target user, or may be obtained from a general case database, without limitation.

[0086] In this embodiment, the corresponding prompt information may be, for example, "Please identify one or more target case data related to a given legal or regulatory text from a given plurality of case data." This prompt information is input into the fourth model to determine at least one target case data corresponding to the legal or regulatory text.

[0087] It should be noted that the fourth large model can be the same large model as the aforementioned first large model, second large model, and third large model, or can be different, and this application does not impose any restrictions on this.

[0088] For the at least one target case data obtained, it can be stored in the target text library along with multiple text features corresponding to the legal and regulatory text, so that similarity calculation can be performed in combination with the target case data and text features to improve the accuracy of text unit search.

[0089] Optionally, for at least one target case data obtained, a feature extraction model, such as an embedding model, can be used to extract the corresponding feature vector, and the feature vector and multiple text features corresponding to the legal and regulatory text can be stored in the target text library so that the feature vector and text features can be combined to perform similarity calculations subsequently to improve the accuracy of text unit search.

[0090] By using the fourth model to determine at least one target case data corresponding to the legal and regulatory text, and storing it together with the text features in the target text library, it helps target users better understand the application of the legal and regulatory text, facilitates the search and determination of at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text, further improves the accuracy of text unit determination, and thus improves the accuracy of legal and regulatory text generation.

[0091] Optionally, the feature vectors, semantic features and multiple text features corresponding to the target case data can also be stored in the target text library, so as to combine the feature vectors, semantic features and text features to perform similarity calculations, thereby improving the accuracy of text unit determination and the accuracy of legal and regulatory text generation. This will not be elaborated here.

[0092] In practical applications, a knowledge graph of legal and regulatory texts can also be constructed to further improve the accuracy and comprehensiveness of text unit determination. In some embodiments, the above method can also include:

[0093] A legal and regulatory knowledge graph is constructed according to the correlation between legal and regulatory texts corresponding to multiple legal and regulatory types.

[0094] Among them, the legal and regulatory knowledge graph is used to use the first large model to determine the associated legal and regulatory texts associated with at least one text unit, and the associated legal and regulatory texts are used to use the second large model to adjust the first legal and regulatory text to generate the second legal and regulatory text.

[0095] Among them, the method for obtaining the legal and regulatory text corresponding to each legal and regulatory type is the same as the method for obtaining the multiple legal and regulatory texts corresponding to the target legal and regulatory type mentioned above, and will not be repeated here.

[0096] In this embodiment, the association relationship may refer to, for example, a superordinate relationship, a subordinate relationship, etc. That is, the constructed legal and regulatory knowledge graph can support the use of a first large model to determine associated legal and regulatory texts that have an association relationship with at least one text unit, and support the use of a second large model to combine the at least one text unit and the associated legal and regulatory texts to adjust the first legal and regulatory text to generate a second legal and regulatory text.

[0097] By constructing a knowledge graph of laws and regulations, the accuracy and comprehensiveness of legal and regulatory text generation have been further improved.

[0098] In actual applications, legal and regulatory texts are usually updated. Therefore, in some embodiments, the above method may further include:

[0099] For any legal text, the fifth model is used to search the target text library for historical legal texts that are consistent with the legal text, and to determine the comparison results between the legal text and the historical legal texts;

[0100] Provide comparison results to target users.

[0101] Specifically, when constructing the target text library, the content, version identifier, and generation time of the corresponding legal and regulatory texts can be stored. Thus, for newly acquired legal and regulatory texts, the fifth model can be used to search for historical versions of legal and regulatory texts with consistent content. Consistent content can refer to the same legal and regulatory items, but the specific textual descriptions may differ. These versions can then be compared with the newly acquired legal and regulatory texts, and the comparison results provided to the target user, allowing them to clearly understand the updates to the legal and regulatory texts.

[0102] In this embodiment, the corresponding prompt information may be, for example: please search the target text library for historical legal and regulatory texts that are consistent with the given legal and regulatory texts, and output the comparison results between the two.

[0103] It should be noted that the fifth large model can be the same large model as the aforementioned first large model, second large model, third large model, and fourth large model, or can be different, and this application does not impose any restrictions on this.

[0104] By using the fifth model, we can search the target text library for historical versions of legal and regulatory texts that are consistent with the content of the legal and regulatory texts, and compare them to obtain comparison results, so that target users can clearly understand the update status of the legal and regulatory texts.

[0105] Optionally, the historical version of the legal and regulatory text can be replaced with the new version of the legal and regulatory text to update the target text library, improve the accuracy of the target text library, and further improve the accuracy of text unit determination and the accuracy of second legal and regulatory text generation.

[0106] like Figure 2 FIG. 1 is a flowchart of an embodiment of a text generation method provided by the present application, which can be executed by a server. The method may include the following steps:

[0107] 201: In response to a text generation request, obtain a first legal and regulatory text.

[0108] The text generation request may be generated by the client in response to a text generation operation triggered by a target user, such as an input operation of a first legal and regulatory text, and sent to the server.

[0109] 202: Determine a text library corresponding to the legal and regulatory type of the first legal and regulatory text.

[0110] The text library is constructed by the text features corresponding to multiple legal and regulatory texts of the corresponding legal and regulatory types. The text features are extracted from text units, which are obtained by segmenting the legal and regulatory texts. The segmentation of text units and the construction of the text library are described in Figure 1 The corresponding descriptions have been given in the illustrated embodiments and will not be repeated here.

[0111] 203: Using the first large model, determine from the text library at least one text unit whose text features meet similarity requirements with the first legal and regulatory text.

[0112] The prompt information may, for example, be: Please calculate the similarity between multiple text features in a given text library and the text features of the first legal document, and select text units corresponding to one or more text features that meet the similarity requirements. This prompt information is input into the first large model to obtain at least one text unit.

[0113] Meeting the similarity requirement may, for example, mean that the similarity is greater than a similarity threshold. The similarity threshold may be set according to actual needs, for example, 90%, 95%, etc.

[0114] 204: Using the second largest model, adjust the first legal and regulatory text according to at least one text unit to generate a second legal and regulatory text.

[0115] The prompt information may be, for example, "Please adjust the first legal text according to at least one given text unit to generate a second legal text." The prompt information is input into the second large model to generate the second legal text.

[0116] In this embodiment, for a first legal text provided by a target user, a first large model can be used to determine, from a text library of the corresponding legal text type, text units corresponding to at least one text feature whose text features meet the similarity requirements. Furthermore, a second large model can be used to adjust the first legal text based on at least one text unit to generate a second legal text. By constructing a target text library corresponding to the target legal text type, the large model can be used to determine, based on the first legal text of the target legal text type, at least one text unit from the corresponding target text library whose text features meet the similarity requirements with the first legal text. Furthermore, the second large model can be used to adjust the first legal text based on at least one text unit to generate the second legal text, thereby improving the accuracy of legal text generation. Furthermore, by obtaining multiple legal texts of the target legal text type, segmenting the obtained legal texts into multiple text units, extracting text features, and constructing a target text library based on the text features, similarity calculations can be performed based on the text features to determine the at least one text unit. This improves computational efficiency, thereby improving the efficiency of legal text generation.

[0117] In some embodiments, determining, from the text library using the first large model, at least one text unit that meets similarity requirements with text features of the first legal and regulatory text may include:

[0118] Extracting first text features of the first legal and regulatory text using a feature extraction model;

[0119] The first large model is used to calculate similarity between the first text feature and multiple text features in the text library to obtain text units corresponding to at least one text feature having a similarity greater than a similarity threshold.

[0120] Specifically, the distance between two text feature vectors, or the cosine of the angle, etc., can be calculated to obtain the similarity between the two text features, which can be set according to actual needs.

[0121] Optionally, similarity calculation may be performed based on the first text feature and the semantic features and text features in the text library, and weighted summation may be performed to obtain at least one text unit having a similarity greater than a similarity threshold.

[0122] Optionally, similarity may be calculated based on the first text feature and the feature vector, semantic feature and text feature of the target case data in the text library, and weighted summed to obtain at least one text unit having a similarity greater than a similarity threshold.

[0123] Optionally, multiple text units corresponding to at least one text unit, obtained by dividing the same legal text, can be determined from the text library based on the text unit identifier, the legal text identifier corresponding to the text unit, and the division order of the legal text corresponding to the text unit. The multiple text units can then be merged according to their corresponding division order to obtain at least one complete legal text. Based on this, the first legal text can be adjusted based on the at least one complete legal text to generate a second legal text, thereby improving the accuracy of the second legal text.

[0124] Optionally, multiple candidate legal texts whose similarities meet similarity requirements can be obtained, and at least one target legal text can be determined from the multiple candidate legal texts using a reranking model. The reranking model (rerank model) can use methods such as an attention mechanism or a matching network to calculate a matching score between the first legal text and each candidate legal text, which can be called a relevance score. This score can reflect the semantic similarity and relevance between the candidate legal text and the first legal text. The multiple candidate legal texts can be sorted in descending order of relevance scores, and based on the sorting results, at least one target legal text can be determined from the multiple candidate legal texts. For example, at least one candidate legal text block with a high ranking can be determined as at least one target legal text.

[0125] By determining the candidate legal and regulatory texts, calculating the relevance scores using the reranking model, sorting multiple candidate legal and regulatory texts in descending order of relevance scores, and determining the target legal and regulatory text based on the sorting results, the accuracy of determining the target legal and regulatory text can be further improved, thereby improving the accuracy of generating the second legal and regulatory text.

[0126] For ease of understanding, the following Figure 3 The system architecture diagram shown in FIG. 1 illustrates the technical solution of this application. Figure 3 As shown, the system architecture may include a user end 301 and a server end 302 .

[0127] Among them, the server end 302 can obtain multiple legal and regulatory texts of the target legal and regulatory type provided by the target user through the user end 301. For any legal and regulatory text, the legal and regulatory text is divided into multiple text units, and the text features corresponding to the multiple text units are extracted. Based on the text features corresponding to the multiple legal and regulatory texts, a target text library corresponding to the target legal and regulatory type is constructed.

[0128] In addition, the server 302 can also respond to the text generation request sent by the user end 301, obtain the first legal and regulatory text, determine the text library corresponding to the legal and regulatory type of the first legal and regulatory text, use the first large model to determine at least one text unit from the text library that meets the similarity requirements with the text features of the first legal and regulatory text, and use the second large model to adjust the first legal and regulatory text according to at least one text unit to generate a second legal and regulatory text.

[0129] By constructing a target text library corresponding to the target legal and regulatory type, a large model can be used to identify at least one text unit from the corresponding target text library based on a first legal and regulatory text of the target legal and regulatory type, whose text features meet the similarity requirements with the first legal and regulatory text. A second large model can then be used to adjust the first legal and regulatory text based on the at least one text unit to generate a second legal and regulatory text, thereby improving the accuracy of legal and regulatory text generation. Furthermore, by obtaining multiple legal and regulatory texts of the target legal and regulatory type, segmenting the obtained legal and regulatory texts into multiple text units, and extracting text features, the target text library is constructed based on the text features, enabling similarity calculation based on the text features to determine the at least one text unit. This improves computational efficiency, thereby improving the efficiency of legal and regulatory text generation.

[0130] like Figure 4 FIG. 1 is a schematic diagram of a structure of an embodiment of a text processing device provided by the present application. The device may include the following modules:

[0131] The first acquisition module 401 is used to acquire multiple legal and regulatory texts of the target legal and regulatory type;

[0132] A segmentation module 402 is used to segment any legal or regulatory text into multiple text units;

[0133] A first extraction module 403 is used to extract text features corresponding to a plurality of text units;

[0134] A construction module 404 is used to construct a target text library corresponding to a target legal and regulatory type based on text features corresponding to the plurality of legal and regulatory texts;

[0135] Among them, the target text library is used to use the first large model to determine at least one text unit whose text features meet the similarity requirements with the first legal and regulatory text of the target legal and regulatory type, and at least one text unit is used to use the second large model to adjust the first legal and regulatory text to generate a second legal and regulatory text.

[0136] In some embodiments, the segmentation module 402 may include:

[0137] A first segmentation unit is used to segment the legal and regulatory text into a plurality of first text segments corresponding to the text structures according to the document format of the legal and regulatory text;

[0138] A second segmentation unit is configured to segment the first text segment corresponding to the same text structure into a plurality of second text segments according to preset symbols;

[0139] The determination unit is used to take any second text segment as a text unit.

[0140] In some embodiments, the apparatus may further comprise:

[0141] a third segmentation unit, configured to segment the second text segment into a plurality of third text segments when the number of bytes in the second text segment is greater than a preset first threshold, wherein the number of bytes in the third text segments is less than the first threshold;

[0142] an adding unit configured to add, for two adjacent third text segments, preset bytes from the third text segment at the earlier segmentation position to the third text segment at the later segmentation position, to obtain a plurality of fourth text segments; wherein the number of the preset bytes is less than a second threshold;

[0143] The determination unit can be specifically used to take any fourth text segment as a text unit.

[0144] In some embodiments, the apparatus may further comprise:

[0145] The parsing module is used to perform semantic parsing on legal and regulatory texts using the third model to obtain semantic data corresponding to the legal and regulatory texts;

[0146] A second extraction module is used to extract semantic features of semantic data;

[0147] Construction module 404 can be specifically used to store semantic features and text features in a target text library in correspondence; wherein the semantic features are used to use the first large model to determine at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text of the target legal and regulatory type.

[0148] In some embodiments, the apparatus may further comprise:

[0149] The second acquisition module is used to acquire multiple case data;

[0150] A first determination module is used to determine at least one target case data corresponding to the legal and regulatory text from the plurality of case data using the fourth model;

[0151] Construction module 404 can be specifically used to store at least one target case data in a target text library; wherein, at least one target case data is used to use the first large model to determine at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text of the target legal and regulatory type.

[0152] In some embodiments, the apparatus may further comprise:

[0153] A comparison module is used to use the fifth model to search for historical legal texts that are consistent with the content of any legal text from the target text library, and to determine the comparison results between the legal text and the historical legal texts;

[0154] Provides a module for providing comparison results to target users.

[0155] Figure 4 The text processing device can execute Figure 1 The implementation principle and technical effects of the text processing method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the text processing device in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.

[0156] like Figure 5 FIG. 1 is a schematic diagram of a structure of an embodiment of a text generation device provided by the present application. The device may include the following modules:

[0157] The third acquisition module 501 is configured to acquire a first legal and regulatory text in response to a text generation request;

[0158] A second determining module 502 is configured to determine a text library corresponding to the legal and regulatory type of the first legal and regulatory text; wherein the text library is constructed from text features corresponding to a plurality of legal and regulatory texts corresponding to the legal and regulatory type, wherein the text features are extracted from text units obtained by segmenting the legal and regulatory text;

[0159] A third determining module 503 is configured to use the first large model to determine at least one text unit from the text library that meets the similarity requirement with the text features of the first legal and regulatory text;

[0160] The generating module 504 is configured to adjust the first legal and regulatory text according to at least one text unit using the second large model to generate a second legal and regulatory text.

[0161] Figure 5 The text generation device can execute Figure 2 The implementation principle and technical effects of the text generation method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the text generation device in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0162] like Figure 6 , which is a schematic structural diagram of an embodiment of a computing device provided by the present application, the device may include a storage component 601 and a processing component 602;

[0163] The storage component 601 can be used to store one or more computer program instructions, wherein the one or more computer program instructions are called and executed by the processing component 602 to implement Figure 1 The text processing method shown, or Figure 2 The text generation method shown.

[0164] Of course, the above computing device may also include other components, such as input / output interfaces, communication components, etc.

[0165] The input / output interface provides an interface between the processing component and peripheral interface modules, which may be output devices, input devices, etc. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices.

[0166] It should be noted that the above computing device implements Figure 1 The text processing method shown or Figure 4In the case of the text generation method shown, it can be a physical device or an elastic computing host provided by a cloud computing platform, etc. It can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device.

[0167] The computing device described above can also be implemented as an electronic device. An electronic device can refer to a device used by a user that has the necessary functions for internet access, computing, and communication, such as a mobile phone, tablet computer, personal computer, or wearable device. It is understood that the electronic device described above may also include a display component, input / output interface, communication component, and other components, which will not be detailed here.

[0168] In one or more of the above embodiments, the processing component may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.

[0169] The storage component is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0170] The display component may be an electroluminescent (EL) element, a liquid crystal display or a micro display having a similar structure, or a direct retinal display or a similar laser scanning display.

[0171] The present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed, the computer program can realize Figure 1 The text processing method shown or Figure 4 The computer-readable medium may be included in the computing device described in the above embodiment, or may exist independently without being assembled into the computing device.

[0172] The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof.

[0173] The present application also provides a computer program product, which includes a computer program carried on a computer-readable storage medium, and when the computer program is executed by a computer, it can achieve Figure 1 The text processing method shown or Figure 4 The text generation method shown.

[0174] In such an embodiment, the computer program may be downloaded and installed from a network, and / or installed from a removable medium. When the computer program is executed by a processor, various functions defined in the system of the present application are performed.

[0175] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0176] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0177] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.

[0178] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A text processing method, characterized in that: include: Obtain multiple legal and regulatory texts of the target legal and regulatory type; For any legal or regulatory text, the legal or regulatory text is divided into multiple text units; Extracting text features corresponding to the multiple text units respectively; Building a target text library corresponding to the target legal and regulatory type based on the text features corresponding to the plurality of legal and regulatory texts; Among them, the target text library is used to use the first large model to determine at least one text unit whose text features meet the similarity requirements with the first legal and regulatory text of the target legal and regulatory type, and the at least one text unit is used to use the second large model to adjust the first legal and regulatory text to generate a second legal and regulatory text.

2. The method according to claim 1, characterized in that The step of dividing the legal text into a plurality of text units includes: According to the document format of the legal and regulatory text, the legal and regulatory text is divided into a plurality of first text segments corresponding to the respective text structures; The first text segment corresponding to the same text structure is divided into multiple second text segments according to preset symbols, and any second text segment is regarded as a text unit.

3. The method according to claim 2, characterized in that Also includes: If the number of bytes in the second text segment is greater than a first threshold, dividing the second text segment into a plurality of third text segments, where the number of bytes in the third text segments is less than the first threshold; For two adjacent third text fragments, adding preset bytes in the previous third text fragment to the latter third text fragment to obtain a plurality of fourth text fragments; wherein the number of the preset bytes is less than the second threshold; The taking any second text segment as a text unit includes: Any fourth text segment is regarded as a text unit.

4. The method according to claim 1, wherein Also includes: Using the third model to perform semantic analysis on the legal and regulatory text to obtain semantic data corresponding to the legal and regulatory text; extracting semantic features of the semantic data; The semantic features and the text features are stored in the target text library in correspondence; wherein, the semantic features are used to determine at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text of the target legal and regulatory type using the first large model.

5. The method according to claim 4, characterized in that Also includes: Get multiple case data; Determining at least one target case data corresponding to the legal and regulatory text from the plurality of case data using a fourth model; The at least one target case data is stored in the target text library; wherein, the at least one target case data is used to use the first large model to determine at least one text unit that meets the similarity requirements with the text features of the first legal and regulatory text of the target legal and regulatory type.

6. The method according to claim 1, characterized in that Also includes: For any legal text, use the fifth model to search the target text library for historical legal texts that are consistent with the content of the legal text, and determine the comparison results between the legal text and the historical legal texts; The comparison result is provided to the target user.

7. A text generation method, characterized in that: include: Responding to the text generation request, obtaining a first legal and regulatory text; Determining a text library corresponding to the legal and regulatory type of the first legal and regulatory text; wherein the text library is constructed by text features corresponding to a plurality of legal and regulatory texts corresponding to the legal and regulatory type, the text features being extracted from text units obtained by segmenting the legal and regulatory text; Determining, from the text library using the first large model, at least one text unit that meets similarity requirements with text features of the first legal and regulatory text; The first legal and regulatory text is adjusted according to the at least one text unit using the second large model to generate a second legal and regulatory text.

8. A computing device, characterized in that It includes a storage component and a processing component; the storage component stores one or more computer program instructions, the computer program instructions are called and executed by the processing component, and the processing component executes the one or more computer program instructions to implement the text processing method according to any one of claims 1 to 6, or the text generation method according to claim 7.

9. A computer-readable storage medium, characterized in that A computer program is stored, and the computer program is executed by a computer to implement the text processing method according to any one of claims 1 to 6, or the text generation method according to claim 7.

10. A computer program product, characterized in that A computer program is stored, and when the computer program is executed by a computer, the text processing method according to any one of claims 1 to 6 or the text generation method according to claim 7 is implemented.

Citation Information

Cited By

  • Legal text processing method, computer readable storage medium and electronic equipment

    CN122047228A