Rapid text generation method and device of LLM model based on key information extraction
Through the methods of key information extraction and trend calculation, the problem of user input content exceeding the limit when LLM model generates text, achieving efficient text generation and strong correlation.
Patent Information
- Application Number
- CN202510435463.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When generating text, the user input content is prone to exceeding the limit, resulting in reduced generation efficiency and poor correlation between generated text.
Using a method based on key information extraction, by obtaining user input text, preprocessing and extracting key information, calculating the trend between keywords and tokens sets, determining the correlation token, and then calling the LLM model to generate text.
It effectively avoids the problem of user input content exceeding the limit, improves the efficiency of text generation, and ensures strong correlation between the generated text and the user input content.
Smart Images

Figure CN119961437A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to text generation, and in particular to a method and device for rapid text generation based on an LLM model for key information extraction. Background Art
[0002] Large Language Model (LLM) is an artificial intelligence model based on deep learning technology, which is specially used for natural language processing (NLP) tasks such as understanding, generation, translation, and summarization. Such models usually contain billions to tens of billions of parameters and are trained with large-scale text data to capture the structure, grammar, semantics, and contextual relationships of the language.
[0003] Currently, when generating text through the LLM model, there is often a problem that the content generated directly by the large model deviates from the actual needs of the user.
[0004] At present, the solution to the above problem is often to generate text by fusing the user's own data, that is, by obtaining the user's input content, fusing private data based on the input content, and then generating text. However, this method still has certain disadvantages, because the LLM model has input tokens restrictions, and the user's input content often exceeds the input limit. In this case, the LLM model will remind the user to repeatedly delete and modify the input content, which not only affects the efficiency of text generation, but also affects the relevance of the final generated text because the user deletes too much content. Summary of the invention
[0005] The purpose of the present invention is to solve at least one of the deficiencies of the prior art and to provide a method and device for quickly generating text based on an LLM model of key information extraction.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: Specifically, a text rapid generation method based on the LLM model of key information extraction is proposed, including the following: Get the user's input text information; Preprocessing the input text information to obtain processed text; Extract key information from the processed text to obtain a keyword sequence; Calculating the degree of orientation between each keyword in the keyword sequence and an element in a tokens set; wherein the tokens set is obtained by pre-extracting tokens from a preset range of materials; For any keyword in the keyword sequence, randomly select one element from the tokens set whose tendency is higher than the first threshold as its associated token, and then obtain the associated tokens of all keywords in the keyword sequence; The associated tokens of all keywords in the keyword sequence call the LLM model for text generation.
[0007] Further, specifically, the input text information is preprocessed to obtain a processed text, including Perform word segmentation on the input text information to decompose the text into words, subwords or characters; Then remove stop words; Finally, the processed text is obtained by standardization.
[0008] Further, specifically, the processed text is subjected to key information extraction to obtain a keyword sequence, including: The processed text is vectorized through the BERT embedding algorithm, and the deep semantic features of the processed text are extracted to generate a semantic embedding vector; Based on the semantic embedding vector, the K-means clustering method is used to automatically group key information, divide the same type of key information according to the distance in the vector space, and generate key information classification labels; based on the key information classification labels, the TF-IDF extraction method is used to extract the keywords in the key information classification labels in turn to obtain the keyword sequence.
[0009] Furthermore, tokens are pre-extracted for the preset range of materials, including: Convert the acquired preset range material into a string to obtain string data, and record the string set composed of the string data corresponding to the preset range material as str_set, and record str_set(t) as the string data of the t-th preset range material in str_set; Execute the word segmentation algorithm on any str_set(t) to obtain multiple substrings corresponding to str_set(t), namely, word segments. The word segments are organized into sub-word sets in order. Any str_set(t) corresponds to a sub-word set, and all sub-word sets form a tokens set.
[0010] Further, specifically, calculating the degree of orientation between each keyword in the keyword sequence and the elements in the tokens set includes: Extracting features of each keyword in the keyword sequence and each element in the tokens set respectively by using a preset feature extraction algorithm to obtain keyword features and features of each element in the tokens set; Calculate the degree of trend between any keyword and each element in the tokens set based on the keyword features and the features of each element in the tokens set; Among them, the preset feature extraction algorithm is as follows: Add each keyword in the keyword sequence to the tokens set to obtain the updated tokens set, denoted as W_seg_set, where the sequence number of each keyword in W_seg_set is predetermined, W_seg_set(i) is the i-th element in W_seg_set, and the value range of i is 1-n. Each character of W_seg_set(i) is then converted into binary to obtain bin(i). If there are w characters in bin(i), then bin(i) is string-segmented to obtain w binary numbers, denoted as w(i). The w-dimensional array of W_seg_set(i) is denoted as w(i), and the set composed of all w(i) is w_set; w(i)_q is denoted as the element with sequence number q in w(i), q∈[1,w], then the array processing function is executed on w(i) get as follows, , The array Recorded as (i) and (i) represents the feature corresponding to W_seg_set(i), the serial number of the element in the feature is q, and the set of features corresponding to each word in the set W_seg_set is called Btqset. (i) The sequence number in the set Btqset is i.
[0011] Further, specifically, the degree of orientation between any keyword and each element in the tokens set is calculated based on the keyword features and the features of each element in the tokens set, including: Let variable s represent the serial number of the element in the set W_seg_set_q, the element with serial number s in W_seg_set_q is recorded as W_seg_set_q(s), k represents the total number of elements in the set W_seg_set_q, s∈[1,k], and the feature of the segmentation W_seg_set_q(s) in the set Btqset obtained by the serial number s of W_seg_set_q(s) is recorded as Btq(s); Set Seq_set. The number of elements in Seq_set is the same as the number of elements in str_set. The sequence number of the elements in Seq_set is t. The element with sequence number t in Seq_set is denoted as Seq_set(t). The trend between each participle is calculated as follows: S110, open a program; obtain W_seg_set_q; obtain the element Seq_set(t) with sequence number t in the set Seq_set, and clear the elements in Seq_set(t); S120, set the value of s to 1; set variable s2; set variable u, set the value of variable u to 0; S130, obtain Btq(s) through s; S141, let the value of s2 be the value of s; S142, increase the value of s2 by 1; S143, obtain Btq(s2) through s2; S144, define the function for calculating the degree of trend between the features of two word segments as ,but To calculate the degree of trend between Btq(s) and Btq(s2), record it as Rel(s,s2), record the element with sequence number t in the array Btq(s) as Btq(s)(t), and the element with sequence number t in the array Btq(s2) as Btq(s2)(t), the calculation formula of Rel(s,s2) is as follows: , Where t is the element number, w is the total number of dimensions of the array segmentation, and s is a predefined variable; S145, calculate the trend between W_seg_set_q(s) and other word segments in the set W_seg_set_q, set the variable s3 to represent the sequence number of the s2th element to the kth element in the set W_seg_set_q, and record the trend between W_seg_set_q(s) and the s2th to the kth word segments in the set W_seg_set_q as Rel(s,s2,k), and the calculation formula is: , The obtained Rel(s,s2,k) is the tendency between W_seg_set_q(s) and the s2th to kth word segments in the set W_seg_set_q; S146, the calculation formula for calculating the threshold u,u is as follows: , The above formula is the threshold value u; S147, let the trend degree of W_seg_set_q(s) and other word segments in the set W_seg_set_q be Seq(W_seg_set_q(s)), and the calculation formula of the trend degree Seq(W_seg_set_q(s)) is: , The obtained trend degree Seq (W_seg_set_q(s)) is recorded as Seq_q_s and added to the set Seq_set(t) as the element with sequence number s of Seq_set(t); S148, set the value of u to the value of Seq_q_s; go to S151; S151, determine whether the value of s is greater than or equal to k, if yes, go to S152, if no, increase the value of s by 1; go to S130; S152, taking the set Seq_set(q) as the element with sequence number q in the set Seq_set and saving it; ending the program; The element Seq_set(q) with sequence number q in the set Seq_set corresponds to the segmented word set W_seg_set_q obtained by segmenting the material data text str_set(q) with sequence number q in the set str_set. The element in the set Seq_set(q) is the tendency of the segmented word with the sequence number corresponding to the element.
[0012] The present invention also proposes a text rapid generation device based on the LLM model of key information extraction, comprising the following: A text information acquisition module is used to acquire the user's input text information; A preprocessing module, used for preprocessing the input text information to obtain processed text; A key information extraction module is used to extract key information from the processed text to obtain a keyword sequence; A trend degree calculation module, used to calculate the trend degree between each keyword in the keyword sequence and the elements in the tokens set; wherein the tokens set is obtained by pre-extracting tokens from a preset range of materials; An associated token determination module is used to randomly select one of the elements in the tokens set with a tendency higher than a first threshold as its associated token for any keyword in the keyword sequence, and then obtain the associated tokens of all the keywords in the keyword sequence; The text generation module is used to call the LLM model for text generation based on the associated tokens of all keywords in the keyword sequence.
[0013] The beneficial effects of the present invention are: The present invention proposes a method and device for rapid text generation based on an LLM model of key information extraction. When obtaining the user's input text information, the key information is first extracted from the input text information to obtain a keyword sequence, thereby avoiding the situation where the user's input content exceeds the limit; then, the associated token is determined according to the extracted keyword sequence through a preset trend degree calculation algorithm, and then the LLM model is called according to the associated token to complete text generation. The text generation method proposed by the present invention not only ensures the generation efficiency but also ensures the strong correlation between the generated text content and the user input content. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and other features of the present disclosure will become more apparent by describing in detail the embodiments shown in the accompanying drawings. The same reference numerals in the accompanying drawings of the present disclosure represent the same or similar elements. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure. For those skilled in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative work. In the accompanying drawings: Figure 1 Shown is a flow chart of the text rapid generation method of the LLM model based on key information extraction of the present invention. DETAILED DESCRIPTION
[0015] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention, so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments in this application and the features in the embodiments can be combined with each other without conflict. The same reference numerals used throughout the drawings indicate the same or similar parts.
[0016] Example 1, reference Figure 1 The present invention proposes a method for quickly generating text based on an LLM model of key information extraction, comprising the following steps: Step 110: Obtaining user input text information; Step 120: pre-process the input text information to obtain processed text; Step 130, extracting key information from the processed text to obtain a keyword sequence; Step 140: Calculate the degree of similarity between each keyword in the keyword sequence and an element in a tokens set; wherein the tokens set is obtained by pre-extracting tokens from a preset range of materials; Step 150: for any keyword in the keyword sequence, randomly select one element from the tokens set whose tendency is higher than the first threshold as its associated token, and then obtain the associated tokens of all keywords in the keyword sequence; Step 160: Call the LLM model to generate text based on the associated tokens of all keywords in the keyword sequence.
[0017] In this embodiment 1, when the user's input text information is obtained, the key information of the input text information is first extracted to obtain a keyword sequence, thereby avoiding the situation where the user's input content exceeds the limit; then, the associated token is determined according to the extracted keyword sequence through a preset trend calculation algorithm, and then the LLM model is called according to the associated token to complete the text generation. The text generation method proposed by the present invention not only ensures the generation efficiency but also ensures the strong correlation between the generated text content and the user input content.
[0018] As a preferred embodiment of the present invention, specifically, the input text information is preprocessed to obtain a processed text, including Perform word segmentation on the input text information to decompose the text into words, subwords or characters; Then remove stop words; Finally, the processed text is obtained by standardization.
[0019] As a preferred embodiment of the present invention, specifically, extracting key information from the processed text to obtain a keyword sequence includes: The processed text is vectorized through the BERT embedding algorithm, and the deep semantic features of the processed text are extracted to generate a semantic embedding vector; Based on the semantic embedding vector, the K-means clustering method is used to automatically group key information, divide the same type of key information according to the distance in the vector space, and generate key information classification labels; based on the key information classification labels, the TF-IDF extraction method is used to extract the keywords in the key information classification labels in turn to obtain the keyword sequence.
[0020] As a preferred embodiment of the present invention, tokens are pre-extracted from the preset range of materials, including: Convert the acquired preset range material into a string to obtain string data, and record the string set composed of the string data corresponding to the preset range material as str_set, and record str_set(t) as the string data of the t-th preset range material in str_set; Execute the word segmentation algorithm on any str_set(t) to obtain multiple substrings corresponding to str_set(t), namely, word segments. The word segments are organized into sub-word sets in order. Any str_set(t) corresponds to a sub-word set, and all sub-word sets form a tokens set.
[0021] In this preferred embodiment, the above method can be used to divide the text of the preset range material into multiple substrings for tokens pre-extraction.
[0022] As a preferred embodiment of the present invention, specifically, calculating the degree of orientation between each keyword in the keyword sequence and the elements in the tokens set includes: Extracting features of each keyword in the keyword sequence and each element in the tokens set respectively by using a preset feature extraction algorithm to obtain keyword features and features of each element in the tokens set; Calculate the degree of trend between any keyword and each element in the tokens set based on the keyword features and the features of each element in the tokens set; Among them, the preset feature extraction algorithm is as follows: Add each keyword in the keyword sequence to the tokens set to obtain the updated tokens set, denoted as W_seg_set, where the sequence number of each keyword in W_seg_set is predetermined, W_seg_set(i) is the i-th element in W_seg_set, and the value range of i is 1-n. Each character of W_seg_set(i) is then converted into binary to obtain bin(i). If there are w characters in bin(i), then bin(i) is string-segmented to obtain w binary numbers, denoted as w(i). The w-dimensional array of W_seg_set(i) is denoted as w(i), and the set composed of all w(i) is w_set; w(i)_q is denoted as the element with sequence number q in w(i), q∈[1,w], then the array processing function is executed on w(i) get as follows, , The array Recorded as (i) and (i) represents the feature corresponding to W_seg_set(i), the serial number of the element in the feature is q, and the set of features corresponding to each word in the set W_seg_set is called Btqset. (i) The sequence number in the set Btqset is i.
[0023] In this preferred embodiment, the above method can be used to extract features of each keyword in the keyword sequence and elements in the tokens set to obtain keyword features and features of each element in the tokens set.
[0024] As a preferred embodiment of the present invention, specifically, the trend degree between any keyword and each element in the tokens set is calculated based on the keyword feature and the feature of each element in the tokens set, including: Let variable s represent the serial number of the element in the set W_seg_set_q, the element with serial number s in W_seg_set_q is recorded as W_seg_set_q(s), k represents the total number of elements in the set W_seg_set_q, s∈[1,k], and the feature of the segmentation W_seg_set_q(s) in the set Btqset obtained by the serial number s of W_seg_set_q(s) is recorded as Btq(s); Set Seq_set. The number of elements in Seq_set is the same as the number of elements in str_set. The sequence number of the elements in Seq_set is t. The element with sequence number t in Seq_set is denoted as Seq_set(t). The trend between each participle is calculated as follows: S110, open a program; obtain W_seg_set_q; obtain the element Seq_set(t) with sequence number t in the set Seq_set, and clear the elements in Seq_set(t); S120, set the value of s to 1; set variable s2; set variable u, set the value of variable u to 0; S130, obtain Btq(s) through s; S141, let the value of s2 be the value of s; S142, increase the value of s2 by 1; S143, obtain Btq(s2) through s2; S144, define the function for calculating the degree of trend between the features of two word segments as ,but To calculate the degree of trend between Btq(s) and Btq(s2), record it as Rel(s,s2), record the element with sequence number t in the array Btq(s) as Btq(s)(t), and the element with sequence number t in the array Btq(s2) as Btq(s2)(t), the calculation formula of Rel(s,s2) is as follows: , Where t is the element number, w is the total number of dimensions of the array segmentation, and s is a predefined variable; S145, calculate the trend between W_seg_set_q(s) and other word segments in the set W_seg_set_q, set the variable s3 to represent the sequence number of the s2th element to the kth element in the set W_seg_set_q, and record the trend between W_seg_set_q(s) and the s2th to the kth word segments in the set W_seg_set_q as Rel(s,s2,k), and the calculation formula is: , The obtained Rel(s,s2,k) is the tendency between W_seg_set_q(s) and the s2th to kth word segments in the set W_seg_set_q; S146, the calculation formula for calculating the threshold u,u is as follows: , The above formula is the threshold value u; S147, let the trend degree of W_seg_set_q(s) and other word segments in the set W_seg_set_q be Seq(W_seg_set_q(s)), and the calculation formula of the trend degree Seq(W_seg_set_q(s)) is: , The obtained trend degree Seq (W_seg_set_q(s)) is recorded as Seq_q_s and added to the set Seq_set(t) as the element with sequence number s of Seq_set(t); S148, set the value of u to the value of Seq_q_s; go to S151; S151, determine whether the value of s is greater than or equal to k, if yes, go to S152, if no, increase the value of s by 1; go to S130; S152, taking the set Seq_set(q) as the element with sequence number q in the set Seq_set and saving it; ending the program; The element Seq_set(q) with sequence number q in the set Seq_set corresponds to the segmented word set W_seg_set_q obtained by segmenting the material data text str_set(q) with sequence number q in the set str_set. The element in the set Seq_set(q) is the tendency of the segmented word with the sequence number corresponding to the element.
[0025] In this preferred embodiment, the above method can accurately calculate the degree of orientation between each keyword in the keyword sequence and the elements in the tokens set.
[0026] The present invention also proposes a text rapid generation device based on the LLM model of key information extraction, comprising the following: A text information acquisition module is used to acquire the user's input text information; A preprocessing module, used for preprocessing the input text information to obtain processed text; A key information extraction module is used to extract key information from the processed text to obtain a keyword sequence; A trend degree calculation module, used to calculate the trend degree between each keyword in the keyword sequence and the elements in the tokens set; wherein the tokens set is obtained by pre-extracting tokens from a preset range of materials; An associated token determination module is used to randomly select one of the elements in the tokens set with a tendency higher than a first threshold as its associated token for any keyword in the keyword sequence, and then obtain the associated tokens of all the keywords in the keyword sequence; The text generation module is used to call the LLM model for text generation based on the associated tokens of all keywords in the keyword sequence.
[0027] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of software functional modules.
[0028] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or system that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0029] Although the description of the present invention has been quite detailed and specifically describes several described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be regarded as providing a broad possible interpretation of these claims in view of the prior art by reference to the appended claims, thereby effectively covering the intended scope of the present invention. In addition, the above description of the present invention is based on the embodiments foreseeable by the inventor, and its purpose is to provide a useful description, and those non-substantial changes to the present invention that have not yet been foreseen may still represent equivalent changes to the present invention.
[0030] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation. As long as the technical effect of the present invention is achieved by the same means, it should belong to the protection scope of the present invention. Within the protection scope of the present invention, its technical scheme and / or implementation method can have various modifications and changes.
Claims
1. A text rapid generation method based on the LLM model of key information extraction, characterized in that: Includes the following: Get the user's input text information; Preprocessing the input text information to obtain processed text; Extract key information from the processed text to obtain a keyword sequence; Calculating the degree of orientation between each keyword in the keyword sequence and an element in a tokens set; wherein the tokens set is obtained by pre-extracting tokens from a preset range of materials; For any keyword in the keyword sequence, randomly select one element from the tokens set whose tendency is higher than the first threshold as its associated token, and then obtain the associated tokens of all keywords in the keyword sequence; The associated tokens of all keywords in the keyword sequence call the LLM model for text generation.
2. The method for quickly generating text based on the LLM model of key information extraction according to claim 1 is characterized in that: Specifically, the input text information is preprocessed to obtain a processed text, including Perform word segmentation on the input text information to decompose the text into words, subwords or characters; Then remove stop words; Finally, the processed text is obtained by standardization.
3. The method for quickly generating text based on the LLM model of key information extraction according to claim 1 is characterized in that: Specifically, the processed text is subjected to key information extraction to obtain a keyword sequence, including: The processed text is vectorized through the BERT embedding algorithm, and the deep semantic features of the processed text are extracted to generate a semantic embedding vector; Based on the semantic embedding vector, the K-means clustering method is used to automatically group key information, divide the same type of key information according to the distance in the vector space, and generate key information classification labels; based on the key information classification labels, the TF-IDF extraction method is used to extract the keywords in the key information classification labels in turn to obtain the keyword sequence.
4. The method for quickly generating text based on the LLM model of key information extraction according to claim 1 is characterized in that: Pre-extract tokens from preset range materials, including: Convert the acquired preset range material into a string to obtain string data, and record the string set composed of the string data corresponding to the preset range material as str_set, and record str_set(t) as the string data of the t-th preset range material in str_set; Execute the word segmentation algorithm on any str_set(t) to obtain multiple substrings corresponding to str_set(t), namely, word segments. The word segments are organized into sub-word sets in order. Any str_set(t) corresponds to a sub-word set, and all sub-word sets form a tokens set.
5. The method for quickly generating text based on the LLM model of key information extraction according to claim 4 is characterized in that: Specifically, the degree of orientation between each keyword in the keyword sequence and the elements in the tokens set is calculated, including: Using a preset feature extraction algorithm, feature extraction is performed on each keyword in the keyword sequence and each element in the tokens set to obtain keyword features and features of each element in the tokens set; Based on the keyword features and the features of each element in the tokens set, the degree of trend between any keyword and each element in the tokens set is calculated; Among them, the preset feature extraction algorithm is as follows: Add each keyword in the keyword sequence to the tokens set to obtain the updated tokens set, denoted as W_seg_set, where the sequence number of each keyword in W_seg_set is predetermined, W_seg_set(i) is the i-th element in W_seg_set, and the value range of i is 1-n. Each character of W_seg_set(i) is then converted into binary to obtain bin(i). If there are w characters in bin(i), then bin(i) is string-segmented to obtain w binary numbers, denoted as w(i). The w-dimensional array of W_seg_set(i) is denoted as w(i), and the set composed of all w(i) is w_set; denoted as w(i)_q is the element with sequence number q in w(i), q∈[1,w], then the array processing function is executed on w(i) get as follows, , The array Recorded as (i) and (i) represents the feature corresponding to W_seg_set(i), the serial number of the element in the feature is q, and the set of features corresponding to each word in the set W_seg_set is called Btqset. (i) The sequence number in the set Btqset is i.
6. The method for quickly generating text based on the LLM model of key information extraction according to claim 5 is characterized in that: Specifically, the degree of orientation between any keyword and each element in the tokens set is calculated based on the keyword features and the features of each element in the tokens set, including: Let variable s represent the serial number of the element in the set W_seg_set_q, the element with serial number s in W_seg_set_q is recorded as W_seg_set_q(s), k represents the total number of elements in the set W_seg_set_q, s∈[1,k], and the feature of the segmentation W_seg_set_q(s) in the set Btqset obtained by the serial number s of W_seg_set_q(s) is recorded as Btq(s); Set Seq_set. The number of elements in Seq_set is the same as the number of elements in str_set. The sequence number of the elements in Seq_set is t. The element with sequence number t in Seq_set is denoted as Seq_set(t). The trend between each word is calculated as follows: S110, open a program; obtain W_seg_set_q; obtain the element Seq_set(t) with sequence number t in the set Seq_set, and clear the elements in Seq_set(t); S120, set the value of s to 1; set variable s2; set variable u, set the value of variable u to 0; S130, obtain Btq(s) through s; S141, let the value of s2 be the value of s; S142, increase the value of s2 by 1; S143, obtain Btq(s2) through s2; S144, define the function for calculating the degree of trend between the features of two word segments as ,but To calculate the degree of trend between Btq(s) and Btq(s2), record it as Rel(s,s2), record the element with sequence number t in the array Btq(s) as Btq(s)(t), and the element with sequence number t in the array Btq(s2) as Btq(s2)(t), the calculation formula of Rel(s,s2) is as follows: , Where t is the element number, w is the total number of dimensions of the array segmentation, and s is a predefined variable; S145, calculate the trend between W_seg_set_q(s) and other word segments in the set W_seg_set_q, set the variable s3 to represent the sequence number of the s2th element to the kth element in the set W_seg_set_q, and record the trend between W_seg_set_q(s) and the s2th to the kth word segments in the set W_seg_set_q as Rel(s,s2,k), and the calculation formula is: , The obtained Rel(s,s2,k) is the tendency between W_seg_set_q(s) and the s2th to kth word segments in the set W_seg_set_q; S146, the calculation formula for calculating the threshold u,u is as follows: , The above formula is the threshold value u; S147, let the trend degree of W_seg_set_q(s) and other word segments in the set W_seg_set_q be Seq(W_seg_set_q(s)), and the calculation formula of the trend degree Seq(W_seg_set_q(s)) is: , The obtained trend degree Seq (W_seg_set_q(s)) is recorded as Seq_q_s and added to the set Seq_set(t) as the element with sequence number s of Seq_set(t); S148, set the value of u to the value of Seq_q_s; go to S151; S151, determine whether the value of s is greater than or equal to k, if yes, go to S152, if no, increase the value of s by 1; go to S130; S152, taking the set Seq_set(q) as the element with sequence number q in the set Seq_set and saving it; ending the program; The element Seq_set(q) with sequence number q in the set Seq_set corresponds to the segmented word set W_seg_set_q obtained by segmenting the material data text str_set(q) with sequence number q in the set str_set. The element in the set Seq_set(q) is the tendency of the segmented word with the sequence number corresponding to the element.
7. A text rapid generation device based on the LLM model of key information extraction, characterized in that: Includes the following: A text information acquisition module is used to acquire the user's input text information; A preprocessing module, used for preprocessing the input text information to obtain processed text; A key information extraction module is used to extract key information from the processed text to obtain a keyword sequence; A trend degree calculation module, used to calculate the trend degree between each keyword in the keyword sequence and the elements in the tokens set; wherein the tokens set is obtained by pre-extracting tokens from a preset range of materials; An associated token determination module is used to randomly select one of the elements in the tokens set with a tendency higher than a first threshold as its associated token for any keyword in the keyword sequence, and then obtain the associated tokens of all the keywords in the keyword sequence; The text generation module is used to call the LLM model for text generation based on the associated tokens of all keywords in the keyword sequence.
Citation Information
Patent Citations
Threat intelligence early warning text analysis method and system based on big data
CN113627179A
Word and sentence vector-based online annotation insurance policy term viewing method, device and system
CN117993393A
Question answering method, device and equipment for business handling, medium and program product
CN118569874A
Speech abstract extraction method and device and computer readable storage medium
CN119400185A
Advertisement-based training method, device and equipment for generative recall large model
CN119646188A