Content tracing method and device based on key information extraction
By introducing key information extraction and content traceability mechanisms into the content generated by large-scale model, the problem of insufficient interpretability and traceability of content generated by large-scale model is solved, and the credibility and transparency of the model are improved.
Patent Information
- Application Number
- CN202510604767.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Large models have problems with interpretability and traceability when generating content, making it difficult for users to understand the source and basis of the content generated by the model, resulting in low credibility and transparency of the model.
The content traceability method based on key information extraction is adopted to trace and mark the source information of the content by obtaining the key information in the generated content (such as complete paragraphs, indicator data and key names) and matching it with the preset range material.
It improves the credibility and transparency of the model, enables users to clearly understand the source and basis of generated content, and enhances users' enthusiasm for using it.
Smart Images

Figure CN120197610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of content traceability, and particularly to a content traceability method and device based on key information extraction. Background Art
[0002] Large models generally refer to deep learning models with a large number of parameters (usually hundreds of millions to tens of billions of parameters). These models are trained with a large amount of data, enabling them to exhibit very strong generalization ability and transfer learning ability in various tasks. Large models are not only suitable for the optimization of specific tasks but can also be applied to other related fields through fine-tuning.
[0003] Currently, there are problems with the interpretability and traceability of large models when generating content. Since users cannot clearly understand the source, basis, and relevant evidence of the content generated by the model, the credibility and transparency of the model are not high, thereby affecting the user's enthusiasm for use. Summary of the Invention
[0004] The purpose of the present invention is to at least solve one of the deficiencies of the prior art, and provide a content traceability method and device based on key information extraction.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: Specifically, a content traceability method based on key information extraction is proposed, including the following: Obtain the generated content to be traced, and extract key information from the generated content. The key information includes complete paragraphs, metric data, and key names. The key names include personal names, place names, and organization names; Match the key names with the materials within a preset range to find the paragraph or sentence position of the materials that match the key names, and obtain the key name traceability result; Match the metric data with the materials within a preset range to find the paragraph or sentence position of the materials that match the metric data, and obtain the metric data traceability result; Perform semantic matching of the complete paragraphs with the materials within a preset range to find the materials and paragraphs that meet the preset conditions, and obtain the paragraph traceability result; Output the key name traceability result, the metric data traceability result, and the paragraph traceability result.
[0006] Further, specifically, the process of obtaining the key name traceability result includes Call the structured data extraction interface to extract personal names, place names, and organization names from the generated content to obtain the extracted key names; Mark the extracted key names with red circles; Perform numerical matching of the extracted key names with the materials within a preset range; If the match is successful, find the location of the material that matches the key name in the paragraph or sentence and mark the source information as the traceability result of the key name. If the match is unsuccessful, mark the traceability result of the key name as empty.
[0007] Further, specifically, obtain the traceability result of the indicator data, including Step 310: Call the structured data extraction interface to extract indicator data from the generated content. The indicator data includes the indicator name, the indicator value, and the text location in the generated content. Step 320: Determine whether indicator data is extracted. If so, go to Step 330; if not, mark the traceability result of the indicator data as empty. Step 330: Determine whether there is a cached indicator in the resource content of the preset range of materials. If so, go to Step 340; if not, extract indicators from the resource content of the preset range of materials. If indicators are extracted, cache the indicators according to the resource ID of the resource content and then go to Step 340. If no indicators are extracted, mark the traceability result of the indicator data as empty. Step 340: Match the indicator data with the cached indicators, and mark the corresponding resource ID of the cached indicators that match the indicator data as the source information as the traceability result of the indicator data.
[0008] Further, the method further includes Set a preset statement recognition interface, call the preset statement recognition interface, perform specific sentence recognition on the generated content to obtain a specific sentence recognition result. The specific sentences include preset statements and ancient poems, and mark the source information of the specific sentences according to the specific sentence recognition result.
[0009] Further, specifically, the process of obtaining the paragraph traceability result includes Perform semantic matching between the complete paragraph and the preset range of materials. Specifically, Loop through the comparison between the complete paragraph and the preset range of materials for paragraph similarity matching to obtain the first similarity matching result. Mark the paragraphs with the first similarity matching result exceeding the first threshold as similar paragraphs and mark the paragraph source information. If there is no first similarity matching result exceeding the first threshold, extract sentences from the complete paragraph to obtain complete sentences, perform sentence similarity matching between the complete sentences and the preset range of materials to obtain the second similarity matching result. Mark the sentences with the second similarity matching result exceeding the second threshold as similar sentences and mark the sentence source information. If there is no second similarity matching result exceeding the second threshold, perform user preselected word tracing to obtain the user preselected word tracing result.
[0010] Further, specifically, user preselected word tracing is performed to obtain a user preselected word tracing result, including Convert the preset range of materials into word forms to obtain multiple different words, and jointly form a set Ptc with these words; According to the user preselected words, calculate multiple word tendency sets of the user preselected words in the set Ptc; Calculate the average tendency degree of all elements in each word tendency set, and select the word tendency sets with the top N rankings in terms of the average tendency degree in the word tendency sets as the traced word tendency sets; Then the content in the preset range of materials corresponding to the traced word tendency sets is the user preselected word tracing result.
[0011] Further, specifically, converting the preset range of materials into word forms to obtain multiple different words, and jointly forming a set Ptc with these words, including Extract the text content in the preset range of materials by removing punctuation marks, and perform word segmentation on the text content through a classic word segmentation algorithm to obtain multiple different words, and jointly form a set Ptc with these words.
[0012] Further, specifically, according to the user preselected words, calculate multiple word tendency sets of each word in the set Ptc for the user preselected words, including Denote the set of user preselected words as set Querys, denote the number of elements in set Querys as n, and denote the serial number of the elements in set Querys as i, i ∈ [1, n], and there is Querys = {Q(1), Q(2), …, Q(n - 1), Q(n)}; Denote the number of elements in set Ptc as m, and denote the serial number of the elements in set Ptc as j, j ∈ [1, m], and there is Ptc = {Ptc(1), Ptc(2), …, Ptc(m - 1), Ptc(m)}; Denote the variable k as the array length of each word Ptc(j) in the set Ptc, denote the variable h as the serial number of the string in the word Ptc(j), Ptc(j) represents Ptc(j, h) represents the string with serial number h in the element with serial number j in the set Ptc, h ∈ [1, k], and there is Ptc(j) = [Ptc(j, 1), Ptc(j, 1), …, Ptc(j, k - 1), Ptc(j, k)]; Let the function Glv() be the function that calculates the word vector of the input string through the word embedding algorithm. Glv(Ptc(j,h)) represents the word vector obtained by the word embedding algorithm for the string with the serial number h in the element with the serial number j in the set Ptc. Denote G(j,h)=Glv(Ptc(j,h)). Glv(Q(i)) represents the word vector obtained by the word embedding algorithm for the string of the element with the serial number i in the set Querys. Denote Gq(i)= Glv(Q(i)). The variable q represents the q-th dimension of the word vector, the variable p represents the number of dimensions of the word vector, G(j,h)[q] represents the value of the q-th dimension of the word vector G(j,h), and Gq(i)[q] represents the value of the q-th dimension of the word vector Gq(i); The function Sim() represents calculating the tendency degree between the two input vectors. The function Sim(Gq(i),G(j,h)) represents calculating the tendency degree between the word vectors Gq(i) and G(j,h) through the function Sim(). The calculation formula for the tendency degree Sim(Gq(i),G(j,h)) is: ; Calculating multiple word tendency sets of each user preselected word in the set Querys among the words in the set Ptc includes the following steps S401, set the value of the variable i to 1; create an empty set Chianset, and the set Chianset has uniqueness and orderliness; go to S402; S402, obtain the element Q(i) with the serial number i in Querys; obtain Gq(i) of Q(i) through the function Glv(); go to S403; S403, set the value of the variable j to 1; go to S404; S404, obtain the element Ptc(j) with the serial number j in Ptc; create an empty array Simset; go to S405; S405, let the value of the variable h be 1; go to S406; S406, obtain the element Ptc(j,h) with the serial number h in Ptc(j); obtain G(j,h) of Ptc(j,h) through the function Glv(); go to S407; S407, obtain the tendency degree Sim(Gq(i),G(j,h)); add the tendency degree Sim(Gq(i),G(j,h)) to the array Simset; go to S408; S408, judge whether the constraint condition h≧k is satisfied. If so, go to S4081; if not, go to S4082; S4081. Calculate the arithmetic mean sim_avg of each element in the array Simset, and use the set of the sequence numbers of the elements in the array Simset that are greater than sim_avg as the set Seq. Use each element in the set Seq as the target sequence number, extract the elements with the target sequence numbers in Ptc(j) as the array Chain, and add the array Chain to the set Chianset. Then go to S409; S4082. Increase the value of h by 1. Then go to S406; S409. Set the value of h to 1. Then go to S410; S410. Determine whether the constraint condition j≥m is satisfied. If so, go to S411; if not, go to S4101; S4101. Increase the value of j by 1. Then go to S404; S411. Set the value of j to 1. Then go to S412; S412. Determine whether the constraint condition i≥n is satisfied. If so, go to S413; if not, go to S4121; S4121. Increase the value of i by 1. Then go to S402; S413. Obtain the set Chianset; Each array in the set Chianset is the word tendency set of each search ranking keyword in the corresponding set Querys. Denote the set of the obtained multiple word tendency sets as the set Litset.
[0013] The present invention also provides a device for content traceability based on key information extraction, which is characterized by including the following: A data acquisition module, configured to acquire the generated content to be subjected to content traceability, and extract key information from the generated content, where the key information includes a complete paragraph, index data, and key names, and the key names include personal names, place names, and institution names; A key name traceability module, configured to match the key name with the preset range of materials, find the position of the paragraph or sentence where the material consistent with the key name is located, and obtain the key name traceability result; An index data traceability module, configured to match the index data with the preset range of materials, find the position of the paragraph or sentence where the material consistent with the index data is located, and obtain the index data traceability result; A paragraph traceability module, configured to perform semantic matching on the complete paragraph with the preset range of materials, find the materials and paragraphs that meet the preset conditions, and obtain the paragraph traceability result; A result output module, configured to output the key name traceability result, the index data traceability result, and the paragraph traceability result.
[0014] The beneficial effects of the present invention are as follows: The present invention provides a content traceability method and device based on key information extraction. By extracting key information from the generated content and tracing the key names, index data, and complete paragraphs based on the extracted key information, a traceability mechanism is added to the content generated by the AI model, enabling users to clearly understand the source, basis, and relevant evidence of the model-generated content, thereby improving the credibility and transparency of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By describing the embodiments shown in the accompanying drawings in detail, the above and other features of the present disclosure will become more apparent. The same reference numerals in the drawings of the present disclosure denote the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 The flowchart of the content traceability method based on key information extraction according to the present invention is shown; Figure 2 The flowchart of complete paragraph traceability in an embodiment of the present invention is shown; Figure 3 The flowchart of preset statement traceability in an embodiment of the present invention is shown; Figure 4 The flowchart of key name traceability in an embodiment of the present invention is shown; Figure 5 The flowchart of index traceability in an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The concept, specific structure, and technical effects of the present invention will be clearly and completely described below in combination with the embodiments and the accompanying drawings to fully understand the purpose, solution, and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The same reference numerals used throughout the drawings denote the same or similar parts.
[0017] Embodiment 1, referring to Figure 1 , the present invention provides a content traceability method based on key information extraction, including the following: Step 110: Obtain the generated content to be traced for content, and extract key information from the generated content. The key information includes complete paragraphs, index data, and key names, and the key names include personal names, place names, and institution names; Step 120: Match the key names with the preset range of materials, find the paragraph or sentence position where the materials consistent with the key names are located, and obtain the key name traceability result; Step 130: Match the indicator data with the preset range of materials, find the location of the paragraph or sentence where the material consistent with the indicator data is located, and obtain the result of tracing the indicator data; Step 140: Perform semantic matching on the complete paragraph with the preset range of materials, find the materials and paragraphs that meet the preset conditions, and obtain the result of tracing the paragraph; Step 150: Output the result of tracing the key name, the result of tracing the indicator data, and the result of tracing the paragraph.
[0018] In this Embodiment 1, by extracting key information from the generated content and tracing the key name, indicator data, and complete paragraph based on the extracted key information, and adding a tracing mechanism to the content generated by the AI model, users can clearly understand the source, basis, and relevant evidence of the content generated by the model, thereby improving the credibility and transparency of the model. In addition, for the traced content, it is marked at the position of the original material content in a highlighted manner for users to consult.
[0019] As a preferred embodiment of the present invention, specifically, the process of obtaining the result of tracing the key name includes Calling the structured data extraction interface to extract personal names, place names, and organization names from the generated content to obtain the extracted key names; Marking the extracted key names with red circles; Performing numerical matching on the extracted key names with the preset range of materials; If the matching is successful, find the location of the paragraph or sentence where the material consistent with the key name is located and mark the source information as the result of tracing the key name; If the matching is unsuccessful, mark the result of tracing the key name as empty.
[0020] In this preferred embodiment, first call the structured data extraction interface to extract personal names, place names, and organization names from the generated content; mark the extracted data with red circles; based on the extracted data, match the resource content of the specified range of materials, and if it exists, mark the source information, and the marked content includes: resource title, resource library identifier, and release date.
[0021] As a preferred embodiment of the present invention, specifically, obtaining the result of tracing the indicator data includes Step 310: Call the structured data extraction interface to extract indicator data from the generated content, where the indicator data includes the indicator name, indicator value, and text position in the generated content; Step 320: Determine whether the indicator data is extracted. If so, go to Step 330; if not, mark the result of tracing the indicator data as empty; Step 330: Determine whether there is a cache index for the resource content in the preset range of materials. If so, go to Step 340; if not, extract the index for the resource content in the preset range of materials. If an index is extracted, perform caching based on the resource ID of the resource content to obtain a cache index and then go to Step 340. If no index is extracted, mark the index data traceability result as empty. Step 340: Match the index data with the cache index, and mark the corresponding resource ID of the cache index that matches the index data successfully as the index data traceability result.
[0022] In this preferred embodiment, first call the large model interface to extract the indexes of the generated content. The extracted information includes the index name, index value, and the text position where it is located. Then extract the indexes of the material resource content. If an index is extracted, perform caching based on the resource ID. Compare and match the indexes of the generated content and the material resources. If the index name and index value are the same, mark the source information of the index, including: resource ID, resource title, resource library identifier, release date, and the text position where the index is located.
[0023] In addition, considering the time-consuming problem of the large model interface, a caching mechanism is adopted to cache the index data of the specified materials (persist to the database). That is, first obtain the cache index according to the resource ID of the material. If it cannot be obtained, then call the large model interface to extract the index. After extracting the index, perform caching according to the resource ID. For the problem of multiple specified materials, an extraction and comparison method is adopted one by one. That is, first pass in the generated content and the first material, perform index extraction and comparison, and render the source of the matched index. The second time, pass in the indexes of the generated content that have been extracted and the second material, perform index extraction and comparison, and then render the source of the newly matched index. Then continue with the next one until the comparison of all specified materials is completed.
[0024] In a specific application, the structure of the resource index cache table is shown in Table 1 below.
[0025] Table 1 Refer to Figure 3 , as a preferred embodiment of the present invention, the method further includes Set a preset statement recognition interface, call the preset statement recognition interface, perform specific sentence recognition on the generated content to obtain a specific sentence recognition result. The specific sentences include preset statements and ancient poems, and mark the source information of the specific sentences according to the specific sentence recognition result.
[0026] In this preferred embodiment, a preset statement recognition interface is called to identify specific sentences in the generated content, including preset statements, ancient Chinese poems, etc., and the source information of the specific sentences is marked. The marked content includes: resource ID, resource title, resource library identifier, and release date.
[0027] As a preferred embodiment of the present invention, specifically, the process of obtaining the paragraph traceability result includes Semantically matching the complete paragraph with the preset range of materials. Specifically, Circularly compare the complete paragraph with the preset range of materials for paragraph similarity matching to obtain the first similarity matching result. Mark the paragraphs with the first similarity matching result exceeding the first threshold as similar paragraphs and mark the paragraph source information. If there is no first similarity matching result exceeding the first threshold, extract sentences from the complete paragraph to obtain complete sentences, perform sentence similarity matching between the complete sentences and the preset range of materials to obtain the second similarity matching result, and mark the sentences with the second similarity matching result exceeding the second threshold as similar sentences and mark the sentence source information; If there is no second similarity matching result exceeding the second threshold, perform user preselected word tracing to obtain the user preselected word tracing result.
[0028] In this preferred embodiment, first, extract paragraphs (separated by line breaks) from the generated content, and paragraphs with more than 20 characters form a list of generated paragraphs; then extract paragraphs from the resource content of the specified range of materials, and paragraphs with more than 20 characters form a list of material paragraphs; extract sentences (separated by commas and periods) from the resources of the specified range of materials, and sentences with more than 20 characters form a list of material sentences; circularly generate the list of paragraphs, and perform similarity comparison with the list of resource paragraphs to obtain the paragraph with the highest similarity and exceeding the specified threshold (tentatively 80%), then mark the source information of the paragraph. The marked content includes: resource ID, resource title, resource library identifier, release date, and content of the similar paragraph; if no similar paragraph is found, then extract sentences from the generated paragraph content and perform similarity comparison with the material sentences to find the sentence with the highest similarity and exceeding the specified threshold, then mark the source information of the sentence. The marked content: resource ID, resource title, resource library identifier, release date, and content of the similar sentence. In addition, considering that even if the traceable content cannot be found through the similarity comparison method, the user may still have doubts. Therefore, a traceability algorithm for user preselected words is added to perform fuzzy traceability of related words through user preselected words to find the traceability result as much as possible.
[0029] The following is a preferred embodiment of the partial interface design for paragraph traceability, as shown in Table II below.
[0030] Table II As a preferred embodiment of the present invention, specifically, obtaining the user preselected word trace result through user preselected word tracing includes: Converting the preset range of materials into word forms to obtain a plurality of different words, and jointly forming a set Ptc with these words; According to the user preselected words, calculating a plurality of word tendency sets of the user preselected words in the set Ptc; Calculating the average tendency degree of all elements in each word tendency set, and selecting the word tendency sets with the top N rankings in terms of the average tendency degree in the word tendency sets as the traced word tendency sets; Then the content in the preset range of materials corresponding to the traced word tendency sets is the user preselected word trace result.
[0031] In this preferred embodiment, by converting the tokens in the preset range of materials into word forms and finding the words that are most likely to have a trace relationship with the user preselected words by calculating the word tendency degree and forming word tendency sets based on the association relationship, and then using the content in the preset range of materials corresponding to the word tendency sets with the top N rankings in terms of the average tendency degree as the user preselected word trace result for jeguo feedback, it can accurately trace the user's doubtful keywords and meet the user's usage requirements.
[0032] As a preferred embodiment of the present invention, specifically, converting the preset range of materials into word forms to obtain a plurality of different words, and jointly forming a set Ptc with these words, includes: Extracting the text content in the preset range of materials by removing punctuation marks, and performing word segmentation processing on the text content through a classic word segmentation algorithm to obtain a plurality of different words, and jointly forming a set Ptc with these words.
[0033] In this preferred embodiment, the preset range of materials is quickly converted into word form tokens through an electrostatic word segmentation algorithm.
[0034] As a preferred embodiment of the present invention, specifically, according to the user preselected words, calculating a plurality of word tendency sets of each word in the set Ptc for the user preselected words, includes: Denoting the set of user preselected words as set Querys, denoting the number of elements in set Querys as n, and the serial number of the elements in set Querys as i, i ∈ [1, n], and having Querys = {Q(1), Q(2), …, Q(n - 1), Q(n)}; Denoting the number of elements in set Ptc as m, and the serial number of the elements in set Ptc as j, j ∈ [1, m], and having Ptc = {Ptc(1), Ptc(2), …, Ptc(m - 1), Ptc(m)}; Let the variable k represent the array length of each word Ptc(j) in the set Ptc, the variable h represent the serial number of the string in the word Ptc(j), Ptc(j) represents Ptc(j,h) which represents the string with serial number h in the element with serial number j in the set Ptc, h ∈ [1, k], and Ptc(j) = [Ptc(j,1), Ptc(j,1), …, Ptc(j,k - 1), Ptc(j,k)]; Let the function Glv() be the function that calculates the word vector of the input string through the word embedding algorithm. Glv(Ptc(j,h)) represents the word vector obtained by the word embedding algorithm for the string with serial number h in the element with serial number j in the set Ptc. Denote G(j,h) = Glv(Ptc(j,h)). Glv(Q(i)) represents the word vector obtained by the word embedding algorithm for the string of the element with serial number i in the set Querys. Denote Gq(i) = Glv(Q(i)). The variable q represents the q-th dimension of the word vector, the variable p represents the number of dimensions of the word vector, G(j,h)[q] represents the value of the q-th dimension of the word vector G(j,h), and Gq(i)[q] represents the value of the q-th dimension of the word vector Gq(i); The function Sim() represents calculating the tendency degree between two input vectors. The function Sim(Gq(i),G(j,h)) represents calculating the tendency degree between the word vectors Gq(i) and G(j,h) through the function Sim(). The calculation formula for the tendency degree Sim(Gq(i),G(j,h)) is: ; Calculating multiple word tendency sets of each user preselected word in the set Querys in each word in the set Ptc includes the following steps S401, set the value of the variable i to 1; create an empty set Chianset, and the set Chianset has uniqueness and orderliness; go to S402; S402, obtain the element Q(i) with serial number i in Querys; obtain Gq(i) of Q(i) through the function Glv(); go to S403; S403, set the value of the variable j to 1; go to S404; S404, obtain the element Ptc(j) with serial number j in Ptc; create an empty array Simset; go to S405; S405, let the value of the variable h be 1; go to S406; S406, obtain the element Ptc(j,h) with serial number h in Ptc(j); obtain G(j,h) of Ptc(j,h) through the function Glv(); go to S407; S407, Obtain the tendency degree Sim(Gq(i), G(j, h)); Add the tendency degree Sim(Gq(i), G(j, h)) to the array Simset; Go to S408; S408, Determine whether the constraint condition h ≧ k is satisfied. If so, go to S4081; if not, go to S4082; S4081, Calculate the arithmetic mean sim_avg of each element in the array Simset. Use the set of the indices of the elements in the array Simset that are greater than sim_avg as the set Seq. Use each element in the set Seq as the target index, and extract the elements with the target indices from Ptc(j) as the array Chain. Add the array Chain to the set Chianset; Go to S409; S4082, Increase the value of h by 1; Go to S406; S409, Set the value of h to 1; Go to S410; S410, Determine whether the constraint condition j ≥ m is satisfied. If so, go to S411; if not, go to S4101; S4101, Increase the value of j by 1; Go to S404; S411, Set the value of j to 1; Go to S412; S412, Determine whether the constraint condition i ≥ n is satisfied. If so, go to S413; if not, go to S4121; S4121, Increase the value of i by 1; Go to S402; S413, Obtain the set Chianset; Each array in the set Chianset is the word tendency set of each search ranking keyword in the corresponding set Querys. Denote the set of the obtained multiple word tendency sets as the set Litset.
[0035] In addition, some of the content is explained as follows: 1. Complete paragraph: Refers to the expression of a complete thought or idea, usually with Chinese full stops, commas, and generally more than 20 characters.
[0036] 2. Specific sentence: Sentences within a specified range, such as preset statements, ancient poems, etc.
[0037] 3. Index data: Refers to quantitative data used to measure, evaluate, or compare the state, performance, effect, etc. of things in a specific field or industry. For example, in economics, GDP growth rate, unemployment rate, inflation rate, etc. are all important economic index data; in medicine, blood pressure, heart rate, blood sugar, etc. are all important health index data.
[0038] 4. Key names: personal names, place names, and organization names.
[0039] The present invention also provides a content traceability device based on key information extraction, which is characterized by including the following: A data acquisition module, configured to acquire the generated content to be traced for content, and extract key information from the generated content, where the key information includes complete paragraphs, index data, and key names, and the key names include personal names, place names, and organization names; A key name traceability module, configured to match the key names with the materials within a preset range, find the position of the paragraph or sentence where the material consistent with the key name is located, and obtain the key name traceability result; An index data traceability module, configured to match the index data with the materials within a preset range, find the position of the paragraph or sentence where the material consistent with the index data is located, and obtain the index data traceability result; A paragraph traceability module, configured to perform semantic matching between the complete paragraphs and the materials within a preset range, find the materials and paragraphs that meet the preset conditions, and obtain the paragraph traceability result; A result output module, configured to output the key name traceability result, the index data traceability result, and the paragraph traceability result.
[0040] In addition, in each embodiment of the present invention, each functional module can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0041] If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or system, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0042] Although the description of the present invention has been quite detailed and particularly describes several of the described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but rather should be regarded as providing a broad interpretation of these claims in light of the prior art by reference to the appended claims, thereby effectively covering the intended scope of the present invention. In addition, the present invention is described above in terms of embodiments foreseeable by the inventors for the purpose of providing a useful description, and non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.
[0043] As described above, these are only the preferred embodiments of the present invention. The present invention is not limited to the above-described embodiments. As long as the same means are used to achieve the technical effects of the present invention, they should fall within the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and variations may be made to its technical solutions and / or embodiments.
Claims
1. A content tracing method based on key information extraction, characterized in that: Includes the following: Obtaining generated content to be traced, and extracting key information from the generated content, wherein the key information includes complete paragraphs, indicator data, and key names, wherein the key names include names of people, places, and institutions; Match the key name with the preset range of materials, find the paragraph or sentence position of the material that matches the key name, and obtain the key name tracing result; Match the indicator data with the preset range of materials, find the paragraph or sentence position of the material consistent with the indicator data, and obtain the indicator data tracing result; Semantically match the complete paragraph with the preset range of materials, find the materials and paragraphs that meet the preset conditions, and obtain the paragraph tracing results; Output key name tracing results, indicator data tracing results and paragraph tracing results.
2. The content tracing method based on key information extraction according to claim 1 is characterized in that: Specifically, the process of obtaining the key name tracing result includes: Calling a structured data extraction interface to extract names of people, places, and institutions from the generated content to obtain extracted key names; Mark the extracted key names with red circles; Match the extracted key names with the preset range of materials; If the match is successful, the paragraph or sentence position of the material that matches the key name is found and the source information is marked as the key name tracing result; If the match fails, the key name tracing result will be marked as empty.
3. The content tracing method based on key information extraction according to claim 1 is characterized in that: Specifically, the indicator data traceability results are obtained, including: Step 310: Calling a structured data extraction interface to extract index data from the generated content, wherein the index data includes an index name, an index value, and a text position of the generated content; Step 320, determine whether the indicator data is extracted, if so, go to step 330, if not, mark the indicator data tracing result as empty; Step 330, determine whether there is a cache indicator for the resource content in the preset range material, if yes, go to step 340, if no, extract the indicator for the resource content in the preset range material, if the indicator is extracted, cache it according to the resource ID of the resource content to obtain the cache indicator, then go to step 340, if no indicator is extracted, mark the indicator data tracing result as empty; Step 340: Match the indicator data with the cache indicator, and mark the corresponding resource ID of the cache indicator that successfully matches the indicator data with source information as the indicator data tracing result.
4. The content tracing method based on key information extraction according to claim 1 is characterized in that: The method further comprises, A preset sentence recognition interface is set, the preset sentence recognition interface is called, and specific sentence recognition is performed on the generated content to obtain a specific sentence recognition result, wherein the specific sentence includes preset sentences and ancient poems, and the source information of the specific sentence is marked according to the specific sentence recognition result.
5. The content tracing method based on key information extraction according to claim 1 is characterized in that: Specifically, the process of obtaining the paragraph tracing result includes: Semantically match the complete paragraph with the preset range of materials. Specifically, Circularly compare the complete paragraph with the preset range of materials to perform paragraph similarity matching to obtain a first similarity matching result, mark the paragraphs whose first similarity matching results exceed a first threshold as similar paragraphs and mark the paragraph source information, If there is no first similarity matching result exceeding the first threshold, extract sentences from the complete paragraph to obtain complete sentences, perform sentence similarity matching on the complete sentences and the preset range of materials to obtain second similarity matching results, and mark sentences whose second similarity matching results exceed the second threshold as similar sentences and mark the sentence source information; If there is no second similarity matching result exceeding the second threshold, the user pre-selected word tracing is performed to obtain the user pre-selected word tracing result.
6. The content tracing method based on key information extraction according to claim 5 is characterized in that: Specifically, the user pre-selected word tracing is performed to obtain the user pre-selected word tracing result, including: Convert the preset range material into word form to obtain multiple different words, and use these words to form a set Ptc; According to the user's pre-selected words, multiple word tendency sets of the user's pre-selected words in the set Ptc are calculated; Calculate the average tendency degree of all elements in each word tendency set, and select the word tendency set whose average tendency degree is ranked in the top N items in the word tendency set as the traceable word tendency set; Then the content in the preset range material corresponding to the traced word tendency set is the user's pre-selected word tracing result.
7. The content tracing method based on key information extraction according to claim 6 is characterized in that: Specifically, the preset range material is converted into a word form to obtain a plurality of different words, and these words together form a set Ptc. include, The text content in the preset range material is extracted by removing punctuation marks, and the text content is segmented by a classic word segmentation algorithm to obtain multiple different words, which together form the set Ptc.
8. The content tracing method based on key information extraction according to claim 7 is characterized in that: Specifically, according to the user's pre-selected words, multiple word tendency sets of the user's pre-selected words in the set Ptc are calculated, including: Assume that the user selects n pre-selected words The set of user pre-selected words is denoted as set Querys, the number of elements in set Querys is n, the sequence number of elements in set Querys is i, i∈[1,n], and Querys={Q(1),Q(2),…,Q(n-1),Q(n)}; Let the number of elements in set Ptc be m, the sequence number of elements in set Ptc be j, j∈[1,m], and Ptc={Ptc(1),Ptc(2),…,Ptc(m-1),Ptc(m)}; Let variable k represent the array length of each word Ptc(j) in the set Ptc, variable h represent the serial number of the string in the word Ptc(j), Ptc(j) represents Ptc(j,h) represents the string with serial number h in the element with serial number j in the set Ptc, h∈[1,k], and Ptc(j)=[Ptc(j,1),Ptc(j,1),…, Ptc(j,k-1), Ptc(j,k)]; Let function Glv() be a function that calculates the word vector of an input string through a word embedding algorithm, Glv(Ptc(j,h)) represents the word vector of the string with sequence number h in the element with sequence number j in the set Ptc obtained through the word embedding algorithm, let G(j,h)=Glv(Ptc(j,h)), Glv(Q(i)) represents the word vector of the string with sequence number i in the set Querys obtained through the word embedding algorithm, let Gq(i)=Glv(Q(i)), variable q represents the qth dimension of the word vector, variable p represents the number of dimensions of the word vector, G(j,h)[q] represents the value of the qth dimension of the word vector G(j,h), and Gq(i)[q] represents the value of the qth dimension of the word vector Gq(i); The function Sim() indicates the calculation of the tendency degree between the two input vectors. The function Sim(Gq(i), G(j, h)) indicates the calculation of the tendency degree between the word vectors Gq(i) and G(j, h) through the function Sim(). The calculation formula of the tendency degree Sim(Gq(i), G(j, h)) is: ; Calculate multiple word preference sets of each user pre-selected word in the set Querys among the words in the set Ptc, including the following steps S401, set the value of variable i to 1; create an empty set Chianset, which has mutual difference and order; go to S402; S402, obtain the element Q(i) with the serial number i in Querys; obtain Gq(i) of Q(i) through the function Glv(); go to S403; S403, set the value of variable j to 1; go to S404; S404, obtaining the element Ptc(j) with sequence number j in Ptc; Create an empty array Simset; go to S405; S405, set the value of variable h to 1; go to S406; S406, obtain the element Ptc(j,h) with sequence number h in Ptc(j); obtain G(j,h) of Ptc(j,h) through function Glv(); go to S407; S407, obtain the tendency degree Sim(Gq(i), G(j, h)); add the tendency degree Sim(Gq(i), G(j, h)) to the array Simset; go to S408; S408, determine whether the constraint condition h≧k is satisfied, if yes, go to S4081, if no, go to S4082; S4081, calculate the arithmetic mean sim_avg of each element in the array Simset, take the set of sequence numbers of each element in the array Simset whose value is greater than sim_avg as the set Seq; take each element in the set Seq as the target sequence number, extract the element with the target sequence number in Ptc(j) as the array Chain, and add the array Chain to the set Chianset; Go to S409; S4082, increase the value of h by 1; go to S406; S409, set the value of h to 1; go to S410; S410, determine whether the constraint condition j≥m is satisfied, if yes, go to S411, if not, go to S4101; S4101, increase the value of j by 1; go to S404; S411, set the value of j to 1; Go to S412; S412, determine whether the constraint condition i≥n is satisfied, if yes, go to S413, if no, go to S4121; S4121, increase the value of i by 1; go to S402; S413, obtain the set Chinat; Each array in the set Chinat is the word tendency set of each search ranking keyword in the corresponding set Querys, and the set of multiple word tendency sets obtained is recorded as the set Litset.
9. A device for content tracing based on key information extraction, characterized in that: Includes the following: A data acquisition module, used to acquire generated content to be traced, and extract key information from the generated content, wherein the key information includes complete paragraphs, indicator data, and key names, wherein the key names include names of people, places, and institutions; The key name tracing module is used to match the key name with the preset range of materials, find the paragraph or sentence position of the material consistent with the key name, and obtain the key name tracing result; The indicator data tracing module is used to match the indicator data with the preset range of materials, find the paragraph or sentence position of the material consistent with the indicator data, and obtain the indicator data tracing result; The paragraph tracing module is used to semantically match the complete paragraph with the preset range of materials, find the materials and paragraphs that meet the preset conditions, and obtain the paragraph tracing results; The result output module is used to output the key name tracing results, indicator data tracing results and paragraph tracing results.
Citation Information
Patent Citations
Article key information tracing method and system, readable medium and equipment
CN114661868A
Multi-source data difference traceability retrieval method based on knowledge graph
CN115809345A
Question and answer result tracing method and device, equipment, medium and program product
CN117909451A
Document comparison tracing method and device and computer storage medium
CN119808752A
Credit assessment method for construction enterprises
WO2024108973A1