A method and system for tracing unstructured documents

By identifying and comparing the structured features of unstructured documents, the problem of difficult to trace unstructured documents in the prior art is solved, and effective traceability and management of documents are achieved.

CN119670751BActive Publication Date: 2025-06-17JIANGSU DAOYUNYIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411740159.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-06-17
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively trace unstructured documents, mainly due to their complexity and diversity, which cannot be effectively managed and stored.

Method used

By identifying structured features of unstructured documents, including document topics and subtopics, using named entity recognition models and semantic similarity calculations, the features are compared to determine whether the document comes from the target system.

Benefits of technology

It realizes effective traceability of unstructured documents, can accurately determine whether the document comes from the target system, and improves the efficiency of document management and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670751B_ABST
    Figure CN119670751B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for tracing unstructured documents. Among them, the method for tracing unstructured documents includes the following steps: S1, identifying the first structured feature of the first unstructured document to be traced, and identifying the second structured feature of the second unstructured document in the target system; S2, comparing the first structured feature with the second structured feature; S3, judging whether the first unstructured document is derived from the target system according to the result of the feature comparison. According to the method for tracing unstructured documents of the present invention, by identifying the structured features of unstructured documents, the unstructured documents can be effectively traced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unstructured document traceability, and particularly to an unstructured document traceability method and an unstructured document traceability system. Background Art

[0002] In related technologies, most traceability is for structured documents. And because unstructured documents are more complex to manage than structured documents, have more diverse storage media, and different data forms, effective traceability cannot be carried out for unstructured documents. Summary of the Invention

[0003] To solve the above technical problems, the present invention provides an unstructured document traceability method, which can effectively trace unstructured documents by identifying the structured features of unstructured documents.

[0004] The technical solution adopted by the present invention is as follows:

[0005] An unstructured document traceability method includes the following steps: S1, identifying the first structured features of a first unstructured document to be traced and the second structured features of a second unstructured document in a target system; S2, comparing the first structured features with the second structured features; S3, judging whether the first unstructured document is derived from the target system according to the result of the feature comparison.

[0006] In an embodiment of the present invention, the first structured features include a first document theme and a first document sub-theme, and the second structured features include a second document theme and a second document sub-theme. Wherein, step S2 specifically includes: S21, calculating a first similarity between the first document theme and the second document theme; S22, judging whether the first similarity is greater than or equal to a first threshold; S23, if the first similarity is greater than or equal to the first threshold, calculating a second similarity between the first document sub-theme and the second document sub-theme; S24, judging whether the second similarity is greater than or equal to a second threshold.

[0007] In an embodiment of the present invention, identifying the first structured features of the first unstructured document in step S1 includes the following steps: S11, obtaining a named entity recognition model and identifying the named entities of the first unstructured document according to the named entity recognition model; S12, identifying a first document theme containing the named entities in the first structured document based on the named entities; S13, obtaining the corresponding first document sub-theme according to the first document theme.

[0008] In one embodiment of the present invention, obtaining the named entity recognition model in step S11 includes: S101, obtaining a set of training documents containing named entity features from a historical document database; S102, performing text segmentation processing on each training document in the set of training documents to obtain a first vocabulary set corresponding to each training document, and classifying the first vocabulary set into a named entity set and a non-named entity set according to the named entity features; S103, inputting the set of training documents into a named entity recognition network, and outputting a first probability value of predicting each vocabulary in the named entity set corresponding to each training document in the set of training documents as a named entity, and a second probability value of predicting each vocabulary in the non-named entity set as a named entity; S104, obtaining a first cost function according to the first probability value and the second probability value, and adjusting each parameter in the named entity network based on the first cost function to obtain the named entity recognition model.

[0009] In one embodiment of the present invention, step S104 specifically includes: S141, generating a second cost function according to the first probability value and the second probability value through the following formula:

[0010]

[0011] where D m represents the second cost function, x ka is the first probability value corresponding to the a-th vocabulary in the named entity set of the k-th training document, y kb is the second probability value corresponding to the b-th vocabulary in the non-named entity set of the k-th training document, M is the number of training documents in the set of training documents, A is the number of vocabularies in the named entity set of the k-th training document, and B is the number of vocabularies in the non-named entity set of the k-th training document;

[0012] S142, calibrating the second cost function, and generating the calibrated first cost function through the following formula:

[0013]

[0014] where D n represents the first cost function, x k_f is the mean value of the first probability values of the f vocabulary with the smallest values in the named entity set of the k-th training document, x k_g is the mean value of the second probability values of the g vocabulary with the largest values in the named entity set of the k-th training document, and c and d are constants.

[0015] An unstructured document traceability system includes: an identification module for identifying the first structured features of a first unstructured document to be traced and the second structured features of a second unstructured document in a target system; a feature comparison module for comparing the first structured features with the second structured features; and a traceability module for determining whether the first unstructured document is derived from the target system according to the feature comparison result.

[0016] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned unstructured document traceability method is implemented.

[0017] A non-transitory computer-readable storage medium stores a computer program, and when the program is executed by a processor, the above-mentioned unstructured document traceability method is implemented.

[0018] Advantages of the present invention:

[0019] By identifying the structured features of unstructured documents, the present invention can effectively trace unstructured documents. Description of the drawings

[0020] Figure 1 It is a flowchart of the unstructured document traceability method according to an embodiment of the present invention;

[0021] Figure 2 It is a block diagram of the unstructured document traceability system according to an embodiment of the present invention. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] As Figure 1 shown, the unstructured document traceability method according to an embodiment of the present invention may include the following steps:

[0024] S1. Identify the first structured features of a first unstructured document to be traced and the second structured features of a second unstructured document in a target system.

[0025] Among them, the first structured features include a first document theme and a first document sub-theme, and the second structured features include a second document theme and a second document sub-theme.

[0026] In one embodiment of the present invention, the steps of identifying the first structured features of the first unstructured document in step S1 include the following steps:

[0027] S11. Obtain a named entity recognition model, and identify the named entities of the first unstructured document according to the named entity recognition model.

[0028] Among them, the named entities may include personal names, institutional names, organization names, place names, products, events, currencies, etc.

[0029] In one embodiment of the present invention, obtaining the named entity recognition model in step S11 includes:

[0030] S101. Obtain a set of training documents containing named entity features from the historical document database.

[0031] S102. Perform text segmentation processing on each training document in the set of training documents to obtain a first vocabulary set corresponding to each training document, and classify the first vocabulary set into a named entity set and a non-named entity set according to the named entity features.

[0032] S103. Input the set of training documents into the named entity recognition network, and output a first probability value for predicting each vocabulary in the named entity set corresponding to each training document in the set of training documents as a named entity, and a second probability value for predicting each vocabulary in the non-named entity set as a named entity.

[0033] S104. Obtain a first cost function according to the first probability value and the second probability value, and adjust each parameter in the named entity network based on the first cost function to obtain the named entity recognition model.

[0034] In one embodiment of the present invention, step S104 specifically includes:

[0035] S141. Generate a second cost function according to the first probability value and the second probability value through the following formula:

[0036]

[0037] Where D m represents the second cost function, x ka is the first probability value corresponding to the a-th vocabulary in the named entity set of the k-th training document, y kb is the second probability value corresponding to the b-th vocabulary in the non-named entity set of the k-th training document, M is the number of training documents in the set of training documents, A is the number of vocabularies in the named entity set of the k-th training document, and B is the number of vocabularies in the non-named entity set of the k-th training document.

[0038] Among them, the second cost function is constructed based on the fact that the second probability value of each word in the non-named entity set corresponding to each to-be-trained document in the to-be-trained document set is less than the first probability value of each word in the named entity set.

[0039] S142. Calibrate the second cost function and generate the calibrated first cost function through the following formula:

[0040]

[0041] Among them, D n represents the first cost function, and x k_f is the mean value of the first probability values of the f words with the smallest values in the named entity set of the kth to-be-trained document, and x k_g is the mean value of the second probability values of the g words with the largest values in the named entity set of the kth to-be-trained document, and c and d are constants.

[0042] Specifically, in the actual training process, when identifying the to-be-trained documents, there is a situation of identification deviation, resulting in the fact that the identified to-be-trained documents cannot meet the above-mentioned second cost function. Therefore, when constructing the cost function corresponding to the named entity recognition model in the present invention, the second cost function is calibrated based on the fact that the second probability values of the first proportion (for example, 90%) of the words in the non-named entity set corresponding to the same to-be-trained document are less than the first probability values of the second proportion (for example, 85%) of the words in the corresponding named entity set, so as to generate the first cost function.

[0043] It should be noted that 0.9c + 0.1x k_f represents the minimum limit value of the first probability values of the words in the second proportion in the named entity set of the kth to-be-trained document, and 0.9d + 0.1x k_g represents the maximum limit value of the second probability values of the words in the first proportion in the non-named entity set of the kth to-be-trained document.

[0044] Furthermore, adjust the parameters in the named entity network according to the variables corresponding to the minimum value of the first cost function to obtain the named entity recognition model.

[0045] After obtaining the named entity recognition model, input the first unstructured document into the named entity recognition model to output the third probability value of each word in the first unstructured document as a named entity. Then, select the corresponding number of words to be extracted as named entities from the words according to the number of named entities to be extracted preset in the first unstructured document.

[0046] S12. Based on the named entity recognition, determine the first document theme including the named entity in the first structured document.

[0047] S13. Summarize each first document topic feature to obtain the corresponding first document sub-topic.

[0048] Specifically, after identifying the named entities in the first unstructured document, multi-level induction of the document topic and sub-topics can be achieved through title hierarchy analysis and combined with the semantic understanding ability of the large model. Specifically, the core topic can be identified by analyzing the document title hierarchy and structure based on the position information of the named entities, and relevant sub-topics can be extracted on this basis. Sub-topics are areas or topics that are related to but secondary to the core topic and can enrich the semantic expression of the document. Through context analysis, semantic similarity calculation (such as Word2Vec or Sentence-BERT), and key sentence extraction, the sub-topics are further expanded to ensure the consistency and accuracy of the sub-topics with the core topic. Among them, if the document is complex, it can also be refined through multi-level sub-topic extraction (such as based on dimensions of time, location, people, etc.) and combined with machine learning techniques (such as topic modeling and clustering analysis) to optimize the extraction process. Finally, the sub-topics are integrated with the core topic to improve the structuring and semantic accuracy of the document and assist subsequent analysis and traceability work.

[0049] Similarly, the second unstructured document can be input into the named entity recognition model, and the named entities of the second unstructured document can be obtained according to the output result. Then, the second document topic and the second document sub-topic are obtained based on the named entities of the second unstructured document. This process is similar to the above-mentioned first unstructured document. To avoid redundancy, it will not be elaborated here.

[0050] S2. Compare the first structured feature with the second structured feature.

[0051] In an embodiment of the present invention, step S2 specifically includes the following steps:

[0052] S21. Calculate the first similarity between the first document topic and the second document topic.

[0053] Specifically, the first similarity between the first document topic and the second document topic can be calculated through the following formula:

[0054] TS = w1·CS(T1, T2) + w2·JS(K1, K2) + w3·SS(E1, E2), (1)

[0055] Among them, TS represents the first similarity, CS(T1, T2) represents the cosine similarity between the first document topic vector T1 and the second document topic vector T2, JS(K1, K2) represents the Jaccard similarity coefficient between the named entity set K1 of the first unstructured document and the named entity set K2 of the second unstructured document, SS(E1, E2) represents the semantic similarity between the semantic embedding vector E1 of the first unstructured document and the semantic embedding vector E2 of the second unstructured document, w1 represents the first weight, w2 represents the second weight, w3 represents the third weight, and w1 + w2 + w3 = 1. For example, w1 = 0.3, w2 = 0.1, and w3 = 0.6.

[0056] Among them,

[0057]

[0058] Among them, in a deep learning-based model, the semantic embedding vector E1 of the first unstructured document and the semantic embedding vector E2 of the second unstructured document are usually generated by a pre-trained language model (such as BERT, Sentence-BERT). Among them, each unstructured document can be converted into an n-dimensional vector representation through the above pre-trained language model.

[0059] S22, determine whether the first similarity is greater than or equal to the first threshold.

[0060] S23, if the first similarity is greater than or equal to the first threshold, then calculate the second similarity between the first document sub-topic and the second document sub-topic.

[0061] Among them, the method for calculating the second similarity is similar to the method for calculating the first similarity above, and can be calculated with reference to the above method. To avoid redundancy, it will not be elaborated here.

[0062] S24, determine whether the second similarity is greater than or equal to the second threshold.

[0063] S3, determine whether the first unstructured document is from the target system according to the feature comparison result.

[0064] Specifically, if it is determined that the first similarity is less than the first threshold (which can be calibrated according to the actual situation, for example, it can be 0.7), then it is considered that the first unstructured document is not from the target system; if it is determined that the first similarity is greater than or equal to the first threshold, and the second similarity is greater than or equal to the second threshold (which can be calibrated according to the actual situation, for example, it can be 0.7), then it is considered that the first unstructured document is from the target system; if it is determined that the first similarity is greater than or equal to the first threshold, and the second similarity is less than the second threshold, then it is considered that the first unstructured document is not from the target system.

[0065] In summary, according to the unstructured document traceability method of the embodiments of the present invention, the first structured feature of the first unstructured document to be traced is identified, the second structured feature of the second unstructured document in the target system is identified, and the first structured feature is compared with the second structured feature, and it is determined whether the first unstructured document is derived from the target system according to the result of the feature comparison. Thus, by identifying the structured features of the unstructured document, the unstructured document can be effectively traced.

[0066] Corresponding to the unstructured document traceability method of the above embodiments, the present invention also proposes an unstructured document traceability system.

[0067] As Figure 2 shown, the unstructured document traceability system of the embodiments of the present invention may include: an identification module 100, a feature comparison module 200, and a traceability module 300.

[0068] Among them, the identification module 100 is used to identify the first structured feature of the first unstructured document to be traced and the second structured feature of the second unstructured document in the target system; the feature comparison module 200 is used to compare the first structured feature with the second structured feature; the traceability module 300 is used to determine whether the first unstructured document is derived from the target system according to the result of the feature comparison.

[0069] In an embodiment of the present invention, the first structured feature includes a first document theme and a first document sub-theme, the second structured feature includes a second document theme and a second document sub-theme, and the feature comparison module 200 is specifically used for: calculating a first similarity between the first document theme and the second document theme; determining whether the first similarity is greater than or equal to a first threshold; if the first similarity is greater than or equal to the first threshold, calculating a second similarity between the first document sub-theme and the second document sub-theme; and determining whether the second similarity is greater than or equal to a second threshold.

[0070] In an embodiment of the present invention, the identification module 100 is specifically used for: obtaining a named entity recognition model, and identifying the named entities of the first unstructured document according to the named entity recognition model; based on the named entities, identifying the first document theme in the first structured document that contains the named entities; and summarizing the features of each first document theme to obtain the corresponding first document sub-theme.

[0071] In one embodiment of the present invention, the recognition module 100 is specifically configured to: obtain a set of documents to be trained containing named entity features from a historical document database; perform text segmentation processing on each document to be trained in the set of documents to be trained, obtain a first vocabulary set corresponding to each document to be trained, and classify the first vocabulary set into a named entity set and a non-named entity set according to the named entity features; input the set of documents to be trained into a named entity recognition network, and output a first probability value of predicting each vocabulary in the named entity set corresponding to each document to be trained in the set of documents to be trained as a named entity, and a second probability value of predicting each vocabulary in the non-named entity set as a named entity; obtain a first cost function according to the first probability value and the second probability value, and adjust each parameter in the named entity network based on the first cost function to obtain the named entity recognition model.

[0072] In one embodiment of the present invention, the recognition module 100 is specifically configured to: generate a second cost function according to the first probability value and the second probability value through the following formula:

[0073]

[0074] where D m represents the second cost function, x ka is the first probability value corresponding to the a-th vocabulary in the named entity set of the k-th document to be trained, y kb is the second probability value corresponding to the b-th vocabulary in the non-named entity set of the k-th document to be trained, M is the number of documents to be trained in the set of documents to be trained, A is the number of vocabularies in the named entity set of the k-th document to be trained, and B is the number of vocabularies in the non-named entity set of the k-th document to be trained;

[0075] Calibrate the second cost function, and generate the calibrated first cost function through the following formula:

[0076]

[0077] where D n represents the first cost function, x k_f is the mean value of the first probability values of the f vocabulary with the smallest values in the named entity set of the k-th document to be trained, x k_g is the mean value of the second probability values of the g vocabulary with the largest values in the named entity set of the k-th document to be trained, and c and d are constants.

[0078] It should be noted that for the details not disclosed in the unstructured document traceability system of the embodiments of the present invention, please refer to the details disclosed in the unstructured document traceability method of the embodiments of the present invention, and will not be elaborated here specifically.

[0079] According to the unstructured document traceability system of the embodiments of the present invention, the first structured feature of the first unstructured document to be traced is identified by an identification module, and the second structured feature of the second unstructured document in the target system is identified, and the first structured feature is compared with the second structured feature by a feature comparison module, and the traceability module determines whether the first unstructured document is derived from the target system according to the feature comparison result. Thus, by identifying the structured features of the unstructured document, the unstructured document can be effectively traced.

[0080] Corresponding to the above embodiments, the present invention also provides a computer device.

[0081] The computer device of the embodiments of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the unstructured document traceability method of the above embodiments is implemented.

[0082] According to the computer device of the embodiments of the present invention, by identifying the structured features of the unstructured document, the unstructured document can be effectively traced.

[0083] Corresponding to the above embodiments, the present invention also provides a non-transitory computer-readable storage medium.

[0084] The non-transitory computer-readable storage medium of the embodiments of the present invention stores a computer program, and when the program is executed by a processor, the above unstructured document traceability method is implemented.

[0085] According to the non-transitory computer-readable storage medium of the embodiments of the present invention, by identifying the structured features of the unstructured document, the unstructured document can be effectively traced.

[0086] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The meaning of "a plurality" is two or more unless otherwise specifically defined.

[0087] In the present invention, unless otherwise clearly specified or limited, the terms "mounted", "connected", "coupled", "fixed", etc. shall be construed broadly. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0088] In the present invention, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "beneath" and "underneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0089] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0090] In addition, in each embodiment of the present invention, the functional units may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0091] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for tracing the provenance of unstructured documents, characterized in that: The following steps are involved: S1, obtaining a named entity recognition model, and identifying a first structured feature of a first unstructured document to be traced according to the named entity recognition model, and identifying a second structured feature of a second unstructured document in a target system; wherein obtaining the named entity recognition model in step S11 includes: S101, obtaining a document set to be trained containing named entity features from a historical document database; S102, performing text segmentation processing on each document to be trained in the set of documents to be trained to obtain a first vocabulary set corresponding to each document to be trained, and classifying the first vocabulary set into a named entity set and a non-named entity set according to the named entity features; S103, inputting the document set to be trained into a named entity recognition network, and outputting a first probability value for predicting each word in the named entity set corresponding to each document to be trained in the document set to be trained as a named entity, and a second probability value for predicting each word in the non-named entity set as a named entity; S104, obtaining a first cost function according to the first probability value and the second probability value, and adjusting each parameter in the named entity recognition network based on the first cost function to obtain the named entity recognition model; wherein step S104 specifically includes: S141: Generate a second cost function according to the first probability value and the second probability value by using the following formula: , in, represents the second cost function, is the first probability value corresponding to the a-th word in the named entity set of the k-th document to be trained, is the second probability value of the bth word in the non-named entity set of the kth document to be trained, is the number of documents to be trained in the document set to be trained, is the number of words in the named entity set of the kth document to be trained, is the number of words in the non-named entity set of the kth document to be trained; S142, calibrating the second cost function, and generating the calibrated first cost function by the following formula: , in, represents the first cost function, is the named entity set of the kth document to be trained The mean of the first probability values ​​of the words with the smallest values, is the named entity set of the kth document to be trained The mean of the second probability values ​​of the words with the largest values, and is a constant; S2, comparing the first structural feature with the second structural feature; S3: Determine whether the first unstructured document originates from the target system according to the feature comparison result.

2. The unstructured document tracing method according to claim 1, characterized in that: The first structural feature includes a first document topic and a first document subtopic, and the second structural feature includes a second document topic and a second document subtopic, wherein step S2 specifically includes: S21, calculating a first similarity between the first document topic and the second document topic; S22, determining whether the first similarity is greater than or equal to a first threshold; S23, if the first similarity is greater than or equal to the first threshold, calculating a second similarity between the first document subtopic and the second document subtopic; S24: Determine whether the second similarity is greater than or equal to a second threshold.

3. The unstructured document tracing method according to claim 2, characterized in that: Identifying the first structured feature of the first unstructured document in step S1 includes the following steps: S11, identifying the named entities of the first unstructured document according to the named entity recognition model; S12, identifying a first document topic containing the named entity in the first unstructured document based on the named entity; S13: Obtain the corresponding first document sub-topic according to the first document topic.

4. An unstructured document traceability system, characterized in that: include: A recognition module, the recognition module is used to obtain a named entity recognition model, and identify the first structured feature of a first unstructured document to be traced according to the named entity recognition model, and identify the second structured feature of a second unstructured document in a target system; wherein the recognition module is specifically used to: obtain a set of documents to be trained containing named entity features from a historical document database; perform text segmentation processing on each document to be trained in the set of documents to be trained to obtain a first vocabulary set corresponding to each document to be trained, and summarize the first vocabulary set into a named entity set and a non-named entity set according to the named entity features; input the set of documents to be trained into a named entity recognition network, and output a first probability value for predicting each vocabulary in the named entity set corresponding to each document to be trained in the set of documents to be trained as a named entity, and a second probability value for predicting each vocabulary in the non-named entity set as a named entity; obtain a first cost function according to the first probability value and the second probability value, and adjust each parameter in the named entity recognition network based on the first cost function to obtain the named entity recognition model; wherein the second cost function is generated according to the first probability value and the second probability value by the following formula: , in, represents the second cost function, is the first probability value corresponding to the a-th word in the named entity set of the k-th document to be trained, is the second probability value of the bth word in the non-named entity set of the kth document to be trained, is the number of documents to be trained in the document set to be trained, is the number of words in the named entity set of the kth document to be trained, is the number of words in the non-named entity set of the kth document to be trained; the second cost function is calibrated, and the calibrated first cost function is generated by the following formula: , in, represents the first cost function, is the named entity set of the kth document to be trained The mean of the first probability values ​​of the words with the smallest values, is the named entity set of the kth document to be trained The mean of the second probability values ​​of the words with the largest values, and is a constant; A feature comparison module, the feature comparison module is used to compare the first structured feature with the second structured feature; A tracing module is used to determine whether the first unstructured document originates from the target system according to a feature comparison result.

5. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, an unstructured document tracing method according to any one of claims 1 to 3 is implemented.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, an unstructured document tracing method according to any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Method and apparatus for word segmentation

    CN109190124A

  • Divulgence traceability method, system and platform suitable for unstructured document

    CN116910713A