Document analysis method, electronic device and program product

By automatically identifying the standard core paragraphs and similarity calculations of documents, the problems of low efficiency and low accuracy of manual processing of documents are solved, and efficient and accurate extraction of core content of documents are achieved.

CN120386875APending Publication Date: 2025-07-29KE COM (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510387409.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the extraction of core content of the document relies on manual processing, resulting in limited processing speed, low efficiency and low accuracy, especially when facing a large number of documents, errors are prone to occur.

Method used

By obtaining the similarity between the standard core paragraphs and paragraph units of the document, the core content of the document is automatically identified using similarity calculation and filtering rules to reduce manual intervention.

Benefits of technology

Improves the efficiency and accuracy of document processing, reduces manual processing time and error rate, and significantly improves processing efficiency and accuracy when processing large amounts of documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386875A_ABST
    Figure CN120386875A_ABST
Patent Text Reader

Abstract

The invention provides a document analysis method, electronic equipment and a program product. The document analysis method comprises the following steps: in response to a received analysis instruction of a to-be-analyzed document, obtaining a plurality of standard core paragraphs related to the to-be-analyzed document and a plurality of paragraph units included in the to-be-analyzed document; obtaining the similarity between each paragraph unit and each standard core paragraph; determining a plurality of similar paragraph units from the plurality of paragraph units according to the similarity between each paragraph unit and each standard core paragraph; and determining a target paragraph unit from the plurality of similar paragraph units according to a screening rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of data processing and the like, and particularly relates to a document analysis method, an electronic device, a storage medium, and a program product. Background Art

[0002] With the rapid development of information technology, the quantity of document data has shown an explosive growth. Whether it is internal reports, contracts, emails within an enterprise, or external news, academic papers, market analysis, documents have become the main carriers of information transmission and storage. However, in the face of a vast amount of document data, how to quickly and accurately extract the core content has become an urgent problem to be solved.

[0003] The traditional extraction of the core content of documents mainly relies on manual processing. Specifically, usually, professional personnel need to read the documents and manually mark the core content.

[0004] However, the speed of manual document processing is limited. Especially when faced with a large number of documents, the processing time will increase significantly, and the processing efficiency is low; moreover, the processing personnel may make incorrect markings due to fatigue, distraction, or neglect of certain details, and the processing accuracy is not high. Summary of the Invention

[0005] The present disclosure provides a document analysis method, an electronic device, a storage medium, and a program product.

[0006] According to one aspect of the present disclosure, there is provided a document analysis method, including: In response to an analysis instruction for a document to be analyzed received, obtaining a plurality of standard core paragraphs related to the document to be analyzed and a plurality of paragraph units included in the document to be analyzed; Obtaining the similarity between each paragraph unit and each standard core paragraph; Determining a plurality of similar paragraph units from the plurality of paragraph units according to the similarity between each paragraph unit and each standard core paragraph; Determining a target paragraph unit from the plurality of similar paragraph units according to a screening rule.

[0007] According to the document analysis method of at least one embodiment of the present disclosure, after determining the target paragraph unit from the plurality of similar paragraph units according to the screening rule, it further includes: In response to an analysis instruction based on the target paragraph unit received, analyzing the target paragraph unit according to the analysis instruction to obtain an analysis result.

[0008] According to the document analysis method of at least one embodiment of the present disclosure, after determining the target paragraph unit from the plurality of similar paragraph units according to the screening rule, it further includes: Perform a reliability assessment on the target paragraph unit according to the multiple similar paragraph units and the screening rules to obtain an assessment result.

[0009] According to the document analysis method of at least one embodiment of the present disclosure, the obtaining of multiple standard core paragraphs related to the document to be analyzed and multiple paragraph units included in the document to be analyzed includes: Obtain the attributes of the document to be analyzed; Obtain multiple standard core paragraphs related to the document to be analyzed according to the attributes of the document to be analyzed; Perform paragraph segmentation on the document to be analyzed to obtain multiple paragraph units.

[0010] According to the document analysis method of at least one embodiment of the present disclosure, the obtaining of the similarity between each paragraph unit and each standard core paragraph includes: Perform paragraph grouping on the multiple paragraph units to obtain multiple paragraph groups; Obtain the similarity between each paragraph group and each standard core paragraph.

[0011] According to the document analysis method of at least one embodiment of the present disclosure, the obtaining of the similarity between each paragraph group and each standard core paragraph includes: Respectively obtain the feature vectors of each paragraph group and the standard vectors of each standard core paragraph; Perform a similarity comparison between each feature vector and each standard vector to obtain the similarity between each paragraph group and each standard core paragraph.

[0012] According to the document analysis method of at least one embodiment of the present disclosure, the performing of a similarity comparison between each feature vector and each standard vector includes: Perform a similarity comparison between each feature vector and each standard vector based on a single-stage similarity calculation method; or, Perform a similarity comparison between each feature vector and each standard vector based on a multi-stage similarity calculation method.

[0013] According to the document analysis method of at least one embodiment of the present disclosure, the determining of the target paragraph unit from the multiple similar paragraph units according to the screening rules includes: Obtain the screening rules of the multiple similar paragraph units; Input the screening rules and the multiple similar paragraph units into a screening model to obtain the target paragraph unit.

[0014] The document analysis method according to at least one embodiment of the present disclosure, before inputting the screening rule and the multiple similar paragraph units into the screening model, further includes: obtaining task knowledge related to the multiple similar paragraph units; The inputting the screening rule and the multiple similar paragraph units into the screening model includes: inputting the screening rule, the multiple similar paragraph units, and the task knowledge into the screening model.

[0015] The document analysis method according to at least one embodiment of the present disclosure, before inputting the screening rule and the multiple similar paragraph units into the screening model, further includes: obtaining guiding examples related to screening the multiple similar paragraph units; The inputting the screening rule and the multiple similar paragraph units into the screening model includes: inputting the screening rule, the multiple similar paragraph units, and the guiding examples into the screening model.

[0016] The document analysis method according to at least one embodiment of the present disclosure, the guiding examples include positive examples; or, The guiding examples include positive examples and negative examples.

[0017] According to another aspect of the present disclosure, there is provided an electronic device, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the document analysis method according to any one of the embodiments of the present disclosure.

[0018] According to still another aspect of the present disclosure, there is provided a computer program product, including a computer program that, when executed by a processor, implements the document analysis method according to any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.

[0020] Figure 1 is a schematic diagram of an application scenario of the document analysis method according to an embodiment of the present disclosure.

[0021] Figure 2 is the flow of the document analysis method according to an embodiment of the present disclosure Figure 1 .

[0022] Figure 3 is the flow of the document analysis method according to an embodiment of the present disclosure Figure 2 .

[0023] Figure 4 is the flowchart of the document analysis method according to an embodiment of the present disclosure Figure 3 .

[0024] Figure 5 is Figure 2 the flowchart of the paragraph acquisition method in the document analysis method shown

[0025] Figure 6 is Figure 2 the flowchart of the similarity acquisition method in the document analysis method shown

[0026] Figure 7 is Figure 6 the flowchart of the vector comparison method in the similarity acquisition method shown Figure 1 .

[0027] Figure 8 is Figure 6 the flowchart of the vector comparison method in the similarity acquisition method shown Figure 2 .

[0028] Figure 9 is Figure 6 the flowchart of the vector comparison method in the similarity acquisition method shown Figure 3 .

[0029] Figure 10 is Figure 2 the flowchart of the decision analysis method in the document analysis method shown Figure 1 .

[0030] Figure 11 is Figure 2 the flowchart of the decision analysis method in the document analysis method shown Figure 2 .

[0031] Figure 12 is Figure 2 the flowchart of the decision analysis method in the document analysis method shown Figure 3 .

[0032] Figure 13 is the schematic flowchart of the document analysis method according to an embodiment of the present disclosure

[0033] Figure 14 is the schematic block diagram of the structure of the document analysis device according to an embodiment of the present disclosure

[0034] Figure 15 is the schematic block diagram of the structure of the electronic device according to an embodiment of the present disclosure Detailed Embodiments

[0035] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and examples. It can be understood that the specific examples described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that for ease of description, only the parts related to the present disclosure are shown in the accompanying drawings.

[0036] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.

[0037] Taking the document to be analyzed as a contract as an example, if the contract content is relatively complex, such as involving multiple basic terms and additional terms, it takes a lot of time for humans to read and understand this content; especially when faced with numerous contracts, the time required for humans to analyze each one increases exponentially; taking the contract volume of a medium-sized enterprise as an example, it may take hundreds of hours of human effort to process these contract documents every year, and the processing efficiency is low. Additionally, in contract processing, contract terms are usually relatively complex and require careful checking and marking; and during the manual processing process, the processing personnel may miss or mismark important terms due to long hours of fatigue reading or distracted attention; for example, the payment conditions or liability for breach of contract involved in a certain contract may be overlooked, thus affecting subsequent contract performance or company decisions, and further causing serious economic losses or legal risks.

[0038] For this reason, the present disclosure proposes a document analysis method, an electronic device, a storage medium, and a program product. The present disclosure can be implemented through document analysis software installed on an electronic device such as a server.

[0039] Figure 1 The application scenario of the document analysis method according to an embodiment of the present disclosure is shown. In this application scenario, it may include a user terminal 100 and a document analysis device 200. The user terminal 100 is connected to the document analysis device 200 through a network; the user terminal 100 is used to upload the document to be analyzed; the document analysis device 200 is used to receive the document to be analyzed sent by the user terminal 100 and analyze it.

[0040] For ease of description and to make the technical solutions of the specific embodiments of the present disclosure easier to understand, the technical terms involved in the present disclosure are explained as follows: Paragraph segmentation refers to the process of splitting a continuous text into relatively independent paragraph units.

[0041] Paragraph grouping refers to the process of combining multiple paragraphs into different groups according to set criteria or rules.

[0042] The single-stage similarity calculation method refers to directly obtaining the similarity result through one-time calculation.

[0043] The multi-stage similarity calculation method divides the similarity calculation process into multiple stages. Usually, a relatively simple and efficient similarity measurement method is first used to preliminarily screen a large number of objects to obtain a smaller candidate set; then a more accurate and complex similarity measurement method is used in the candidate set for secondary calculation to obtain the final similarity result.

[0044] Figure 2 FIG. 4 shows an overall flowchart of a document analysis method M100 according to an embodiment of the present disclosure. As Figure 2 shown, the document analysis method includes steps S110 to S140. Among them, the document analysis method can be executed by an electronic device such as a server.

[0045] Specifically, Figure 2 the shown document analysis method includes: Step S110, in response to an analysis instruction of a document to be analyzed received, obtain a plurality of standard core paragraphs related to the document to be analyzed and a plurality of paragraph units included in the document to be analyzed.

[0046] In some embodiments of the present disclosure, the analysis instruction of the document to be analyzed received in step S110 may be a specific analysis instruction, that is, an instruction directly indicating which document to analyze; the analysis instruction of the document to be analyzed received in step S110 may also be a fuzzy analysis instruction, that is, providing a plurality of documents and instructing to perform document analysis, but not specifying which specific document to analyze.

[0047] When the received analysis instruction is a specific analysis instruction, step S110 may directly obtain a plurality of standard core paragraphs and a plurality of paragraph units; when the received analysis instruction is a fuzzy analysis instruction, step S110 may first use a preset keyword, a preset screening model, a preset screening rule, etc. to determine the final document to be analyzed from a plurality of documents, and then obtain a plurality of standard core paragraphs and a plurality of paragraph units of the final document to be analyzed.

[0048] The plurality of standard core paragraphs in step S110 may be paragraphs obtained from pre-set standard paragraphs. The pre-set standard paragraphs may be fixed or may be adjusted by the user according to needs. The plurality of standard core paragraphs may be core paragraphs determined by the user based on a plurality of types related to the document to be analyzed from the pre-set standard paragraphs; any one of the plurality of standard core paragraphs may include one or more paragraph units.

[0049] In some embodiments of the present disclosure, since the purposes, structures, audiences, included contents, etc. of different types of documents vary, the core contents of different types of documents are different; for different types of documents, different standard core paragraphs can be set. At this time, step S110 may specifically include: in response to an analysis instruction of a document to be analyzed received, obtaining the category of the document to be analyzed; obtaining a plurality of standard core paragraphs according to the category.

[0050] Step S120, obtaining the similarity between each paragraph unit and each standard core paragraph.

[0051] In some embodiments of the present disclosure, step S120 may obtain the similarity between each paragraph unit and each standard core paragraph based on vocabulary, term frequency-inverse document frequency (TF-IDF), etc.

[0052] Step S130, determining a plurality of similar paragraph units from the plurality of paragraph units according to the similarity between each paragraph unit and each standard core paragraph.

[0053] In some embodiments of the present disclosure, step S130 may obtain a preset first number of similar paragraph units with higher similarity from the plurality of paragraph units according to the level of similarity between each paragraph unit and each standard core paragraph.

[0054] Step S140, determining a target paragraph unit from the plurality of similar paragraph units according to a screening rule.

[0055] In some embodiments of the present disclosure, through step S140, decision analysis is performed on the plurality of similar paragraph units according to the screening rule, and the most similar target paragraph unit can be selected from the plurality of similar paragraph units; through decision analysis, abnormal paragraph units can also be identified and removed from the plurality of similar paragraph units to improve the accuracy of the decision, etc. Step S140 may use a screening model to determine the target paragraph unit.

[0056] For the document analysis method provided by the present disclosure, since the similar paragraph units are determined according to the standard core paragraphs, the similar paragraph units can represent the core content to a certain extent; by screening the similar paragraph units, the core content of the document to be analyzed can be accurately identified. This document analysis method can reduce manual intervention and solve the problems in the prior art that the speed of manual document processing is limited, especially when facing a large number of documents, the processing time will increase significantly and the processing efficiency is low; and the processing personnel may make incorrect markings due to fatigue, distraction or neglect of certain details, and the processing accuracy is not high.

[0057] Further, the document analysis method provided by the present disclosure, after determining the target paragraph unit through step S140, may further include step S150 as Figure 3 shown.

[0058] Step S150: In response to the received analysis instruction based on the target paragraph unit, analyze the target paragraph unit according to the analysis instruction to obtain an analysis result.

[0059] In some embodiments of the present disclosure, when the user needs to conduct a more in-depth exploration based on the target paragraph unit, an analysis instruction for the target paragraph unit may be sent, and the analysis instruction may indicate a more refined analysis or a supplementary analysis of the target paragraph unit.

[0060] Step S150 may analyze the target paragraph unit through corresponding analysis rules, analysis models, etc. of the analysis instruction. The analysis model and the screening model used in step S140 may be the same model or different models.

[0061] Through step S150, after determining the target paragraph unit, a further in-depth analysis can be performed based on the target paragraph unit to obtain a more accurate and comprehensive understanding and provide more valuable information.

[0062] Further, the document analysis method provided by the present disclosure, after determining the target paragraph unit through step S140, may further include step S160 as Figure 4 shown.

[0063] Step S160: Evaluate the reliability of the target paragraph unit according to multiple similar paragraph units and screening rules to obtain an evaluation result.

[0064] In some embodiments of the present disclosure, step S160 may evaluate the reliability of the target paragraph unit according to multiple similar paragraph units and screening rules based on evaluation indicators to evaluate whether the target paragraph unit meets the expected standard; step S160 may also use an evaluation model to evaluate the reliability of the target paragraph unit, and the evaluation model and the screening model used in step S140 may be the same model or different models.

[0065] After determining the target paragraph unit through step S140, evaluating the reliability of the target paragraph unit through step S160 can effectively judge the effectiveness, accuracy, and applicability of the screening process.

[0066] Regarding step S110, in some embodiments of the present disclosure, it may include steps S111 to S113 as Figure 5 shown.

[0067] Step S111: Obtain the attributes of the document to be analyzed.

[0068] In some embodiments of the present disclosure, due to the different uses, structures, and information expression methods of documents with different attributes, the core contents of documents with different attributes are different. The attribute of the document to be analyzed obtained through step S111 generally refers to the type or classification attribute of the document to be analyzed, etc.

[0069] Step S112: Obtain multiple standard core paragraphs related to the document to be analyzed according to the attribute of the document to be analyzed.

[0070] In some embodiments of the present disclosure, the multiple standard core paragraphs related to the document to be analyzed obtained through step S112 are the criteria for identifying and determining which paragraphs in the document to be analyzed belong to the core content. The multiple standard core paragraphs related to the document to be analyzed in step S112 can specifically be multiple standard core paragraphs of the same type as the document to be analyzed.

[0071] Step S113: Perform paragraph segmentation on the document to be analyzed to obtain multiple paragraph units.

[0072] In some embodiments of the present disclosure, the document to be analyzed can be in a plain text format, such as TXT (.txt), Markdown (.md), CSV (.csv), etc.; the format of the document to be analyzed can also be a document in a structured document format, such as Word (.doc or.docx), PDF (.pdf), etc.

[0073] When the format of the document to be analyzed is.docx, since the.docx file is an open file format based on XML and has standard paragraph information, step S113 can perform paragraph segmentation on the document to be analyzed through processes such as document reading, XML parsing, and paragraph processing.

[0074] When the format of the document to be analyzed is.doc, since the paragraph information of this format document is not clear and cannot be effectively split through the information of the document itself, step S113 can perform paragraph segmentation on the document to be analyzed through methods such as win32com parsing or parsing after format conversion.

[0075] When the format of the document to be analyzed is.pdf, step S113 can first perform image rendering on the document to be analyzed, then perform character recognition through Optical Character Recognition (OCR), erase interference information (such as numbers, superscripts, headers, etc.), and finally perform paragraph segmentation on the content obtained through the above processes.

[0076] When the format of the document to be analyzed is.doc or.pdf, artificial intelligence (AI) (such as rule-based models, ordinary machine learning models, or deep learning models, etc.) can also be used to achieve paragraph segmentation.

[0077] Through steps S111 to S112, the obtained multiple standard core paragraphs can be made relevant to the document to be analyzed, ensuring that the core content of the document to be analyzed can be correctly extracted; through step S113, the paragraph boundaries in the document to be analyzed can be automatically recognized, reducing the workload of manual annotation and improving the document analysis efficiency.

[0078] Regarding step S120, in some embodiments of the present disclosure, it may include steps S121 to S122 as Figure 6 shown.

[0079] Step S121, perform paragraph grouping on multiple paragraph units to obtain multiple paragraph groups.

[0080] In some embodiments of the present disclosure, step S121 may perform paragraph grouping on multiple paragraph units based on a preset grouping word count; in particular, to prevent being truncated in the middle of a paragraph, when performing paragraph grouping on multiple paragraph units based on the preset grouping word count, it may also be determined whether the content of the last paragraph in the paragraph grouping is complete; if it is complete, no processing may be performed on this paragraph grouping; if it is incomplete, the content of the last paragraph may be supplemented, or the content of the last paragraph may be deleted, etc.

[0081] Step S121 may also perform paragraph grouping based on the structured information of the document (such as chapter titles, subheadings, etc.), semantic similarity between paragraph units, etc.

[0082] Step S122, obtain the similarity between each paragraph group and each standard core paragraph.

[0083] In some embodiments of the present disclosure, step S122 may obtain the similarity between each paragraph group and each standard core paragraph based on vocabulary, term frequency-inverse document frequency (TF-IDF), etc.

[0084] To intuitively represent the similarity between each paragraph group and each standard core paragraph, the similarity obtained through step S122 may be represented as a similarity matrix.

[0085] Steps S121 to S122 can improve the calculation efficiency of similarity, reduce noise interference, and enhance semantic consistency by grouping multiple paragraph units and then obtaining the similarity.

[0086] At this time, step S130 can be specifically: obtaining a plurality of similar paragraph units according to the similarity between each paragraph group and each standard core paragraph. The process of obtaining the plurality of similar paragraph units may include: obtaining at least one paragraph group with a higher similarity according to the similarity between each paragraph group and each standard core paragraph; splitting the at least one paragraph group with a higher similarity into paragraphs to obtain a plurality of similar paragraph units.

[0087] Regarding step S122, in some embodiments of the present disclosure, it may include steps S1221 to S1222 as Figure 7 shown.

[0088] Step S1221, respectively obtaining the feature vectors of each paragraph group and the standard vectors of each standard core paragraph.

[0089] In some embodiments of the present disclosure, step S1221 may respectively obtain the feature vectors of each paragraph group and the standard vectors of each standard core paragraph based on methods such as the Bag of Words (BoW), TF-IDF, and the word vector Word2Vec.

[0090] Step S1222, comparing the similarity of each feature vector and each standard vector to obtain the similarity between each paragraph group and each standard core paragraph.

[0091] In some embodiments of the present disclosure, step S1222 may compare the similarity of each feature vector and each standard vector based on cosine similarity, Euclidean distance, dot product, Mahalanobis distance, etc.

[0092] Calculating the similarity after vectorization through steps S1221 to S1222 can improve the efficiency and accuracy of document analysis.

[0093] Regarding step S1222, in some embodiments of the present disclosure, it may be replaced with the following steps as Figure 8 shown: Comparing the similarity of each feature vector and each standard vector based on a single-stage similarity calculation method.

[0094] In some embodiments of the present disclosure, this step may use a similarity calculation method to compare the similarity of each feature vector and each standard vector.

[0095] Comparing the similarity based on the single-stage similarity calculation method can reduce the calculation complexity and improve the calculation efficiency.

[0096] Regarding step S1222, in some embodiments of the present disclosure, it may be replaced with the following steps asFigure 9 The following steps shown: Based on the multi-stage similarity calculation method, compare the similarity of each feature vector and each standard vector.

[0097] In some embodiments of the present disclosure, this step can adopt multiple similarity calculation methods and be carried out step by step in stages. For example: In the first stage, a simple similarity calculation method such as TF-IDF can be used for quick similarity comparison to obtain the Top-N feature vectors among multiple feature vectors; in the second stage, more complex similarity calculations can be performed on the Top-N feature vectors based on a preset weight algorithm or BERT semantic matching, etc., and further screen out the Top-M feature vectors from the Top-N feature vectors.

[0098] This step performs similarity comparison based on the multi-stage similarity calculation method, which can improve the accuracy of similarity calculation.

[0099] Regarding step S140, in some embodiments of the present disclosure, it may include steps S141 to S142 as Figure 10 shown.

[0100] Step S141, obtain the screening rules for multiple similar paragraph units.

[0101] In some embodiments of the present disclosure, the screening rules in step S141 can guide how to perform decision analysis on multiple similar paragraph units and screen out the most suitable target paragraph unit. The screening rules can be obtained based on the analysis of the business characteristics of relevant businesses.

[0102] Step S142, input the screening rules and multiple similar paragraph units into the screening model to obtain the target paragraph unit.

[0103] In some embodiments of the present disclosure, by inputting the screening rules and multiple similar paragraph units into the screening model in step S142, the screening model can be guided by the screening rules to perform decision analysis on multiple similar paragraph units.

[0104] In the process of guiding the screening model to perform decision analysis, the multi-round dialogue ability can be used to guide the screening model to think more deeply; at the same time, the idea of Chain-of-Thought (CoT) is used in each dialogue to guide the screening model to think comprehensively in each decision analysis.

[0105] The screening model in step S142 can be a simple semantic matching model, a scoring and weighting model, a large language model, etc.

[0106] Steps S141 to S142 use screening rules to guide the screening model to perform decision analysis on multiple similar passage units, which can improve the consistency and compliance of the decision analysis process.

[0107] Before step S142, it may also include step S143 as Figure 11 shown.

[0108] Step S143: Obtain task knowledge related to multiple similar passage units.

[0109] In some embodiments of the present disclosure, the task knowledge in step S143 is an information resource used to support and guide the decision analysis process. By providing data and professional insights for the decision analysis, it helps the screening model make more accurate, compliant, and effective selections. The task knowledge may include content such as laws and regulations, business logics, and tacit experiences.

[0110] At this time, step S142 is specifically: input the screening rules, multiple similar passage units, and task knowledge into the screening model.

[0111] Introducing task knowledge into the input of the screening model through step S143 can help the screening model improve the accuracy, controllability, and compliance of the decision analysis.

[0112] Before step S142, it may also include step S144 as Figure 12 shown.

[0113] Step S144: Obtain guidance examples related to screening multiple similar passage units.

[0114] In some embodiments of the present disclosure, the guidance examples obtained through step S144 may be instances that have occurred in the past and are related to the current document to be analyzed.

[0115] The guidance examples obtained through step S144 may specifically be positive examples, that is, correct and expected examples, which show the targets that the screening model should imitate; positive examples can help the screening model understand and confirm the demonstration data of the target state or correct result; through positive examples, the screening model can know what to do under what circumstances.

[0116] In particular, in addition to positive examples, the guidance examples may also include negative examples, that is, incorrect and unexpected examples, which can help the screening model identify errors and avoid their occurrence.

[0117] In particular, for the convenience of users, when the guiding examples include positive examples and negative examples, the positive examples can be fixed, and the negative examples can be dynamic. Users can adjust the specific content of the negative examples in real time according to their needs. Through the fixed positive examples and dynamic negative examples, a dynamic prompt can be realized, making the decision-making analysis process of the screening model more intelligent.

[0118] At this time, step S142 is specifically: inputting the screening rules, multiple similar paragraph units, and guiding examples into the screening model.

[0119] By introducing guiding examples into the input of the screening model through step S144, successful practices can be learned from positive examples, and similar mistakes can be helped to avoid recurring through negative examples, helping the screening model to make more accurate decision-making analysis and improving the accuracy of decision-making analysis.

[0120] In some embodiments of the present disclosure, the input of the screening model can also integrate the screening rules, multiple similar paragraph units, task knowledge, and guiding examples, so that the screening model makes decision-making analysis based on the above content.

[0121] In some embodiments of the present disclosure, in order to protect privacy and sensitive data and improve compliance, before obtaining multiple paragraph units through step S110, the document to be analyzed can also be desensitized, that is, by processing the sensitive information in the document to be analyzed, so that it can still be effectively analyzed on the premise of protecting personal privacy, business secrets or other sensitive content. The desensitized document is usually achieved by replacing, removing or obscuring specific data to avoid unauthorized persons obtaining sensitive information. These sensitive information can include personal identity information, financial data, medical records, business secrets, etc.

[0122] The document analysis method provided by the present disclosure is mainly used for documents containing important information and structured content, such as contract documents, legal instruments, financial reports and audit documents, tender documents, policy documents and regulations, academic papers and research reports, etc., various documents containing key terms, information or data. This document analysis method can be used in document entry, document review, document comparison, clause retrieval, risk warning and other processing processes.

[0123] After the document analysis method determines the core content of the document to be analyzed through steps S110 to S140, the reviewer can review the core content, which can significantly reduce the workload of manual review and reduce the error rate. And in the process of document analysis, the core risks of the document to be analyzed can also be determined, the target risk paragraphs obtained by evaluation can be obtained, and the document quality can be improved. And when the document to be analyzed is specifically a contract, the analysis results can also be used in the subsequent cooperation process, reducing the problem points of disputes between the two parties and reducing the risk of later litigation.

[0124] Taking the real estate field as an example, this document analysis method can be used to analyze contracts such as "Second-hand House Sales Contract", "New House Sales Contract", "Service Contract", etc., build a contract clause dashboard (visualization tool or platform), and guide the user to sign more and better clause contracts.

[0125] Figure 13 An exemplary flowchart implemented based on the document analysis method of the present disclosure is shown.

[0126] Figure 13 In the shown flowchart, taking the document to be analyzed as the "Service Contract" as an example, the document analysis process may include: Step S210, in response to the analysis instruction of the document to be analyzed received, obtain a plurality of standard core paragraphs related to the document to be analyzed.

[0127] In some embodiments of the present disclosure, when a user needs to analyze commission information based on the "Service Contract", the user can input the analysis instruction of the "Service Contract" at the user terminal and send the analysis instruction to the document analysis software.

[0128] Since the commission clause is directly related to remuneration and economic interests, it is one of the core contents of the "Service Contract", that is, the plurality of standard core paragraphs obtained through step S210 may specifically include a plurality of standard core paragraphs related to the commission clause.

[0129] The plurality of standard core paragraphs obtained through step S210 may be obtained from a preset standard library according to the type of the document to be analyzed. The standard core paragraphs stored in the preset standard library may be representative paragraphs extracted by the user from sample contracts in advance; the standard core paragraphs stored in the preset standard library may include one or more paragraph units.

[0130] Step S220, perform paragraph segmentation on the document to be analyzed to obtain a plurality of paragraph units.

[0131] In some embodiments of the present disclosure, the "Service Contract" may include the basic information of both parties, service fees and payment methods, obligations of both parties, liability for breach of contract, contract term, dispute resolution method, supplementary clauses, etc. Performing paragraph segmentation on the "Service Contract" through step S220 may specifically be to perform natural paragraph segmentation on the "Service Contract", so as to take each natural paragraph as a paragraph unit.

[0132] Specifically, to prevent information leakage, before segmenting the document to be analyzed in step S220, the document to be analyzed may also be desensitized, such as replacing sensitive information such as the information of the parties involved and trade secret information (e.g., replacing real sensitive information with fictional but similarly formatted data) or masking to obtain a desensitized document, and then segmenting the desensitized document in step S220.

[0133] Step S230: Group the multiple paragraph units to obtain multiple paragraph groups.

[0134] In some embodiments of the present disclosure, step S230 may specifically group the multiple paragraph units according to a preset number of words; taking the terms of service fees and payment methods in the "Service Contract" as an example: Article 2 Service Fees and Payment Methods 1. Service Fees: Party A agrees to pay service fees to Party B. The specific fee is X% of the transaction amount completed by Party A through Party B, but not less than RMB XXXX yuan.

[0135] The transaction amount refers to the total transaction amount of transactions, business cooperations, etc. facilitated between Party A and Party B, unless otherwise clearly stipulated in the contract.

[0136] 2. Payment Time and Method: Party A shall pay the service fees in proportion within 5 working days after the transaction is successful and the formal contract is signed.

[0137] The payment of the fees can be made by bank transfer to the designated account of Party B. The account information is as follows: Account Name: Beijing XXXX Co., Ltd. Opening Bank: XX Sub-branch of Beijing XX Bank Bank Account: XXXXXXXXXX Party A shall indicate the purpose of payment as "service fees" when making the payment.

[0138] 3. Installment Payment: If the transaction reached between Party A and Party B involves installment payment, Party A may pay the service fees in proportion according to the actual receipt of the transaction funds.

[0139] The service fees for each payment shall be paid to Party B within 3 working days after Party A receives the transaction funds.

[0140] 4. Delayed Payment: If Party A fails to pay the service fees as agreed, it shall pay a late payment penalty of one-thousandth of the outstanding amount per day.

[0141] Party B has the right to suspend the provision of subsequent services in case of Party A's overdue payment until Party A has paid all the amounts.

[0142] 5. Fee Adjustment: If the transaction amount or other conditions involved by Party A change, the two parties shall negotiate and adjust the service fees and implement them after written confirmation.

[0143] If a force majeure event occurs during the contract period, resulting in the postponement or failure of the transaction, the two parties may negotiate to adjust the payment time of the service fees.

[0144] Taking the preset number of words as 90 as an example, the content of the paragraph grouping is "Article 2 Service Fees and Payment Methods 1. Service Fees: Party A agrees to pay Party B service fees. The specific fee is X% of the transaction amount completed by Party A through Party B, but not less than RMB XXXX yuan.

[0145] The transaction amount refers to the total transaction amount of the transactions, business cooperations, etc. facilitated by Party A and Party B. Since the content of the last paragraph is incomplete, the content of this paragraph can be supplemented. The final content of the paragraph grouping obtained is "Article 2 Service Fees and Payment Methods 1. Service Fees: Party A agrees to pay Party B service fees. The specific fee is X% of the transaction amount completed by Party A through Party B, but not less than RMB XXXX yuan.

[0146] The transaction amount refers to the total transaction amount of the transactions, business cooperations, etc. facilitated by Party A and Party B, unless otherwise clearly stipulated in the contract."

[0147] The document analysis software can group paragraphs of multiple paragraph units in the above manner. The paragraph groups obtained through step S230 may include one or more paragraphs.

[0148] Step S240, obtain the similarity between each paragraph group and each standard core paragraph.

[0149] In some embodiments of the present disclosure, step S240 can convert each paragraph group and each standard core paragraph into vectors respectively, and obtain the similarity according to the distance between vectors, etc.

[0150] Step S250, obtain multiple similar paragraph units according to the similarity between each paragraph group and each standard core paragraph.

[0151] In some embodiments of the present disclosure, step S250 may select N similar paragraph groups with relatively high similarity from multiple paragraph groups, where N is a positive integer and the value of N can be preset by the user according to needs; then, perform natural paragraph segmentation on the N similar paragraph groups to obtain multiple similar paragraph units.

[0152] Step S260, determine a target paragraph unit from multiple similar paragraph units according to a screening rule.

[0153] In some embodiments of the present disclosure, through step S260, decision analysis is performed on multiple similar paragraph units according to a screening rule, and terms related to commission can be extracted from the "Service Contract", that is, the target paragraph unit is the paragraph unit where the commission-related terms are located.

[0154] Step S270, in response to the received analysis instruction based on the target paragraph unit, analyze the target paragraph unit according to the analysis instruction to obtain an analysis result.

[0155] In some embodiments of the present disclosure, through step S270, the commission-related terms can be analyzed to understand the commission calculation method and obtain the commission jump point (referring to a specific condition or value that triggers commission payment or commission level change), so as to realize commission settlement according to the commission jump point.

[0156] Steps S260 and S270 can use an artificial intelligence model for screening and analysis; the artificial intelligence models relied on by steps S260 and S270 can be the same model or different models.

[0157] Based on any of the above embodiments, the present disclosure also provides a document analysis device.

[0158] Figure 14 It is a structural schematic block diagram of a document analysis device according to an embodiment of the present disclosure.

[0159] As Figure 14 shown, the document analysis device includes: A paragraph acquisition module 110, configured to acquire multiple standard core paragraphs related to the document to be analyzed and multiple paragraph units included in the document to be analyzed in response to the received analysis instruction for the document to be analyzed.

[0160] A similarity acquisition module 120, configured to acquire the similarity between each paragraph unit and each standard core paragraph.

[0161] A paragraph screening module 130, configured to determine multiple similar paragraph units from multiple paragraph units according to the similarity between each paragraph unit and each standard core paragraph.

[0162] A paragraph analysis module 140 is configured to determine a target paragraph unit from multiple similar paragraph units according to a screening rule.

[0163] The above-mentioned document analysis device may be in the form of computer software, and each module of the above-mentioned document analysis device may be implemented by computer software modules.

[0164] For the specific implementation process of the functions and roles of each module in the above-mentioned device, please refer to the implementation process of the corresponding steps in the above-mentioned method for details, and will not be elaborated here.

[0165] The execution subject of the document analysis method in the specific implementation manner of the present disclosure may be an electronic device such as a server.

[0166] Therefore, based on any one of the above embodiments, the present disclosure further provides an electronic device, which can execute the document analysis method of any one of the above embodiments described in the present disclosure.

[0167] Figure 15 It is a structural schematic diagram of an electronic device 1000 according to an embodiment of the present disclosure.

[0168] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application of the hardware and overall design constraints. The bus 1100 connects various circuits including one or more processors 1200, a memory 1300, and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripheral devices, voltage regulators, power management circuits, external antennas, etc.

[0169] The bus 1100 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only one connecting line is shown in this figure, but it does not mean that there is only one bus or one type of bus.

[0170] The present disclosure also provides a readable storage medium storing a computer program which, when executed by a processor, is used to implement the above method. The "readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples of the readable storage medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.

[0171] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed, the processes or functions of the present disclosure are executed in whole or in part.

[0172] The computer program or instructions can be stored in a readable storage medium, or transmitted from one readable storage medium to another. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The readable storage medium can be any available medium that can be accessed, or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc; or it can be a semiconductor medium, such as a solid state drive. The computer-readable storage medium can be a volatile or non-volatile storage medium, or can include both volatile and non-volatile types of storage media.

[0173] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system, or a computer program product. Therefore, the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0174] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0175] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the blocks.

[0177] In the description of this specification, the description with reference to terms such as "an embodiment / way", "some embodiments / ways", "example", "specific example", or "some examples", etc., means that the specific features, structures, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of this disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.

[0178] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0179] Those skilled in the art should understand that the above-described embodiments are merely for clearly illustrating the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.

Claims

1. A document analysis method, characterized in that including: In response to an analysis instruction of a document to be analyzed, obtaining a plurality of standard core paragraphs related to the document to be analyzed and a plurality of paragraph units included in the document to be analyzed; obtaining the similarity between each paragraph unit and each standard core paragraph; determining a plurality of similar paragraph units from the plurality of paragraph units according to the similarity between each paragraph unit and each standard core paragraph; and determining a target paragraph unit from the plurality of similar paragraph units according to a screening rule.

2. The document analysis method according to claim 1, characterized in that: After determining the target paragraph unit from the plurality of similar paragraph units according to the screening rule, it further includes: In response to an analysis instruction based on the target paragraph unit, analyzing the target paragraph unit according to the analysis instruction to obtain an analysis result.

3. The document analysis method according to claim 1, wherein: After determining the target paragraph unit from the plurality of similar paragraph units according to the screening rule, it further includes: evaluating the reliability of the target paragraph unit according to the plurality of similar paragraph units and the screening rule to obtain an evaluation result.

4. The document analysis method according to any one of claims 1 to 3, characterized in that The obtaining a plurality of standard core paragraphs related to the document to be analyzed and a plurality of paragraph units included in the document to be analyzed includes: obtaining the attributes of the document to be analyzed; obtaining a plurality of standard core paragraphs related to the document to be analyzed according to the attributes of the document to be analyzed; and performing paragraph segmentation on the document to be analyzed to obtain a plurality of paragraph units.

5. The document analysis method according to any one of claims 1 to 3, characterized in that: The obtaining the similarity between each paragraph unit and each standard core paragraph includes: grouping the plurality of paragraph units into a plurality of paragraph groups; and obtaining the similarity between each paragraph group and each standard core paragraph.

6. The document analysis method according to claim 5, characterized in that The obtaining the similarity between each paragraph group and each standard core paragraph includes: respectively obtaining the feature vectors of each paragraph group and the standard vectors of each standard core paragraph; and performing similarity comparison between each feature vector and each standard vector to obtain the similarity between each paragraph group and each standard core paragraph.

7. The document analysis method according to claim 6, characterized in that, The performing similarity comparison between each feature vector and each standard vector includes: performing similarity comparison between each feature vector and each standard vector based on a single-stage similarity calculation method; or performing similarity comparison between each feature vector and each standard vector based on a multi-stage similarity calculation method.

8. The document analysis method according to any one of claims 1 to 3, characterized in that: The determining a target paragraph unit from the plurality of similar paragraph units according to the screening rule includes: obtaining the screening rule of the plurality of similar paragraph units; and inputting the screening rule and the plurality of similar paragraph units into a screening model to obtain a target paragraph unit.

9. An electronic device, characterized in that, including: a memory storing execution instructions; and a processor that executes the execution instructions stored in the memory, so that the processor executes the document analysis method according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the document analysis method according to any one of claims 1 to 8.