Document key information fusion analysis method based on vectorization technology

By considering keyword position and context information in the document key information fusion analysis method, and calculating row and column similarity to infer document logical relationships, the deviations and limitations of traditional methods when building vector representation matrix and judging document logical relationships are solved, and more accurate document content representation and more refined logical relationship recognition are achieved.

CN120181081APending Publication Date: 2025-06-20INST OF ECONOMIC & TECH STATE GRID HEBEI ELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510246550.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The traditional document key information fusion analysis method based on vectorization technology ignores keyword position and context relationships when constructing vector representation matrix, and has deviations and limitations when judging document logical relationships and generating summary description graphs.

Method used

By collecting documents, performing Chinese word segmentation and filtering stop words, a vector representation matrix of the document is constructed, the position and context information of the keywords are considered, and the logical relationship of the document is inferred by calculating row and column similarity, and a summary description diagram of the label priority and hierarchy is generated.

Benefits of technology

The generated vector representation matrix reflects the document content more accurately, can more accurately identify complex logical relationships between documents, and improve the readability and ease of use of summary description diagrams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181081A_ABST
    Figure CN120181081A_ABST
Patent Text Reader

Abstract

The invention discloses a document key information fusion analysis method based on a vectorization technology, which comprises the following steps of: collecting a document, and performing Chinese word segmentation on the document; filtering words in the document to obtain a keyword K in the document; constructing a vector representation matrix VM of the document by utilizing the keyword K; according to the vector representation matrix VM, judging a logic relationship between the documents to obtain an abstract description graph based on the logic relationship; and querying and analyzing according to the abstract description graph to obtain a text analysis result. According to the method, when the vector representation matrix VM is constructed, not only is the word frequency of the keyword K considered, but also the vector representation QK of the keyword K is obtained by obtaining the position set of the keyword and combining the word vector T, so that the generated vector representation matrix can better reflect the real content of the document; because the position of the keyword and the context information are crucial to understanding the meaning of the document, the logic relationship between the documents is deduced by calculating the row similarity SMm and the column similarity SNn.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document information processing, and in particular to a method for fusing and analyzing key information of documents based on vectorization technology. Background Art

[0002] Today, with the rapid development of information technology, document processing and analysis technologies have increasingly become an important part of the information processing field. Especially in the big data era, how to quickly and accurately extract key information from a large number of documents is of great significance for improving information processing and decision-making efficiency. The method for fusing and analyzing key information of documents based on vectorization technology is an efficient and advanced document processing technology born to meet this need.

[0003] This method realizes the effective extraction and fusion of key information of documents through steps such as collecting documents, performing Chinese word segmentation, filtering stop words to obtain keywords, constructing a vector representation matrix of documents, judging the logical relationship between documents, and generating a summary description graph. This method can not only quickly identify the core information in the document, but also reveal the internal logical relationship between documents, thus providing strong support for the further analysis and utilization of documents.

[0004] However, the traditional method for fusing and analyzing key information of documents based on vectorization technology still has some deficiencies in specific use. First, when constructing the vector representation of documents, the traditional method often only considers simple statistical features such as the word frequency of keywords, while ignoring important information such as the position and context relationship of keywords in the document, which leads to a deviation in the generated vector representation matrix in reflecting the true content of the document. Second, when judging the logical relationship between documents, the traditional method usually relies on simple similarity calculations, such as cosine similarity, etc. This calculation method is unable to handle complex logical relationships and is difficult to accurately reveal complex relationships such as superior-subordinate and parallel relationships between documents. In addition, when generating a summary description graph, the traditional method often only marks the association relationship between keywords, while ignoring the priority and hierarchical structure between relationships, which makes the generated summary description graph have certain limitations in readability and usability. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a method for fusing and analyzing key information of documents based on vectorization technology.

[0006] To solve the above problems, the technical solution of the present invention is as follows: including the following steps:

[0007] S1. Collect documents and perform Chinese word segmentation on the documents;

[0008] S2. Filter the words in the documents to obtain the keywords K in the documents;

[0009] S3. Construct a vector representation matrix VM of the document using the keyword K;

[0010] S4. Determine the logical relationship between the documents according to the vector representation matrix VM, and obtain an abstract description graph based on the logical relationship;

[0011] S5. Perform query and analysis according to the abstract description graph to obtain a text analysis result.

[0012] Further, the step S2 specifically includes the following steps:

[0013] S2.1. Construct a preset word library and obtain the corresponding stop words in the preset word library;

[0014] S2.2. Filter the words in the document according to the stop words to obtain the keyword K in the document.

[0015] Further, the step S3 specifically includes the following steps:

[0016] S3.1. Create a matrix of the document and obtain the row vectors and column vectors in the matrix;

[0017] S3.2. Replace the row vectors and column vectors in the matrix with the keyword K to obtain the vector representation of the keyword K;

[0018] S3.3. Integrate the vector representations of the keyword K to obtain the vector representation matrix VM of the document.

[0019] Further, in the step S3.3, the vector representations QK of the keyword K are integrated according to the formula to obtain the vector representation matrix of the document, and the formula is as follows:

[0020]

[0021] where M ∈ R m×n , n represents the length of the matrix, m represents the height of the matrix, R represents the set of positive real numbers, and [X1, X2, … X n represents the set of vector representations corresponding to the keyword K.

[0022] Further, the step S4.1 includes the following steps:

[0023] S4.1.1. Calculate the number of words in two rows;

[0024] S4.1.2. Obtain a word similarity matrix EK according to the number of words;

[0025] S4.1.3. Put the word similarity matrix EK into a word similarity queue;

[0026] S4.1.4. Traverse the word similarity queue and calculate the similarity between two rows.

[0027] Further, step S5 queries and analyzes according to the abstract description graph to obtain a text analysis result, which specifically includes the following:

[0028] Construct upper and lower level keyword pairs and upper and lower level association relationships according to the upper and lower level relationships;

[0029] Construct multiple parallel keyword pairs and parallel association relationships according to the parallel relationships;

[0030] Construct the abstract description graph according to the upper and lower level keyword pairs, upper and lower level association relationships, multiple parallel keyword pairs and multiple parallel association relationships;

[0031] Query whether the input information appears at a preset position in the document according to the abstract description graph, and output the text analysis result.

[0032] Further, step S2 specifically includes the following steps:

[0033] S2.1. Construct a preset word library and obtain the corresponding stop words in the preset word library;

[0034] S2.2. Filter the words in the document according to the stop words to obtain the keyword K in the document.

[0035] Further, S3.2 specifically includes the following steps:

[0036] S3.2.1. Obtain the word vector T of the document;

[0037] S3.2.2. Construct the vectorized representation ZK of the keyword K;

[0038] S3.2.3. Obtain the position set corresponding to the keyword K according to the vectorized representation ZK;

[0039] S3.2.3. Obtain the vector representation QK of the keyword K according to the position set and the word vector T.

[0040] Further, step S4 specifically includes the following steps:

[0041] S4.1. Calculate the row similarity SM of the vector representation matrix VM m , m = 1, 2, 3,..., M;

[0042] S4.2. Calculate the column similarity SN of the vector representation matrix VM n , n = 1, 2, 3,..., N;

[0043] S4.3. Obtain the abstract description graph according to the row similarity SM m and the column similarity SN n .

[0044] Further, the operations included in the step S4.3 are as follows:

[0045] S4.3.1. Determine whether the row similarity is less than the column similarity. If so, execute step S4.3.2; otherwise, execute step S4.3.3;

[0046] S4.3.2. If the row similarity is less than the column similarity, obtain the association matrix FM of the matrix. The association matrix FM is the lower triangular matrix of the matrix diagonal, and obtain the superior-inferior relationship;

[0047] S4.3.3. If the row similarity is greater than the column similarity, obtain the association matrix FM of the matrix. The association matrix FM is the upper and lower triangular matrix of the matrix, and obtain the parallel relationship;

[0048] S4.3.4. If the row similarity is equal to the column similarity, obtain the association matrix FM of the matrix as a symmetric matrix structure, and obtain the parallel relationship.

[0049] Further, the construction criteria of the abstract description graph further include:

[0050] If the document contains only one relationship, draw an abstract description graph based on the logical relationship and label the relationship;

[0051] If the document contains multiple relationships, draw an abstract description graph based on the logical relationship and label multiple relationships and the priorities between the relationships.

[0052] The advantages of the present invention compared with the existing technologies are as follows:

[0053] 1. The present invention provides a method for fusing and analyzing key information of a document based on a vectorization technology. When constructing the vector representation matrix VM, not only the word frequency of the keyword K is considered, but also the position set of the keyword is obtained, and the vector representation QK of the keyword K is obtained by combining the word vector T. This makes the generated vector representation matrix better reflect the true content of the document, because the position and context information of the keyword are crucial for understanding the meaning of the document;

[0054] 2. The present invention provides a method for fusing and analyzing key information of documents based on vectorization technology. By calculating the row similarity SMm and the column similarity SNn to infer the logical relationship between documents, it does not simply rely on a single index such as cosine similarity. By comparing the similarities of rows and columns, complex logical relationships such as superior-subordinate relationships and parallel relationships can be accurately distinguished, thereby providing a more refined and accurate relationship description;

[0055] 3. The present invention provides a method for fusing and analyzing key information of documents based on vectorization technology. When generating a summary description graph, this technical solution not only marks the association relationships between keywords, but also further introduces the concepts of priority and hierarchical structure between relationships. When a document contains multiple relationships, these relationships and their priorities will be clearly marked. This approach improves the readability and usability of the summary description graph, making it easier for users to understand and utilize the analysis results. Brief Description of the Drawings

[0056] Figure 1 is a flowchart of a method for fusing and analyzing key information of documents based on vectorization technology according to the present invention. Detailed Description of the Embodiments

[0057] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices consistent with some aspects of the present disclosure as detailed in the appended claims.

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments; based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] As Figure 1 , this embodiment proposes a method for fusing and analyzing key information of documents based on vectorization technology, including the following steps:

[0060] S1. Collect documents and perform Chinese word segmentation on the documents;

[0061] S2. Filter the words in the documents to obtain the keywords K in the documents;

[0062] S3. Use the keywords K to construct a vector representation matrix VM of the documents;

[0063] S4. Determine the logical relationship between documents based on the vector representation matrix VM, and obtain a summary description graph based on the logical relationship;

[0064] S5. Perform queries and analysis based on the summary description graph to obtain the text analysis results.

[0065] Furthermore, step S2 specifically includes the following steps:

[0066] S2.1. Construct a preset word library and obtain the corresponding stop words in the preset word library;

[0067] S2.2. Filter the words in the document according to the stop words to obtain the keyword K in the document.

[0068] Furthermore, step S3 specifically includes the following steps:

[0069] S3.1. Create a matrix of the document and obtain the row vectors and column vectors in the matrix;

[0070] S3.2. Replace the row vectors and column vectors in the matrix with the keyword K to obtain the vector representation of the keyword K;

[0071] S3.3. Integrate the vector representation of the keyword K to obtain the vector representation matrix VM of the document.

[0072] Furthermore, S3.2 specifically includes the following steps:

[0073] S3.2.1. Obtain the word vector T of the document;

[0074] S3.2.2. Construct the vectorized representation ZK of the keyword K;

[0075] S3.2.3. Obtain the position set corresponding to the keyword K according to the vectorized representation ZK;

[0076] S3.2.3. Obtain the vector representation QK of the keyword K according to the position set and the word vector T.

[0077] Furthermore, in step S3.3, the vector representation QK of the keyword K is integrated according to the formula to obtain the vector representation matrix of the document. The formula is as follows:

[0078]

[0079] where M ∈ R m×n , n represents the length of the matrix, m represents the height of the matrix, R represents the set of positive real numbers, and [X1, X2, … X n represents the set of vector representations corresponding to the keyword K.

[0080] Furthermore, step S4 specifically includes the following steps:

[0081] S4.1. Calculate the row similarity SM of the vector representation matrix VM m , where m = 1, 2, 3, …, M;

[0082] S4.2. Calculate the column similarity SN of the vector representation matrix VM n , where n = 1, 2, 3, …, N;

[0083] S4.3. Obtain the abstract description graph based on the row similarity SM m and the column similarity SN n .

[0084] Furthermore, step S4.1 includes the following steps:

[0085] S4.1.1. Calculate the number of words in two rows;

[0086] S4.1.2. Obtain the word similarity matrix EK based on the number of words;

[0087] S4.1.3. Put the word similarity matrix EK into the word similarity queue;

[0088] S4.1.4. Traverse the word similarity queue and calculate the similarity between two rows.

[0089] Furthermore, step S4.3 includes the following operations:

[0090] S4.3.1. Determine whether the row similarity is less than the column similarity. If so, execute step S4.3.2; otherwise, execute step S4.3.3;

[0091] S4.3.2. If the row similarity is less than the column similarity, obtain the association relationship matrix FM of the matrix. The association relationship matrix FM is the lower triangular matrix of the matrix diagonal, and obtain the superior-subordinate relationship;

[0092] S4.3.3. If the row similarity is greater than the column similarity, obtain the association relationship matrix FM of the matrix. The association relationship matrix FM is the upper and lower triangular matrix of the matrix, and obtain the parallel relationship;

[0093] S4.3.4. If the row similarity is equal to the column similarity, obtain the association relationship matrix FM of the matrix as a symmetric matrix structure, and obtain the parallel relationship.

[0094] Furthermore, step S5 queries and analyzes according to the abstract description graph to obtain the text analysis result, which specifically includes the following contents:

[0095] Construct superior-subordinate keyword pairs and superior-subordinate association relationships according to the superior-subordinate relationship;

[0096] Construct multiple parallel keyword pairs and parallel association relationships according to the parallel relationship;

[0097] Construct an abstract description graph based on upper and lower level keyword pairs, upper and lower level association relationships, multiple parallel keyword pairs, and multiple parallel association relationships;

[0098] Query whether the input information appears at the preset position in the document according to the abstract description graph, and output the text analysis result.

[0099] Furthermore, the construction criteria for the abstract description graph also include:

[0100] If the document contains only one type of relationship, draw an abstract description graph based on the logical relationship and label the relationship;

[0101] If the document contains multiple types of relationships, draw an abstract description graph based on the logical relationship and label multiple relationships and the priority between relationships.

[0102] Specifically, when in use, refer to Figure 1 As shown, the system obtains the document to be analyzed from the database or user input, uses a mature Chinese word segmentation tool (such as Jieba segmentation) to perform word segmentation on the document, divides the continuous text into independent lexical units for more refined operations in subsequent steps, constructs a preset word library, a preset word library containing common stop words (such as "de", "shi", etc.), applies stop word filtering to the word segmentation result to remove meaningless words, and retains the keyword K. Suppose we have a simple sentence: "The weather is really nice today and it is suitable to go out for a walk." After word segmentation, the result is "today / weather / really / nice / , / suitable / go out / walk / ." After applying stop word filtering, the keyword K may be "weather / nice / suitable / walk".

[0103] Create a matrix to represent the document, where the rows represent different document fragments or sentences, and the columns represent all different words that may appear in these fragments. Initialize an empty matrix M, where each row represents a document fragment or sentence, and each column represents a potential word. Use a pre-trained word embedding model (such as Word2Vec, GloVe, etc.) to find the corresponding word vector T for each keyword K, determine the position of each keyword in the document, and calculate the final vector representation QK of the keyword according to the position set and the word vector TT. The formula for QK is as follows:

[0104] QK = f(ZK, T)

[0105] Here, f(.) represents a certain function, such as weighted average or direct value extraction.

[0106] Integrate the vector representation QK of the keyword K according to the formula to obtain the vector representation matrix of the document. The formula is as follows:

[0107]

[0108] where \(M\in R\) m×n , \(n\) represents the length of the matrix, \(m\) represents the height of the matrix, \(R\) represents the set of positive real numbers, and \([X_1, X_2, \ldots, X\) n represents the set of vector representations corresponding to the keyword \(K\).

[0109] If we have three keywords and their vector representations \(QK_1\), \(QK_2\), \(QK_3\) respectively, then the integrated matrix \(M\) will be a linear combination of these vectors.

[0110] Calculate the number of words between two rows to form a word similarity matrix \(EK\), put \(EK\) into a queue, and traverse the queue to calculate the similarity between two rows. The formula for calculating the row similarity is as follows:

[0111] \(SM\) m \(= g(EK)\)

[0112] Here, \(g(\cdot)\) represents the function used to calculate the similarity, such as cosine similarity.

[0113] Similarly, calculate the number of corresponding words between two columns to form a word similarity matrix. The formula for calculating the column similarity is as follows:

[0114] \(SN\) n \(= h(EK\) ′ )

[0115] Here, \(h(\cdot)\) also represents the function used to calculate the similarity, and \((EK\) ′ ) represents the word similarity matrix between columns. According to the results of the row similarity \(SM\) m and the column similarity \(SN\) n , construct an association relationship matrix \(FM\). If the row similarity is less than the column similarity, it is considered that there is a superior-subordinate relationship; otherwise, it is considered a parallel relationship; when the two are equal, it is defined as a parallel relationship. Suppose the row similarity is \(0.7\) and the column similarity is \(0.9\), then according to the rule, this will be regarded as a parallel relationship.

[0116] Based on the generated summary description graph, efficient text analysis can be performed for specific queries. For example, check whether the input information appears in the preset position of the document and output the corresponding text analysis results.

[0117] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device.

[0118] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

[0119] The above describes the present invention and its implementation manners. Such description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of the present invention, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present invention.

Claims

1. A document key information fusion analysis method based on vectorization technology, characterized in that: The following steps are involved: S1. Collect documents and perform Chinese word segmentation on the documents; S2, filtering the words in the document to obtain the keyword K in the document; S3, constructing a vector representation matrix VM of the document using the keyword K; S4, judging the logical relationship between the documents according to the vector representation matrix VM, and obtaining a summary description graph based on the logical relationship; S5. Query and analyze according to the summary description graph to obtain text analysis results.

2. According to the method for fusion analysis of key information of documents based on vectorization technology in claim 1, it is characterized in that: The step S2 specifically includes the following steps: S2.

1. Build a preset word library and obtain the corresponding stop words in the preset word library; S2.

2. Filter the words in the document according to the stop words to obtain the keyword K in the document.

3. According to the method for fusion analysis of key information of documents based on vectorization technology in claim 2, it is characterized in that: The step S3 specifically comprises the following steps: S3.

1. Create a matrix of the document and obtain row vectors and column vectors in the matrix; S3.

2. Use the keyword K to replace the row vector and column vector in the matrix to obtain a vector representation of the keyword K; S3.

3. Integrate the vector representation of the keyword K to obtain the vector representation matrix VM of the document.

4. According to the method for fusion analysis of key information of documents based on vectorization technology as claimed in claim 3, it is characterized in that: The S3.2 specifically includes the following steps: S3.2.

1. Obtain the word vector T of the document; S3.2.2, constructing a vectorized representation ZK of the keyword K; S3.2.3, obtaining a position set corresponding to the keyword K according to the vectorized representation ZK; S3.2.

3. According to the position set and the word vector T, obtain the vector representation QK of the keyword K.

5. According to the method for fusion analysis of key information of documents based on vectorization technology as claimed in claim 4, it is characterized in that: In step S3.3, the vector representation QK of the keyword K is integrated according to the formula to obtain the vector representation matrix of the document, and the formula is as follows: Where M∈R m×n , n represents the length of the matrix, m represents the height of the matrix, R represents the positive real number set, [X1,X2,…X n ] represents the vector representation set corresponding to the keyword K.

6. The document key information fusion analysis method based on vectorization technology according to claim 5 is characterized in that: The step S4 specifically comprises the following steps: S4.

1. Calculate the row similarity SM of the vector representation matrix VM m , m=1,2,3,…,M; S4.

2. Calculate the column similarity SN of the vector representation matrix VM n , n=1,2,3,…,N; S4.3, according to the row similarity SM m and column similarity SN n , and obtain the summary description graph.

7. The document key information fusion analysis method based on vectorization technology according to claim 1 is characterized in that: The step S4.1 includes the following steps: S4.1.

1. Count the number of words in two lines; S4.1.2, according to the number of words, obtain a word similarity matrix EK; S4.1.3, putting the word similarity matrix EK into the word similarity queue; S4.1.

4. Traverse the word similarity queue and calculate the similarity between two rows.

8. The method for fusion analysis of key information of a document based on vectorization technology according to claim 7 is characterized in that: The step S4.3 includes the following operations: S4.3.1, determine whether the row similarity is less than the column similarity, if so, execute step S4.3.2, otherwise execute step S4.3.3; S4.3.

2. If the row similarity is less than the column similarity, then the association matrix FM of the matrix is ​​obtained, and the association matrix FM is a lower triangular matrix of the diagonal of the matrix, and the superior-subordinate relationship is obtained; S4.3.

3. If the row similarity is greater than the column similarity, then the association matrix FM of the matrix is ​​obtained, and the association matrix FM is the upper and lower triangular matrix of the matrix, and a parallel relationship is obtained; S4.3.

4. If the row similarity is equal to the column similarity, the association relationship matrix FM of the matrix is ​​a symmetric matrix structure, and a parallel relationship is obtained.

9. The method for fusion analysis of key information of a document based on vectorization technology according to claim 8 is characterized in that: The step S5 performs query and analysis based on the summary description graph to obtain text analysis results, which specifically include the following contents: According to the superior-subordinate relationship, construct superior-subordinate keyword pairs and superior-subordinate association relationships; According to the parallel relationship, construct a plurality of parallel keyword pairs and parallel association relationships; Constructing the summary description graph according to the upper-lower keyword pairs, the upper-lower association relationships, the plurality of parallel keyword pairs and the plurality of parallel association relationships; According to the summary description graph, it is queried whether the input information appears at a preset position of the document, and the text analysis result is output.

10. The document key information fusion analysis method based on vectorization technology according to claim 9 is characterized in that: The construction criteria of the summary description diagram also include: If the document contains only one relationship, draw a summary description diagram based on the logical relationship and mark the relationship; If the document contains multiple relationships, draw a summary description diagram based on the logical relationships and mark the multiple relationships and the priorities between the relationships.