Information comparison method, device, electronic device and storage medium

By extracting multiple text information and metadata features from files, fusing and processing the text features, and determining file similarity, the problem of insufficient accuracy in file screening and similarity comparison in existing technologies is solved, and more efficient file comparison is achieved.

CN115329850BActive Publication Date: 2025-09-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210920358.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-02
Publication Date
2025-09-12
Estimated Expiration
2042-08-02

AI Technical Summary

Technical Problem

It is difficult to effectively filter out desired files from massive files with existing technologies, and the accuracy of file similarity comparison is insufficient.

Method used

By extracting multiple text information and metadata features from the reference file, the text features of each text information are extracted separately, and fused together to obtain comprehensive text features. The similarity between the reference file and the file to be compared is determined based on the metadata features and the comprehensive text features.

Benefits of technology

The accuracy of file similarity comparison is improved, and the similarity between files can be described more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115329850B_ABST
    Figure CN115329850B_ABST
Patent Text Reader

Abstract

The present disclosure provides an information comparison method, device, electronic device and storage medium, which relate to the field of artificial intelligence technology, and in particular to the field of intelligent search technology. The specific implementation scheme is: extracting multiple text information from the text content of the reference file, and extracting metadata features based on the metadata of the reference file; extracting the text features of each text information separately, and fusing each text feature to obtain a comprehensive text feature; based on the metadata feature and the comprehensive text feature, determining the similarity between the reference file and the file to be compared. In the embodiment of the present disclosure, multiple text information are extracted from the reference file, and text features are extracted separately, which is conducive to refining the ideological features independently expressed by each text information. The comprehensive text features obtained by combining multiple text features can represent the overall text features. Further combined with metadata features, it is possible to achieve feature description of the reference file from multiple dimensions, thereby improving the accuracy of file similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the field of intelligent search technology. Background Art

[0002] The accumulation of documents over the years, coupled with the constant creation of new ones, has resulted in a massive volume of documents. While databases can be used to manage these documents, and fuzzy search techniques can be used to retrieve relevant documents, the challenge remains to identify the desired documents from this vast volume. Summary of the Invention

[0003] The present disclosure provides an information comparison method, device, electronic device and storage medium.

[0004] According to a first aspect of the present disclosure, there is provided an information comparison method, comprising:

[0005] extracting multiple pieces of text information from the text content of the reference file, and extracting metadata features based on the metadata of the reference file;

[0006] Extract the text features of each text information separately, and fuse the text features of each text information to obtain comprehensive text features;

[0007] Based on metadata features and comprehensive text features, the similarity between the reference file and the file to be compared is determined.

[0008] According to a second aspect of the present disclosure, there is provided an information comparison device, comprising:

[0009] an acquisition module, configured to extract multiple pieces of text information from the text content of the reference file, and extract metadata features based on the metadata of the reference file;

[0010] The extraction module is used to extract the text features of each text information separately, and fuse the text features of each text information to obtain the comprehensive text features;

[0011] The comparison module is used to determine the similarity between the reference file and the file to be compared based on metadata features and comprehensive text features.

[0012] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect mentioned above.

[0016] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the aforementioned first aspect.

[0017] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method of the aforementioned first aspect when executed by a processor.

[0018] The solution provided in this embodiment extracts multiple pieces of text information from a reference file and extracts text features from each piece of text information separately, facilitating the refinement of the conceptual characteristics independently expressed by each piece of text information. Combining multiple text features yields comprehensive text features that represent the overall text characteristics of the reference file. Furthermore, by integrating the metadata features of the reference file, it is possible to describe the characteristics of the reference file from multiple dimensions, thereby improving the accuracy of file similarity.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0021] Figure 1 is a flowchart of an information comparison method according to an embodiment of the present disclosure;

[0022] Figure 2 is another flowchart of an information comparison method according to another embodiment of the present disclosure;

[0023] Figure 3 Schematic diagram of a model structure for extracting text features in an information comparison method according to an embodiment of the present disclosure;

[0024] Figure 4 is another flowchart of an information comparison method according to another embodiment of the present disclosure;

[0025] Figure 5 2 is a schematic diagram of another model structure for extracting text features in an information comparison method according to another embodiment of the present disclosure;

[0026] Figure 6 is another flowchart of an information comparison method according to another embodiment of the present disclosure;

[0027] Figure 7 is another flowchart of an information comparison method according to another embodiment of the present disclosure;

[0028] Figure 8 Schematic diagram of a model structure for extracting comprehensive text features in an information comparison method according to an embodiment of the present disclosure;

[0029] Figure 9 This is another schematic diagram of a model structure for extracting comprehensive text features in an information comparison method according to another embodiment of the present disclosure;

[0030] Figure 10 2 is a schematic diagram of a model structure for extracting comprehensive text features and metadata features in an information comparison method according to another embodiment of the present disclosure;

[0031] Figure 11 1 is a schematic diagram of the structure of an information comparison device according to an embodiment of the present disclosure;

[0032] Figure 12 is another structural diagram of an information comparison device according to another embodiment of the present disclosure;

[0033] Figure 13 It is a block diagram of an electronic device used to implement the information comparison method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0035] The terms "first" and "second" and the like in the embodiments and claims of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0036] The present disclosure provides an information comparison method, such as Figure 1 The process flow diagram of the method is shown, which includes the following steps:

[0037] S101 , extracting multiple pieces of text information from the text content of a reference file, and extracting metadata features based on the metadata of the reference file.

[0038] S102 , extracting text features of each piece of text information respectively, and fusing the text features of each piece of text information to obtain comprehensive text features.

[0039] Taking patent application documents as an example, multiple text information may include at least two of the following: abstract, title, claims, and technical effect description. The abstract can reflect the core invention of the patent application document and complete the overview of the technical solution; the title of the patent application document is a thematic summary of the content involved in the patent application document; the claims can concisely and directly express the core solution of the patent application, and the technical effect description can deeply explain the technical effects brought about by the implementation scheme. All of the above text information can reflect the core content of the patent application document. Therefore, using at least one of the above text information can extract the core idea of ​​the patent application document, so as to accurately retrieve similar patents to the patent application document.

[0040] In addition, other types of documents, such as papers, journals, magazines, and novels, generally have abstracts and titles. For other types of documents, the multiple text items can be selected as abstracts and / or titles. Alternatively, appropriate multiple text items can be selected based on actual needs, such as the template requirements of a journal. This is not limited in the present embodiment.

[0041] S103: Determine the similarity between the reference file and the file to be compared based on the metadata features and the comprehensive text features.

[0042] In the embodiment of the present disclosure, multiple text information are extracted from the reference file, and text features are extracted from each text information respectively, which is conducive to separately refining the ideological features independently expressed by each text information. Then, the text features of each text information are combined to obtain the comprehensive text features of the reference file, so that the comprehensive text features contain the ideological features independently expressed by each text information, while being able to represent the overall text features of the reference file. The embodiment of the present disclosure further combines the metadata features of the reference file, and can realize the feature description of the reference file from multiple dimensions. When comparing with the file to be compared, the feature description based on multiple dimensions can accurately describe the similarity between the reference file and the file to be compared.

[0043] In some embodiments, in order to better extract the text features of the reference file, the embodiments of the present disclosure may perform the following operations for each text information when extracting the text features of each text information separately: Figure 2 The operations shown include:

[0044] S201 : extracting initial text features of the text information using a first language model corresponding to the text information.

[0045] In the disclosed embodiments, the language model may be a pre-trained BERT (Bidirectional Encoder Representation from Transformers) model. Alternatively, a RoBERTa (Robustly Optimized BERT Pretraining Approach) model may be selected, or NEZHA (Nezha), a Chinese pre-trained language model based on BERT, may be selected. The specific model may be determined based on actual needs and is not limited in the disclosed embodiments.

[0046] S202 : Input the initial text features of the text information into a first fully connected layer corresponding to the text information to obtain text features of the text information output by the first fully connected layer.

[0047] After extracting multiple pieces of text information from the reference document, each piece of text information is converted into a text sequence. The resulting text sequence may contain or exclude punctuation, as is not limited in the present disclosure. For example, a document focused on literature may need to express emotion, so punctuation may be retained in the text sequence. However, a patent application document may focus more on technical solutions, so punctuation may be excluded.

[0048] After obtaining the text sequence for each piece of text information, each character in the text sequence can be converted into a word vector to obtain the word vector representation of the text sequence. In practice, word embedding technologies such as FastText (fast text classifier) ​​and word2vec (word to vector) can be used to convert the text sequence into a word vector representation.

[0049] Then the word vector representation of the text sequence is input into the respective BERT (i.e. the first language model) to extract the initial text features of each text information. Figure 3 As shown in the figure, assuming that the multiple text information extracted from the reference file include text item 1, text item 2 and text item 3, the word vector representation of each text item is input into the corresponding BERT model respectively to obtain the initial text features, and then processed by FC (fully connected layers) (i.e., the first fully connected layer) to output the text features of each text information.

[0050] In the embodiment of the present disclosure, each text information corresponds to its own first language model, and multiple language models are used in parallel to extract the initial text features of each text information, which can improve the efficiency of extracting the initial text features. Then, the initial text features are input into the fully connected layer to obtain text features. Among them, since each node of the fully connected layer is connected to all the nodes of the previous layer, the features extracted by the language model can be integrated. The initial text features extracted by the language model are processed by the fully connected layer, and the fully connected layer can play the role of a "classifier". The fully connected layer also plays the role of mapping the "distributed feature representation" learned by the language model to other feature spaces, so that the features extracted by the language model can be further sublimated and transformed into features that can be distinguished from other files, and the consistency of similar features can be guaranteed. Therefore, the embodiment of the present disclosure, based on the language model and the fully connected layer, can extract accurate features for text comparison to improve the accuracy of file comparison.

[0051] In some embodiments, some text items (i.e., text information) may be more complex. For example, the abstract in a patent application document includes a paragraph, the title also includes a paragraph, and the claims include multiple claims. In the embodiment of the present disclosure, each claim is regarded as a paragraph. Obviously, the content of the claims is more complex. In view of this, in the embodiment of the present disclosure, text information including multiple paragraphs can be defined as a complex text item. For complex text items, since each paragraph will express a content separately, such as each claim in the claims expresses a content separately, each paragraph of the complex text item can be processed separately in the hope of obtaining a comprehensive feature expression of the complex text item. In the embodiment of the present disclosure, the text information containing a single paragraph can be processed as follows Figure 2 and Figure 3 The described method extracts text features. For complex text items, in addition to Figure 2 and Figure 3 In addition to extracting text features in the manner described above, the present disclosure also provides the following Figure 4 The method shown in the figure extracts text features of a complex text item, including the following steps:

[0052] S401 : extracting sub-text features of each text segment in the complex text item based on the second language model corresponding to the complex text item.

[0053] S402 , performing dimensionality reduction processing on the sub-text features of each text segment to obtain dimensionality reduction features of the complex text item.

[0054] S403: Input the dimension reduction features of the complex text item into the second fully connected layer corresponding to the complex text item to obtain the text features of the complex text item output by the second fully connected layer.

[0055] exist Figure 3Based on this, the operation of extracting text features of complex text items is as follows Figure 5 As shown. Figure 5 In the example, assuming that text item 3 is a complex text item, each paragraph in the complex text item corresponds to a text sequence. Each text sequence is represented by a word vector using word embedding technology. The word vector representation of the complex text item forms a multidimensional matrix, which is input into the BERT model to obtain the initial text features of each paragraph.

[0056] Because complex text items have a large number of initial text features, in order to avoid focusing too heavily on the text features of complex text items and neglecting the text features of text items with fewer paragraphs, the disclosed embodiments perform dimensionality reduction on the text features of multiple paragraphs within a complex text item. Furthermore, dimensionality reduction further enables the extraction of deeper features within the text features of the complex text item, thereby improving the accuracy of the comparison results between the reference document and the document to be compared.

[0057] The dimensionality reduction method in the disclosed example can be selected as follows: Figure 5 The AVG (Word Averaging) model shown here averages the initial text features of each paragraph to achieve feature dimensionality reduction.

[0058] In other embodiments, linear dimensionality reduction methods such as subset selection, principal component analysis, etc. may also be selected for dimensionality reduction.

[0059] The above describes how to extract text features of various text information. The following describes how to fuse the text features of various text information to obtain comprehensive text features.

[0060] One possible implementation involves processing the text features of each piece of text information based on a hierarchical attention mechanism to generate comprehensive text features. This mechanism can focus on key features, making the resulting comprehensive text features more suitable for information comparison and improving the accuracy of information comparison.

[0061] In another embodiment, the text features of each piece of text information can be processed based on a hierarchical attention mechanism to obtain semantic features of the reference file. Afterwards, the semantic features are spliced ​​with the text features of at least one piece of text information to obtain comprehensive text features. In the disclosed embodiment, the hierarchical attention mechanism can help extract deep semantic features of the reference file. By splicing semantic features with text features of other text information, the comprehensive text features can not only express the semantics of the reference file but also integrate other text features, thereby being able to comprehensively describe the features of the reference file from multiple dimensions and improve the accuracy of text comparison.

[0062] In some embodiments, the text features of each text information are processed based on the hierarchical attention mechanism to obtain comprehensive text features, which can be implemented as follows: Figure 6 As shown:

[0063] S601, determining a complex text item containing multiple paragraphs in the multiple text information, and determining that the text information other than the complex text item in the multiple text information is simple text information;

[0064] S602, based on the complex text item, determining the key features, value features, and query features of the hierarchical attention mechanism; wherein the subtext features of each paragraph in the text features of the complex text item are the key features and value features, and the text features of the simple text information are the query features;

[0065] S603, determining optimized text features of the complex text item based on the key features, value features, and query features;

[0066] S604: Concatenate the optimized text features of the complex text item and the text features of the simple text item to obtain comprehensive text features.

[0067] It can also be understood as taking the sub-text features of each paragraph in the text features of the complex text item as the key features and value features of the hierarchical attention mechanism, and taking the text features of each simple text information as the query features, to obtain the sub-semantic features corresponding to each simple text information (which can also be understood as the sub-semantic features in the semantic features of the previous text); based on the sub-semantic features, the optimized text features of the complex text item are determined.

[0068] Taking patent application documents as an example, the extracted text information includes title, abstract and claims. Query vector (i.e. query feature), subtext feature of each claim in the text representation of the claims They are Key vector (i.e. key feature) and Value vector (i.e. value feature). is the subtext feature of claim i in the claims list, where the total number of claims is s. Each claim can be considered as a paragraph.

[0069] The Attention weight is calculated by the Query vector, Key vector, and Value vector, and then the weighted sum of each sub-text feature in the claim is performed. The calculation process is shown in formula (1):

[0070]

[0071] in, for Attention weight;

[0072] Used to measure titles and claims The similarity can be calculated by methods such as dot product and cosine similarity. The dimensions of the query vector, key vector and value vector are all d k , so √d k are the preset hyperparameters.

[0073] Similarly, you can also summarize Query vector, with multiple claims The Key vector and the Value vector are used to calculate the Attention weight. Then, the weighted sum of the claims is performed. The calculation process is shown in formula (2):

[0074]

[0075] in for Attention weight;

[0076] Used to measure titles and claims The similarity can be calculated by point multiplication, cosine similarity, etc., √d k are the preset hyperparameters.

[0077] Afterwards, the obtained sub-semantic features with title and summary as query are averaged to obtain the semantic features of the reference document, which is recorded as The calculation process is shown in formula (3):

[0078]

[0079] In summary, the disclosed embodiments of a complex text item can consider the differences in content conveyed by each paragraph within it. Using the subtext features of each paragraph as a benchmark, sub-semantic features are derived using the text features of other text information, thereby enabling the extracted sub-semantic features to better focus on key features. Derivation of semantic features for a reference document based on sub-semantic features can comprehensively describe the information characteristics conveyed by a complex text item, improving the accuracy of semantic feature extraction.

[0080] Furthermore, still taking the patent application document as an example, the text representation and semantic feature representation of the title and abstract are spliced ​​together to obtain a spliced ​​representation. The spliced ​​feature can be expressed as Then, the concatenated representation is processed by the fully connected layer to obtain the comprehensive text features of the reference file. The processing process of the concatenated features by the fully connected layer can be expressed as shown in formula (4):

[0081]

[0082] Among them, W o and b o are the parameters of the fully connected layer.

[0083] In some embodiments, semantic features can be combined with the text features of the abstract, or they can be combined with the text features of the title alone. Of course, they can also be combined with the text features of the abstract, title, and claims. That is, in the embodiments of the present disclosure, semantic features can be combined with the text features of some of the text information of multiple pieces of text information, or they can be combined with the text features of all the text information, and the embodiments of the present disclosure are not limited to this.

[0084] In some embodiments, in addition to using a hierarchical attention mechanism to obtain comprehensive text features, the text features of each piece of text information can also be spliced ​​together to obtain comprehensive text features. The splicing process is simple and easy to operate, and can obtain comprehensive text features while improving the efficiency of extracting comprehensive text features.

[0085] In the disclosed embodiment, the text features of each piece of text information are extracted separately, and the text features of each piece of text information are fused to obtain a comprehensive text feature. This can be achieved based on a comprehensive text feature network model, that is, the text features of each piece of text information are obtained through the comprehensive text feature network model, and are fused to obtain a comprehensive text feature.

[0086] The comprehensive text feature network model can use artificial intelligence technology to mine the text features of the reference file to improve the accuracy of the comparison results between the reference file and the file to be compared.

[0087] The comprehensive text feature network model can use the aforementioned hierarchical attention mechanism to obtain comprehensive text features, or it can be obtained by splicing the text features of various text information. Regardless of the method used, the embodiment of the present disclosure can use the following method to train the comprehensive text feature network model, such as Figure 7 As shown, the following steps are included:

[0088] S701, extract multiple text information from the same file to construct a positive sample, and extract multiple text information from different files to construct a negative sample.

[0089] For example, multiple pieces of text information include the title, abstract, and claims. If these three pieces of text information all come from the same patent application document, it is a positive sample. If at least one of the pieces of text information comes from another document, it is a negative sample. For example, extracting the title, abstract, and claims from patent application document A creates a positive sample. Extracting the title and abstract from patent application document A, but extracting the claims from patent application document B, creates a negative sample from these three pieces of text information.

[0090] S702: Input the positive sample and the negative sample into the initial text feature network respectively to obtain the comprehensive text features of the positive sample and the comprehensive text features of the negative sample output by the initial text feature network.

[0091] S703 , using a classifier to perform classification processing on the comprehensive text features of the positive sample and the comprehensive text features of the negative sample, respectively, to obtain classification processing results, wherein the classification categories of the classifier include positive samples and negative samples.

[0092] S704 : Determine a classification loss value based on the classification processing result, the category label of the positive sample, and the category label of the negative sample.

[0093] In the embodiment of the present disclosure, positive and negative samples are input into a classifier for discrimination prediction to obtain classification processing results of positive and negative samples, and the loss function can adopt a commonly used cross entropy loss function.

[0094] S705: Based on the classification loss value, adjust the model parameters of the initial text feature network to obtain a comprehensive text feature network model.

[0095] In the disclosed embodiments, training samples do not require manual labeling, improving sample acquisition efficiency. Furthermore, a classification model is used to combine positive and negative samples for training. Training only requires classifying positive and negative samples. This makes model training simple and efficient, enabling the rapid acquisition of a converged comprehensive text feature network model.

[0096] The following describes two comprehensive text feature network model structures: one that obtains comprehensive text features by concatenating text features and the other that obtains comprehensive text features by using a hierarchical attention mechanism.

[0097] like Figure 8 As shown in the figure, it is a structural diagram of the comprehensive text feature network model that obtains comprehensive text features by concatenating text features. Figure 8 The network model includes a language model (BERT), a fully connected layer (FC), a dimensionality reduction layer, and a concatenation layer.

[0098] Taking patent application documents as an example, we extract the text sequences of title, abstract and claims from the patent application text. Each claim in the claims is considered as a paragraph, and the text sequence is extracted from each claim separately. The word vector of the text sequence of the title is defined as The word vector representation of the abstract is The word vector representation of the claims is in l is the length of the word sequence, s is the number of claims;

[0099] like Figure 8 As shown, the word vector representations of each text information are processed by their respective language models and then passed through their respective fully connected layers to obtain the text features of the title. Text features of abstracts and the textual features of the claims Figure 8 In the claims, each claim is extracted by the language model, and then the features of all claims are reduced in dimension by the AVG dimensionality reduction layer and then input into the fully connected layer FC.

[0100] The process of processing by the fully connected layer is shown in formula (5):

[0101]

[0102] In formula (5), W t 、b t 、W a 、b a 、W c 、b c These are the parameters that need to be trained for each fully connected layer; represents the result obtained after the i-th word vector in the title is processed by the BERT model; l t Indicates the length of the text sequence in the title; similarly, Represents the result obtained after the i-th word vector in the summary is processed by the BERT model; l a Indicates the length of the text sequence in the summary; The result obtained by processing the word vector of claim i in the claims through the BERT model; Indicates the number of claims in the claims.

[0103] Finally, the text features of the title Text features of abstracts and the textual features of the claims Comprehensive text features are obtained by concatenating the concatenation layer Concat

[0104] like Figure 9 As shown in the figure, it is a structural diagram of the comprehensive text feature network model that obtains comprehensive text features using the hierarchical attention mechanism. Figure 9 The network model includes a language model (BERT), a third fully connected layer (FC), a hierarchical attention mechanism network, a splicing layer, and a fourth fully connected layer.

[0105] Taking patent application documents as an example, we extract the text sequences of title, abstract and claims from the patent application text. Each claim in the claims is considered as a paragraph and the text sequence is extracted separately. The word vector of the text sequence of the title is defined as The word vector of the abstract is represented as The word vector representation of the claims is in l is the length of the word sequence and s is the number of claims.

[0106] After being processed by their respective language models and then passing through the first fully connected layer, the text features of the title are obtained Text features of abstracts and the textual features of the claims

[0107] The process of processing each text information through its respective third fully connected layer is shown in formula (5), which will not be repeated here. Figure 9 The features of the claims may not be subjected to dimensionality reduction.

[0108] Then with the title Query vector, with the claim For Key vector and Value vector, through the hierarchical attention mechanism network ( Figure 9 Attention in , and get the first sub-semantic feature.

[0109] Similarly, the abstract Query vector, with the claim For Key vector and Value vector, through the hierarchical attention mechanism network ( Figure 9 Attention in , and obtain the second sub-semantic feature.

[0110] Then the semantic feature is obtained by averaging the first and second sub-semantic features. Finally, the semantic features Text features of the title and text features of the abstract After processing through the Concat layer and then through the fourth fully connected layer, the comprehensive text features are obtained

[0111] exist Figure 9 Based on the comprehensive text features, the metadata features of the reference files can be The fusion is performed to obtain the feature representation of the reference document. Figure 10 A schematic diagram of a citation representation network for extracting metadata features is shown in FIG. In the embodiment of the present disclosure, a patent embedding network can be defined, which includes a comprehensive text feature network model for extracting comprehensive text features and a citation representation network for extracting metadata features.

[0112] For ease of understanding, refer to Figure 10 How to obtain the citation representation network is described.

[0113] In the embodiment of the present disclosure, metadata is first extracted from the reference document. Taking the patent application document as an example, the metadata content is shown in Table 1. It should be noted that Table 1 is only used to illustrate the embodiment of the present disclosure, and does not limit the embodiment of the present disclosure.

[0114] Table 1

[0115]

[0116]

[0117] The forward citation trend can be determined based on the forward citations of Patent A in recent years. For example, if the number of forward citations in 2017 is a and the number of forward citations in 2018 is b, the forward citation trend can be expressed based on the number of forward citations in recent years.

[0118] After obtaining the metadata, a reference representation network F is established. Each node in the reference representation network F is a file, and the attribute representation of each node is the metadata of the file.

[0119] Reference network G={V,E,X}, where V={v k |k=1,2,…,N} represents all file nodes, Indicates the reference relationship between file nodes, X={X k |k∈S P} represents the metadata of the file node. If the file node References the file node Then we can observe an edge This edge represents the reference relationship between file nodes, and nodes with a reference relationship become neighbor nodes. Subsequently, through network embedding learning, the attribute representations of the nodes in the high-dimensional, unstructured reference representation network are converted into low-dimensional representations, thereby achieving a more accurate characterization of file features.

[0120] Taking patent application documents as an example, we first integrate the attribute representations of all patent application documents to form an attribute matrix Where N is the number of patent application documents, d m The dimension of each node's attribute representation. Attribute matrix X m The kth row of k , represents the attribute representation of patent k.

[0121] According to the random walk strategy, taking each node in the patent citation network G as the root node and randomly sampling its neighbor nodes, different paths can be generated, such as:

[0122] <root,neighborhood1,neighborhood2,…>

[0123] For each of the above paths, the set of neighbor nodes of patent k is also called the context information of patent k. The context information expression is shown in formula (6):

[0124] context(v k )={v k-s ,…,v k+s}\{v k} (6)

[0125] Formula (6) indicates that the extracted context information of patent k needs to exclude k itself; s is used to limit the length of the context information.

[0126] Afterwards, the reference representation network F is trained by maximizing the following conditional probability (as shown in Equation (7)), which expresses the probability of inferring k as the central node based on the neighbor nodes:

[0127]

[0128] Among them, the element value range j represents the neighbor node of k; and They represent the metadata features of patent k output through the citation representation network F and the metadata features of context information output through the citation representation network F, It is defined as shown in formula (8):

[0129]

[0130] in, Represents the metadata features output by the reference representation network F of neighbor node j in the context information of patent k.

[0131] In order to convert the patent node X m Embedded into the network learning process, defining the reference representation network For each patent k, its attribute representation The conversion process is

[0132] When training the reference representation network F, in the embodiment of the present disclosure, non-adjacent nodes of patent node k are used as negative samples of patent node k. The objective function is approximated by the negative sampling strategy and parameters are optimized. The objective function is shown in formula (9):

[0133]

[0134] Wherein, the parameter of the σ() function in Equation 9 is simplified by x, σ(x)=1 / (1+exp(-x)) is the Sigmoid function, neg is the number of negative samples sampled for each positive sample k, E is the mathematical expectation, P n (v)∝d v 3 / 4 is the noise distribution of Layer negative sample nodes, d v Represents the out-degree of node v.

[0135] After the citation representation network is trained, for each patent k, its attribute matrix is ​​input to obtain its metadata features.

[0136] like Figure 10 As shown in the figure, the input layer of the citation representation network inputs the patent citation network, with node k as the center. The context information and attribute representation of node k are extracted from the patent citation network through sequence generation. The context information and attribute representation are input into the citation representation network F for training the citation representation network. The trained citation representation network can obtain the metadata features of patent k based on the attribute representation and context information of patent k. Figure 10 The reference representation network is used to project the attribute matrix into the feature space of metadata features through embedding learning. The metadata features extracted by the reference representation network and comprehensive text features extracted by the comprehensive text feature network model In the fusion layer, the data is first concatenated by the Concat layer and then processed by the fully connected layer (FC). This yields the comparison features of the patent application document. This comparison feature can be used to calculate similarity with the comparison features of other documents, thereby enabling the search for similar patents to the patent application document.

[0137] Based on the same technical concept, the present disclosure also provides an information comparison device, such as Figure 11 Shown, including:

[0138] An acquisition module 1101 is configured to extract multiple pieces of text information from the text content of a reference file, and extract metadata features based on the metadata of the reference file;

[0139] Extraction module 1102, for extracting text features of each piece of text information respectively, and fusing the text features of each piece of text information to obtain comprehensive text features;

[0140] The comparison module 1103 is used to determine the similarity between the reference file and the file to be compared based on metadata features and comprehensive text features.

[0141] In some embodiments, Figure 11 On the basis of Figure 12 As shown, the extraction module 1102 is used to process the text features of each text information based on the hierarchical attention mechanism to obtain comprehensive text features.

[0142] In some embodiments, Figure 11 Based on Figure 12 As shown, the extraction module 1102 includes:

[0143] The text item determining unit 1201 is configured to determine a complex text item including multiple paragraphs in the multiple text information items, and to determine the text information other than the complex text item in the multiple text information items as simple text information;

[0144] A feature determination unit 1202 is configured to determine, based on the complex text item, key features, value features, and query features of a hierarchical attention mechanism; wherein the subtext features of each paragraph in the text features of the complex text item are used as key features and value features, and the text features of the simple text information are used as query features;

[0145] A feature optimization unit 1203 is configured to determine optimized text features of a complex text item based on key features, value features, and query features;

[0146] The splicing unit 1204 is used to splice the optimized text features of the complex text item and the text features of the simple text item to obtain comprehensive text features.

[0147] In some embodiments, the extraction module 1102 is used to perform splicing processing on the text features of each piece of text information to obtain a comprehensive text feature.

[0148] In some embodiments, the extraction module 1102 is configured to perform the following operations on each item of text information:

[0149] Extracting initial text features of the text information using a first language model corresponding to the text information;

[0150] The initial text features of the text information are input into the first fully connected layer corresponding to the text information to obtain the text features of the text information output by the first fully connected layer.

[0151] In some embodiments, for a complex text item containing multiple paragraphs in multiple text messages, the extraction module 1102 extracts text features of the complex text item based on the following method:

[0152] Based on the second language model corresponding to the complex text item, the sub-text features of each paragraph of the complex text item are extracted respectively;

[0153] Perform dimensionality reduction processing on the sub-text features of each text segment to obtain the dimensionality reduction features of the complex text item;

[0154] The dimensionality reduction features of the complex text item are input into the second fully connected layer corresponding to the complex text item to obtain the text features of the complex text item output by the second fully connected layer.

[0155] In some embodiments, the extraction module 1102 is used to extract text features of each piece of text information based on a comprehensive text feature network model, and fuse the text features of each piece of text information to obtain a comprehensive text feature.

[0156] In some embodiments, a training module 1203 is further included for training a comprehensive text feature network model based on the following method:

[0157] Extract multiple text information from the same file to construct positive samples, and extract multiple text information from different files to construct negative samples;

[0158] The positive samples and negative samples are input into the initial text feature network respectively, and the comprehensive text features of the positive samples and the comprehensive text features of the negative samples are obtained as outputs of the initial text feature network;

[0159] A classifier is used to classify the comprehensive text features of the positive sample and the comprehensive text features of the negative sample respectively to obtain a classification processing result, wherein the classification categories of the classifier include positive samples and negative samples;

[0160] Determine the classification loss value based on the classification processing results, the category labels of the positive samples, and the category labels of the negative samples;

[0161] Based on the classification loss value, the model parameters of the initial text feature network are adjusted to obtain a comprehensive text feature network model. In the disclosed embodiment, based on the first ranking value, the first ranking value is further adjusted using the adjustment parameters of the user's reference recommendation information of the same category, thereby adjusting the first ranking value based on the same category information, thereby improving the accuracy of the ranking.

[0162] According to another embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0163] Figure 13 A schematic block diagram of an example electronic device 1300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0164] like Figure 13 As shown, the electronic device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. Various programs and data required for the operation of the electronic device 1300 can also be stored in the RAM 1303. The computing unit 1301, the ROM 1302, and the RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0165] Multiple components in electronic device 1300 are connected to I / O interface 1305, including: an input unit 1306, such as a keyboard, mouse, etc.; an output unit 1307, such as various types of displays, speakers, etc.; a storage unit 1308, such as a magnetic disk, optical disk, etc.; and a communication unit 1309, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1309 allows electronic device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0166] The computing unit 1301 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1301 performs the information comparison method described above. In some embodiments, the information comparison method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 1308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by the computing unit 1301, one or more steps of the information comparison method can be performed. Alternatively, in other embodiments, the computing unit 1301 can be configured to perform the information comparison method by any other appropriate means (e.g., by means of firmware).

[0167] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0168] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0169] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0170] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0171] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0172] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0173] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0174] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. An information comparison method, comprising: extracting a plurality of text information from the text content of the reference file, and extracting metadata features based on the metadata of the reference file; Extracting text features of each item of text information separately, and fusing the text features of each item of text information to obtain comprehensive text features, including: determining a complex text item containing multiple paragraphs in the multiple items of text information, and determining that the text information other than the complex text item in the multiple items of text information is simple text information; based on the complex text item, determining the key feature, value feature and query feature of the hierarchical attention mechanism; wherein the sub-text feature of each paragraph in the text feature of the complex text item is the key feature and the value feature, and the text feature of the simple text information is the query feature; determining the optimized text feature of the complex text item based on the key feature, the value feature and the query feature; and concatenating the optimized text feature of the complex text item with the text feature of the simple text information to obtain the comprehensive text feature; Based on the metadata features and the comprehensive text features, the similarity between the reference file and the file to be compared is determined.

2. The method according to claim 1, wherein the fusing of the text features of each piece of text information to obtain a comprehensive text feature further comprises: The text features of each piece of text information are spliced ​​together to obtain the comprehensive text features.

3. The method according to claim 1 or 2, wherein extracting text features of each item of text information comprises: For each text message, perform the following actions: extracting initial text features of the text information using a first language model corresponding to the text information; The initial text features of the text information are input into a first fully connected layer corresponding to the text information to obtain text features of the text information output by the first fully connected layer.

4. The method according to claim 1, wherein for a complex text item including multiple paragraphs in the plurality of text information, extracting text features of the complex text item comprises: extracting sub-text features of each text segment in the complex text item based on the second language model corresponding to the complex text item; Performing dimensionality reduction processing on the sub-text features of each text segment to obtain the dimensionality reduction features of the complex text item; The dimensionality reduction features of the complex text item are input into a second fully connected layer corresponding to the complex text item to obtain text features of the complex text item output by the second fully connected layer.

5. The method according to claim 1, wherein the extracting of text features of each piece of text information and fusing the text features of each piece of text information to obtain a comprehensive text feature further comprises: Based on the comprehensive text feature network model, the text features of each piece of text information are extracted respectively, and the text features of each piece of text information are fused to obtain the comprehensive text feature.

6. The method according to claim 5, further comprising training the comprehensive text feature network model based on the following method: Extract multiple text information from the same file to construct positive samples, and extract multiple text information from different files to construct negative samples; Inputting the positive sample and the negative sample into an initial text feature network respectively, and obtaining comprehensive text features of the positive sample and the negative sample output by the initial text feature network; A classifier is used to classify the comprehensive text features of the positive sample and the comprehensive text features of the negative sample to obtain classification results, wherein: The classification categories of the classifier include positive samples and negative samples; Determining a classification loss value based on the classification processing result, the category label of the positive sample, and the category label of the negative sample; Based on the classification loss value, the model parameters of the initial text feature network are adjusted to obtain the comprehensive text feature network model.

7. An information comparison device comprising: an acquisition module, configured to extract multiple pieces of text information from the text content of a reference file, and extract metadata features based on the metadata of the reference file; The extraction module is used to extract the text features of each text information respectively and fuse the text features of each text information to obtain a comprehensive text feature; the extraction module includes: a text item determining unit, configured to determine a complex text item including multiple paragraphs in the plurality of text information items, and to determine text information other than the complex text item in the plurality of text information items as simple text information; a feature determination unit, configured to determine, based on the complex text item, a key feature, a value feature, and a query feature of a hierarchical attention mechanism; wherein the subtext features of each paragraph in the text features of the complex text item are the key features and the value features, and the text features of the simple text information are the query features; a feature optimization unit, configured to determine an optimized text feature of the complex text item based on the key feature, the value feature, and the query feature; a concatenation unit, configured to concatenate the optimized text features of the complex text item and the text features of the simple text information to obtain the comprehensive text features; The comparison module is used to determine the similarity between the reference file and the file to be compared based on the metadata features and the comprehensive text features.

8. The device according to claim 7, wherein the extraction module is further configured to perform concatenation processing on the text features of each piece of text information to obtain the comprehensive text features.

9. The apparatus according to claim 7 or 8, wherein the extraction module is configured to perform the following operations for each item of text information: extracting initial text features of the text information using a first language model corresponding to the text information; The initial text features of the text information are input into a first fully connected layer corresponding to the text information to obtain text features of the text information output by the first fully connected layer.

10. The apparatus according to claim 7, wherein for a complex text item including multiple paragraphs in the plurality of text information, the extraction module extracts text features of the complex text item based on the following method: extracting sub-text features of each text segment in the complex text item based on the second language model corresponding to the complex text item; Performing dimensionality reduction processing on the sub-text features of each text segment to obtain the dimensionality reduction features of the complex text item; The dimensionality reduction features of the complex text item are input into a second fully connected layer corresponding to the complex text item to obtain text features of the complex text item output by the second fully connected layer.

11. The device according to claim 7, wherein the extraction module is further used to extract text features of each piece of text information based on a comprehensive text feature network model, and fuse the text features of each piece of text information to obtain the comprehensive text feature.

12. The apparatus according to claim 11, further comprising a training module for training the comprehensive text feature network model based on the following method: Extract multiple text information from the same file to construct positive samples, and extract multiple text information from different files to construct negative samples; Inputting the positive sample and the negative sample into an initial text feature network respectively, and obtaining comprehensive text features of the positive sample and the negative sample output by the initial text feature network; A classifier is used to classify the comprehensive text features of the positive sample and the comprehensive text features of the negative sample to obtain classification results, wherein: The classification categories of the classifier include positive samples and negative samples; Determining a classification loss value based on the classification processing result, the category label of the positive sample, and the category label of the negative sample; Based on the classification loss value, the model parameters of the initial text feature network are adjusted to obtain the comprehensive text feature network model.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Retrieval sorting method and device, electronic equipment and storage medium

    CN118551036A