Enterprise patent demand text matching method and system based on multi-modal data fusion

By using multimodal data fusion technology and multimodal data from enterprise databases, word triples and demand intensity values ​​are obtained, which solves the problem of inaccurate division of patent demand matching scope and improves matching efficiency and accuracy.

CN121188185BActive Publication Date: 2026-02-24QINGDAO ASIDUN ENG TECH TRANSFER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511350239.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-02-24
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately define the matching scope of patent requirements based on multimodal data, leading to abnormal scoring results and affecting the accuracy of patent text matching.

Method used

By fusing multimodal data from the enterprise database, multiple word triples and demand intensity values ​​are obtained. Then, using word segmentation and dependency parsing algorithms, the salience of demand words is calculated, a matching range is generated, and text matching is performed.

Benefits of technology

It improves the efficiency and accuracy of matching patent demands with enterprise operational data, ensuring that the matching results better reflect users' actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121188185B_ABST
    Figure CN121188185B_ABST
Patent Text Reader

Abstract

The application discloses an enterprise patent demand text matching method and system based on multi-modal data fusion, and relates to the technical field of data processing. The method comprises the following steps: obtaining multi-modal operation data from an enterprise database; based on the operation data, obtaining a plurality of first lexical triplets; based on a patent demand text, obtaining a plurality of demand words and a plurality of second lexical triplets; based on each first lexical triplet and each second lexical triplet, obtaining a demand intensity value corresponding to each demand word; and based on each demand intensity value, obtaining a matching range of the patent demand text. By calculating the demand intensity value of the demand word, the application divides the matching range of the patent demand based on the demand intensity value, and dynamically adjusts the final matching range by combining the topological distance and the intensity difference value, so that the matching result is more in line with the real demand of the user, and the matching efficiency and accuracy of the patent demand and the enterprise operation data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a method and system for matching enterprise patent requirements text based on multimodal data fusion. Background Technology

[0002] A company's patents are often technical solutions specifically designed to protect its own product capabilities and technological development. However, the methods for integrating and quantifying a company's multimodal data to assess the conformity between the company and its patent text often rely on unstructured methods such as expert review and weighted voting. This approach is prone to errors in scoring due to the reviewer's own biases.

[0003] To avoid anomalies in scoring results caused by human factors, existing technologies extract entities from multimodal data of enterprises to analyze the status of enterprise technology keywords represented in the multimodal data, thereby completing the description of the enterprise's technical characteristics and dynamically processing the corresponding result text. However, in the actual process of mining technical characteristics through multimodal data, different changing characteristics represent different themes with varying real-time significance in the current data. That is, a certain technical theme may have undergone technological progress, but the significance of the technological progress may not be significant, and the process of technological progress is relatively difficult. However, by using text keyword matching, the keywords representing the current technological progress may be similar to the keywords representing the original technological level, leading to an incorrect judgment of the significance of the current technological progress. This results in the loss of information in the keyword extraction and analysis process of multimodal data, affecting the matching accuracy of patent text. Summary of the Invention

[0004] The purpose of this application is to provide a method and system for matching enterprise patent demand texts based on multimodal data fusion, so as to solve the technical problem that it is difficult to accurately divide the matching scope from multimodal data based on patent demand.

[0005] To achieve the above objectives, this application provides the following technical solution:

[0006] Firstly, this application proposes a technical solution for a method of matching enterprise patent demand text based on multimodal data fusion, which includes:

[0007] The data consists of multimodal operational data from an enterprise database; the operational data includes any one or a combination of multiple types of technical text, R&D text, images, logs, videos, and audio; the enterprise database is pre-established.

[0008] Based on the aforementioned operational data, multiple first-word triples are obtained;

[0009] Based on the patent requirement text, multiple requirement terms and multiple second term triples are obtained; the patent requirement text is obtained in advance; the requirement terms are any words in the patent requirement text.

[0010] Based on each first word triplet and each second word triplet, obtain the demand intensity value corresponding to each demand word; the demand intensity value is used at least to characterize the salience of the corresponding demand word in the patent demand text and operational data;

[0011] Based on each demand intensity value, the matching range of the patent demand text is obtained;

[0012] Based on the specified matching range, text matching is performed on the required words to obtain the matching results.

[0013] As a specific solution in the technical solution of this application, the step of obtaining multiple first-word triples based on the operational data includes:

[0014] Based on the operational data, first text data and non-text data are obtained; the first text data is the text data in the operational data.

[0015] Based on the non-text data, obtain the second text data;

[0016] Based on the word segmentation algorithm, multiple operational terms are obtained from the first text data and the second text data;

[0017] Based on the dependency parsing algorithm, multiple first-word triples are obtained from each operational word.

[0018] As a specific solution in this application, the step of obtaining the demand intensity value corresponding to each demand word based on each first word triplet and each second word triplet includes:

[0019] Based on each demand term, obtain the first demand term; the first demand term is any demand term among the demand terms for which no corresponding demand intensity value has been obtained.

[0020] Based on the first required vocabulary, multiple third vocabulary triples are obtained from each first vocabulary triple and each second vocabulary triple; the third vocabulary triple is any first vocabulary triple or second vocabulary triple, and the primary or secondary position in the third vocabulary triple is the first required vocabulary.

[0021] Based on each third-word triplet, a first quantity and a second quantity are obtained; the first quantity is the number of triplets in each third-word triplet where the first demand word is in a primary position; the second quantity is the number of triplets in each third-word triplet where the first demand word is in a secondary position.

[0022] Based on the first quantity and the second quantity, obtain the demand intensity value corresponding to the first demand term.

[0023] As a specific solution in this application, the step of obtaining the demand intensity value corresponding to the first demand term based on the first quantity and the second quantity includes:

[0024] Based on each first word triplet, obtain a node tree to be concatenated that corresponds one-to-one with each first word triplet; the root node in the node tree to be concatenated is the word corresponding to the main position in the first word triplet; the child nodes in the node tree to be concatenated are the words corresponding to the secondary positions in the first word triplet.

[0025] Merge identical nodes in each node tree to be spliced ​​to obtain the operational node topology;

[0026] Based on the operational node topology, a third quantity and a fourth quantity are obtained; the third quantity is the number of diffusion nodes in the operational node topology corresponding to the first demand term; the fourth quantity is the number of description nodes in the operational node topology corresponding to the first demand term.

[0027] Based on the first quantity, the second quantity, the third quantity, and the fourth quantity, obtain the demand intensity value corresponding to the first demand term.

[0028] As a specific solution in this application, obtaining the matching range of the patent demand text based on various demand intensity values ​​includes:

[0029] Sort the demand terms in descending order of demand intensity to obtain a demand term sequence;

[0030] Based on the sequence of required terms, an intensity difference value is obtained; the intensity difference value is used at least to characterize the magnitude of the difference in the intensity values ​​of the requirements among the various required terms in the patent requirement text.

[0031] Based on the demand vocabulary sequence, a first sequence and a second sequence are obtained; the demand intensity values ​​corresponding to the demand words in the first sequence are all greater than or equal to the average demand intensity values ​​of each demand word in the demand vocabulary sequence; the demand intensity values ​​corresponding to the demand words in the second sequence are all less than the average demand intensity values ​​of each demand word in the demand vocabulary sequence.

[0032] Based on the first sequence, the first vocabulary topology is obtained from the topology of the operating nodes;

[0033] Based on the second sequence, a second vocabulary topology is obtained from the operational node topology;

[0034] Based on the intensity difference value, the first lexical topology, and the second lexical topology, the matching range of the patent claim text is obtained.

[0035] As a specific solution in this application, obtaining the intensity difference value based on the required vocabulary sequence includes:

[0036] Traverse the sequence of required words to obtain a second required word; the second required word is any required word in the sequence of required words for which a corresponding first ratio has not been obtained.

[0037] Based on the demand vocabulary sequence, the mean and standard deviation are obtained; the mean is the average value of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence; the standard deviation is the standard deviation of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence.

[0038] Based on the second demand vocabulary, obtain the first demand intensity value;

[0039] Based on the first demand intensity value and the average value, a first difference is obtained; the first difference is equal to the first demand intensity value minus the average value;

[0040] Based on the first difference and the standard deviation, a first ratio is obtained; the first ratio is equal to the ratio of the first difference to the standard deviation.

[0041] Based on the first ratio, the intensity difference value is obtained.

[0042] As a specific solution in the technical solution of this application, based on the first sequence, the first vocabulary topology is obtained from the operational node topology, including:

[0043] Based on the first sequence, a third demand term is obtained; the third demand term is any demand term in the first sequence for which no corresponding topological matching range has been obtained.

[0044] Based on the third requirement term, obtain the farthest topological distance in the operation node topology;

[0045] Based on the farthest topological distance, obtain the topological matching range corresponding to the third required vocabulary;

[0046] The topological matching ranges corresponding to each required word in the first sequence are merged to obtain the topology of the first word.

[0047] As a specific solution in this application, obtaining the topological matching range corresponding to the third requirement term based on the farthest topological distance includes:

[0048] Based on the third requirement vocabulary and each first vocabulary triplet, a fifth quantity and a sixth quantity are obtained; the fifth quantity is the number of triplets in each first vocabulary triplet where the third requirement vocabulary is located in a primary or secondary position; the sixth quantity is the number of each first vocabulary triplet.

[0049] Based on the fifth quantity and the sixth quantity, a second ratio is obtained; the second ratio is the ratio of the fifth quantity to the sixth quantity.

[0050] Based on the second ratio and the farthest topological distance, obtain the expansion requirement topological distance;

[0051] Based on the expansion requirement topology distance, the topology matching range corresponding to the third requirement term is obtained.

[0052] As a specific solution in this application, obtaining the matching range of the patent claim text based on the intensity difference value, the first lexical topology, and the second lexical topology includes:

[0053] The first vocabulary topology and the second vocabulary topology are overlapped to obtain multiple extended matching nodes; the extended matching nodes are nodes that are present in the first vocabulary topology but not in the second vocabulary topology.

[0054] Based on the intensity difference value and each extended matching node, the matching range of the patent claim text is obtained.

[0055] Secondly, this application proposes a technical solution for a corporate patent demand text matching system based on multimodal data fusion, which includes:

[0056] A reader is used to process multimodal operational data from an enterprise database; the operational data includes any one or a combination of technical text, R&D text, images, logs, video, and audio; the enterprise database is pre-built.

[0057] A server is used to obtain multiple first-word triples based on the operational data;

[0058] Furthermore, based on the patent requirement text, multiple requirement terms and multiple second term triples are obtained; the patent requirement text is obtained in advance; the requirement terms are any terms in the patent requirement text.

[0059] Furthermore, based on each first word triplet and each second word triplet, the demand intensity value corresponding to each demand word is obtained; the demand intensity value is at least used to characterize the salience of the corresponding demand word in the patent demand text and operational data;

[0060] Furthermore, based on each demand intensity value, the matching range of the patent demand text is obtained;

[0061] Furthermore, based on the matching range, text matching is performed on the required words to obtain the matching results.

[0062] As a specific solution in this application, the server is further configured to acquire first text data and non-text data based on the operational data; the first text data is the text data in the operational data.

[0063] And, based on the non-text data, obtain the second text data;

[0064] Furthermore, based on the word segmentation algorithm, multiple operational terms are obtained from the first text data and the second text data;

[0065] Furthermore, based on the dependency parsing algorithm, multiple first-word triples are obtained from each operational word.

[0066] As a specific solution in the technical solution of this application, the server is further configured to obtain a first demand term based on each demand term; the first demand term is any demand term among the various demand terms for which a corresponding demand intensity value has not been obtained.

[0067] Furthermore, based on the first required vocabulary, multiple third vocabulary triples are obtained from each first vocabulary triple and each second vocabulary triple; the third vocabulary triple is any first vocabulary triple or second vocabulary triple, and the primary or secondary position in the third vocabulary triple is the first required vocabulary.

[0068] Furthermore, based on each third-word triplet, a first quantity and a second quantity are obtained; the first quantity is the number of triplets in each third-word triplet where the first demand word is in a primary position; the second quantity is the number of triplets in each third-word triplet where the first demand word is in a secondary position.

[0069] Furthermore, based on the first quantity and the second quantity, the demand intensity value corresponding to the first demand term is obtained.

[0070] As a specific solution in the technical solution of this application, the server is further configured to obtain a node tree to be concatenated that corresponds one-to-one with each of the first word triplets; the root node in the node tree to be concatenated is the word corresponding to the main position in the first word triplet; and the child nodes in the node tree to be concatenated are the words corresponding to the secondary positions in the first word triplets.

[0071] Additionally, merge identical nodes in each node tree to be spliced ​​to obtain the operational node topology;

[0072] Furthermore, based on the operational node topology, a third quantity and a fourth quantity are obtained; the third quantity is the number of diffusion nodes in the operational node topology corresponding to the first demand term; the fourth quantity is the number of description nodes in the operational node topology corresponding to the first demand term.

[0073] Furthermore, based on the first quantity, the second quantity, the third quantity, and the fourth quantity, the demand intensity value corresponding to the first demand term is obtained.

[0074] As a specific solution in the technical solution of this application, the server is also used to sort the various demand words in descending order of demand intensity value to obtain a demand word sequence;

[0075] Furthermore, based on the sequence of required terms, an intensity difference value is obtained; the intensity difference value is used at least to characterize the magnitude of the difference in required intensity values ​​between various required terms in the patent requirement text;

[0076] Furthermore, based on the demand vocabulary sequence, a first sequence and a second sequence are obtained; the demand intensity values ​​corresponding to the demand words in the first sequence are all greater than or equal to the average demand intensity values ​​of each demand word in the demand vocabulary sequence; the demand intensity values ​​corresponding to the demand words in the second sequence are all less than the average demand intensity values ​​of each demand word in the demand vocabulary sequence.

[0077] And, based on the first sequence, the first vocabulary topology is obtained from the operational node topology;

[0078] And, based on the second sequence, a second vocabulary topology is obtained from the operational node topology;

[0079] Furthermore, based on the intensity difference value, the first lexical topology, and the second lexical topology, the matching range of the patent claim text is obtained.

[0080] As a specific solution in the technical solution of this application, the server is further configured to traverse the demand word sequence to obtain a second demand word; the second demand word is any demand word in the demand word sequence that has not obtained a corresponding first ratio;

[0081] Furthermore, based on the demand vocabulary sequence, the mean and standard deviation are obtained; the mean is the average value of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence; the standard deviation is the standard deviation of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence.

[0082] And, based on the second demand vocabulary, obtain the first demand intensity value;

[0083] Furthermore, based on the first demand intensity value and the average value, a first difference is obtained; the first difference is equal to the first demand intensity value minus the average value;

[0084] Furthermore, based on the first difference and the standard deviation, a first ratio is obtained; the first ratio is equal to the ratio of the first difference to the standard deviation.

[0085] And, based on the first ratio, the intensity difference value is obtained.

[0086] As a specific solution in the technical solution of this application, the server is further configured to obtain a third demand term based on the first sequence; the third demand term is any demand term in the first sequence for which no corresponding topology matching range has been obtained.

[0087] Furthermore, based on the third requirement term, the farthest topological distance is obtained in the operational node topology;

[0088] And, based on the farthest topological distance, obtain the topological matching range corresponding to the third requirement word;

[0089] Furthermore, the topological matching ranges corresponding to each required word in the first sequence are merged to obtain the first word topology.

[0090] As a specific solution in the technical solution of this application, the server is further configured to obtain a fifth quantity and a sixth quantity based on the third requirement word and each first word triplet; the fifth quantity is the number of triplets in each first word triplet where the third requirement word is located in a primary or secondary position; the sixth quantity is the number of each first word triplet.

[0091] Furthermore, based on the fifth quantity and the sixth quantity, a second ratio is obtained; the second ratio is the ratio of the fifth quantity to the sixth quantity.

[0092] And, based on the second ratio and the farthest topological distance, obtain the expansion requirement topological distance;

[0093] Furthermore, based on the expansion requirement topology distance, the topology matching range corresponding to the third requirement term is obtained.

[0094] Compared with the prior art, the beneficial effects of this application are:

[0095] This application calculates the demand intensity value of demand terms and divides the matching range of patent demand based on the demand intensity value, so that the matching results are more in line with the user's real needs, thereby improving the matching efficiency and accuracy of patent demand and enterprise operation data. Attached Figure Description

[0096] Figure 1 This is a flowchart illustrating a method for matching enterprise patent demand text based on multimodal data fusion proposed in an embodiment of this application.

[0097] Figure 2 This is a schematic diagram of the structure of an enterprise patent demand text matching system based on multimodal data fusion proposed in an embodiment of this application;

[0098] Figure 3 This is a schematic diagram of an operating node topology proposed in an embodiment of this application. Detailed Implementation

[0099] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0100] The terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. For example, the first text data and the second text data mentioned below are different types of text data. It should be understood that such names can be used interchangeably where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The division of modules in the embodiments of this application is merely a logical division. In actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not performed. In addition, the shown or discussed mutual coupling or direct coupling or communication connection may be through some interface, and the indirect coupling or communication connection between modules may be electrical or other similar forms. None of these are limited in the embodiments of this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed among multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.

[0101] To address the technical problem mentioned in the background art of the difficulty in accurately defining the matching scope from multimodal data based on patent requirements, this application proposes an embodiment of a method for matching enterprise patent requirement texts based on multimodal data fusion. For example... Figure 1 As shown, the enterprise patent demand text matching method based on multimodal data fusion includes steps S100 to S600.

[0102] Step S100: Multimodal operational data from the enterprise database.

[0103] In this embodiment, the operational data includes any one or a combination of multiple types of technical text, R&D text, images, logs, videos, and audio. The enterprise database is pre-established.

[0104] Step S200: Based on the operational data, obtain multiple first-word triplets.

[0105] In this embodiment, step S200 involves obtaining multiple first word triples based on the operational data, including steps S210 to S240.

[0106] Step S210: Based on the operational data, obtain the first text data and non-text data.

[0107] In this embodiment, the first text data is the text data in the operational data.

[0108] Step S220: Based on the non-text data, obtain the second text data.

[0109] It's important to understand that converting non-text data into text data (secondary text data) is a mature technology, which won't be elaborated on here. For example, Optical Character Recognition (OCR) technology can convert text in image data into text data; sound event recognition and text conversion technology can convert audio data into text data; and video data can be converted into text data through text extraction technology or video content description generation technology.

[0110] Step S230: Based on the word segmentation algorithm, obtain multiple operational terms from the first text data and the second text data.

[0111] In this embodiment, the word segmentation algorithm can be any commercially available word segmentation algorithm, such as Forward Maximum Matching (FMM), Backward Maximum Matching (BMM), and Hidden Markov Model (HMM).

[0112] Step S240: Based on the dependency parsing algorithm, obtain multiple first-word triples from each operational word.

[0113] It is important to understand that using dependency parsing algorithms to divide each word (i.e., each operational word) into multiple triples (i.e., multiple first-word triples) is a mature technology, which will not be elaborated here.

[0114] Step S300: Based on the patent requirement text, obtain multiple requirement words and multiple second word triplets.

[0115] In this embodiment, the patent requirement text is obtained in advance. The requirement vocabulary can be any word in the patent requirement text.

[0116] In this embodiment, obtaining multiple requirement terms based on the patent requirement text can be referred to step S230. Obtaining multiple second term triples based on the patent requirement text can be referred to step S240.

[0117] Step S400: Based on each first word triplet and each second word triplet, obtain the demand intensity value corresponding to each demand word.

[0118] In this embodiment, the demand intensity value is used at least to characterize the salience of the corresponding demand terms in the patent demand text and operational data.

[0119] It's important to understand that the core logic of patent demand text matching is to analyze the scale of presentation of different technical fields within the enterprise's multimodal data (i.e., operational data). This "scale of presentation" includes two aspects: the number of descriptions of a particular field in the data, and the depth of those descriptions. These two aspects together reflect the enterprise's level of data accumulation in that field: the more numerous and in-depth the descriptions, the richer the enterprise's data reserves in that field. Based on this analysis, when performing patent demand text matching for that field, the actual needs of the enterprise in that field can be more accurately identified, thereby improving the quality of the matching results. Based on this, step S400, based on each first-word triplet and each second-word triplet, obtains the demand intensity value corresponding to each demand word, including steps S410 to S440.

[0120] Step S410: Based on each required vocabulary, obtain the first required vocabulary.

[0121] In this embodiment, the first demand term is any demand term among the various demand terms for which no corresponding demand intensity value has been obtained. That is, in this embodiment, the method for obtaining the demand intensity value corresponding to each demand term is the same as the method for obtaining the demand intensity value corresponding to the first demand term.

[0122] Step S420: Based on the first required vocabulary, obtain multiple third vocabulary triples from each first vocabulary triple and each second vocabulary triple.

[0123] In this embodiment, the third word triplet can be any first word triplet or second word triplet, and the primary or secondary position in the third word triplet is the first required word.

[0124] Step S430: Based on each third-word triplet, obtain the first and second counts.

[0125] In this embodiment, the first quantity is the number of triples in each third-word triplet where the first demand word is in a primary position. The second quantity is the number of triples in each third-word triplet where the first demand word is in a secondary position.

[0126] Step S440: Based on the first quantity and the second quantity, obtain the demand intensity value corresponding to the first demand term.

[0127] It is important to understand that the more primary keywords appear in prominent positions (i.e., the more prominent they are), the more significant they are; conversely, the fewer primary keywords appear in prominent positions (i.e., the less prominent they are), the less significant they are.

[0128] It should be noted that in this embodiment, assuming a triple is (a, b, c), then a represents the word in the primary position; b represents the word in the secondary position; and c represents the relationship between the word in the secondary position and the word in the primary position.

[0129] In this embodiment, the demand intensity value corresponding to the first demand term can be obtained based on the first quantity and the second quantity in any reasonable manner. For example, the demand intensity value can be the ratio of the first quantity to the second quantity.

[0130] It is important to understand that the more dispersed the distribution of a demand term (e.g., the first demand term) and its corresponding descriptive terms across different domains, the greater the differences between the domains corresponding to that demand term. Therefore, the initial range should be set wider during matching. Based on this, step S440, based on the first quantity and the second quantity, obtains the demand intensity value corresponding to the first demand term, including steps S450 to S480.

[0131] Step S450: Based on each first word triplet, obtain the node tree to be spliced ​​that corresponds one-to-one with each first word triplet.

[0132] In this embodiment, the root node in the node tree to be spliced ​​is the word corresponding to the major position in the first word triplet; the child nodes in the node tree to be spliced ​​are the words corresponding to the minor positions in the first word triplet.

[0133] Step S460: Merge identical nodes in each node tree to be spliced ​​to obtain the operational node topology.

[0134] It is important to understand that merging identical nodes in various node trees to be spliced ​​to form a complete node topology (i.e., operational node topology) is a mature technology, which will not be elaborated here.

[0135] Step S470: Based on the operational node topology, obtain the third and fourth quantities.

[0136] In this embodiment, the third quantity is the number of diffusion nodes in the operation node topology corresponding to the first demand term. The fourth quantity is the number of description nodes in the operation node topology corresponding to the first demand term.

[0137] In this embodiment, a description node refers to a node in the operation node topology whose reachability distance to the first demand term is 1; a diffusion node refers to a node whose reachability distance to the description node of the first demand term is 1, but which is not directly connected to the first demand term. For example, Figure 3 This is a schematic diagram of an operational node topology. Assume node 1 is the node corresponding to the first demand term. Since nodes 2 and 4 have a reachability distance of 1 from the first demand term (i.e., node 1), nodes 2 and 4 are both description nodes. Since node 3 has a reachability distance of 1 from node 2 (i.e., the description node), and node 3 is not directly connected to the first demand term (i.e., node 1), node 3 is a diffusion node.

[0138] Step S480: Based on the first quantity, the second quantity, the third quantity, and the fourth quantity, obtain the demand intensity value corresponding to the first demand term.

[0139] In this embodiment, any reasonable method can be used to obtain the demand intensity value corresponding to the first demand term based on the first quantity, the second quantity, the third quantity, and the fourth quantity. For example, in one embodiment of this application, step S480, the calculation formula for obtaining the demand intensity value corresponding to the first demand term based on the first quantity, the second quantity, the third quantity, and the fourth quantity, can be as follows:

[0140] ;

[0141] in, This indicates the demand intensity value corresponding to the first demand term; Indicates the first quantity; Indicates the second quantity; Indicates the third quantity; Indicates the fourth quantity; This indicates a non-zero function, used to change the value inside the parentheses to 1 if the value is 0.

[0142] In another embodiment of this application, step S480, based on the first quantity, the second quantity, the third quantity, and the fourth quantity, the formula for calculating the demand intensity value corresponding to the first demand term can be as follows:

[0143] ;

[0144] in, This indicates the demand intensity value corresponding to the first demand term; Indicates the first quantity; Indicates the second quantity; Indicates the third quantity; Indicates the fourth quantity; This indicates a non-zero function, used to change the value inside the parentheses to 1 if the value is 0.

[0145] Step S500: Based on each demand intensity value, obtain the matching range of the patent demand text.

[0146] It is important to understand that the different emphases of the demand terms in the patent demand text input by the user reflect the user's actual needs. Therefore, it is necessary to determine the matching range of the patent demand text based on the demand intensity value corresponding to each demand term and the descriptive relationship between the demand terms. Based on this, step S500, based on each demand intensity value, obtains the matching range of the patent demand text, including steps S510 to S560.

[0147] Step S510: Sort the demand words in descending order of demand intensity value to obtain the demand word sequence.

[0148] It is important to understand that sorting various demand terms according to rules (i.e., in descending order of demand intensity) to form a sequence (i.e., a demand term sequence) is a mature technology, which will not be elaborated here.

[0149] Step S520: Based on the required word sequence, obtain the intensity difference value.

[0150] It is important to note that the importance (i.e., demand intensity value) of different demand terms varies in the patent demand text input by the user. Some terms have relatively high demand intensity values. By further analyzing the descriptive synergy between these terms, the core focus of the user-input demand (i.e., the patent demand text) can be extracted, and the range of matching text can be adaptively set based on this core focus. In other words, in this embodiment, the intensity difference value is used to characterize at least the magnitude of the difference in demand intensity values ​​between the various demand terms in the patent demand text.

[0151] In embodiments of this application, the intensity difference value can be obtained based on the demand word sequence in any reasonable manner. For example, the intensity difference value can be the variance or standard deviation of the demand intensity values ​​corresponding to each demand word in the demand word sequence. Alternatively, in one embodiment of this application, step S520, obtaining the intensity difference value based on the demand word sequence, includes steps S521 to S526.

[0152] Step S521: Traverse the sequence of required words to obtain the second required words.

[0153] In this embodiment, the second demand term is any demand term in the demand term sequence for which no corresponding first ratio has been obtained. That is, in this embodiment, the method for obtaining the first ratio corresponding to any demand term in the demand term sequence is the same as the method for obtaining the first ratio corresponding to the second demand term.

[0154] Step S522: Based on the required vocabulary sequence, obtain the mean and standard deviation.

[0155] In this embodiment, the average value is the average value of the demand intensity corresponding to each demand term in the demand term sequence. The standard deviation is the standard deviation of the demand intensity value corresponding to each demand term in the demand term sequence.

[0156] Step S523: Based on the second demand vocabulary, obtain the first demand intensity value.

[0157] In this embodiment, the first demand intensity value is the demand intensity value corresponding to the second demand term.

[0158] Step S524: Obtain a first difference based on the first demand intensity value and the average value.

[0159] In this embodiment, the first difference is equal to the first demand intensity value minus the average value. In the computer field, obtaining the difference (i.e., the first difference) between two values ​​(i.e., the first demand intensity value and the average value) is a mature technology and will not be elaborated upon here.

[0160] Step S525: Obtain the first ratio based on the first difference and the standard deviation.

[0161] In this embodiment, the first ratio is equal to the ratio of the first difference to the standard deviation. In the computer field, obtaining the ratio of two values ​​(i.e., the first difference and the standard deviation) is a mature technique, and will not be elaborated upon here.

[0162] Step S526: Based on the first ratio, obtain the intensity difference value.

[0163] Specifically, in step S526, based on the first ratio, the calculation formula for the intensity difference value is as follows:

[0164] ;

[0165] in, Indicates the intensity difference value; n represents the number of demand words in the demand word sequence; This represents the demand intensity value (i.e., the first demand intensity value) corresponding to the i-th demand term (i.e., the second demand term). This represents the average value; It represents the standard deviation.

[0166] In this embodiment, the intensity difference value The numerical value tends to reflect the magnitude of the difference in demand intensity values ​​between various demand terms in the patent demand text. If the intensity difference value... The larger the value, the more demand terms with higher demand intensity values ​​there are, meaning more demand points can be extracted from the patent demand text; if the intensity difference value... The smaller the value, the fewer the corresponding demand words with high demand intensity, which means that the demand points entered by the user in the patent demand text are more concentrated, which is conducive to the subsequent extraction of the user's precise demand points.

[0167] Step S530: Based on the required vocabulary sequence, obtain the first sequence and the second sequence.

[0168] In this embodiment, the demand intensity values ​​corresponding to the demand words in the first sequence are all greater than or equal to the average demand intensity values ​​of all demand words in the demand word sequence. The demand intensity values ​​corresponding to the demand words in the second sequence are all less than the average demand intensity values ​​of all demand words in the demand word sequence.

[0169] Step S540: Based on the first sequence, obtain the first vocabulary topology from the topology of the operating nodes.

[0170] In this embodiment, the first vocabulary topology can be the topology formed by each demand vocabulary in the first sequence in the topology of the operation node.

[0171] In a specific embodiment of this application, step S540, based on the first sequence, obtains the first vocabulary topology from the operating node topology, including steps S541 to S544.

[0172] Step S541: Based on the first sequence, obtain the third required vocabulary.

[0173] In this embodiment, the third demand term is any demand term in the first sequence for which no corresponding topology matching range has been obtained. That is, in this embodiment, the method for obtaining the topology matching range corresponding to any demand term in the first sequence is the same as the method for obtaining the topology matching range corresponding to the third demand term.

[0174] Step S542: Based on the third requirement term, obtain the farthest topological distance in the operation node topology.

[0175] It should be clear that the farthest topological distance of the third requirement term in the operational node topology refers to the maximum number of edges traversed from the node corresponding to the third requirement term to reach each of its associated nodes in the operational node topology.

[0176] Step S543: Based on the farthest topological distance, obtain the topological matching range corresponding to the third requirement word.

[0177] In embodiments of this application, the topological matching range corresponding to the third requirement term can be obtained based on the farthest topological distance using any reasonable method. For example, the range formed by all nodes in the operating node topology whose distance from the node corresponding to the third requirement term is less than or equal to the farthest topological distance can be used as the topological matching range corresponding to the third requirement term.

[0178] It is important to understand that for the third requirement word, the more times the third requirement word appears in each of the first word triplets, the more significant the third requirement word is, and the greater the need to increase the topological matching range corresponding to the third requirement word. Based on this, step S543, based on the farthest topological distance, obtains the topological matching range corresponding to the third requirement word, including steps S543a to S543d.

[0179] Step S543a: Based on the third requirement vocabulary and each first vocabulary triplet, obtain the fifth quantity and the sixth quantity.

[0180] In this embodiment, the fifth quantity refers to the number of triples in each first word triplet where the third demand word occupies a primary or secondary position. The sixth quantity refers to the number of each first word triplet.

[0181] Step S543b: Obtain a second ratio based on the fifth quantity and the sixth quantity.

[0182] In this embodiment, the second ratio is the ratio of the fifth quantity to the sixth quantity.

[0183] Step S543c: Based on the second ratio and the farthest topology distance, obtain the expansion requirement topology distance.

[0184] In this embodiment, step S543c: Based on the second ratio and the farthest topology distance, the calculation formula for the expansion requirement topology distance is as follows:

[0185] ;

[0186] in, This indicates the topological distance of the extended demand corresponding to the third demand term; Indicates the fifth quantity; Indicates the sixth quantity; Indicates the farthest topological distance corresponding to the third requirement term; This represents the floor function, used to round the value within the parentheses up to the nearest integer.

[0187] Step S543d: Based on the expansion requirement topology distance, obtain the topology matching range corresponding to the third requirement term.

[0188] In this embodiment, the range formed by all nodes in the operating node topology whose distance from the node corresponding to the third demand term is less than or equal to the distance of the expansion demand topology can be used as the topology matching range corresponding to the third demand term.

[0189] Step S544: Merge the topological matching ranges corresponding to each required word in the first sequence to obtain the first word topology.

[0190] It is important to understand that merging the various topological matching ranges to form a complete topology (i.e., the first-word topology) is a mature technology, which will not be elaborated here.

[0191] Step S550: Based on the second sequence, obtain the second vocabulary topology from the operating node topology.

[0192] In this embodiment, the method for obtaining the second vocabulary topology can refer to the method for obtaining the first vocabulary topology, and will not be described in detail here.

[0193] Step S560: Based on the intensity difference value, the first lexical topology, and the second lexical topology, obtain the matching range of the patent claim text.

[0194] In this embodiment, nodes present in the first vocabulary topology (i.e., demand words) but not present in the second vocabulary topology are highly likely to be key demand words. Based on this, step S560, based on the intensity difference value, the first vocabulary topology, and the second vocabulary topology, obtains the matching range of the patent demand text, including steps S561 and S562.

[0195] Step S561: Overlap the first vocabulary topology and the second vocabulary topology to obtain multiple extended matching nodes.

[0196] In this embodiment, the extended matching node is a node that exists in the first vocabulary topology but not in the second vocabulary topology.

[0197] Step S562: Based on the intensity difference value and each extended matching node, obtain the matching range of the patent claim text.

[0198] In this embodiment, the topological matching range formed by each extended matching node is the matching range of the patent claim text. In this embodiment, the method for obtaining the topological matching range formed by the extended matching nodes can refer to step S543d.

[0199] In one embodiment of this application, the topology matching range formed by the extended matching nodes is obtained as follows:

[0200] Step S562a: Based on the extended matching node, obtain the extended topology distance.

[0201] Specifically, based on the extended matching nodes, the formula for calculating the extended topology distance is as follows:

[0202] ;

[0203] in, Indicates the extended topological distance; Indicates the difference in intensity; Indicates the number of extended matching nodes; This represents a normalization function used to map the values ​​within the parentheses to the range [0, 1]. This represents the floor function, used to round the value within the parentheses up to the nearest integer.

[0204] Step S562b: Based on the extended topology distance, obtain the topology matching range of the extended matching node.

[0205] In this embodiment, the range formed by all nodes whose distance from the extended matching node is less than or equal to the extended topology distance can be used as the topology matching range of the extended matching node in the operating node topology.

[0206] Step S600: Based on the matching range, perform text matching on the required words to obtain the matching results.

[0207] It should be clear that in this technical field, performing text matching on required words within a certain range (i.e., the matching range) and obtaining the corresponding matching results is a mature technology, which will not be elaborated here.

[0208] The embodiment of the enterprise patent demand text matching method based on multimodal data fusion proposed in this application calculates the demand intensity value of demand words, divides the matching range of patent demand based on the demand intensity value, and dynamically adjusts the final matching range by combining topological distance and intensity difference value, so that the matching result is more in line with the user's real needs, and improves the matching efficiency and accuracy of patent demand and enterprise operation data.

[0209] Having introduced the enterprise patent demand text matching method based on multimodal data fusion proposed in the embodiments of this application, the following describes an embodiment of an enterprise patent demand text matching system based on multimodal data fusion proposed in this application. For example... Figure 2 As shown, the enterprise patent demand text matching system 10 based on multimodal data fusion includes:

[0210] Reader 11 is used to process multimodal operational data from an enterprise database; the operational data includes any one or a combination of technical text, R&D text, images, logs, videos, and audio; the enterprise database is pre-established.

[0211] Server 12 is used to obtain multiple first-word triples based on the operational data;

[0212] Furthermore, based on the patent requirement text, multiple requirement terms and multiple second term triples are obtained; the patent requirement text is obtained in advance; the requirement terms are any terms in the patent requirement text.

[0213] Furthermore, based on each first word triplet and each second word triplet, the demand intensity value corresponding to each demand word is obtained; the demand intensity value is at least used to characterize the salience of the corresponding demand word in the patent demand text and operational data;

[0214] Furthermore, based on each demand intensity value, the matching range of the patent demand text is obtained;

[0215] Furthermore, based on the matching range, text matching is performed on the required words to obtain the matching results.

[0216] As a specific embodiment of this application, the server 12 is further configured to obtain first text data and non-text data based on the operational data; the first text data is the text data in the operational data;

[0217] And, based on the non-text data, obtain the second text data;

[0218] Furthermore, based on the word segmentation algorithm, multiple operational terms are obtained from the first text data and the second text data;

[0219] Furthermore, based on the dependency parsing algorithm, multiple first-word triples are obtained from each operational word.

[0220] As a specific embodiment of this application, the server 12 is further configured to obtain a first demand term based on each demand term; the first demand term is any demand term among the various demand terms that has not obtained a corresponding demand intensity value.

[0221] Furthermore, based on the first required vocabulary, multiple third vocabulary triples are obtained from each first vocabulary triple and each second vocabulary triple; the third vocabulary triple is any first vocabulary triple or second vocabulary triple, and the primary or secondary position in the third vocabulary triple is the first required vocabulary.

[0222] Furthermore, based on each third-word triplet, a first quantity and a second quantity are obtained; the first quantity is the number of triplets in each third-word triplet where the first demand word is in a primary position; the second quantity is the number of triplets in each third-word triplet where the first demand word is in a secondary position.

[0223] Furthermore, based on the first quantity and the second quantity, the demand intensity value corresponding to the first demand term is obtained.

[0224] As a specific embodiment of this application, the server 12 is further configured to obtain a node tree to be concatenated that corresponds one-to-one with each of the first word triples based on each first word triple; the root node in the node tree to be concatenated is the word corresponding to the main position in the first word triple; the child nodes in the node tree to be concatenated are the words corresponding to the secondary positions in the first word triples.

[0225] Additionally, merge identical nodes in each node tree to be spliced ​​to obtain the operational node topology;

[0226] Furthermore, based on the operational node topology, a third quantity and a fourth quantity are obtained; the third quantity is the number of diffusion nodes in the operational node topology corresponding to the first demand term; the fourth quantity is the number of description nodes in the operational node topology corresponding to the first demand term.

[0227] Furthermore, based on the first quantity, the second quantity, the third quantity, and the fourth quantity, the demand intensity value corresponding to the first demand term is obtained.

[0228] As a specific embodiment of this application, the server 12 is further configured to sort the demand words in descending order of demand intensity values ​​to obtain a demand word sequence;

[0229] Furthermore, based on the sequence of required terms, an intensity difference value is obtained; the intensity difference value is used at least to characterize the magnitude of the difference in required intensity values ​​between various required terms in the patent requirement text;

[0230] Furthermore, based on the demand vocabulary sequence, a first sequence and a second sequence are obtained; the demand intensity values ​​corresponding to the demand words in the first sequence are all greater than or equal to the average demand intensity values ​​of each demand word in the demand vocabulary sequence; the demand intensity values ​​corresponding to the demand words in the second sequence are all less than the average demand intensity values ​​of each demand word in the demand vocabulary sequence.

[0231] And, based on the first sequence, the first vocabulary topology is obtained from the operational node topology;

[0232] And, based on the second sequence, a second vocabulary topology is obtained from the operational node topology;

[0233] Furthermore, based on the intensity difference value, the first lexical topology, and the second lexical topology, the matching range of the patent claim text is obtained.

[0234] As a specific embodiment of this application, the server 12 is further configured to traverse the demand word sequence to obtain a second demand word; the second demand word is any demand word in the demand word sequence that has not obtained a corresponding first ratio;

[0235] Furthermore, based on the demand vocabulary sequence, the mean and standard deviation are obtained; the mean is the average value of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence; the standard deviation is the standard deviation of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence.

[0236] And, based on the second demand vocabulary, obtain the first demand intensity value;

[0237] Furthermore, based on the first demand intensity value and the average value, a first difference is obtained; the first difference is equal to the first demand intensity value minus the average value;

[0238] Furthermore, based on the first difference and the standard deviation, a first ratio is obtained; the first ratio is equal to the ratio of the first difference to the standard deviation.

[0239] And, based on the first ratio, the intensity difference value is obtained.

[0240] As a specific embodiment of this application, the server 12 is further configured to obtain a third demand term based on the first sequence; the third demand term is any demand term in the first sequence for which no corresponding topological matching range has been obtained.

[0241] Furthermore, based on the third requirement term, the farthest topological distance is obtained in the operational node topology;

[0242] And, based on the farthest topological distance, obtain the topological matching range corresponding to the third requirement word;

[0243] Furthermore, the topological matching ranges corresponding to each required word in the first sequence are merged to obtain the first word topology.

[0244] As a specific embodiment of this application, the server 12 is further configured to obtain a fifth quantity and a sixth quantity based on the third demand word and each first word triplet; the fifth quantity is the number of triplets in each first word triplet where the third demand word is located in a primary or secondary position; the sixth quantity is the number of each first word triplet.

[0245] Furthermore, based on the fifth quantity and the sixth quantity, a second ratio is obtained; the second ratio is the ratio of the fifth quantity to the sixth quantity.

[0246] And, based on the second ratio and the farthest topological distance, obtain the expansion requirement topological distance;

[0247] Furthermore, based on the expansion requirement topology distance, the topology matching range corresponding to the third requirement term is obtained.

[0248] As a specific embodiment of this application, the server 12 is further configured to overlap the first vocabulary topology and the second vocabulary topology to obtain multiple extended matching nodes; the extended matching nodes are nodes that are present in the first vocabulary topology but not in the second vocabulary topology.

[0249] Furthermore, based on the intensity difference value and each extended matching node, the matching range of the patent claim text is obtained.

[0250] The embodiment of the enterprise patent demand text matching system based on multimodal data fusion proposed in this application calculates the demand intensity value of demand words, divides the matching range of patent demand based on the demand intensity value, and dynamically adjusts the final matching range by combining topological distance and intensity difference value, so that the matching result is more in line with the user's real needs, and improves the matching efficiency and accuracy of patent demand and enterprise operation data.

[0251] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0252] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the methods, apparatuses, and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0253] In the several embodiments provided in this application, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between devices or modules may be electrical, mechanical, or other forms.

[0254] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0255] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0256] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.

[0257] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video optical disc), or a semiconductor medium (e.g., solid-state drive (SSD)).

[0258] Although embodiments of this application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles of this application.

Claims

1. A method for matching enterprise patent demand text based on multimodal data fusion, characterized in that, include: The data consists of multimodal operational data from an enterprise database; the operational data includes any one or a combination of multiple types of technical text, R&D text, images, logs, videos, and audio; the enterprise database is pre-established. Based on the aforementioned operational data, multiple first-word triples are obtained; Based on the patent requirement text, multiple requirement terms and multiple second term triples are obtained; the patent requirement text is obtained in advance; the requirement terms are any words in the patent requirement text. Based on each first-word triplet and each second-word triplet, obtain the demand intensity value corresponding to each demand word; The demand intensity value is used at least to characterize the salience of the corresponding demand terms in the patent demand text and operational data; Based on each demand intensity value, the matching range of the patent demand text is obtained; Based on the matching range, perform text matching on the required words to obtain the matching results; The process of obtaining the demand intensity value corresponding to each demand word based on each first word triplet and each second word triplet includes: Based on each demand term, obtain the first demand term; the first demand term is any demand term among the demand terms for which no corresponding demand intensity value has been obtained. Based on the first required vocabulary, multiple third vocabulary triples are obtained from each first vocabulary triple and each second vocabulary triple; the third vocabulary triple is any first vocabulary triple or second vocabulary triple, and the primary or secondary position in the third vocabulary triple is the first required vocabulary. Based on each third-word triplet, a first quantity and a second quantity are obtained; the first quantity is the number of triplets in each third-word triplet where the first demand word is in a primary position; the second quantity is the number of triplets in each third-word triplet where the first demand word is in a secondary position. Based on the first quantity and the second quantity, obtain the demand intensity value corresponding to the first demand term.

2. The enterprise patent demand text matching method based on multimodal data fusion according to claim 1, characterized in that, Based on the operational data, multiple first-word triples are obtained, including: Based on the operational data, first text data and non-text data are obtained; the first text data is the text data in the operational data. Based on the non-text data, obtain the second text data; Based on the word segmentation algorithm, multiple operational terms are obtained from the first text data and the second text data; Based on the dependency parsing algorithm, multiple first-word triples are obtained from each operational word.

3. The enterprise patent demand text matching method based on multimodal data fusion according to claim 1, characterized in that, The step of obtaining the demand intensity value corresponding to the first demand term based on the first quantity and the second quantity includes: Based on each first word triplet, obtain a node tree to be concatenated that corresponds one-to-one with each first word triplet; the root node in the node tree to be concatenated is the word corresponding to the main position in the first word triplet; the child nodes in the node tree to be concatenated are the words corresponding to the secondary positions in the first word triplet. Merge identical nodes in each node tree to be spliced ​​to obtain the operational node topology; Based on the operational node topology, a third quantity and a fourth quantity are obtained; the third quantity is the number of diffusion nodes in the operational node topology corresponding to the first demand term; the fourth quantity is the number of description nodes in the operational node topology corresponding to the first demand term. Based on the first quantity, the second quantity, the third quantity, and the fourth quantity, obtain the demand intensity value corresponding to the first demand term.

4. The enterprise patent demand text matching method based on multimodal data fusion according to claim 3, characterized in that, The process of obtaining the matching range of the patent demand text based on each demand intensity value includes: Sort the demand terms in descending order of demand intensity to obtain a demand term sequence; Based on the sequence of required terms, an intensity difference value is obtained; the intensity difference value is used at least to characterize the magnitude of the difference in the intensity values ​​of the requirements among the various required terms in the patent requirement text. Based on the demand vocabulary sequence, a first sequence and a second sequence are obtained; the demand intensity values ​​corresponding to the demand words in the first sequence are all greater than or equal to the average demand intensity values ​​of each demand word in the demand vocabulary sequence; the demand intensity values ​​corresponding to the demand words in the second sequence are all less than the average demand intensity values ​​of each demand word in the demand vocabulary sequence. Based on the first sequence, the first vocabulary topology is obtained from the topology of the operating nodes; Based on the second sequence, a second vocabulary topology is obtained from the operational node topology; Based on the intensity difference value, the first lexical topology, and the second lexical topology, the matching range of the patent claim text is obtained.

5. The enterprise patent demand text matching method based on multimodal data fusion according to claim 4, characterized in that, The step of obtaining the intensity difference value based on the required vocabulary sequence includes: Traverse the sequence of required words to obtain a second required word; the second required word is any required word in the sequence of required words for which a corresponding first ratio has not been obtained. Based on the demand vocabulary sequence, the mean and standard deviation are obtained; the mean is the average value of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence; the standard deviation is the standard deviation of the demand intensity value corresponding to each demand vocabulary in the demand vocabulary sequence. Based on the second demand vocabulary, obtain the first demand intensity value; Based on the first demand intensity value and the average value, a first difference is obtained; the first difference is equal to the first demand intensity value minus the average value; Based on the first difference and the standard deviation, a first ratio is obtained; the first ratio is equal to the ratio of the first difference to the standard deviation. Based on the first ratio, the intensity difference value is obtained.

6. The enterprise patent demand text matching method based on multimodal data fusion according to claim 4, characterized in that, Based on the first sequence, the first vocabulary topology is obtained from the operational node topology, including: Based on the first sequence, a third demand term is obtained; the third demand term is any demand term in the first sequence for which no corresponding topological matching range has been obtained. Based on the third requirement term, obtain the farthest topological distance in the operation node topology; Based on the farthest topological distance, obtain the topological matching range corresponding to the third required vocabulary; The topological matching ranges corresponding to each required word in the first sequence are merged to obtain the topology of the first word.

7. The enterprise patent demand text matching method based on multimodal data fusion according to claim 6, characterized in that, The step of obtaining the topological matching range corresponding to the third requirement term based on the farthest topological distance includes: Based on the third requirement vocabulary and each first vocabulary triplet, a fifth quantity and a sixth quantity are obtained; the fifth quantity is the number of triplets in each first vocabulary triplet where the third requirement vocabulary is located in a primary or secondary position; the sixth quantity is the number of each first vocabulary triplet. Based on the fifth quantity and the sixth quantity, a second ratio is obtained; the second ratio is the ratio of the fifth quantity to the sixth quantity. Based on the second ratio and the farthest topological distance, obtain the expansion requirement topological distance; Based on the expansion requirement topology distance, the topology matching range corresponding to the third requirement term is obtained.

8. The enterprise patent demand text matching method based on multimodal data fusion according to claim 7, characterized in that, The step of obtaining the matching range of the patent claim text based on the intensity difference value, the first lexical topology, and the second lexical topology includes: The first vocabulary topology and the second vocabulary topology are overlapped to obtain multiple extended matching nodes; the extended matching nodes are nodes that are present in the first vocabulary topology but not in the second vocabulary topology. Based on the intensity difference value and each extended matching node, the matching range of the patent claim text is obtained.

9. A text matching system for enterprise patent requirements based on multimodal data fusion, characterized in that: include: A reader is used to process multimodal operational data from an enterprise database; the operational data includes any one or a combination of technical text, R&D text, images, logs, video, and audio; the enterprise database is pre-built. A server is used to obtain multiple first-word triples based on the operational data; Furthermore, based on the patent requirement text, multiple requirement terms and multiple second term triples are obtained; the patent requirement text is obtained in advance; the requirement terms are any terms in the patent requirement text. In addition, based on each first-word triplet and each second-word triplet, the demand intensity value corresponding to each demand word is obtained; The demand intensity value is used at least to characterize the salience of the corresponding demand terms in the patent demand text and operational data; Furthermore, based on each demand intensity value, the matching range of the patent demand text is obtained; Furthermore, based on the matching range, text matching is performed on the required words to obtain matching results; The process of obtaining the demand intensity value corresponding to each demand word based on each first word triplet and each second word triplet includes: Based on each demand term, obtain the first demand term; the first demand term is any demand term among the demand terms for which no corresponding demand intensity value has been obtained. Based on the first required vocabulary, multiple third vocabulary triples are obtained from each first vocabulary triple and each second vocabulary triple; the third vocabulary triple is any first vocabulary triple or second vocabulary triple, and the primary or secondary position in the third vocabulary triple is the first required vocabulary. Based on each third-word triplet, a first quantity and a second quantity are obtained; the first quantity is the number of triplets in each third-word triplet where the first demand word is in a primary position; the second quantity is the number of triplets in each third-word triplet where the first demand word is in a secondary position. Based on the first quantity and the second quantity, obtain the demand intensity value corresponding to the first demand term.

Citation Information

Patent Citations

  • Method and device for requirement identification

    CN102855251A

  • Patent evaluation system and evaluation method based on big data analysis

    CN115935061A