Semantic Recognition-Based Associated Technology Text Novelty Search Method and System
Through the associated technical text new search method based on semantic recognition, the shortcomings of traditional search methods are solved, efficient and accurate technical content new search is achieved, and users' innovative judgment ability and efficiency are improved.
Patent Information
- Application Number
- CN202411919693.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-12-25
AI Technical Summary
Traditional text-based keyword-based direct search of knowledge resources cannot meet users' daily needs and cannot obtain accurate information from a large number of resources.
The correlation technology text new search method based on semantic recognition is adopted, and the key index of the technical keyword to be retrieved is obtained, and the search matching and association algorithm calculation is performed, and the correlation technical text content is output with the technical content to be retrieved.
It improves the correlation between the retrieved related technical content and its own technology, provides accurate existing technology text, and improves users' accuracy in new search and judgment and innovative evaluation capabilities.
Smart Images

Figure CN119739848B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of patent novelty search, and in particular, to a method and system for novelty search of associated technical texts based on semantic recognition. Background Art
[0002] Intelligent knowledge novelty search is an important part of patent retrieval innovation. An excellent intelligent resource novelty search system can perform intelligent technical novelty search according to different technical needs of different users, helping users accurately search for relevant knowledge materials with less query content and master the target knowledge and skills. At present, traditional novelty search methods are generally static, divided into novelty search based on collaborative filtering and novelty search based on content. In recent years, novelty search based on deep query and reinforcement query has also emerged. Although many novelty search algorithms have been proposed, due to the continuous innovation and iteration of technologies, directly retrieving knowledge resources based on text keywords can no longer meet the daily needs of users and cannot obtain accurate information from a large number of resources. Therefore, a better novelty search scheme for associated technical texts needs to be found. Summary of the Invention
[0003] To solve the above technical problems, a method and system for novelty search of associated technical texts based on semantic recognition are provided. This technical solution solves the problem that directly retrieving knowledge resources based on text keywords can no longer meet the daily needs of users and cannot obtain accurate information from a large number of resources.
[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A method for novelty search of associated technical texts based on semantic recognition, comprising:
[0006] Based on the technical content to be retrieved, obtain at least one keyword of the technical content to be retrieved;
[0007] Based on the keyword of the technical content to be retrieved, perform a retrieval match on the existing technical content to be retrieved, and obtain the key index of each keyword of the technical content to be retrieved in the text of the technical content to be retrieved;
[0008] Use the keyword of the technical content to be retrieved to perform a retrieval match in the existing prior art database, obtain all the existing associated technical content corresponding to each keyword of the technical content to be retrieved, and form a set of existing associated technical content corresponding to each keyword of the technical content to be retrieved;
[0009] Take the union of the total sets of the existing associated technical content corresponding to all the keywords of the technical content to be retrieved to obtain the total set of associated technical content of the technical content to be retrieved;
[0010] Based on the technical association algorithm, calculate the association index between each existing associated technology in the total set of associated technical content and the technical content to be retrieved respectively;
[0011] Output the associated technical text content with the technical content to be retrieved based on the association index with the technical content to be retrieved.
[0012] Preferably, the method of retrieving and matching the existing technical content to be retrieved based on the technical keywords to be retrieved, and obtaining the key index of each technical keyword to be retrieved in the text of the technical content to be retrieved specifically includes:
[0013] Use the word segmentation technology to segment the text of the technical content to be retrieved to obtain the word set of the text of the technical content to be retrieved;
[0014] Take the ratio of the number of occurrences of the technical keyword to be retrieved in the word set to the total number of words in the word set to obtain the TF value of the technical keyword to be retrieved in the text of the technical content to be retrieved;
[0015] According to the key indexes of all technical keywords to be retrieved in the text of the technical content to be retrieved, perform normalization standard processing to obtain the key indexes of the technical keywords to be retrieved in the text of the technical content to be retrieved;
[0016] The formula for the normalization standard processing is:
[0017]
[0018] where G i is the key index of the i-th technical keyword to be retrieved in the text of the technical content to be retrieved, TF i is the TF value of the i-th technical keyword to be retrieved in the text of the technical content to be retrieved, and n is the total number of technical keywords to be retrieved.
[0019] Preferably, the method of retrieving and matching with the existing prior art database using the technical keywords to be retrieved, obtaining all the existing associated technical content corresponding to each technical keyword to be retrieved, and forming the set of existing associated technical content corresponding to each technical keyword to be retrieved specifically includes:
[0020] Based on the retrieval and matching of the technical keywords to be retrieved with the existing prior art database, record the existing technical data in all content texts that contain the technical keywords to be retrieved as the existing associated technology corresponding to the technical keywords to be retrieved;
[0021] Obtain all the existing associated technologies corresponding to the technical keywords to be retrieved, and form the set of existing associated technical content corresponding to the technical keywords to be retrieved.
[0022] Preferably, the technical association algorithm specifically includes:
[0023] Based on the keyword matching algorithm, determine the initial technical association index of each existing related technology in the total set of existing related technology content;
[0024] Based on the citation relationship of the existing related technologies in the total set of existing related technology content, determine the citation technology set and the cited technology set of each existing related technology;
[0025] Based on the association analysis formula, combine the initial technical association index, citation technology set and cited technology set of each existing related technology in the total set of existing related technology content, and calculate the updated value of the technical association index of each existing related technology in the total set of existing related technology content;
[0026] Judge whether the deviation value between the updated value of the technical association index of each existing related technology in the total set of existing related technology content and the initial technical association index of each existing related technology in the total set of existing related technology content is less than the deviation threshold. If so, use the updated value of the technical association index of each existing related technology in the total set of existing related technology content as the association index between each existing related technology in the total set of related technology content and the technology content to be retrieved. If not, normalize the updated value of the technical association index of each existing related technology in the total set of existing related technology content and use it as the initial technical association index of each existing related technology in the total set of existing related technology content and substitute it into the association analysis formula for calculation again;
[0027] Among them, the association analysis formula is specifically:
[0028]
[0029] In the formula, is the updated value of the technical association index of the u-th existing related technology in the total set of existing related technology content;
[0030] is the initial technical association index of the u-th existing related technology in the total set of existing related technology content;
[0031] α is the damping coefficient;
[0032] B u is the cited technology set of the u-th existing related technology in the total set of existing related technology content;
[0033] bv is the v-th element in the cited technology set of the u-th existing related technology in the total set of existing related technology content;
[0034] is the initial technical association index of the v-th element in the cited technology set of the u-th existing related technology in the total set of existing related technology content;
[0035] N v is the total number of elements in the reference technology set of the v-th element in the referenced technology set of the u-th existing related technology in the total set of existing related technology content.
[0036] Preferably, the keyword matching algorithm specifically includes:
[0037] Obtain at least one to-be-retrieved technology keyword corresponding to the existing related technology in the total set of existing related technology content, and construct a technology keyword set for this existing related technology;
[0038] Use word segmentation technology to segment the existing related technology text to obtain a word set of the existing related technology text;
[0039] Take the ratio of the number of occurrences of the elements in the technology keyword set of the existing related technology in the word set of the existing related technology text to the total number of words in the word set of the existing related technology text to obtain the TF value of the elements in the technology keyword set of the existing related technology in the existing related technology text;
[0040] Combine the key index of the elements in the technology keyword set of the existing related technology and the TF value of the elements in the technology keyword set of the existing related technology in the existing related technology text, and calculate the initial technology correlation index of the existing related technology through the initial correlation formula;
[0041] Among them, the specific initial correlation formula is:
[0042]
[0043] In the formula, a j is the j-th element in the technology keyword set of the existing related technology, A u is the technology keyword set of the u-th existing related technology, G j is the key index of the j-th element in the technology keyword set of the existing related technology, TF j is the TF value of the j-th element in the technology keyword set of the existing related technology in the existing related technology text, and m is the total number of existing related technologies in the total set of existing related technology content.
[0044] Preferably, the calculation formula for the deviation value between the updated value of the technology correlation index of each existing related technology in the total set of existing related technology content and the initial technology correlation index of each existing related technology in the total set of existing related technology content is:
[0045]
[0046] Among them, P is the deviation value;
[0047] The calculation formula for normalizing the updated value of the technical association index of each existing related technology in the total set of existing related technology content is as follows:
[0048]
[0049] Where is the initial technical association index of the u-th existing related technology after update.
[0050] Furthermore, a novelty search system for related technology texts based on semantic recognition is proposed to implement the novelty search method for related technology texts based on semantic recognition as described above, including:
[0051] A keyword extraction module, where the keyword extraction unit is used to obtain at least one keyword of the technology to be retrieved based on the technology content to be retrieved;
[0052] A self-matching module, which is electrically connected to the keyword extraction module. The self-matching module is used to retrieve and match the existing technology content to be retrieved based on the keyword of the technology to be retrieved, and obtain the key index of each keyword of the technology to be retrieved in the text of the technology content to be retrieved;
[0053] A retrieval module, which is electrically connected to the keyword extraction module. The retrieval module is used to retrieve and match the existing technology database with the keyword of the technology to be retrieved, obtain all the existing related technology content corresponding to each keyword of the technology to be retrieved, and form a set of existing related technology content corresponding to each keyword of the technology to be retrieved, and take the union of the total sets of existing related technology content corresponding to all the keywords of the technology to be retrieved to obtain the total set of related technology content of the technology content to be retrieved;
[0054] An association matching module, which is electrically connected to the keyword extraction module and the retrieval module. The association matching module is used to calculate the association index between each existing related technology in the total set of related technology content and the technology content to be retrieved based on the technical association algorithm, and output the related technology text content of the technology content to be retrieved based on the association index with the technology content to be retrieved.
[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0056] The present invention proposes an associated technology text novelty search solution based on semantic recognition. By performing multiple associated technology searches on the relevant weights between technical keywords and the technology itself, the correlation between technical keywords and the retrieved text, and the mutual citation and association relationships between the retrieved technical texts, the relevance between the retrieved associated technology content and the self-technology is effectively improved. Furthermore, accurate prior art texts are provided for the innovative evaluation of technical content, thereby improving the accuracy of the novelty search judgment of users for technical content. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of the method for novelty search of associated technology text based on semantic recognition proposed by this solution;
[0058] Figure 2 It is a flowchart of the method for the key index of each technical keyword to be retrieved in the text of the technical content to be retrieved in this solution;
[0059] Figure 3 It is a flowchart of the method for forming the set of existing associated technology content corresponding to the technical keywords to be retrieved in this solution;
[0060] Figure 4 It is a flowchart of the technical association algorithm in this solution;
[0061] Figure 5 It is a flowchart of the keyword matching algorithm in this solution;
[0062] Figure 6 It is an architecture diagram of the electronic device of this application;
[0063] Figure 7 It is a schematic diagram of the structure of the computer-readable storage medium of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variations.
[0065] Referring to Figure 1 as shown, a method for novelty search of associated technology text based on semantic recognition includes:
[0066] Based on the technical content to be retrieved, obtain at least one technical keyword to be retrieved;
[0067] Based on the technical keywords to be retrieved, perform a retrieval match on the existing technical content to be retrieved, and obtain the key index of each technical keyword to be retrieved in the text of the technical content to be retrieved;
[0068] Retrieve and match the keywords of the technology to be retrieved in the existing prior art database, obtain all the existing related technology content corresponding to each keyword of the technology to be retrieved, and form a set of existing related technology content corresponding to each keyword of the technology to be retrieved;
[0069] Take the union of the total sets of existing related technology content corresponding to all keywords of the technology to be retrieved to obtain the total set of related technology content of the technology to be retrieved;
[0070] Based on the technology association algorithm, calculate the association index between each existing related technology in the total set of related technology content and the technology content to be retrieved respectively;
[0071] Based on the association index with the technology content to be retrieved, output the related technology text content of the technology content to be retrieved.
[0072] The present invention is not limited to the keyword analysis and retrieval at a single level, but adopts a multi-dimensional and multi-level comprehensive evaluation method. It deeply explores the internal connection between technology keywords and technology entities, carefully quantifies the correlation degree between technology keywords and the retrieved text, and carefully traces the mutual citation and association context between the retrieved technology texts. Through this all-round and deep-level retrieval strategy, this solution greatly improves the fitness of the retrieved technology content with the user's own technology field. It provides rich and accurate prior art text resources for evaluating the innovation of technology content, and these resources are like the lighthouses of wisdom, illuminating the road of technological innovation. In addition, this solution also significantly enhances the novelty search and judgment ability of users when facing a large amount of technological information. It can not only help users quickly locate the most relevant and valuable technology texts, but also guide users to discover potential technological breakthrough points and innovation opportunities through intelligent analysis and suggestions. Thus, it greatly improves the efficiency and success rate of users in the process of technological research and development and innovation.
[0073] Refer to Figure 2 As shown, based on the keywords of the technology to be retrieved, retrieve and match the existing technology content to be retrieved, and the key indexes of each keyword of the technology to be retrieved in the text of the technology content to be retrieved specifically include:
[0074] Use the word segmentation technology to segment the text of the technology content to be retrieved to obtain the word set of the text of the technology content to be retrieved;
[0075] Take the ratio of the number of occurrences of the keyword of the technology to be retrieved in the word set to the total number of words in the word set to obtain the TF value of the keyword of the technology to be retrieved in the text of the technology content to be retrieved;
[0076] According to the key index of all technical keywords to be retrieved in the technical content text to be retrieved, perform normalization standard processing to obtain the key index of the technical keywords to be retrieved in the technical content text to be retrieved.
[0077] The formula for normalization standard processing is:
[0078]
[0079] Among them, G i is the key index of the i-th technical keyword to be retrieved in the technical content text to be retrieved, TF i is the TF value of the i-th technical keyword to be retrieved in the technical content text to be retrieved, and n is the total number of technical keywords to be retrieved.
[0080] In this solution, the word frequency analysis method is used to analyze the correlation between technical keywords and technical content. The more times a technical keyword appears in the technical text, the more it can represent the core technical content.
[0081] Refer to Figure 3 As shown, use the technical keywords to be retrieved to perform retrieval and matching in the existing prior art database, obtain all the existing related technical content corresponding to each technical keyword to be retrieved, and form the existing related technical content set corresponding to each technical keyword to be retrieved, which specifically includes:
[0082] Based on the retrieval and matching of the technical keywords to be retrieved and the existing prior art database, record the existing technical data in all content texts where the technical keywords to be retrieved appear as the existing related technologies corresponding to the retrieval and matching of the technical keywords.
[0083] Obtain all the existing related technologies corresponding to the retrieval and matching of the technical keywords to be retrieved, and form the existing related technical content set corresponding to the technical keywords to be retrieved.
[0084] By directly retrieving, screen out all the existing technologies that may be related to the technical content to be retrieved, which can ensure the integrity of the retrieved content, thereby reducing the missed detection rate and avoiding novelty search omissions.
[0085] Refer to Figure 4 As shown, the technical association algorithm specifically includes:
[0086] Based on the keyword matching algorithm, determine the initial technical association index of each existing related technology in the total set of existing related technical content.
[0087] Based on the citation relationship of the existing related technologies in the total set of existing related technical content, determine the citation technology set and the cited technology set of each existing related technology.
[0088] Based on the association analysis formula, combined with the initial technical association index, cited technology set, and cited-by technology set of each existing related technology in the total set of existing related technology content, calculate the updated value of the technical association index of each existing related technology in the total set of existing related technology content;
[0089] Judge whether the deviation value between the updated value of the technical association index of each existing related technology in the total set of existing related technology content and the initial technical association index of each existing related technology in the total set of existing related technology content is less than the deviation threshold. If so, use the updated value of the technical association index of each existing related technology in the total set of existing related technology content as the association index between each existing related technology in the total set of related technology content and the technology content to be retrieved and output it. If not, normalize the updated value of the technical association index of each existing related technology in the total set of existing related technology content and use it as the initial technical association index of each existing related technology in the total set of existing related technology content and substitute it back into the association analysis formula for calculation;
[0090] Among them, the association analysis formula is specifically:
[0091]
[0092] In the formula, is the updated value of the technical association index of the \(u\)-th existing related technology in the total set of existing related technology content;
[0093] is the initial technical association index of the \(u\)-th existing related technology in the total set of existing related technology content;
[0094] α is the damping coefficient;
[0095] B u is the cited-by technology set of the \(u\)-th existing related technology in the total set of existing related technology content;
[0096] \(b_v\) is the \(v\)-th element in the cited-by technology set of the \(u\)-th existing related technology in the total set of existing related technology content;
[0097] is the initial technical association index of the \(v\)-th element in the cited-by technology set of the \(u\)-th existing related technology in the total set of existing related technology content;
[0098] N v is the total number of elements in the cited technology set of the \(v\)-th element in the cited-by technology set of the \(u\)-th existing related technology in the total set of existing related technology content.
[0099] The calculation formula for the deviation value between the updated value of the technical association index of each existing related technology in the total set of existing related technologies and the initial technical association index of each existing related technology in the total set of existing related technologies is as follows:
[0100]
[0101] Where P is the deviation value;
[0102] The calculation formula for normalizing the updated value of the technical association index of each existing related technology in the total set of existing related technologies is as follows:
[0103]
[0104] Where is the initial technical association index of the u-th existing related technology after update.
[0105] Referring to Figure 5 shown, the keyword matching algorithm specifically includes:
[0106] Obtain at least one technical keyword to be retrieved corresponding to the existing related technology in the total set of existing related technologies, and construct the technical keyword set of this existing related technology;
[0107] Use the word segmentation technology to segment the existing related technology text to obtain the word set of the existing related technology text;
[0108] Take the number of occurrences of the elements in the technical keyword set of the existing related technology in the word set of the existing related technology text and divide it by the total number of words in the word set of the existing related technology text to obtain the TF value of the elements in the technical keyword set of the existing related technology in the existing related technology text;
[0109] Combine the key index of the elements in the technical keyword set of the existing related technology and the TF value of the elements in the technical keyword set of the existing related technology in the existing related technology text, and calculate the initial technical association index of the existing related technology through the initial association formula;
[0110] Where the initial association formula is specifically:
[0111]
[0112] In the formula, a j is the j-th element in the technical keyword set of the existing related technology, A u is the technical keyword set of the u-th existing related technology, G j is the key index of the j-th element in the technical keyword set of the existing related technology, TF jThe TF value of the j-th element in the set of technical keywords of the existing related technology in the text of the existing related technology, and m is the total number of existing related technologies in the total set of existing related technology content.
[0113] In this solution, the present invention assigns an initial association index to the existing related technologies by comprehensively considering two key dimensions: the correlation weight between the technical keywords and the technology itself, and the correlation between the technical keywords and the retrieved text. This index not only intuitively reflects the direct relevance between the existing related technologies and the technology to be retrieved, but also lays a solid foundation for subsequent in-depth exploration. On this basis, the present invention ingeniously combines the relevant citation relationships between the existing related technologies and draws on the core idea of the PageRank algorithm to further screen and refine the more authoritative and influential technical content in the existing related technologies. This process aims to capture the technical content that truly has the significance of novelty search. Through this series of precise and efficient multiple association technology retrieval operations, the present invention can quickly and accurately output the technical content highly relevant to the innovation subject. This not only provides a shortcut for users to quickly locate the most relevant and valuable technical texts, but also inspires users' innovation inspiration through intelligent analysis and suggestions, guiding them to discover potential technical breakthrough points and innovation opportunities, greatly improving the work efficiency and success rate of users in the process of technical research and development and innovation.
[0114] Furthermore, based on the same inventive concept as the above-mentioned method for novelty search of related technology texts based on semantic recognition, this solution proposes a system for novelty search of related technology texts based on semantic recognition, including:
[0115] A keyword extraction module, and the keyword extraction unit is used to obtain at least one technical keyword to be retrieved based on the technical content to be retrieved;
[0116] A self-matching module, the self-matching module is electrically connected to the keyword extraction module, and the self-matching module is used to perform retrieval and matching on the existing technical content to be retrieved based on the technical keywords to be retrieved, and obtain the key index of each technical keyword to be retrieved in the text of the technical content to be retrieved;
[0117] A retrieval module, the retrieval module is electrically connected to the keyword extraction module, and the retrieval module is used to perform retrieval and matching in the existing prior art database using the technical keywords to be retrieved, obtain all the existing related technical content corresponding to each technical keyword to be retrieved, and form a set of existing related technical content corresponding to each technical keyword to be retrieved, and take the union of the total sets of existing related technical content corresponding to all technical keywords to be retrieved to obtain the total set of related technical content of the technical content to be retrieved;
[0118] The associated matching module is electrically connected to the keyword extraction module and the retrieval module. The associated matching module is used to calculate the association index between each existing associated technology in the total set of associated technical content and the technical content to be retrieved based on the technical association algorithm, and output the associated technical text content of the technical content to be retrieved based on the association index with the technical content to be retrieved.
[0119] The usage process of the above system is as follows:
[0120] Step 1: The keyword extraction module obtains at least one keyword of the technology to be retrieved based on the technical content to be retrieved.
[0121] Step 2: The self-matching module retrieves and matches the existing technical content to be retrieved based on the keyword of the technology to be retrieved, and obtains the key index of each keyword of the technology to be retrieved in the text of the technical content to be retrieved.
[0122] Step 3: The retrieval module uses the keyword of the technology to be retrieved to perform retrieval and matching in the existing prior art database, obtains all the existing associated technical content corresponding to each keyword of the technology to be retrieved, and forms a set of existing associated technical content corresponding to each keyword of the technology to be retrieved, and takes the union of the total sets of existing associated technical content corresponding to all keywords of the technology to be retrieved to obtain the total set of associated technical content of the technical content to be retrieved.
[0123] Step 4: The associated matching module calculates the association index between each existing associated technology in the total set of associated technical content and the technical content to be retrieved based on the technical association algorithm, and outputs the associated technical text content of the technical content to be retrieved based on the association index with the technical content to be retrieved.
[0124] Furthermore, the method according to the embodiment of the present application can also be implemented by means of Figure 6 the architecture of the electronic device shown. As Figure 6 shown, the electronic device 500 may include a bus 501, one or more CPUs 502, a read-only memory (ROM) 503, a random access memory (RAM) 504, a communication port 505 connected to the network, an input / output component 506, a hard disk 507, etc. The storage device in the electronic device 500, such as the ROM 503 or the hard disk 507, may store a method for novelty search of associated technical texts based on semantic recognition provided by the present application. The electronic device 500 may also include a user interface 508. Of course, Figure 6 the architecture shown is only exemplary. When implementing different devices, one or more components shown in the Figure 6 electronic device may be omitted according to actual needs.
[0125] Figure 7It is a schematic structural diagram of a computer-readable storage medium provided by an embodiment of the present application. As Figure 7 shown, it is a computer-readable storage medium 600 according to an embodiment of the present application. Computer-readable instructions are stored on the computer-readable storage medium 600. When the computer-readable instructions are run by a processor, a method for novelty search of related technology texts based on semantic recognition according to an embodiment of the present application described with reference to the above drawings can be executed. The storage medium 600 includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0126] In summary, the advantages of the present invention are as follows: It can quickly and accurately output technical content highly relevant to the innovation entity, greatly improving the work efficiency and success rate of users in the process of technological research and development and innovation.
[0127] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A novelty search method for related technology texts based on semantic recognition, characterized in that, Including: Based on the technical content to be retrieved, obtain at least one technical keyword to be retrieved; Based on the technical keywords to be retrieved, conduct a retrieval and matching on the existing technical content to be retrieved, and obtain the key index of each technical keyword to be retrieved in the text of the technical content to be retrieved; Use the technical keywords to be retrieved to conduct a retrieval and matching in the existing prior art database, obtain all the existing related technical content corresponding to each technical keyword to be retrieved, and form a set of existing related technical content corresponding to each technical keyword to be retrieved; Take the union of the total sets of existing related technical content corresponding to all the technical keywords to be retrieved to obtain the total set of related technical content of the technical content to be retrieved; Based on the technical association algorithm, calculate the association index between each existing related technology in the total set of related technical content and the technical content to be retrieved respectively; Based on the association index with the technical content to be retrieved, output the text content of the related technology of the technical content to be retrieved; The technical association algorithm specifically includes: Based on the keyword matching algorithm, determine the initial technical association index of each existing related technology in the total set of existing related technical content; Based on the citation relationship of the existing related technologies in the total set of existing related technical content, determine the set of cited technologies and the set of technologies being cited for each existing related technology; Based on the association analysis formula, combine the initial technical association index, the set of cited technologies and the set of technologies being cited of each existing related technology in the total set of existing related technical content, and calculate the updated value of the technical association index of each existing related technology in the total set of existing related technical content; Judge whether the deviation value between the updated value of the technical association index of each existing related technology in the total set of existing related technical content and the initial technical association index of each existing related technology in the total set of existing related technical content is less than the deviation threshold. If so, use the updated value of the technical association index of each existing related technology in the total set of existing related technical content as the association index between each existing related technology in the total set of related technical content and the technical content to be retrieved and output it. If not, normalize the updated value of the technical association index of each existing related technology in the total set of existing related technical content and use it as the initial technical association index of each existing related technology in the total set of existing related technical content and substitute it into the association analysis formula for calculation again; Among them, the association analysis formula is specifically: In the formula, is the updated value of the technical correlation index of the u-th existing related technology in the total set of existing related technology content; is the initial technical association index of the u-th existing related technology in the total set of existing related technology content; α is the damping coefficient; B u It is the cited technology set of the u-th existing related technology in the total set of existing related technology content; bv is the v-th element in the set of technologies being cited of the u-th existing related technology in the total set of existing related technical content; is the initial technical association index of the v-th element in the cited technology set of the u-th existing related technology in the total set of existing related technology content; N v It is the total number of elements in the set of cited technologies of the v-th element in the set of cited technologies of the u-th existing related technology in the total set of existing related technology content.
2. The novelty search method for related technology texts based on semantic recognition according to claim 1, characterized in that The specific process of obtaining the key index of each technical keyword to be retrieved in the text of the technical content to be retrieved by conducting a retrieval and matching on the existing technical content to be retrieved based on the technical keywords to be retrieved includes: Use the word segmentation technology to segment the text of the technical content to be retrieved to obtain the word set of the text of the technical content to be retrieved; Take the ratio of the number of occurrences of the technical keyword to be retrieved in the word set to the total number of words in the word set to obtain the TF value of the technical keyword to be retrieved in the text of the technical content to be retrieved; According to the key indexes of all the technical keywords to be retrieved in the technical content text to be retrieved, perform normalization standard processing to obtain the key indexes of the technical keywords to be retrieved in the technical content text to be retrieved; The formula for the normalization standard processing is: Among them, G i is the key index of the i-th technical keyword to be retrieved in the text of the technical content to be retrieved, TF i is the TF value of the i-th technical keyword to be retrieved in the text of the technical content to be retrieved, and n is the total number of technical keywords to be retrieved.
3. The novelty search method for related technology texts based on semantic recognition according to claim 2, characterized in that, The process of using the technical keywords to be retrieved to perform retrieval matching in the existing prior art database, obtaining all the existing related technical content corresponding to each technical keyword to be retrieved, and forming the set of existing related technical content corresponding to each technical keyword to be retrieved specifically includes: Based on the retrieval matching of the technical keywords to be retrieved and the existing prior art database, the existing technical data in all content texts where the technical keywords to be retrieved appear are recorded as the existing related technologies corresponding to the technical keywords to be retrieved; Obtain all the existing related technologies corresponding to the technical keywords to be retrieved, and form the set of existing related technical content corresponding to the technical keywords to be retrieved.
4. The novelty search method for related technology texts based on semantic recognition according to claim 3, characterized in that The keyword matching algorithm specifically includes: Obtain at least one technical keyword to be retrieved corresponding to the existing related technology in the total set of existing related technical content, and construct the set of technical keywords of this existing related technology; Use the word segmentation technology to perform word segmentation processing on the existing related technology text to obtain the word set of the existing related technology text; Take the ratio of the number of occurrences of the elements in the set of technical keywords of the existing related technology in the word set of the existing related technology text to the total number of words in the word set of the existing related technology text to obtain the TF value of the elements in the set of technical keywords of the existing related technology in the existing related technology text; Combine the key indexes of the elements in the set of technical keywords of the existing related technology and the TF values of the elements in the set of technical keywords of the existing related technology in the existing related technology text, and calculate the initial technical association index of the existing related technology through the initial association formula; Among them, the specific initial association formula is: where a j is the j-th element in the technical keyword set of the existing related technology, A u is the technical keyword set of the u-th existing related technology, G j is the key index of the j-th element in the technical keyword set of the existing related technology, TF j is the TF value of the j-th element in the technical keyword set of the existing related technology in the text of the existing related technology, and m is the total number of existing related technologies in the total set of existing related technology content.
5. The novelty search method for related technology texts based on semantic recognition according to claim 4, characterized in that, The calculation formula for the deviation value between the updated value of the technical association index of each existing related technology in the total set of existing related technical content and the initial technical association index of each existing related technology in the total set of existing related technical content is: Among them, P is the deviation value; The calculation formula for normalizing the updated value of the technical association index of each existing related technology in the total set of existing related technical content is: Among them, is the initial technical association index of the updated u-th existing related technology.
6. An associated technology text novelty search system based on semantic recognition, characterized in that, Used to implement the related technology text novelty search method based on semantic recognition as described in any one of claims 1-5, including: A keyword extraction module, where the keyword extraction unit is used to obtain at least one technical keyword to be retrieved based on the technical content to be retrieved; A self-matching module, the self-matching module is electrically connected to the keyword extraction module, and the self-matching module is used to perform retrieval matching on the existing technical content to be retrieved based on the technical keywords to be retrieved, and obtain the key index of each technical keyword to be retrieved in the technical content text to be retrieved; A retrieval module, electrically connected to the keyword extraction module, for retrieving and matching the to-be-retrieved technical keywords in an existing prior art database, obtaining all the existing related technical contents corresponding to each to-be-retrieved technical keyword, forming a set of existing related technical contents corresponding to each to-be-retrieved technical keyword, and taking the union of the total sets of existing related technical contents corresponding to all to-be-retrieved technical keywords to obtain the total set of related technical contents of the to-be-retrieved technical contents; An association matching module, electrically connected to the keyword extraction module and the retrieval module, for calculating the association index between each existing related technology in the total set of related technical contents and the to-be-retrieved technical contents respectively based on a technology association algorithm, and outputting the related technical text content of the to-be-retrieved technical contents based on the association index with the to-be-retrieved technical contents.
7. An electronic device, characterized in that, Comprising: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute an association technology text novelty search method based on semantic recognition as described in any one of claims 1-5.
8. A computer-readable storage medium storing computer-readable instructions, characterized in that, When the computer-readable instructions are executed by a processor, an association technology text novelty search method based on semantic recognition as described in any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Semantic-based approximate text search method and device, computer equipment and medium
CN113434636A
Information entropy-based quotation recommendation method, device and terminal
CN117076658A