Research and development project repetition analysis method and device, electronic equipment and storage medium

By obtaining the vectors of R&D project documents, calculating the repetition rate using semantic substructure similarity, and combining optical character recognition and technical feature knowledge graphs, we solved the accuracy and efficiency issues of repeated reviews of R&D projects and achieved automated repetition analysis.

CN120706430APending Publication Date: 2025-09-26中国联合网络通信有限公司广东省分公司
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510801431.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the problem of duplicate R&D projects relies on the subjective judgment of experts, resulting in a lack of consistency and accuracy in the review results. As the number of projects increases, manual review becomes difficult to fully identify duplicate R&D projects.

Method used

By obtaining the vectors of the R&D project documents that are pending and already approved, the repetition rate is calculated using the semantic substructure similarity, and word vectors are constructed by combining optical character recognition, cleaning, named entity recognition and technical feature knowledge graphs for automated repetition analysis.

Benefits of technology

It realizes the automated repeated analysis of R&D projects, improves the accuracy and efficiency of the review, and reduces omissions in manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706430A_ABST
    Figure CN120706430A_ABST
Patent Text Reader

Abstract

The invention discloses a research and development project repetition analysis method and device, electronic equipment and a storage medium. The method comprises the steps that a vector of a research and development project document to be approved and a vector of an approved research and development project document are obtained, the vector of the research and development project document to be approved comprises a plurality of first semantic substructures, and the vector of the approved research and development project document comprises a plurality of second semantic substructures; for any project-approved research and development project document, based on a plurality of first semantic substructures corresponding to the research and development project document to be approved and a plurality of second semantic substructures corresponding to the project-approved research and development project document, determining a repetition rate of a vector of the research and development project document to be approved and a vector of the project-approved research and development project document; and determining a repetition analysis result of the research and development project document to be subjected to project establishment based on the repetition rate of the vector of the research and development project document to be subjected to project establishment and the vector of each research and development project document subjected to project establishment. According to the technical scheme, whether research and development projects are repeated or not is automatically analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a method, device, electronic device and storage medium for repeat analysis of R&D projects. Background Art

[0002] In the process of R&D project application and management, the issue of whether there is duplicate R&D in the technical solutions of R&D projects has gradually become prominent, and it mainly depends on the subjective judgment of experts based on their own professional knowledge and experience.

[0003] The expert review method has significant limitations. On the one hand, there are differences in personal cognition among experts, and different experts may have different criteria for judging whether the same technical solution is repeated, resulting in a lack of consistency and objectivity in the review results; on the other hand, with the rapid increase in the number of R&D projects, the amount of information that experts need to process is growing exponentially. Manual review is bound to have omissions, making it difficult to comprehensively and accurately identify duplicate R&D projects. Summary of the Invention

[0004] The present invention provides a method, device, electronic device and storage medium for analyzing duplication of R&D projects, so as to realize automatic analysis of whether R&D projects are duplicated.

[0005] According to one aspect of the present invention, a method for analyzing duplication of R&D projects is provided, comprising:

[0006] Obtaining a vector of a document of a pending R&D project and a vector of a document of an approved R&D project, wherein the vector of the document of the pending R&D project includes a plurality of first semantic substructures, and the vector of the document of the approved R&D project includes a plurality of second semantic substructures;

[0007] For any approved R&D project document, based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the approved R&D project document, determine the repetition rate of the vector of the R&D project document to be approved and the vector of the approved R&D project document;

[0008] The repetition analysis result of the research and development project document to be established is determined based on the repetition rate of the vector of the research and development project document to be established and the vector of each established research and development project document.

[0009] According to another aspect of the present invention, there is provided a device for repeated analysis of R&D projects, comprising:

[0010] A vector acquisition module for R&D project documents, configured to acquire vectors of R&D project documents to be approved and vectors of approved R&D project documents, wherein the vectors of the R&D project documents to be approved include multiple first semantic substructures, and the vectors of the approved R&D project documents include multiple second semantic substructures;

[0011] A module for determining the repetition rate of vectors of R&D project documents, for determining, for any approved R&D project document, the repetition rate of the vectors of the R&D project document to be approved and the vectors of the approved R&D project document based on a plurality of first semantic substructures corresponding to the R&D project document to be approved and a plurality of second semantic substructures corresponding to the R&D project document to be approved;

[0012] The module for determining the repeated analysis results of the R&D project documents to be established is used to determine the repeated analysis results of the R&D project documents to be established based on the repetition rates of the vectors of the R&D project documents to be established and the vectors of each established R&D project document.

[0013] According to another aspect of the present invention, an electronic device is provided, comprising:

[0014] at least one processor;

[0015] and a memory communicatively coupled to the at least one processor;

[0016] In which, the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the R&D project repeated analysis method described in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the R&D project duplication analysis method described in any embodiment of the present invention when executed.

[0018] The technical solution of the embodiment of the present invention is to obtain the vector of the R&D project document to be established and the vector of the R&D project document that has been established, wherein the vector of the R&D project document to be established includes multiple first semantic substructures, and the vector of the R&D project document that has been established includes multiple second semantic substructures; for any R&D project document that has been established, based on the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document that has been established, determine the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document that has been established; based on the repetition rate of the vector of the R&D project document to be established and the vector of each R&D project document that has been established, determine the repetition analysis result of the R&D project document to be established. The above technical solution realizes the automatic analysis of whether the R&D project is repeated, and improves the accuracy and efficiency of R&D project review.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0021] Figure 1 This is a flow chart of a method for repeated analysis of R&D projects provided according to the first embodiment of the present invention;

[0022] Figure 2 This is a flow chart of a method for repeated analysis of R&D projects provided according to the second embodiment of the present invention;

[0023] Figure 3 This is a flow chart of a method for repeated analysis of R&D projects provided according to the third embodiment of the present invention;

[0024] Figure 4 This is a flow chart of a method for repeated analysis of R&D projects provided according to a fourth embodiment of the present invention;

[0025] Figure 5 This is a schematic diagram of the structure of a repeated analysis device for R&D projects provided according to a fifth embodiment of the present invention;

[0026] Figure 6 It is a structural diagram of an electronic device for implementing the repeated analysis method of R&D projects according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a method for analyzing duplicate R&D projects provided in the first embodiment of the present invention. This embodiment is applicable to the situation where the R&D project documents are automatically judged as duplicate R&D. The method can be executed by a R&D project duplication analysis device, which can be implemented in the form of hardware and / or software. The R&D project duplication analysis device can be configured in electronic devices such as terminals and / or servers. Figure 1 As shown, the method includes:

[0031] S110. Obtain a vector of a document of a research and development project to be established and a vector of a document of an established research and development project, wherein the vector of the document of the research and development project to be established includes multiple first semantic substructures, and the vector of the document of the established research and development project includes multiple second semantic substructures.

[0032] R&D project documentation refers to documents related to the R&D project, including but not limited to project proposals, opening reports, milestone reports, and final reports. R&D project documentation can be in PDF, Word, or image formats, with no specific restrictions here. Semantic substructure refers to the semantic and structural information within R&D project documentation. For example, a semantic substructure could be a technical operational process fragment or a set of related technical terms.

[0033] For example, the documents of R&D projects to be established and the documents of R&D projects that have been established can be read from the R&D project management system, and then the documents of R&D projects to be established and the documents of R&D projects that have been established can be vectorized to obtain the vectors of the documents of R&D projects to be established and the vectors of the documents of R&D projects that have been established.

[0034] S120. For any R&D project document that has been approved, determine the repetition rate between the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the R&D project document that has been approved.

[0035] Among them, the duplication rate refers to the degree of duplication between two R&D project documents, which can be used to determine whether the R&D project document to be established is a duplicate R&D project.

[0036] Specifically, the similarity between the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document that has been established can be calculated, and then the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document that has been established can be determined based on the similarity between the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document that has been established.

[0037] S130 : Determine a repetition analysis result of the document of the research and development project to be established based on the repetition rate between the vector of the document of the research and development project to be established and the vector of each established research and development project document.

[0038] The duplication analysis result refers to the judgment result of whether the R&D project document to be established is a duplicate R&D project, which can be a duplicate R&D project or a non-duplicate R&D project.

[0039] In some optional embodiments, for any R&D project document that has been approved, if the repetition rate between the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved is greater than a preset repetition rate threshold, then the repetition analysis result of the R&D project document to be approved is determined to be a repeated R&D project; otherwise, the repetition analysis result of the R&D project document to be approved is determined to be a non-repetitive R&D project.

[0040] In some optional embodiments, for any R&D project document that has been approved, if the repetition rate between the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved is within a preset repetition rate range, then the repetition analysis result of the R&D project document to be approved is determined to be a repeated R&D project; otherwise, the repetition analysis result of the R&D project document to be approved is determined to be a non-repetitive R&D project.

[0041] The technical solution of the embodiment of the present invention is to obtain the vector of the R&D project document to be established and the vector of the R&D project document that has been established, wherein the vector of the R&D project document to be established includes multiple first semantic substructures, and the vector of the R&D project document that has been established includes multiple second semantic substructures; for any R&D project document that has been established, based on the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document that has been established, determine the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document that has been established; based on the repetition rate of the vector of the R&D project document to be established and the vector of each R&D project document that has been established, determine the repetition analysis result of the R&D project document to be established. The above technical solution realizes the automatic analysis of whether the R&D project is repeated, and improves the accuracy and efficiency of R&D project review.

[0042] Example 2

[0043] Figure 2 A flow chart of a method for repeated analysis of R&D projects provided in the second embodiment of the present invention, the method of this embodiment can be combined with the various optional schemes in the method for repeated analysis of R&D projects provided in the above embodiments. The method for repeated analysis of R&D projects provided in this embodiment has been further optimized. Optionally, the method of determining the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document to be established based on the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document to be established includes: determining the semantic substructure similarity between the first semantic substructure and the second semantic substructure based on the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document to be established; and determining the repetition rate between the vector of the R&D project document to be established and the vector of the R&D project document to be established based on the semantic substructure similarity between the first semantic substructure and the second semantic substructure.

[0044] like Figure 2 As shown, the method includes:

[0045] S210. Obtain a vector of a document of a research and development project to be established and a vector of a document of an established research and development project, wherein the vector of the document of the research and development project to be established includes multiple first semantic substructures, and the vector of the document of the established research and development project includes multiple second semantic substructures.

[0046] S220. For any approved R&D project document, determine the semantic substructure similarity between the first semantic substructure and the second semantic substructure based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the approved R&D project document.

[0047] The semantic substructure similarity is used to measure the similarity between the first semantic substructure and the second semantic substructure. A higher score corresponding to the semantic substructure similarity indicates a higher similarity.

[0048] Specifically, the semantic substructure similarity between the first semantic substructure and the second semantic substructure may be calculated using distance metrics such as Euclidean distance and Manhattan distance.

[0049] On the basis of the above embodiments, optionally, based on multiple first semantic substructures corresponding to the R&D project documents to be established and multiple second semantic substructures corresponding to the R&D project documents that have been established, the semantic substructure similarity between the first semantic substructure and the second semantic substructure is determined, including: for any of the first semantic substructures, determining the word frequency-inverse file frequency vector based on the topic relevance factor corresponding to the first semantic substructure; for any of the second semantic substructures, determining the word frequency-inverse file frequency vector based on the topic relevance factor corresponding to the second semantic substructure; based on the word frequency-inverse file frequency vector based on the topic relevance factor corresponding to the first semantic substructure and the word frequency-inverse file frequency vector based on the topic relevance factor corresponding to the second semantic substructure, determining the semantic substructure similarity between the first semantic substructure and the second semantic substructure.

[0050] In an embodiment of the present invention, the formula for determining the term frequency-inverse document frequency vector based on the topic relevance factor is:

[0051]

[0052] Where t represents the word in the semantic substructure, d represents the document where the semantic substructure is located, TF(t,d) represents the word frequency of word t in document d, IDF(t) represents the inverse document frequency of word t, and β represents the topic relevance adjustment parameter. represents the vector representation of word t, represents the topic vector of document d, express and The cosine similarity of .

[0053] Specifically, Where D represents the total number of R&D project documents. β is used to adjust the influence of topic-related factors on the TF-IDF value and can range from [0, 1]. It can be generated through a word vector model (such as Word2Vec or a semantically enhanced word vector generation model). It can be obtained through topic models such as latent Dirichlet allocation, representing the main topic direction of the R&D project document. It is used to measure the correlation between word t and document topic vector. Its value range is between [-1, 1]. The closer the value is to 1, the stronger the correlation is.

[0054] It should be noted that, in the embodiment of the present invention, based on the traditional TF-IDF, a topic relevance factor is introduced so that the TF-IDF value can better reflect the importance and topic relevance of a word in a document.

[0055] Furthermore, the word frequency-inverse document frequency vector based on the topic relevance factor corresponding to the first semantic substructure and the word frequency-inverse document frequency vector based on the topic relevance factor corresponding to the second semantic substructure can be calculated using distance metrics such as Euclidean distance and Manhattan distance to obtain the semantic substructure similarity between the first semantic substructure and the second semantic substructure.

[0056] S230. Determine a repetition rate between the vector of the to-be-established R&D project document and the vector of the established R&D project document based on the semantic substructure similarity between the first semantic substructure and the second semantic substructure.

[0057] The formula for determining the repetition rate between the vector of the R&D project document to be approved and the vector of the approved R&D project document is:

[0058]

[0059] Where A={a1,a2,...,a m} represents the vector of R&D project documents to be established, B={b1,b2,...,b n} represents the vector of the R&D project document that has been approved, Similarity(A,B) represents the repetition of the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved, a i The first semantic substructure of the i-th vector in the R&D project document to be established, b j represents the jth second semantic substructure in the vector of the R&D project document that has been approved, m represents the number of first semantic substructures, and n represents the number of second semantic substructures; S(a i ,b j ) represents the semantic substructure similarity between the i-th first semantic substructure and the j-th second semantic substructure.

[0060] It should be noted that the embodiment of the present invention uses the above-mentioned repetition rate calculation formula to comprehensively consider the matching situations in two directions, from the semantic substructure of A to the semantic substructure of B, and from the semantic substructure of B to the semantic substructure of A, to realize the repetition rate calculation based on semantic structure matching, thereby making the calculated repetition rate more accurate.

[0061] S240: Determine a repetition analysis result of the document of the research and development project to be established based on the repetition rate between the vector of the document of the research and development project to be established and the vector of each established research and development project document.

[0062] The technical solution of the embodiment of the present invention calculates the repetition rate between the vector of the R&D project document to be established and the vector of the R&D project document that has been established through the semantic substructure similarity between the first semantic substructure and the second semantic substructure, thereby realizing the quantitative calculation of the repetition of the R&D project documents.

[0063] Example 3

[0064] Figure 3 This is a flow chart of a method for analyzing R&D project duplications provided in Example 3 of the present invention. The method of this embodiment can be combined with the various optional solutions in the R&D project duplication analysis methods provided in the above embodiments. The R&D project duplication analysis method provided in this embodiment is further optimized. Optionally, obtaining the vector of the document of the R&D project to be established includes: obtaining the document of the R&D project to be established; performing optical character recognition on the document of the R&D project to be established to obtain text information of the document of the R&D project to be established; cleaning the text information of the document of the R&D project to be established to obtain the cleaned text information of the document of the R&D project to be established; performing named entity recognition on the cleaned text information of the document of the R&D project to be established to obtain the technical features of the document of the R&D project to be established, and the technical features of the document of the R&D project to be established include technical terms, technical parameters and technical operation procedures; constructing a technical feature knowledge graph corresponding to the R&D project document based on the technical features of the document of the R&D project to be established, and the technical feature knowledge graph corresponding to the document of the R&D project to be established includes multiple technical feature vectors and the relationship between the technical feature vectors; updating the word vector of the document of the R&D project to be established based on the technical feature knowledge graph corresponding to the document of the R&D project to be established to obtain the vector of the document of the R&D project to be established.

[0065] like Figure 3 As shown, the method includes:

[0066] S310. Obtain the R&D project documents to be established.

[0067] S320: Perform optical character recognition on the document of the research and development project to be established to obtain text information of the document of the research and development project to be established.

[0068] S330: Clean the text information of the document of the research and development project to be established to obtain the cleaned text information of the document of the research and development project to be established.

[0069] In an embodiment of the present invention, by cleaning the text information of the R&D project document to be established, noise in the R&D project document can be efficiently removed.

[0070] Specifically, the steps for cleaning the text information of the R&D project documents to be approved may include but are not limited to:

[0071] Use regular expressions to clean the text information of the R&D project documents to be established, and remove unnecessary HTML tags, special characters and extra spaces in the text information.

[0072] Convert the text information of R&D project documents to be established to lowercase to avoid data inconsistencies caused by different uppercase and lowercase letters.

[0073] Use a stop word list to remove meaningless words from the text information of the R&D project documents to be established.

[0074] Use spell checking tools or libraries (such as TextBlob or spellchecker in Python) to perform spelling correction to ensure that the words in the text information of the R&D project documents to be established are written correctly.

[0075] S340. Perform named entity recognition on the cleaned text information of the R&D project document to be established to obtain the technical features of the R&D project document to be established. The technical features of the R&D project document to be established include technical terms, technical parameters and technical operation procedures.

[0076] Specifically, named entity recognition technology can be used to extract technical features such as technical terms, technical parameters, and technical operation procedures from the cleaned text information of the R&D project documents to be established.

[0077] S350. Construct a technical feature knowledge graph corresponding to the R&D project document based on the technical features of the R&D project document to be established. The technical feature knowledge graph corresponding to the R&D project document to be established includes multiple technical feature vectors and the relationship between the technical feature vectors.

[0078] It should be noted that, through the technical feature knowledge graph corresponding to the R&D project documents, the technical features identified by named entities can be associated and organized, and the hierarchical relationship and causal relationship between the technical features can be clarified.

[0079] For example, in a software development project, the technical feature knowledge graph corresponding to the R&D project document can clarify the causal relationship between the "algorithm optimization" technical feature and the "performance improvement" technical feature, as well as the specific technical parameters and operation steps involved in the "algorithm optimization" technical feature.

[0080] S360. Based on the technical feature knowledge graph corresponding to the document of the research and development project to be established, the word vector of the document of the research and development project to be established is updated to obtain the vector of the document of the research and development project to be established.

[0081] In this embodiment of the present invention, the formula for updating the word vector is:

[0082]

[0083] Among them, w i Represents the i-th word vector in the R&D project document to be established, represents w i The vector representation before updating is, represents w i The vector before update is represented, α represents the learning rate, and N(i) represents w i The set of adjacent nodes in the technical feature knowledge graph, represents w i In the technical feature knowledge graph j The associated weight of .

[0084] Specifically, the learning rate can control the step size of each word vector update and is used to adjust the learning speed of the word vector generation model. A higher association weight indicates a closer association.

[0085] It should be noted that the formula for updating word vectors is calculated by considering w i and the adjacent node word w j The association weight e ij , and the adjacent node word vector With its own word vector The difference between i Update so that the word vector can better reflect the semantic position of the word in the knowledge graph.

[0086] S370. Obtain a vector of an approved R&D project document, wherein the vector of the R&D project document to be approved includes a plurality of first semantic substructures, and the vector of the approved R&D project document includes a plurality of second semantic substructures.

[0087] Similarly, the documents of approved R&D projects can be subjected to optical character recognition, cleaning, named entity recognition, construction of technical feature knowledge graphs, and word vector updates to obtain the vectors of the approved R&D project documents.

[0088] S380. For any R&D project document that has been approved, determine the repetition rate between the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the R&D project document that has been approved.

[0089] S390. Determine a repetition analysis result of the document of the research and development project to be established based on the repetition rate between the vector of the document of the research and development project to be established and the vector of each established research and development project document.

[0090] The technical solution of the embodiment of the present invention obtains the vector of the R&D project document by performing optical character recognition, cleaning, named entity recognition, building a technical feature knowledge graph and updating the word vector on the R&D project document, so that the word vector can better reflect the semantic position of the word in the knowledge graph.

[0091] Example 4

[0092] Figure 4 A flowchart of a method for repeated analysis of R&D projects provided in Example 4 of the present invention, the method of this embodiment can be combined with the various optional schemes in the method for repeated analysis of R&D projects provided in the above embodiments. The method for repeated analysis of R&D projects provided in this embodiment has been further optimized. Optionally, after determining the repeated analysis results of the R&D project documents to be established based on the repetition rate of the vectors of the R&D project documents to be established and the vectors of each R&D project document that has been established, it also includes: performing density and hierarchical fusion cluster analysis on the vectors of the R&D project documents to be established and the vectors of each R&D project document that has been established to obtain a combination of repeated R&D projects.

[0093] like Figure 4 As shown, the method includes:

[0094] S410. Obtain a vector of a document of an R&D project to be established and a vector of a document of an established R&D project, wherein the vector of the document of the R&D project to be established includes multiple first semantic substructures, and the vector of the document of the established R&D project includes multiple second semantic substructures.

[0095] S420. For any approved R&D project document, determine the repetition rate between the vector of the R&D project document to be approved and the vector of the approved R&D project document based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the R&D project document to be approved.

[0096] S430: Determine a repetition analysis result of the document of the research and development project to be established based on the repetition rate between the vector of the document of the research and development project to be established and the vector of each established research and development project document.

[0097] S440: Perform density and hierarchical fusion cluster analysis on the vectors of the R&D project documents to be approved and the vectors of each approved R&D project document to obtain a combination of repeated R&D projects.

[0098] The density and hierarchy fusion clustering analysis can be a DHC (Divisive Hierarchical Clustering) clustering algorithm. A duplicate R&D project combination refers to R&D project documents clustered in the same group or cluster. In other words, R&D project documents in the same group correspond to R&D projects that are duplicated.

[0099] In some embodiments, a distance decay factor γ can be introduced during density calculation in a density-hierarchical cluster analysis to adjust the contribution of neighboring points to the density calculation of a data point. The value of γ generally ranges from [0 to 1], with a larger γ indicating a more significant effect of distance on density.

[0100] For example, for any data point x in the vector of R&D project documents, the calculation formula for its density β(x) is:

[0101]

[0102] Here, N(x) represents the neighborhood of data point x. That is, N(x) is the set of other data points within a certain distance range from x. The specific distance range can be set based on the data distribution characteristics and clustering requirements. d(x,y) represents the distance between data point x and its neighboring point y. It can be calculated using distance metrics such as Euclidean distance and Manhattan distance to measure the proximity of two data points in data space.

[0103] It should be emphasized that the above density calculation formula introduces a distance attenuation factor so that the neighborhood points closer to the data point x contribute more to its density, thereby more accurately reflecting the local density around the data point and providing a more reasonable density estimation for cluster analysis.

[0104] The technical solution of the embodiment of the present invention obtains a combination of duplicate R&D projects by performing density and hierarchical fusion cluster analysis on the vectors of the R&D project documents to be approved and the vectors of each approved R&D project document, thereby realizing secondary analysis of duplicate R&D projects and further improving the detection accuracy of duplicate R&D projects.

[0105] Example 5

[0106] Figure 5 This is a schematic diagram of the structure of a device for repeated analysis of R&D projects provided in Example 5 of the present invention. Figure 5 As shown, the device includes:

[0107] A R&D project document vector acquisition module 510 is configured to acquire vectors of R&D project documents to be approved and vectors of approved R&D project documents, wherein the vectors of the R&D project documents to be approved include multiple first semantic substructures, and the vectors of the approved R&D project documents include multiple second semantic substructures.

[0108] A module 520 for determining a repetition rate of vectors of R&D project documents, for determining, for any approved R&D project document, a repetition rate between the vectors of the R&D project document to be approved and the vectors of the approved R&D project document based on a plurality of first semantic substructures corresponding to the R&D project document to be approved and a plurality of second semantic substructures corresponding to the approved R&D project document;

[0109] The module 530 for determining the duplication analysis results of the R&D project document to be established is used to determine the duplication analysis results of the R&D project document to be established based on the duplication rates of the vectors of the R&D project document to be established and the vectors of each established R&D project document.

[0110] The technical solution of the embodiment of the present invention is to obtain the vector of the R&D project document to be established and the vector of the R&D project document that has been established, wherein the vector of the R&D project document to be established includes multiple first semantic substructures, and the vector of the R&D project document that has been established includes multiple second semantic substructures; for any R&D project document that has been established, based on the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document that has been established, determine the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document that has been established; and determine the repetition analysis result of the R&D project document to be established based on the repetition rate of the vector of the R&D project document to be established and the vector of each R&D project document that has been established. The above technical solution realizes the automatic analysis of whether the R&D project is repeated.

[0111] In some optional implementations, the repetition rate determination module 520 of the vector of the R&D project document includes:

[0112] a semantic substructure similarity determination unit, configured to determine the semantic substructure similarity between a first semantic substructure and a second semantic substructure based on a plurality of first semantic substructures corresponding to the R&D project document to be established and a plurality of second semantic substructures corresponding to the R&D project document that has been established;

[0113] The repetition rate determination unit of the vector of the R&D project document is used to determine the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document that has been established based on the semantic substructure similarity between the first semantic substructure and the second semantic substructure.

[0114] In some optional implementations, the formula for determining the repetition rate between the vector of the R&D project document to be approved and the vector of the approved R&D project document is:

[0115]

[0116] Among them, A represents the vector of the R&D project document to be approved, B represents the vector of the R&D project document that has been approved, Similarity(A,B) represents the repetition of the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved, a i The first semantic substructure of the i-th vector in the R&D project document to be established, b j represents the jth second semantic substructure in the vector of the R&D project document that has been approved, m represents the number of first semantic substructures, and n represents the number of second semantic substructures; S(a i ,b j ) represents the semantic substructure similarity between the i-th first semantic substructure and the j-th second semantic substructure.

[0117] In some optional implementations, the semantic substructure similarity determination unit may further be specifically configured to:

[0118] For any of the first semantic substructures, determining a word frequency-inverse document frequency vector based on a topic relevance factor corresponding to the first semantic substructure;

[0119] For any of the second semantic substructures, determining a term frequency-inverse document frequency vector based on a topic relevance factor corresponding to the second semantic substructure;

[0120] The semantic substructure similarity between the first semantic substructure and the second semantic substructure is determined based on the word frequency-inverse document frequency vector based on the topic relevance factor corresponding to the first semantic substructure and the word frequency-inverse document frequency vector based on the topic relevance factor corresponding to the second semantic substructure.

[0121] In some optional implementations, the formula for determining the term frequency-inverse document frequency vector based on the topic relevance factor is:

[0122]

[0123] Where t represents the word in the semantic substructure, d represents the document where the semantic substructure is located, TF(t,d) represents the word frequency of word t in document d, IDF(t) represents the inverse document frequency of word t, and β represents the topic relevance adjustment parameter. represents the vector representation of word t, represents the topic vector of document d, express and The cosine similarity of .

[0124] In some optional implementations, the module 530 for determining the duplication analysis results of the R&D project documents to be established may also be specifically used to:

[0125] For any R&D project document that has been approved, if the repetition rate between the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved is greater than the preset repetition rate threshold, the repetition analysis result of the R&D project document to be approved is determined to be a duplicate R&D project.

[0126] In some optional implementations, the R&D project document vector acquisition module 510 may also be specifically configured to:

[0127] Obtain documents for R&D projects to be established;

[0128] Performing optical character recognition on the research and development project document to be established to obtain text information of the research and development project document to be established;

[0129] Cleaning the text information of the research and development project document to be established to obtain the cleaned text information of the research and development project document to be established;

[0130] Performing named entity recognition on the cleaned text information of the R&D project document to be established to obtain technical features of the R&D project document to be established, where the technical features of the R&D project document to be established include technical terms, technical parameters, and technical operation procedures;

[0131] Constructing a technical feature knowledge graph corresponding to the R&D project document based on the technical features of the R&D project document to be established, wherein the technical feature knowledge graph corresponding to the R&D project document to be established includes multiple technical feature vectors and relationships between the technical feature vectors;

[0132] Based on the technical feature knowledge graph corresponding to the research and development project document to be established, the word vector of the research and development project document to be established is updated to obtain the vector of the research and development project document to be established.

[0133] In some optional implementations, the formula for updating the word vector is:

[0134]

[0135] Among them, w i Represents the i-th word vector in the R&D project document to be established, represents w i The vector representation before updating is, represents w iThe vector before update is represented, α represents the learning rate, and N(i) represents w i The set of adjacent nodes in the technical feature knowledge graph, represents w i In the technical feature knowledge graph j The associated weight of .

[0136] In some optional embodiments, the R&D project repeated analysis device further includes:

[0137] The cluster analysis module is used to perform density and hierarchical fusion cluster analysis on the vectors of the R&D project documents to be established and the vectors of each R&D project document that has been established, so as to obtain a combination of repeated R&D projects.

[0138] The R&D project duplication analysis device provided in the embodiment of the present invention can execute the R&D project duplication analysis method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0139] Example 6

[0140] Figure 6 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0141] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 and a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An I / O interface 15 is also connected to the bus 14.

[0142] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0143] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the R&D project duplication analysis method, which includes:

[0144] Obtaining a vector of a document of a pending R&D project and a vector of a document of an approved R&D project, wherein the vector of the document of the pending R&D project includes a plurality of first semantic substructures, and the vector of the document of the approved R&D project includes a plurality of second semantic substructures;

[0145] For any approved R&D project document, based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the approved R&D project document, determine the repetition rate of the vector of the R&D project document to be approved and the vector of the approved R&D project document;

[0146] The repetition analysis result of the research and development project document to be established is determined based on the repetition rate of the vector of the research and development project document to be established and the vector of each established research and development project document.

[0147] In some embodiments, the R&D project duplication analysis method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the R&D project duplication analysis method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the R&D project duplication analysis method in any other appropriate manner (e.g., via firmware).

[0148] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0149] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0150] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0151] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0152] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0153] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0154] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0155] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for repeated analysis of R&D projects, characterized in that: include: Obtaining a vector of a document of a pending R&D project and a vector of a document of an approved R&D project, wherein the vector of the document of the pending R&D project includes a plurality of first semantic substructures, and the vector of the document of the approved R&D project includes a plurality of second semantic substructures; For any approved R&D project document, based on the multiple first semantic substructures corresponding to the R&D project document to be approved and the multiple second semantic substructures corresponding to the approved R&D project document, determine the repetition rate of the vector of the R&D project document to be approved and the vector of the approved R&D project document; The repetition analysis result of the research and development project document to be established is determined based on the repetition rate of the vector of the research and development project document to be established and the vector of each established research and development project document.

2. The method according to claim 1, characterized in that The determining, based on the multiple first semantic substructures corresponding to the R&D project document to be established and the multiple second semantic substructures corresponding to the R&D project document to be established, of the repetition rate of the vector of the R&D project document to be established and the vector of the R&D project document to be established, includes: Determining semantic substructure similarities between the first semantic substructures and the second semantic substructures based on the plurality of first semantic substructures corresponding to the R&D project document to be established and the plurality of second semantic substructures corresponding to the R&D project document that has been established; Based on the semantic substructure similarity between the first semantic substructure and the second semantic substructure, a repetition rate between the vector of the to-be-established R&D project document and the vector of the established R&D project document is determined.

3. The method according to claim 2, characterized in that The formula for determining the repetition rate between the vector of the R&D project document to be approved and the vector of the approved R&D project document is: Among them, A represents the vector of the R&D project document to be approved, B represents the vector of the R&D project document that has been approved, Similarity(A,B) represents the repetition of the vector of the R&D project document to be approved and the vector of the R&D project document that has been approved, a i The first semantic substructure of the i-th vector in the R&D project document to be established, b j represents the jth second semantic substructure in the vector of the R&D project document that has been approved, m represents the number of first semantic substructures, and n represents the number of second semantic substructures; S(a i ,b j ) represents the semantic substructure similarity between the i-th first semantic substructure and the j-th second semantic substructure.

4. The method according to claim 2, characterized in that The determining, based on the plurality of first semantic substructures corresponding to the R&D project document to be established and the plurality of second semantic substructures corresponding to the R&D project document that has been established, the semantic substructure similarity between the first semantic substructure and the second semantic substructure includes: For any of the first semantic substructures, determining a word frequency-inverse document frequency vector based on a topic relevance factor corresponding to the first semantic substructure; For any of the second semantic substructures, determining a term frequency-inverse document frequency vector based on a topic relevance factor corresponding to the second semantic substructure; The semantic substructure similarity between the first semantic substructure and the second semantic substructure is determined based on the word frequency-inverse document frequency vector based on the topic relevance factor corresponding to the first semantic substructure and the word frequency-inverse document frequency vector based on the topic relevance factor corresponding to the second semantic substructure.

5. The method according to claim 4, characterized in that The formula for determining the term frequency-inverse document frequency vector based on the topic relevance factor is: Where t represents the word in the semantic substructure, d represents the document where the semantic substructure is located, TF(t,d) represents the word frequency of word t in document d, IDF(t) represents the inverse document frequency of word t, and β represents the topic relevance adjustment parameter. represents the vector representation of word t, represents the topic vector of document d, express and The cosine similarity of .

6. The method according to claim 1, wherein The vector for obtaining the R&D project documents to be established includes: Obtain documents for R&D projects to be established; Performing optical character recognition on the research and development project document to be established to obtain text information of the research and development project document to be established; Cleaning the text information of the research and development project document to be established to obtain the cleaned text information of the research and development project document to be established; Performing named entity recognition on the cleaned text information of the R&D project document to be established to obtain technical features of the R&D project document to be established, where the technical features of the R&D project document to be established include technical terms, technical parameters, and technical operation procedures; Constructing a technical feature knowledge graph corresponding to the R&D project document based on the technical features of the R&D project document to be established, wherein the technical feature knowledge graph corresponding to the R&D project document to be established includes multiple technical feature vectors and relationships between the technical feature vectors; Based on the technical feature knowledge graph corresponding to the research and development project document to be established, the word vector of the research and development project document to be established is updated to obtain the vector of the research and development project document to be established.

7. The method according to claim 6, characterized in that The formula for updating the word vector is: Among them, w i Represents the i-th word vector in the R&D project document to be established, represents w i The vector representation before updating is, represents w i The vector before update is represented, α represents the learning rate, and N(i) represents w i The set of adjacent nodes in the technical feature knowledge graph, represents w i In the technical feature knowledge graph j The associated weight of .

8. A device for repeated analysis of R&D projects, characterized in that: include: A vector acquisition module for R&D project documents, configured to acquire vectors of R&D project documents to be approved and vectors of approved R&D project documents, wherein the vectors of the R&D project documents to be approved include multiple first semantic substructures, and the vectors of the approved R&D project documents include multiple second semantic substructures; A module for determining the repetition rate of vectors of R&D project documents, for determining, for any approved R&D project document, the repetition rate of the vectors of the R&D project document to be approved and the vectors of the approved R&D project document based on a plurality of first semantic substructures corresponding to the R&D project document to be approved and a plurality of second semantic substructures corresponding to the R&D project document to be approved; The module for determining the repeated analysis results of the R&D project documents to be established is used to determine the repeated analysis results of the R&D project documents to be established based on the repetition rates of the vectors of the R&D project documents to be established and the vectors of each established R&D project document.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the R&D project repeated analysis method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the R&D project duplication analysis method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Intelligent analysis device for repeated declaration of scientific research projects

    CN112214986A

  • Knowledge graph complex problem generation method and system based on graph neural network

    CN113722510A

  • Semantic recognition method and device, equipment, medium and program product

    CN117975946A

  • Method for mining threat figure in chat group and related equipment

    CN118626641A

  • File duplicate checking method and device, equipment, storage medium and program product

    CN118733717A