A patent recommendation method based on multi-domain network division
By building a multi-field patent network and conducting community testing, the problem of designers being difficult for them to obtain accurate patent information from massive patent data is solved, efficient patent recommendation and knowledge space expansion is achieved, and design efficiency and quality are improved.
Patent Information
- Application Number
- CN202310690773.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-06-12
AI Technical Summary
When designers face massive patent data, it is difficult for designers to efficiently obtain accurate patent functional and structural information to assist design innovation.
Using a patent recommendation method based on multi-domain network division, a target functional patent set is constructed and a patent network is built through semantic dictionary and IPC classification code, community detection and division of fields is carried out, and cross-domain patent recommendation is finally carried out based on the core degree of the field and the correlation degree.
It effectively reduces designers' time and energy investment in patent knowledge retrieval, improves designers' efficiency in obtaining design-related knowledge, and improves concept design quality and design efficiency.
Smart Images

Figure CN116881394B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the information field based on semantic network and patent knowledge recommendation, relates to a multi-domain division method of a network, and specifically relates to a patent recommendation method based on multi-domain network division. Background Art
[0002] Conceptual design is a process in which innovative design thinking and design knowledge interact to solve design problems. In this process, designers need to comprehensively use external knowledge and their own experience to find possible conceptual solutions. This not only places high demands on the designer's innovative ability, but also requires the designer to have sufficient knowledge reserves. As an integrated product of innovative knowledge, patent texts have the latest and most innovative knowledge and are an important knowledge resource to assist designers in expanding their design knowledge space[1]. However, manual search and reading of patents by designers brings a great cognitive burden to designers and requires a lot of time and energy. Based on this, patent analysis-related technologies have emerged, and through information technologies such as artificial intelligence, big data, and machine learning, they can efficiently search patents and extract knowledge from them[2]. Although patent analysis technology has reduced the burden of manual knowledge search by designers, in the face of massive patent data, designers still need to rely on their own experience to obtain patent content for specific design problems. How to ensure that designers can more efficiently obtain accurate patent texts to assist design innovation is still an urgent problem to be solved.
[0003] As a scientific and technological document, patent text has a clear and rigorous writing format, usually including the title page, claims, specification, and appendix to the specification. Figure 4 The above four parts elaborate on the detailed technical solutions to solve the technical problems. Such technical solutions are usually reflected in the structural composition of products that realize certain functions, without describing the specific parameters of the structure and the processing technology information. This feature of patents enables them to provide reference design knowledge for conceptual design. The functional knowledge and structural knowledge contained therein reflect the innovative technical features of the patents and can be used to assist designers in product concept design. However, how to obtain accurate patent function and structural information in the massive patent texts still requires designers to spend time and energy to search, read and understand the design knowledge related to the design concept.
[0004] References
[0005] [1] Cao Shujin, Li Ruijing. Construction and application of innovation knowledge graph based on patent document abstracts[J]. Information Theory and Practice, 2022, 45(11): 21-28.
[0006] [2] Liu Chunjiang, Zhu Jiang. Research on the architecture of patent big data service platform for intelligence analysis[J]. Library Work and Research, 2022(04):57-64. Summary of the invention
[0007] In order to assist designers in expanding their knowledge space and solve the problems existing in existing patent knowledge recommendation technologies, the present invention aims to provide a patent recommendation method based on multi-domain network division. The main technical solutions are as follows:
[0008] A patent recommendation method based on multi-domain network division includes the following steps:
[0009] The first step is to obtain the target function patent set: conduct product concept design under the determined design requirements, decompose the design functions according to the design problems, and obtain the design target function concept; obtain the main functional concepts of the patent text set based on the design target function concept; obtain the target function patents based on the similarity of the semantic dictionary and construct the target function patent set;
[0010] The second step is to build a multi-field patent network based on the target functional patent set. The method is as follows:
[0011] (1) Select patent features from a technical perspective based on the IPC classification code of the patent;
[0012] (2) Selecting patent features from a semantic perspective: Semantic features are used to reflect the main semantic information of the patent text content and extract the multi-dimensional patent text vector of each patent text;
[0013] (3) By calculating the weighted correlation between the semantic-technical features of each pair of patents in the target function patent set, the patent network of the target function patent set is constructed as follows:
[0014] The degree of association between patents in terms of semantic features is measured by semantic similarity, which is obtained by calculating the cosine similarity between patent text vectors. The greater the semantic similarity between two patent texts, the greater the degree of association in terms of semantic features.
[0015] The degree of association between patents in terms of technical features can be measured by the co-occurrence ratio of IPC classification codes: the same IPC classification code represents the same technical subject. The more IPC classification codes are the same between two patents, the greater the overlap between the two in terms of technical subject matter. The degree of association between patents in terms of technical features can be measured by the co-occurrence ratio of IPC classification codes.
[0016] According to the calculated correlation degree of the patent in terms of technical features and semantic features, weights are assigned to the correlation degrees from the two perspectives to calculate the semantic-technical weighted correlation degree;
[0017] By calculating the semantic-technical weighted correlation between each pair of patents in the target functional patent set, a patent network of the target functional patent set is constructed. The patent network is a fully connected network. There are edges representing correlation between each pair of patent nodes. The higher the correlation degree of the semantic-technical weighted features on the edge, the better.
[0018] The third step is to divide the patent network domain based on community detection
[0019] Conduct community detection on the patent network and divide the constructed patent network into communities. Patents in the same community have a high degree of correlation, while patents in different communities have a low degree of correlation. Each community is regarded as a field, and each patent has its own field.
[0020] The fourth step is to recommend cross-domain patents based on domain coreness and relevance. The method is as follows:
[0021] Define each patent in the target function patent set as a target patent, and combine the design target function concept and the main function concept in each target patent in pairs to form a concept combination;
[0022] In the target patent, the more times the concept combination formed by the design target function concept and the main function concept co-occurs, and the shorter the context distance, the greater the correlation between the concept combination and the patent. The correlation of the concept combination in a single target patent is calculated;
[0023] In the same field of the patent network, if a patent has a high degree of correlation with other patents in the same field, then the patent can better represent the patent field in which it is located; calculate the coreness of the target patents with concept combinations in each field, and sort the patents in each field according to the coreness. The higher the coreness score, the higher the ranking. The patents with the highest coreness ranking in each field are selected as valuable patents and recommended to designers to broaden the knowledge space for designers to solve design problems;
[0024] The importance of target patents with concept combinations in multiple fields is calculated based on the correlation between the concept combinations in each target patent and the coreness of each target patent in the field. Patents in each field of the target patents are ranked according to their importance. The higher the importance score, the higher the ranking. Patents with high importance ranking in each field are selected as valuable patents and recommended to designers to broaden their knowledge space for solving design problems.
[0025] Furthermore, the first step specifically includes the following steps:
[0026] (1) Decompose the design function of the design problem and obtain the concept of the design target function
[0027] According to the description of the design problem, the hierarchical decomposition method is used to decompose the total functions required by the design, obtain the main functions and auxiliary functions for solving the design problem, and extract the main functions as the target functional concepts in the conceptual design process;
[0028] (2) Obtaining the main functional concepts of the patent text set based on the design target functional concepts
[0029] Extraction of main functional concepts: Decomposing the Subject-Action-Object structure in the patent text through dependency analysis, extracting the Actions from it and using them as the main functional concepts of the patent text;
[0030] (3) Obtain target function patents based on the similarity of semantic dictionaries and construct a target function patent set
[0031] After determining the main functional concept of each patent text, the target functional patent set is obtained according to the similarity between the main functional concept of the patent and the target functional concept. The similarity between the main functional concept and the target functional concept adopts the semantic similarity based on the wordnet path distance. The method is as follows:
[0032] For each patent in the patent text set, in WordNet, nouns, verbs, adjectives, and adverbs are grouped into synonym sets. Each synonym set contains all words with the same meaning. A hierarchical structure is used to express the relationship between synonym sets. In the hierarchical structure, the deeper the level, the more refined the semantics represented by the synonym set. That is, under the same shortest path length, two synonym sets at a deep level are more similar than two synonym sets at a shallow level. In addition to considering the shortest path length, the hierarchical depth of the synonym set is also considered.
[0033] The WordNet semantic similarity between the main functional concept of each patent text in the patent text set and the target functional concept is calculated in turn. If a patent has multiple main functional concepts, the semantic similarity between each main functional concept and the target functional concept is calculated separately and the maximum value is taken. The meanings of the main functional concepts and target functional concepts with different semantic similarities are analyzed, and the semantic similarity threshold T is set. Patents with main functional concepts with semantic similarity higher than T are added to the target functional patent set.
[0034] Furthermore, the semantic similarity threshold T is set to 0.7.
[0035] Furthermore, according to the IPC classification code of the patent, the method for selecting patent features from a technical perspective is as follows: according to the IPC classification code of the patent, regular expressions are used for extraction, and the technical features of the patent are represented in the form of an IPC classification code set.
[0036] Furthermore, the method for extracting a multidimensional patent text vector for each patent text is as follows: input the patent's keywords into the Skip-gram model of Word2Vec to generate a multidimensional word vector for the keywords, and weighted sum all keyword vectors based on TFIDF keyword weights to extract a multidimensional patent text vector for each patent text.
[0037] Furthermore, the method of the third step is as follows:
[0038] (1) Each network node in the patent network of the entire field is regarded as a community, and the number of communities is the same as the number of nodes;
[0039] (2) For each network node, the modularity increment after joining the community of its neighboring nodes is calculated in turn, and the community with the largest modularity increment after joining is selected; if there is no community with an increased modularity increment after joining, no community is added; repeat the above process until the communities to which all network nodes belong no longer change;
[0040] (3) Compress the patent network in the entire field, compress a community into a new node, convert the edge weights between nodes within the community into the weights of the ring of the new node, and convert the sum of the edge weights between communities into the edge weights between new nodes to obtain the compressed network;
[0041] (4) Repeat steps (1) to (3) until the post-community modularity of the full-field patent network no longer changes, and obtain multiple communities, each of which represents a patent field. A full-field patent network is divided into a multi-field patent network consisting of multiple communities.
[0042] Furthermore, the calculation of the correlation of the concept combination in a single patent is shown in Formula 1. The higher the correlation P_cor_score of the concept combination in a target patent, the more relevant the target patent is to the concept combination;
[0043]
[0044] Where P, w1, w2 are the two concepts in the target patent and concept combination respectively, n is the total number of sentences where w1 and w2 co-occur in the target patent, min_context_dist(sent i ,w1,w2) is the minimum context distance between w1 and w2 in the i-th sentence, len(sent i ) is the length of the i-th sentence.
[0045] Furthermore, the core score of the target patent in the field is calculated as shown in Formula 2; the higher the core score P_core_score of a target patent in the field, the more representative the target patent is of the field;
[0046]
[0047] Where P and c are the target patent and field respectively, Cor(P,P i ) is the target patent P and the patent P in the same field i The correlation degree value on the connecting edge between them, |c| is the total number of patents in field c.
[0048] Furthermore, the calculation method for the importance of target patents with concept combinations in multiple fields is as follows:
[0049] Considering that the relevance and coreness have different dimensions, the maximum-minimum standardization is used for standardization, and then the average value of the two is calculated to obtain the importance of the target patent in the field. The calculation method is shown in Formulas 3, 4, and 5:
[0050]
[0051]
[0052]
[0053] Where P_cor_score min and P_cor_score max P_core_score is the minimum and maximum value of the concept combination correlation of the target patent with the concept combination in the field. min and P_core_score max They are the minimum and maximum coreness of the target patents with concept combinations in the field. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 Building the overall framework of the method
[0055] Figure 2 WordNet Hierarchy
[0056] Figure 3 Patent semantic feature extraction process
[0057] Figure 4 Louvain algorithm community detection process
[0058] Figure 5 Patent network multi-field division DETAILED DESCRIPTION
[0059] In order to make the technical solution of the present invention clearer, the present invention is further described below in conjunction with the accompanying drawings.
[0060] This patent aims to analyze the characteristics of patent texts, build a patent network divided into multiple fields by combining the key semantic information of patent texts, recommend patent texts related to design concepts to designers based on specific design problem requirements, and extract the conceptual product structure knowledge to support designers in generating design solutions, accurately expand the knowledge space of designers, effectively reduce the difficulty of designers in searching for patent knowledge, and ultimately achieve the purpose of improving the quality of conceptual design and design efficiency. The present invention is specifically implemented in the following steps:
[0061] 1. Acquisition of target functional patent set
[0062] Carry out product concept design under the condition of certain design requirements, decompose design functions according to design problems, obtain design target function concepts, and obtain target function patents based on the similarity of the design target function concepts in the main functional concepts of the massive patent text set based on the semantic dictionary, and build the target function patent set. Specifically, it includes the following steps:
[0063] (1) Acquisition of target functional concepts based on functional analysis of design problems
[0064] According to the description of the design problem, the hierarchical decomposition method is used to decompose the total functions required for the design, obtain the main functions and auxiliary functions for solving the design problem, and extract the main functions as the target functional concepts in the conceptual design process.
[0065] (2) Obtaining the main functional concepts of the patent text. The expression of functions is usually in the form of "functional verb + functional noun", which respectively expresses the behavior and the object of the behavior that completes the function, such as "recycle waste". However, in order to ensure that as many functional concepts from different fields as possible are included, this patent only uses functional verbs as the main functional concepts of the patent text to limit the patent set. The extraction of the main functional concepts is done by decomposing the SAO (Subject-Action-Object) structure in the patent text through dependency analysis, and extracting the Action from it and using it as the main functional concept of the patent text.
[0066] (3) Acquisition of target functional patents based on semantic similarity
[0067] After determining the main functional concept of each patent, the target functional patent set is obtained based on the similarity between the main functional concept of the patent and the target functional concept. The similarity between the main functional concept and the target functional concept adopts the semantic similarity based on the wordnet path distance. The WordNet path distance calculates the semantic similarity distance between the main functional concept and the target functional concept. For each patent in the patent text set, in WordNet, nouns, verbs, adjectives, and adverbs are grouped into synonym sets (Synsets), and each synonym set contains all words with the same meaning. In WordNet, the following is used: Figure 2 The hierarchical structure shown is used to express the relationship between synsets. Based on the hierarchical structure of WordNet, the semantic similarity is calculated using Formula 1.
[0068]
[0069] Where w1 and w2 are two concept words, S1 and S2 are the synonym sets of w1 and w2, respectively, and path_dist(S1, S2) is the shortest path length between the synonym sets S1 and S2 in the hierarchy.
[0070] In the hierarchical structure, the deeper the level, the more refined the semantics represented by the synonym set. That is, under the same shortest path length, two synonym sets at a deep level are more similar than two synonym sets at a shallow level. Therefore, in addition to considering the shortest path length, the hierarchical depth of the synonym set is also considered.
[0071] Formula 1 is used to calculate the WordNet semantic similarity between the main functional concept and the target functional concept of each patent text in the patent dataset. If a patent has multiple main functional concepts, the semantic similarity between each main functional concept and the target functional concept is calculated and the maximum value is taken. wn The meaning of the functional concept and the target functional concept under Sim wn Threshold, patents with main functional concepts below this threshold are added to the target functional patent set.
[0072] 2. Construction of multi-field patent network
[0073] The patent network reflects the correlation between massive patents and can divide massive patents into different fields according to the degree of correlation. Selecting appropriate patent features to express patent information is a prerequisite for establishing correlation between patents. Patent features used in previous studies, such as keywords, literature citations and citations, and the region to which the right holder belongs, mostly describe patent information from a certain perspective, and the correlation established in this way also has limitations. In order to make the constructed patent network reflect not only the degree of correlation of patents in text content, but also the degree of correlation in technical fields, this patent selects patent features from two perspectives: semantics and technology. The specific method is as follows:
[0074] (1) Selection of patent features from a technical perspective
[0075] Technical features are used to reflect the technology used in the patent and the technical field to which the patent belongs, and are represented by the patent's IPC classification code. IPC classification is an internationally common patent classification method that classifies patents based on technical themes through five classification levels. Each patent has one or more IPC classification codes to reflect one or more technical themes of the patent. Patents with the same IPC classification code have the same technical themes. IPC classification codes are generated according to coding rules and have a formatted expression. Therefore, regular expressions are used for extraction, and the technical features of the patent are represented in the form of an IPC classification code set T = {t1, t2, ..., t n}, where each element t is an IPC classification code of the patent.
[0076] (2) Patent feature selection from a semantic perspective
[0077] Semantic features are used to reflect the main semantic information of patent text content. Based on the keywords extracted by TF-IDF method and the trained Word2Vec Skip-gram model, patent text vectors are used as the semantic features of patents. The extraction process of patent semantic features is as follows: Figure 3 As shown, the keywords of the patent are input into the Skip-gram model of Word2Vec to generate a 300-dimensional word vector for the keywords, and then all keyword vectors are weighted and summed using the keyword weight based on TFIDF to obtain a 300-dimensional patent text vector.
[0078] The patent text vector has the same dimension as the word vector generated by the Skip-gram model, but the difference is that the patent text vector comprehensively considers the importance of the patent's keywords and the semantic information of all keywords. The calculation method of the patent text vector S is shown in formula (3-6).
[0079]
[0080]
[0081] Where W i is the i-th keyword vector of the patent, ω i is the weight of the ith keyword, and n is the number of all keywords in the patent.
[0082] This patent uses the Skip-gram model of Word2Vec to generate 300-dimensional word vectors, and the similarity of vocabulary is calculated using the cosine similarity shown in Formula 4.
[0083]
[0084] In the formula, w1 and w2 are two words, and their word vectors are A(a1,a2,a3,…,a 300 ) and B(b1,b2,b3,…,b 300 ).
[0085] The technical features of a patent are expressed in the form of an IPC classification code set, while the semantic features of a patent are expressed in the form of a 300-dimensional vector. The two cannot use the same correlation calculation method. The correlation between the semantic features of a patent can be measured by semantic similarity. The greater the semantic similarity between two patent texts, the greater the correlation between the semantic features. Therefore, the correlation between the semantic features can be calculated by the cosine similarity between the patent text vectors using Formula 1. The same IPC classification code represents the same technical subject. The more IPC classification codes are the same between two patents, the greater the overlap between the two in terms of technical subject matter. Therefore, the correlation between the technical features of a patent can be measured by the co-occurrence ratio of the IPC classification codes. The calculation method is shown in Formula 5.
[0086]
[0087] In the formula, P1 and P2 are two patents, T1 is the IPC classification code set of patent P1, and T2 is the IPC classification code set of patent P2.
[0088] According to Formula 4 and Formula 5, the correlation degree of patents in technical features Cor T Correlation degree between semantic features S , and the two have the same dimension, giving the same importance to the correlation degree of the two angles, and calculating the semantic-technical weighted correlation degree, as shown in Formula 6.
[0089] Cor(P1,P2)=w T ×Cor T (P1,P2)+w S ×Cor S (P1,P2)#(6)
[0090] In the formula, w T is the weight of the degree of association in technical features, w S Cor is the correlation degree of semantic features S .
[0091] According to the above formula, the semantic-technical weighted correlation between each pair of patents in the target functional patent set is calculated to construct the patent network of the target functional patent set. The patent network is a fully connected network, and there are edges representing the correlation between each pair of patent nodes. The higher the semantic-technical weighted correlation on the edge, the more related the two patents are, and the greater the possibility that they are in the same field.
[0092] 3. Patent network domain division based on community detection. In order to recommend patents in different fields to designers and expand the designer's knowledge space, it is necessary to divide the constructed patent network into several fields. Patents in the patent network have a higher degree of correlation, while patents in different fields have a lower degree of correlation. This patent uses the Louvain algorithm to perform community detection on the patent network. The Louvain algorithm is a community detection algorithm based on modularity. It can hierarchically divide the network by maximizing the modularity of the community network. Modularity is an indicator used to measure the strength of the network community structure. Considering that the patent network is a weighted fully connected network, the modularity used also takes into account the weight of the edge. Its calculation method is shown in Formula 7.
[0093]
[0094] In the formula, a ij is the weight of the edge between node i and node j; c i and c j are the communities where nodes i and j are located respectively, if c i =c j ,δ(c i ,c j )=1, otherwise δ(c i ,c j )=0;k i is the sum of the weights of the edges connected to node i, and m is the sum of the weights of all edges in the entire network.
[0095] The community detection process of Louvain algorithm is as follows Figure 4 As shown, the specific process is as follows:
[0096] 1) Treat each node in the network as a community, and the number of communities is the same as the number of nodes;
[0097] 2) For each node, calculate the modularity increment ΔQ after joining the communities of its neighboring nodes in turn. The calculation method is shown in Formula 8, and select the community with the largest ΔQ to join. If there is no community with an increased ΔQ, do not join any community. Repeat this process until the communities to which all nodes belong no longer change.
[0098] 3) Compress the network, compress a community into a new node, convert the edge weights between nodes within the community into the weights of the ring of the new node, and convert the sum of the edge weights between communities into the edge weights between new nodes to obtain the compressed network;
[0099] 4) Repeat steps 1), 2), and 3) until the modularity of the network no longer changes.
[0100]
[0101] In the formula, is the sum of the edge weights between nodes within community c, is the sum of the weights of all edges connected to the nodes in community c, is the sum of the weights of the edges between node i and the nodes in community c.
[0102] Through the above-mentioned Louvian algorithm community detection process, the patent network can be divided into several communities. Patents divided into the same community have a high degree of correlation. Each community can be regarded as a field, and each patent has its own field.
[0103] 4. Cross-domain patent recommendation based on domain coreness and relevance
[0104] According to the WordNet semantic similarity results, the target functional concept and the main functional concepts in the target functional patent set above the semantic similarity threshold are combined to form a concept combination. However, a single concept combination can only provide conceptual incentives for designers, but cannot expand the knowledge space used by designers to solve design problems. Designers still need to spend time and energy to search for design knowledge related to conceptual incentives. Therefore, it is necessary to recommend patents related to conceptual incentives to designers to help reduce the burden of designers acquiring relevant knowledge.
[0105] In the vast amount of patent documents, a concept combination usually appears in many patent texts in different fields, and these corresponding patent texts can be provided to designers as the expanded design knowledge of the concept combination. 。However, in the actual design process, designers still need to spend a lot of time to screen a small number of more targeted and valuable patents from a large number of related patents as the main reference. Therefore, in order to further reduce the burden of designers screening the main reference patents from related patents, relevant patents containing concept combinations are recommended from the perspectives of the relevance of concept combinations in patents and the coreness of patents in patent networks.
[0106] When a concept combination co-occurs more times in a patent and the context distance is shorter, the concept combination is more important in the patent, the patent is more relevant to the concept combination, and the value to the designer is greater. The calculation of the correlation of the concept combination in a single patent is shown in Formula 9. The higher the correlation P_cor_score, the more relevant the patent is to the concept combination.
[0107]
[0108] Where P, w1, w2 are the two concepts in the target patent and concept combination respectively, n is the total number of sentences where w1 and w2 co-occur in the target patent, min_context_dist(sent i ,w1,w2) is the minimum context distance between w1 and w2 in the i-th sentence, len(sent i ) is the length of the i-th sentence.
[0109] In the patent network, the correlation between patent texts in the same field is greater. If a patent has a high correlation with other patents in the same field, the patent can be regarded as the core patent in the field and can well represent the field. The core degree of a patent in a field is calculated as shown in Formula 10. The higher the core degree P_core_score of a patent in a field, the more representative the patent is of the field.
[0110]
[0111] Where P,c are the target patents and fields, Cor(P,P i ) is the target patent P and the patent P in the same field i The correlation degree value on the connecting edge between them, |c| is the total number of patents in field c.
[0112] The relevance of the concept combination in the patent and the coreness of the patent in the field are comprehensively considered, and the importance of the patents with concept combinations in multiple fields is calculated respectively. Considering that the relevance and coreness have different dimensions, the maximum-minimum normalization is first used for standardization, and then the average of the two is calculated to obtain the importance of the patent in the field. The calculation method is shown in Formulas 11, 12, and 13.
[0113]
[0114]
[0115]
[0116] Where P_cor_score min and P_cor_score max P_core_score is the minimum and maximum value of the concept combination correlation of patents with concept combinations in the field, min and P_core_score max They are the minimum and maximum coreness of patents with concept combinations in the field content.
[0117] The importance of patents with concept combinations in each field is calculated, and patents are sorted according to their importance in each field. The higher the importance score, the higher the ranking. The top-ranked patents in each field are selected as valuable patents and recommended to designers to broaden the knowledge space for designers to solve design problems.
[0118] 5. Example Application
[0119] This case solves a design problem of "crushing masonry waste in construction waste" and generates a corresponding design solution. The concept generation method based on patent knowledge cognition proposed in this patent is applied to the concept generation and solution generation process of this design problem.
[0120] Construction waste refers to the rocks excavated during the construction of roads and bridges, municipal construction and renovation, as well as the concrete, bricks and tiles, abandoned slag and waste materials, etc., with masonry waste being the main component. With the rapid development of my country's economy, large-scale projects such as urban transformation, expansion and renovation of roads and bridges, and collective relocation are continuously carried out, and with it comes an increasing amount of construction waste. Most of the construction waste is transported to rural areas for open-air storage or landfill without any treatment. Not only will the transportation process consume a lot of construction funds, causing environmental pollution problems such as dust, but it will also occupy arable land. Construction waste masonry waste is a useful substance. The bricks, stones, concrete, and waste residues can be crushed to replace sand, used for mortar, masonry mortar, etc., and can also be used to make building materials. Therefore, crushing and reusing masonry waste can not only save construction costs, but also reduce environmental pollution caused by garbage transportation and storage.
[0121] (1) Data description
[0122] The data used in this example are some patent texts published in the Derwent patent database in 2020. A total of 104,214 patents were obtained, covering 7 Derwent classification codes, covering the period from January 1, 2020 to December 31, 2020, including patent information such as invention name, application number, abstract, inventor, IPC classification code, claims, instructions, citations, etc. The invention name, application number, abstract, IPC classification code, claims, and instructions of these 104,214 patents are saved as fields to form a patent database used by this patent later.
[0123] (2) Construction of patent network in all fields and division into multiple fields
[0124] First, semantic features and technical features are obtained from patents. The IPC classification code set of each patent is extracted from the IPC classification code text of patent data using regular expressions. At the same time, the keywords extracted from each patent are input into Word2Vec to obtain the corresponding word vectors. The patent text vectors are calculated based on the word vectors and the TF-IDF values of each keyword. Then, the cosine similarity method is used to calculate the similarity of patent text vectors between two patents as the degree of association of patents in semantic features. The co-occurrence ratio of IPC classification codes is used to calculate the degree of association of patents in technical features. The weights of the two are 0.5 respectively. The weighted sum is calculated to obtain the semantic-technical weighted association degree as the weight of the edge between patent nodes in the patent network, thereby obtaining a fully connected patent network.
[0125] The patent network is divided into fields by community detection. After the Louvain algorithm is used for community detection, the patent network is divided into five communities, that is, five fields. The number of patents in each field is 705, 682, 893, 634, and 261 respectively. The patent network divided into fields is as follows Figure 5 As shown, since the patent network is a fully connected network with a large number of edges, only the edges with larger weights are shown in the figure for clear graphical expression. It can be seen from the figure that there are more large-weight edges within the same field, while there are fewer large-weight edges between fields. This also reflects that patents in the same field are more technically and semantically related.
[0126] (3) Cross-field patent recommendation
[0127] After the design concept is generated, the designer is provided with patents related to the design concept. Patents from five different fields are recommended for the above three groups of design concepts. The importance of each patent in each field is obtained by calculating the sum of the correlation of the concept combination corresponding to each group of design concepts in the patents in each field and the core degree of the patents in each field, and standardizing the correlation and core degree and taking the weighted sum. The higher the importance, the more relevant the corresponding patent is to the design concept, and the more core it is in the field. The calculation results are shown in Table 1, which only lists the top three patents in each field of importance for each group of design concepts, a total of 45 patents.
[0128] Table 1 Recommended patents corresponding to design concepts
[0129]
[0130]
[0131] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A patent recommendation method based on multi-domain network division, comprising the following steps: The first step is to obtain the target function patent set: conduct product concept design under the determined design requirements, decompose the design functions according to the design problems, and obtain the design target function concept; obtain the main functional concepts of the patent text set based on the design target function concept; obtain the target function patent based on the similarity of the semantic dictionary, and construct the target function patent set; The second step is to build a multi-field patent network based on the target functional patent set. The method is as follows: (1) Select patent features from a technical perspective based on the IPC classification code of the patent; (2) Selecting patent features from a semantic perspective: Semantic features are used to reflect the main semantic information of patent text content and extract multi-dimensional patent text vectors for each patent text; (3) By calculating the weighted correlation between the semantic-technical features of each pair of patents in the target function patent set, the patent network of the target function patent set is constructed as follows: The correlation degree of patents in semantic features Cor S The semantic similarity is measured by calculating the cosine similarity between patent text vectors. The greater the semantic similarity between two patent texts, the greater the degree of association in semantic features. The degree of correlation between patents in terms of technical features T Measured by the co-occurrence ratio of IPC classification codes: The same IPC classification code represents the same technical subject. The more IPC classification codes are the same between two patents, the greater the overlap between the two in terms of technical subject matter. The degree of correlation between patents in terms of technical features is measured by the co-occurrence ratio of IPC classification codes. The calculation method is as follows: In the formula, P1 and P2 are two patents, T1 is the IPC classification code set of patent P1, and T2 is the IPC classification code set of patent P2; According to the calculation of the correlation degree of patents in technical features Cor T Correlation between semantic features and S , assign weights to the correlation degrees of the two perspectives respectively, and calculate the semantic-technical weighted correlation degree; By calculating the semantic-technical weighted correlation between each pair of patents in the target functional patent set, a patent network of the target functional patent set is constructed. The patent network is a fully connected network, and there are edges representing correlations between each pair of patent nodes. The higher the semantic-technical weighted correlation on the edge, the more related the two patents are, and the greater the possibility that they are in the same field. The third step is to divide the patent network domain based on community detection Conduct community detection on the patent network and divide the constructed patent network into communities. Patents in the same community have a high degree of correlation, while patents in different communities have a low degree of correlation. Each community is regarded as a field, and each patent has its own field. The fourth step is to recommend cross-domain patents based on domain coreness and relevance. The method is as follows: Define each patent in the target function patent set as a target patent, and combine the design target function concept and the main function concept in each target patent in pairs to form a concept combination; In the target patent, the more times the concept combination formed by the design target function concept and the main function concept co-occurs, and the shorter the context distance, the greater the correlation between the concept combination and the patent. The correlation of the concept combination in a single target patent is calculated; In the same field of the patent network, if a patent has a high degree of correlation with other patents in the same field, then the patent can better represent the patent field in which it is located; calculate the coreness of the target patents with concept combinations in each field, and sort the patents in each field according to the coreness. The higher the coreness score, the higher the ranking. The patents with the highest coreness ranking in each field are selected as valuable patents and recommended to designers to broaden the knowledge space for designers to solve design problems; The importance of target patents with concept combinations in multiple fields is calculated based on the correlation between the concept combinations in each target patent and the coreness of each target patent in the field. Patents in each field of the target patents are ranked according to their importance. The higher the importance score, the higher the ranking. Patents with high importance ranking in each field are selected as valuable patents and recommended to designers to broaden their knowledge space for solving design problems.
2. The patent recommendation method based on multi-domain network division according to claim 1 is characterized in that: The first step specifically includes the following steps: (1) Decompose the design function of the design problem and obtain the concept of the design target function According to the description of the design problem, the hierarchical decomposition method is used to decompose the total functions required by the design, obtain the main functions and auxiliary functions for solving the design problem, and extract the main functions as the target functional concepts in the conceptual design process; (2) Obtaining the main functional concepts of the patent text set based on the design target functional concepts Extraction of main functional concepts: Decomposing the Subject-Action-Object structure in the patent text through dependency analysis, extracting the Actions from it and using them as the main functional concepts of the patent text; (3) Obtain target function patents based on the similarity of semantic dictionaries and construct a target function patent set After determining the main functional concept of each patent text, the target functional patent set is obtained according to the similarity between the main functional concept of the patent and the target functional concept. The similarity between the main functional concept and the target functional concept adopts the semantic similarity based on the wordnet path distance. The method is as follows: For each patent in the patent text set, in WordNet, nouns, verbs, adjectives, and adverbs are grouped into synonym sets. Each synonym set contains all words with the same meaning. A hierarchical structure is used to express the relationship between synonym sets. In the hierarchical structure, the deeper the level, the more refined the semantics represented by the synonym set. That is, under the same shortest path length, the synonym sets at a deeper level are more similar than those at a shallower level. In addition to considering the shortest path length, the hierarchical depth of the synonym set is also considered. The WordNet semantic similarity between the main functional concept of each patent text in the patent text set and the target functional concept is calculated in turn. If a patent has multiple main functional concepts, the semantic similarity between each main functional concept and the target functional concept is calculated separately and the maximum value is taken. The meanings of the main functional concepts and target functional concepts with different semantic similarities are analyzed, and the semantic similarity threshold T is set. Patents with main functional concepts with semantic similarity higher than T are added to the target functional patent set.
3. The patent recommendation method based on multi-domain network division according to claim 2 is characterized in that: The semantic similarity threshold T is set to 0.
7.
4. The patent recommendation method based on multi-domain network division according to claim 1 is characterized in that: According to the IPC classification code of the patent, the method for selecting patent features from a technical perspective is as follows: according to the IPC classification code of the patent, regular expressions are used for extraction, and the technical features of the patent are expressed in the form of a set of IPC classification codes.
5. The patent recommendation method based on multi-domain network division according to claim 1 is characterized in that: The method for extracting the multidimensional patent text vector of each patent text is as follows: input the patent keywords into the Skip-gram model of Word2Vec to generate multidimensional word vectors of the keywords, and extract the multidimensional patent text vector of each patent text by weighting and summing all keyword vectors based on TFIDF keyword weights.
6. The patent recommendation method based on multi-domain network division according to claim 1 is characterized in that: The method for the third step is as follows: (1) Each network node in the patent network of the entire field is regarded as a community, and the number of communities is the same as the number of nodes; (2) For each network node, the modularity increment after joining the communities of its neighboring nodes is calculated in turn, and the community with the largest modularity increment is selected to join; if there is no community with an increased modularity increment, no community is added; the above process is repeated until the communities to which all network nodes belong no longer change; (3) Compress the patent network in the entire field, compress a community into a new node, convert the edge weights between nodes within the community into the weights of the ring of the new node, and convert the sum of the edge weights between communities into the edge weights between new nodes to obtain the compressed network; (4) Repeat steps (1) to (3) until the modularity of the full-field patent network no longer changes, and obtain multiple communities, each of which represents a patent field. A full-field patent network is divided into a multi-field patent network consisting of multiple communities.
7. The patent recommendation method based on multi-domain network division according to claim 1 is characterized in that: The calculation of the correlation of the concept combination in a single patent is shown in Formula 1. The higher the correlation P_cor_score of the concept combination in a target patent, the more relevant the target patent is to the concept combination. Where P is the target patent, w1 and w2 are the two concepts in the concept combination, n is the total number of sentences where w1 and w2 co-occur in the target patent, and min_context_dist(sent i ,w1,w2) is the minimum context distance between w1 and w2 in the i-th sentence, len(sent i ) is the length of the i-th sentence.
8. The patent recommendation method based on multi-domain network division according to claim 7 is characterized in that: The core score of the target patent in the field is calculated as shown in Formula 2. The higher the core score P_core_score of a target patent in the field, the more representative the target patent is of the field. Where P and c are the target patent and field respectively, Cor(P,P i ) is the target patent P and the patent P in the same field i The correlation degree value on the connecting edge between them, |c| is the total number of patents in field c.
9. The patent recommendation method based on multi-domain network division according to claim 8 is characterized in that: The calculation method for the importance of target patents with concept combinations in multiple fields is as follows: Considering that relevance and coreness have different dimensions, the maximum-minimum standardization is used for standardization, and then the average of the two is calculated to obtain the importance of the target patent in the field P_impt_score std (P,c), the calculation method is shown in equations 3, 4, and 5: Where P_cor_score min and P_cor_score max P_core_score is the minimum and maximum value of the concept combination correlation of the target patent with the concept combination in the field. min and P_core_score max They are the minimum and maximum coreness of the target patents with concept combinations in the field.
Citation Information
Patent Citations
Method and terminal for automatically generating ideas based on knowledge network
CN106940726A
Patent recommendation method and device based on semantic link heterogeneous information network embedding
CN110175224A