A method and system for identifying key common technical entities

By combining neural networks and cyclic dilated convolutional neural networks, a quintuple extraction model was constructed, which solved the problem that the LDA model did not consider semantic relationships in the identification of key common technologies, and achieved high-precision identification of technical entities and accurate screening of key common technologies.

CN120562415BActive Publication Date: 2025-09-26NANJING UNIV OF INFORMATION SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511056105.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-26
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

In the existing technology, the LDA topic model fails to effectively consider the semantic relationship between technical entities in the identification of key common technologies, resulting in insufficient recognition accuracy and reliability.

Method used

A neural network entity relationship extraction model is used to identify technical entities through the bertopic topic model. Combined with the recurrent dilated convolutional neural network and social network analysis, a quintuple extraction model is constructed to define the types and semantic relationships of technical entities. A data annotation platform is used to form a training corpus, and identification is carried out in combination with leading and key indicators.

Benefits of technology

It improves the accuracy and reliability of technical entity identification, can extract key common technologies in a fine-grained manner, and improves the precision and effectiveness of key common technology identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120562415B_ABST
    Figure CN120562415B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for identifying key common technical entities, belonging to the technical field of text data recognition; the method comprises: step S1: obtaining a technical text data set in a required field, and performing data cleaning on the technical text data set; step S2: defining entities and semantic relationships, and annotating text summary contents to form a corpus; step S3: using the corpus to train an entity relationship extraction model introduced into a neural network; step S4: performing commonality measurement through universality, efficiency, and relevance, and screening common technical entities; step S5: measuring the importance of technical entities with the help of social network analysis, and combining leading indicators to measure technical criticality, so as to accurately and efficiently identify key common technical entities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for identifying key common technical entities, and belongs to the technical field of text data processing. Background Art

[0002] The sustainable development of future industries is highly dependent on the breakthroughs and support of key common technologies. They are the cornerstone of industrial development and are of great significance for promoting the optimization and upgrading of industrial structure.

[0003] Currently, the identification of key common technologies primarily relies on LDA topic models for technology entity mining. This is achieved by constructing a multidimensional evaluation index system and calculating the criticality and commonality scores of technical entities to identify and screen key common technologies. However, traditional topic mining models have significant limitations: they only extract the probability of potential topics in documents and fail to consider the semantic relationships between technical entities, resulting in a low level of accuracy in technical entity identification. This limitation directly impacts the accuracy of subsequent calculations of criticality and commonality scores, ultimately making it difficult to ensure the reliability and effectiveness of key common technical entity identification. Summary of the Invention

[0004] Purpose of the invention: In response to the problems and shortcomings in the prior art, the present invention provides a method and system for identifying key common technical entities.

[0005] Technical solution: A method for identifying key common technical entities, including the following steps:

[0006] Step S1: Obtain a technical text dataset in the required field and perform data cleaning on the technical text dataset;

[0007] Step S2: define technical entity types and semantic relationships, and annotate the text summary content to form a corpus;

[0008] Step S3: Using the corpus to train the entity relationship extraction model introduced into the neural network;

[0009] Step S4: Measure commonality through versatility, efficiency, and relevance to screen common technical entities;

[0010] Step S5: Use social network analysis to measure the importance of technical entities, and combine leading indicators to measure technical criticality, so as to accurately and efficiently identify key common technical entities.

[0011] Preferably, in step S2, the technical entity types and semantic relationships are defined, and the text summary content (such as the abstract of the technical text) is annotated to form a corpus, including:

[0012] Step S21: Use the bertopic topic model to perform technical entity recognition on the text summary content, and define the technical entity type and semantic relationship based on the technical entity recognition results and in combination with relevant domain expertise.

[0013] Step S22: Based on the technical entity types and semantic relationships, the doccano data annotation platform is used to annotate the technical text with entity semantic relationships to form an entity relationship extraction model training corpus containing technical text, head / tail entities, head / tail entity types, and semantic relationships. The corpus is divided into a training set and a test set in a ratio of 7:3.

[0014] Preferably, the training of the entity relationship extraction model introduced into the neural network in step S3 includes:

[0015] Step S31: First, the technical text in the entity relationship extraction model training corpus is segmented by the BERT built-in word segmenter and converted into a string sequence X=(x1,x2,...,x n ), where x i For characters in technical text, such as text, i=1, 2, ..., n represents the character sequence number, and at the same time generates the vocabulary index input_ids, attention mask attention_mask and character position mapping offset_mapping; then the generated vocabulary index and attention mask are used to generate the context-related word vector H through the pre-trained language model BERT.

[0016] H=BERT(input_ids, attention_mask)∈

[0017] represents a set of real numbers; b represents the batch size; n represents the length of the string sequence of technical text; d represents the hidden layer dimension of BERT.

[0018] Secondly, traverse the character position mapping offset_mapping to map the character-level position of the head / tail entity in the technical text to the string sequence, and mark all strings within the head / tail entity span as the same head / tail entity type, generating the head / tail entity boundary index position and the real label y of the head / tail entity type.

[0019] Step S32: Build a shared semantic feature enhancement layer using a cyclic dilated convolutional neural network (RDC module). The context-sensitive word vector H output by BERT passes through the RDC module to capture long-range dependencies and hierarchical features. Because information in text summaries often has multi-level dependencies, important technical entities and relationship information may be located at the boundaries of the text. Therefore, by adjusting the dilation rate and padding size of the cyclic dilated convolutional neural network, convolutional layers with different dilation rates are used to extract multi-scale feature information:

[0020] H enhanced =RDC(H)∈

[0021] C(H)=Res(H)+GLU(Conv1D(H))

[0022] CP(H)=C k=2 0 (C k=2 1 (…C k=2 L-1 (H)…))

[0023] The RDC module gradually enhances the input features by stacking multiple dilated convolutional layers and residual connections to capture multi-scale information. At the same time, it introduces a gated linear unit to gate the convolution results and filter out unimportant information. enhanced Represents semantically enhanced word vectors; Res(H) represents residual connection; GLU represents gated linear unit; C(H) represents a single-layer symmetric dilated convolution operation; Conv1D represents a one-dimensional convolution layer; CP(H) represents a multi-layer symmetric dilated convolution operation; L represents the convolution parameter; k represents the dilation rate.

[0024] Step S33: Construct a head entity decoding layer to realize the two tasks of head entity recognition and head entity type prediction. The head entity decoding layer is used to enhance the semantic word vector H enhanced Decoding is performed by constructing two binary classifiers to predict the index positions of the start and end of the head entity, for each character x in the technical text X i Calculate the probability of (i=1, 2, ..., n) as the starting and ending positions, and filter out the starting index position and ending index position of the head entity according to the set threshold. Then, use the nearest matching principle to pair the identified starting and ending index positions to obtain a set of candidate head entities:

[0025] p sub_head =σ(W sub_head H enhanced +b sub_head )∈

[0026] p sub_tail =σ(W sub_tail H enhanced +b sub_tail )∈

[0027] Among them, p sub_head is the probability of the starting boundary of the head entity, P sub_tail is the end boundary probability of the head entity, σ is the sigmoid nonlinear activation function, W sub_head 、W sub_tail is the learnable weight, b sub_head 、b sub_tail is the bias term.

[0028] The detection is considered as a multi-classification problem. By adding two new linear layers to the entity relationship extraction model to predict the entity type label, the traditional triple extraction is converted into a five-tuple extraction, including the head entity, the head entity type, the semantic relationship, the tail entity, and the tail entity type. The head entity type prediction is achieved by calling the softmax classification model:

[0029] p sub_type =softmax(W sub_type H enhanced +b sub_type )∈

[0030] Among them, p sub_type is the predicted probability of the head entity type, W sub_type is the learnable weight, b sub_type is the bias term, and |T| is the number of head / tail entity types.

[0031] Step S34: Construct the tail entity decoding layer, and for each semantic relationship, correspond to a tail entity label. For the identified head entity, traverse all relationships. For a certain semantic relationship r∈R (R represents the set of all semantic relationships), the semantic enhancement word vector H enhanced and head entity feature h sub After concatenation, an RDC module and a linear layer are used to decode the start and end index positions of the tail entity of the relation r:

[0032] H r =RDC r ([H enhanced ;h sub ])∈

[0033]

[0034]

[0035] Among them, hsub is the identified head entity feature, i is the starting index position of the head entity, j is the ending index position of the head entity, is the starting boundary probability of the tail entity, is the end boundary probability of the tail entity, 、 are the learnable weights, 、 is the bias term, and σ is the sigmoid nonlinear activation function.

[0036] At the same time, the tail entity type is predicted, the softmax classification model is called, and the tail entity type prediction result is output:

[0037]

[0038] in, is the predicted probability of the tail entity type, are the learnable weights, is the bias term, and |T| is the number of head / tail entity types.

[0039] Step S35: A binary cross entropy loss function is used to predict the start and end positions of the head entity / tail entity:

[0040]

[0041] L sub =BCE(p sub_head ,y sub_head )+BCE(p sub_tail ,y sub_tail )

[0042]

[0043] Among them, BCE represents binary cross entropy loss, p i represents the predicted probability, y i represents the true label, p sub_head represents the predicted probability of the starting boundary of the head entity, y sub_head The ground truth label representing the starting boundary of the head entity, p sub_tail represents the predicted probability of the end boundary of the head entity, y sub_tail The ground truth label representing the end boundary of the head entity, represents the predicted probability of the starting boundary of the tail entity, The true label representing the starting boundary of the tail entity, represents the predicted probability of the end boundary of the tail entity, Indicates the true label of the end boundary of the tail entity, L sub represents the detection loss of the head entity, L objrepresents the detection loss of the tail entity, and |R| represents the number of semantic relations.

[0044] For the prediction of entity type, a multi-classification cross entropy loss function is used:

[0045]

[0046] L sub_type =CE(p sub_type ,y sub_type )

[0047]

[0048] Among them, CE represents multi-classification cross entropy loss, p ij represents the predicted probability, y ij represents the true label, L sub_type represents the prediction loss of the head entity type, L obj_type represents the prediction loss of the tail entity type, p sub_type Denotes the predicted probability of the head entity type, y sub_type The true label representing the head entity type, represents the predicted probability of the tail entity type, represents the true label of the tail entity type, |T| is the number of head / tail entity types, and |R| represents the number of semantic relations.

[0049] Step S36: During the entity relationship extraction model training process, the entity relationship extraction loss value of each epoch is monitored in real time and dynamically compared with the current optimal loss value. The model with the lowest loss value is obtained and saved to achieve continuous optimization of model performance. At the same time, an early stopping mechanism is added to the model training. As the epoch increases, if the entity relationship extraction loss value increases for five consecutive times, the model training is terminated early and the model before the loss value increased is retained.

[0050] Step S37: Select the three indicators of precision, recall, and F1 to evaluate the performance of the entity relationship extraction model finally saved. The calculation formulas are:

[0051] Precision=TP / (TP+FP)

[0052] Recall=TP / (TP+FN)

[0053] F1=2×Precision×Recall / (Precision+Recall)

[0054] Where TP stands for true positives, which indicates the number of entities, entity types, or semantic relations correctly predicted by the entity relationship extraction model. FP stands for false positives, which indicates the number of entities, entity types, or semantic relations incorrectly predicted by the entity relationship extraction model. FN stands for false negatives, which indicates the number of entities, entity types, or semantic relations that should have been predicted by the entity relationship extraction model but were missed.

[0055] Preferably, step S4 includes the following steps:

[0056] Step S41: Utilize an entity relationship extraction model to extract the semantic relationship quintuples of technical entities to construct a technical topic co-occurrence matrix, which serves as the basis for calculating relevant metrics. A co-occurrence relationship between technical topics refers to the existence of a co-occurrence relationship between two or more technical topic entities if the same technical document contains them.

[0057] Step S42: The universality of the technical theme is measured by the technical theme co-occurrence rate, and the average value of the co-occurrence times is set as the threshold for the co-occurrence partner statistics. The calculation formula for the technical co-occurrence rate is:

[0058] R i =coop i / (n-1)

[0059] Where R i represents the technology co-occurrence rate of technology topic i, coop i represents the number of co-occurring partners of technical topic i, and n is the number of technical topics.

[0060] Step S43: The effectiveness of the technical theme is measured by the average number of homologous systems and the average number of technical contribution points of the theme. The calculation formula is:

[0061]

[0062] Where, represents the average number of homologous systems for topic i, represents the average number of technical contribution points of topic i, N F (t) represents the number of different versions or variants derived from the technical text t to which subject i belongs, N c (t) represents the number of innovative elements in the technical text t to which topic i belongs, M i Indicates the number of technical texts belonging to topic i.

[0063] Step S44: Measure the relevance of technical topics using the technical topic co-occurrence intensity index, and the calculation formula is:

[0064] I ij =coo(i,j) / (occ(i)+occ(j)-coo(i,j))

[0065]

[0066] Where, I ij represents the co-occurrence intensity of technical topic i and technical topic j, I i represents the co-occurrence intensity of technical topic i, coo(i,j) represents the co-occurrence frequency of technical topics i and j in the technical text. occ(i) and occ(j) represent the frequencies of occurrence of technical topics i and j in the technical text, respectively.

[0067] Step S45: Calculate the technical theme commonality score according to the entropy method:

[0068]

[0069] Among them, S G (i) is the commonality score of technical topic i, norm represents standardization, and w i , i=1,2,3,4 are the common index weights obtained by the entropy method.

[0070] Step S46: If the commonality score S of the technical topic i G (i)≥avg[S G ] (avg means taking the average value), then technical topic i is determined to be the identified common technical topic.

[0071] Preferably, the method for measuring the importance and leadership of the technical theme in step S5 is:

[0072] Preferably, the measurement of the importance and leadership of the technical topic and the calculation of the criticality score are:

[0073] Step S51: In combination with the social network analysis method, according to the technology theme co-occurrence matrix, the technology entity types are used as nodes, and the co-occurrence times higher than the average co-occurrence times are used as edge weights to establish a technology theme co-occurrence network.

[0074] Step S52: quantify the importance of technology topics using degree centrality, betweenness centrality, and closeness centrality, thereby identifying technology nodes with significant status in the network. The calculation formula is:

[0075] DC i =k i / (N-1)

[0076]

[0077] Where, DC i , BC i 、CC i are degree centrality, betweenness centrality, and closeness centrality; ki represents the number of edges connected to node i, represents the number of paths that pass through node i and are the shortest paths, g st represents the number of shortest paths connecting s and t, d ij represents the distance from node i to node j, and N represents the number of technology nodes.

[0078] Step S53: The leadership of common technologies is measured by the average citation frequency of the topic. The calculation formula is:

[0079]

[0080] Where, represents the average citation frequency of topic i, M i represents the number of technical texts belonging to topic i, and F(t) represents the frequency of technical text t belonging to topic i being cited by other technical texts.

[0081] Step S54: Calculate the technical topic criticality score using the entropy method:

[0082]

[0083] Among them, S K (i) is the critical score of technical topic i, norm represents normalization, w i , i=5,6,7,8 are the key indicator weights obtained by the entropy method.

[0084] Step S55: If the critical score S of the technical topic i K (i)≥avg[S K ] (avg means taking the average value), then the common technical theme i is determined to be the identified key common technical theme.

[0085] A system for identifying key common technical entities includes the following modules:

[0086] Data preprocessing module: obtains technical text datasets in the required fields and performs data cleaning on the technical text datasets;

[0087] Corpus formation module: defines technical entity types and semantic relationships, and annotates text summary content to form a corpus;

[0088] Model training module: Use the corpus to train the entity relationship extraction model introduced into the neural network;

[0089] Screening common technology entity module: Screen common technology entities through commonality measurement, efficiency and relevance;

[0090] Identify key common technology entities module: Use social network analysis to measure the importance of technology entities, and combine leading indicators to measure technology criticality, accurately and efficiently identifying key common technology entities.

[0091] The system implementation process is the same as the method implementation process and will not be repeated here.

[0092] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for identifying key common technical entities as described above are implemented.

[0093] A computer-readable storage medium stores a computer program for executing the method for identifying key common technical entities as described above.

[0094] Beneficial effects: 1. The entity relationship extraction model incorporates a recurrent dilated convolutional neural network based on the existing casrel model, which is suitable for longer technical texts and improves the entity relationship extraction effect.

[0095] 2. Add an entity type prediction module to the entity relationship extraction model, which is suitable for technical texts with complex semantic relationships. Expand the triple extraction to five-tuple extraction including entity types to facilitate the subsequent calculation of key and common indicators.

[0096] 3. Compared with the current method of using LDA to identify key common technologies, the use of entity relationship extraction technology can achieve fine-grained technology entity extraction, thereby improving the accuracy of key common technology identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 is a flow chart of the method principle of an embodiment of the present invention;

[0098] Figure 2 This is a co-occurrence network diagram of the technical themes of an embodiment of the present invention. DETAILED DESCRIPTION

[0099] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0100] like Figure 1 As shown, an embodiment of the present invention provides a key common technology identification method based on deep learning and entity relationship extraction, including the following steps:

[0101] Step S1: Obtain a technical text dataset in a certain field, such as a patent text dataset, and perform data cleaning on the text dataset.

[0102] Specifically, the Incopat global patent database was selected as the experimental data source, and a patent search formula was constructed (TIAB=("humanoid intelligent robot" OR "humanoid robot" OR "humanoid robot" OR "humanoid robot" OR "walking robot" OR "bipedal walking robot" OR "anthropomorphic robot" OR "robot head" OR "robot foot" OR "robot leg" OR "robot foot" OR "humanoid robotic hand" OR "dexterous hand" OR "humanoid intelligent robot" OR "humanoid robot" OR "biped robot" OR "walking robot" OR "anthropomorphic robot" OR "robot head" OR "robot foot" OR "robot leg" OR "dexterous hand" OR "dexterous robot hand")). The search date was December 12, 2024, and the search time span was 1992-2024. A total of 2,356 Chinese patents were retrieved.

[0103] In an embodiment of the present invention, in order to improve the operability of the data source, the original corpus required for constructing the entity relationship extraction model is mainly extracted from the patent abstracts as the text summary content, and the retrieved patent data set is subjected to deduplication and other cleaning operations to obtain 2035 patent abstract texts.

[0104] Step S2: define entities and semantic relationships, and manually annotate technical texts to form a corpus;

[0105] Specifically, we used the bertopic topic model to mine technical topics on the technical text dataset. The results are shown in Table 1. Based on these mining results and combined with relevant expertise in the field of humanoid robots, we defined entities and semantic relationships. We then used the doccano data annotation platform for manual annotation to form a model training corpus, which was divided into training and test sets in a 7:3 ratio.

[0106] Table 1bertopic theme-keyword matrix

[0107]

[0108] Table 2 Definition of technical entity types

[0109]

[0110] Table 3 Relationship type definition

[0111]

[0112] Step S3: training the entity relationship extraction model introduced into the neural network;

[0113] In this embodiment of the present invention, the experimental operating system was Windows 11 64-bit, the central processing unit (CPU) was an Intel(R) Core(TM) i7-14700 2.10GHz, and the memory was 32.0GB. The programming language was Python 3.9.13, and the development tool was Jupyter. The model was built using PyTorch, a Torch-based framework developed by Facebook Artificial Intelligence Research (FAIR). The BERT model used in the improved Casrel model was the pre-trained model BERT-Base-Chinese provided by Google. The specific experimental parameters are shown in Table 4. The experimental parameters for the cyclic dilated convolutional neural network are shown in Table 5.

[0114] Table 4 Hyperparameter settings (BERT)

[0115]

[0116] Table 5 Hyperparameters (CDIL-CNN)

[0117]

[0118] Step 1) First, the technical text in the entity relationship extraction model training corpus formed in step S2 is segmented by BERT's built-in word segmenter and converted into a string sequence X=(x1,x2,...,x n ), where x i For characters in technical text, such as text, i=1, 2, ..., n represents the character sequence number, and at the same time generates the vocabulary index input_ids, attention mask attention_mask and character position mapping offset_mapping; then the generated vocabulary index and attention mask are used to generate the context-related word vector H through the pre-trained language model BERT.

[0119] H=BERT(input_ids, attention_mask)∈

[0120] represents a set of real numbers; b represents the batch size; n represents the length of the string sequence of technical text; d represents the hidden layer dimension of BERT.

[0121] Secondly, the character position mapping offset_mapping is traversed to map the character-level position of the head / tail entity in the technical text to the string sequence. At the same time, all strings within the head / tail entity span are marked as the same head / tail entity type, and the head / tail entity boundary index position and the real label y of the head / tail entity type are generated.

[0122] Step 2): Build a shared semantic feature enhancement layer using a cyclic dilated convolutional neural network (RDC module). The context-sensitive word vector H output by BERT passes through the RDC module to capture long-range dependencies and hierarchical features. Because information in text summaries often has multi-level dependencies, important technical entities and relationship information may be located at the boundaries of the text. Therefore, by adjusting the dilation rate and padding size of the cyclic dilated convolutional neural network, convolutional layers with different dilation rates are used to extract multi-scale feature information:

[0123] H enhanced =RDC(H)∈

[0124] C(H)=Res(H)+GLU(Conv1D(H))

[0125] CP(H)=C k=2 0 (C k=2 1 (…C k=2 L-1 (H)…))

[0126] The RDC module gradually enhances the input features by stacking multiple dilated convolutional layers and residual connections to capture multi-scale information. At the same time, it introduces a gated linear unit to gate the convolution results and filter out unimportant information. enhanced Represents semantically enhanced word vectors; Res(H) represents residual connection; GLU represents gated linear unit; C(H) represents a single-layer symmetric dilated convolution operation; Conv1D represents a one-dimensional convolution layer; CP(H) represents a multi-layer symmetric dilated convolution operation; L represents the convolution kernel size; k represents the dilation rate.

[0127] Step 3) Construct the head entity decoding layer to realize the two tasks of head entity recognition and head entity type prediction. The head entity decoding layer is used to enhance the semantic word vector H enhanced Decoding is performed by constructing two binary classifiers to predict the index positions of the start and end of the head entity, for each character x in the technical text X iCalculate the probability of (i=1, 2, ..., n) as the starting and ending positions, and filter out the starting index position and ending index position of the head entity according to the set threshold. Then, use the nearest matching principle to pair the identified starting and ending index positions to obtain a set of candidate head entities:

[0128] p sub_head =σ(W sub_head H enhanced +b sub_head )∈

[0129] p sub_tail =σ(W sub_tail H enhanced +b sub_tail )∈

[0130] Among them, p sub_head is the probability of the starting boundary of the head entity, P sub_tail is the end boundary probability of the head entity, σ is the sigmoid nonlinear activation function, W sub_head 、W sub_tail is the learnable weight, b sub_head 、b sub_tail is the bias term.

[0131] We treat the prediction of entity types as a multi-classification problem and add two new linear layers to the entity relationship extraction model to predict entity type labels. This converts the traditional triple extraction into a five-tuple extraction, which includes the head entity, head entity type, semantic relationship, tail entity, and tail entity type. The head entity type prediction is implemented by calling the softmax classification model:

[0132] p sub_type =softmax(W sub_type H enhanced +b sub_type )∈

[0133] Among them, p sub_type is the predicted probability of the head entity type, W sub_type is the learnable weight, b sub_type is the bias term, and |T| is the number of head / tail entity types.

[0134] Step 4) Construct the tail entity decoding layer, and for each semantic relationship, correspond to a tail entity label. For the identified head entity, traverse all relationships. For the semantic relationship r∈R (R represents the set of all semantic relationships), the semantic enhancement word vector H enhanced and head entity feature h subAfter concatenation, an RDC module and a linear layer are used to decode the start and end index positions of the tail entity for the relationship:

[0135]

[0136] Among them, h sub is the identified head entity feature, i is the starting index position of the head entity, j is the ending index position of the head entity, is the starting boundary probability of the tail entity, is the end boundary probability of the tail entity, 、 are the learnable weights, 、 is the bias term, which is the sigmoid nonlinear activation function.

[0137] At the same time, the tail entity type is predicted, the softmax classification model is called, and the tail entity type prediction result is output:

[0138]

[0139] in, is the predicted probability of the tail entity type, are the learnable weights, is the bias term, and |T| is the number of head / tail entity types.

[0140] Step 5) For the start and end position prediction of the head entity / tail entity, a binary cross entropy loss function is used:

[0141]

[0142] L sub =BCE(p sub_head , y sub_head )+BCE(p sub_tail , y sub_tail )

[0143]

[0144] Among them, BCE represents binary cross entropy loss, p i represents the predicted probability, y i represents the true label, p sub_head represents the predicted probability of the starting boundary of the head entity, y sub_head The ground truth label representing the starting boundary of the head entity, p sub_tail represents the predicted probability of the end boundary of the head entity, y sub_tail The ground truth label representing the end boundary of the head entity, represents the predicted probability of the starting boundary of the tail entity, The true label representing the starting boundary of the tail entity, represents the predicted probability of the end boundary of the tail entity, Indicates the true label of the end boundary of the tail entity, L sub represents the detection loss of the head entity, L obj represents the detection loss of the tail entity, and |R| represents the number of semantic relations.

[0145] For the prediction of entity type, a multi-classification cross entropy loss function is used:

[0146]

[0147] L sub_type =CE(p sub_type ,y sub_type )

[0148]

[0149] Among them, CE represents multi-classification cross entropy loss, L sub_type represents the prediction loss of the head entity type, L obj_type represents the prediction loss of the tail entity type, p sub_type Denotes the predicted probability of the head entity type, y sub_type The true label representing the head entity type, represents the predicted probability of the tail entity type, represents the true label of the tail entity type, and |R| represents the number of semantic relations.

[0150] Step 6) During entity relationship extraction model training, the entity relationship extraction loss value at each epoch is monitored in real time and dynamically compared with the current optimal loss value. The model with the lowest loss value is obtained and saved to continuously optimize model performance. At the same time, an early stopping mechanism is incorporated into model training. As epochs increase, if the entity relationship extraction loss value increases five times in a row, the model training is terminated early, and the model before the loss increase is retained.

[0151] Step 7) Integrate a cyclic dilated convolutional neural network into the Casrel model to improve the model's entity relationship extraction performance. To evaluate the effectiveness of the improved Casrel model, an ablation experiment was conducted on the technical text test set. The performance of the final saved entity relationship extraction model was evaluated using three metrics: precision, recall, and F1. The calculation formulas are:

[0152] Precision=TP / (TP+FP)

[0153] Recall=TP / (TP+FN)

[0154] F1=2×Precision×Recall / (Precision+Recall)

[0155] Where TP stands for true positives, which indicates the number of entities, entity types, or semantic relations correctly predicted by the entity relationship extraction model. FP stands for false positives, which indicates the number of entities, entity types, or semantic relations incorrectly predicted by the entity relationship extraction model. FN stands for false negatives, which indicates the number of entities, entity types, or semantic relations that should have been predicted by the entity relationship extraction model but were missed.

[0156] The experimental results are shown in Table 6. The results demonstrate that the CASREL model, incorporating a cyclic dilated convolutional neural network (CDCN), achieves superior performance in both entity recognition and relation extraction tasks in technical text, with F1 improvements of 2.95 and 2.57 percentage points, respectively. This demonstrates that the improvements significantly enhance entity relation extraction. This improvement is primarily due to the CDCN's ability to capture long-range dependencies within technical text, making it more suitable for the complex context of entities and relations within it. Furthermore, the accuracy of entity type predictions exceeded 94%, laying the foundation for the subsequent calculation of key commonality indicators.

[0157] Table 6 Experimental results

[0158]

[0159] Step S4: Measure commonality through versatility, efficiency, and relevance to screen common technical entities;

[0160] Specifically, an entity relationship extraction model is used to obtain the five-tuple semantic relationship of technical topics to calculate the common index value, and common technologies are identified by calculating the common technology feature score through the entropy method.

[0161] Step 1) Use the entity relationship extraction model to obtain the technical topic semantic relationship quintuple to construct the technical topic co-occurrence matrix, and then calculate the common technical indicator values ​​of technical topic i;

[0162] Table 7 Technical topic co-occurrence matrix (partial)

[0163]

[0164] Step 2) Measure the universality of technical topics through the technical topic co-occurrence rate, and set the average co-occurrence count as the threshold for co-occurrence partner statistics. The calculation formula for the technical co-occurrence rate is:

[0165] R i =coop i / (n-1)

[0166] Where R i represents the technology co-occurrence rate of technology topic i, coop i represents the number of co-occurring partners of technical topic i, and n is the number of technical topics.

[0167] Step 3) Measure the effectiveness of a technical topic by the average number of homologous systems (such as the number of families) and the average technical contribution points (such as the number of problems solved). The calculation formula is:

[0168]

[0169] Where, represents the average number of homologous systems (such as the number of families) for subject i, represents the average technical contribution points of topic i (such as the number of problems solved), M i Indicates the number of technical documents to which topic i belongs, such as the number of patents.

[0170] Step 4) Use the technical topic co-occurrence intensity index to measure the relevance of technical topics. The calculation formula is:

[0171] I ij =coo(i,j) / (occ(i)+occ(j)-coo(i,j))

[0172]

[0173] Where, I ij represents the co-occurrence intensity of technical topic i and technical topic j, I i represents the co-occurrence intensity of technical theme i, coo(i,j) represents the co-occurrence frequency of technical theme i and technical theme j in the technical text. occ(i) and occ(j) represent the frequency of occurrence of technical theme i and technical theme j in the patent abstract, respectively.

[0174] Step 5) Use the entropy method to determine the indicator weight coefficient and calculate the commonality score of technical topic i:

[0175]

[0176] Where S G (i) is the commonality score of technical topic i, norm means standardizing the index value, R i is the technical topic co-occurrence rate, is the average number of siblings of subject i, is the average number of problems solved for topic i, I i is the technical theme co-occurrence intensity of topic i, w1, w2, w3, and w4 are the indicator weight coefficients determined by the entropy method;

[0177] Step 6) Set the mean of the commonality scores of all technical topics as the threshold. If the commonality score of technical topic i is S G (i)≥avg[S G ], then the technical subject i is determined to be the identified common technical subject.

[0178] Table 8 Common technology identification results

[0179]

[0180] Step S5: Use social network analysis to measure the importance of technical topics, and combine leading indicators to measure the criticality of technology, so as to accurately and efficiently identify key common technologies.

[0181] Specifically, an entity relationship extraction model is used to obtain the five-tuple semantic relationship of technical topics to calculate the key indicator value, and the key common technologies are identified by calculating the key technical feature scores through the entropy method.

[0182] Step 1) Combined with the social network analysis method, according to the technical theme co-occurrence matrix, a technical theme co-occurrence network is established, and then the key technical indicator values ​​of technical theme i are calculated.

[0183] Step 2) Quantify the importance of technology topics using degree centrality, betweenness centrality, and closeness centrality to identify technology nodes with significant positions in the technology topic co-occurrence network. The calculation formula is:

[0184] DC i =k i / (N-1)

[0185]

[0186]

[0187] Where, DC i , BC i 、CC i are degree centrality, betweenness centrality, and closeness centrality; k i represents the number of edges connected to node i, represents the number of paths that pass through node i and are the shortest paths, g st represents the number of shortest paths connecting s and t, d ij represents the distance from node i to node j, and N represents the number of technology nodes.

[0188] Step 3) Measure the leadership of common technologies by the average citation frequency of the topic. The calculation formula is:

[0189]

[0190] Where, represents the average citation frequency of topic i, M i represents the number of patent abstracts belonging to topic i, and F(t) represents the frequency of technical text t belonging to topic i being cited by other technical texts.

[0191] Step 4) Use the entropy method to determine the indicator weight coefficient and calculate the criticality score of technical topic i:

[0192]

[0193] Where S K (i) is the critical score of technical topic i, norm means standardizing the index value, DC i is the degree centrality of topic i, BC i is the betweenness centrality of subject i, CC i is the closeness centrality of topic i, is the average citation frequency of topic i, w5, w6, w7, and w8 are the indicator weight coefficients determined by the entropy method;

[0194] Step 5) Set the mean of all technical topic criticality scores as the threshold. If the criticality score of technical topic i is S K (i)≥avg[S K ], the common technical theme i is determined to be the identified key common technical theme.

[0195] Table 9 Key common technology identification results

[0196]

[0197] Through the method of the embodiment of the present invention, the key common technologies in the field can be accurately and quickly identified based on the field technical text data.

[0198] A system for identifying key common technical entities includes the following modules:

[0199] Data preprocessing module: obtains technical text datasets in the required fields and performs data cleaning on the technical text datasets;

[0200] Corpus formation module: defines technical entity types and semantic relationships, and annotates text summary content to form a corpus;

[0201] Model training module: Use the corpus to train the entity relationship extraction model introduced into the neural network;

[0202] Screening common technology entity module: Screen common technology entities through commonality measurement, efficiency and relevance;

[0203] Identify key common technology entities module: Use social network analysis to measure the importance of technology entities, and combine leading indicators to measure technology criticality, accurately and efficiently identifying key common technology entities.

[0204] Obviously, those skilled in the art should understand that the various steps of the method for identifying key common technical entities or the various modules of the system for identifying key common technical entities of the above-mentioned embodiments of the present invention can be implemented using a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Alternatively, they can be implemented using program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described can be performed in an order different from that shown here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

Claims

1. A method for identifying key common technical entities, characterized in that: The following steps are involved: Step S1: Obtain a technical text dataset in the required field and perform data cleaning on the technical text dataset; Step S2: define technical entity types and semantic relationships, and annotate the text summary content to form a corpus; Step S3: Using the corpus to train the entity relationship extraction model introduced into the neural network; Step S4: Measure commonality through versatility, efficiency, and relevance to screen common technical entities; Step S5: Use social network analysis to measure the importance of technical entities, and combine leading indicators to measure technical criticality, accurately and efficiently identifying key common technical entities; The step S4 comprises the following steps: Step S41: Using an entity relationship extraction model to obtain a five-tuple of semantic relationships between technical entities to construct a technical topic co-occurrence matrix; the co-occurrence relationship between technical topics means that if the same technical document contains two or more technical topic entities at the same time, then there is a co-occurrence relationship between the two or more topics; Step S42: The versatility of the technical theme is measured by the technical theme co-occurrence rate, and the average value of the co-occurrence times is set as the threshold for the co-occurrence partner statistics. The calculation formula for the technical co-occurrence rate is: R i =coop i / (n-1), where R i represents the technology co-occurrence rate of technology topic i, coop i represents the number of co-occurring partners of technical topic i, and n is the number of technical topics; Step S43: measuring the effectiveness of the technical theme by the average number of homologous systems and the average number of technical contribution points of the theme; Step S44: Measure the relevance of technical topics using the technical topic co-occurrence intensity index, and the calculation formula is: I ij =coo(i,j) / (occ(i)+occ(j)-coo(i,j)), , where I ij represents the co-occurrence intensity of technical topic i and technical topic j, I i represents the co-occurrence intensity of technical topic i, coo(i,j) represents the co-occurrence frequency of technical topic i and technical topic j in the technical text; occ(i) and occ(j) represent the frequency of technical topic i and technical topic j in the technical text respectively; Step S45: Calculate the technical theme commonality score according to the entropy method: Among them, S G (i) is the commonality score of technical topic i, norm represents standardization, w1-w4 are the commonality index weights obtained by entropy method, R i represents the technical co-occurrence rate of topic i, represents the average number of homologous systems for topic i, represents the average number of technical contribution points of topic i, I i represents the co-occurrence strength of topic i; Step S46: If the commonality score S of the technical topic i G (i)≥avg[S G ], avg represents the average value, and the technical topic i is determined to be the identified common technical topic.

2. The method for identifying key common technical entities according to claim 1, characterized in that: In step S2, the technical entity types and semantic relationships are defined, and the text summary content is annotated to form a corpus, including: Step S21: Using the bertopic topic model to identify technical entities in the text summary content, and defining the types and semantic relationships of technical entities based on the technical entity identification results and in combination with relevant domain expertise; Step S22: Based on the technical entity types and semantic relationships, the doccano data annotation platform is used to annotate the entity semantic relationships of the technical text to form an entity relationship extraction model training corpus containing technical text, head / tail entities, head / tail entity types, and semantic relationships. The corpus is divided into training set and test set according to the proportion.

3. The method for identifying key common technical entities according to claim 1, characterized in that: The step S3 includes training the entity relationship extraction model introduced into the neural network, including: Step S31: First, the technical text in the entity relationship extraction model training corpus is segmented using BERT's built-in word segmenter and converted into a string sequence. At the same time, the word table index input_ids, attention mask attention_mask, and character position mapping offset_mapping are generated. The generated word table index and attention mask are then used to generate context-related word vectors through the pre-trained language model BERT. Secondly, traverse the character position mapping offset_mapping to map the character-level position of the head / tail entity in the technical text to the string sequence, and mark all the strings within the head / tail entity span as the same head / tail entity type, generating the head / tail entity boundary index position and the real label of the head / tail entity type; Step S32: A shared semantic feature enhancement layer is constructed using a cyclic dilated convolutional neural network (RDC module). The context-dependent word vector H output by BERT passes through the RDC module to capture long-range dependencies and hierarchical features. By adjusting the dilation rate and padding size of the cyclic dilated convolutional neural network, convolutional layers with different dilation rates are used to extract multi-scale feature information. Step S33: Construct a header entity decoding layer to realize the two tasks of header entity recognition and header entity type prediction; the header entity decoding layer decodes the semantic enhancement word vector, and predicts the index position of the start and end of the header entity by constructing two binary classifiers. i Calculate the probability of it being the starting and ending position, and filter out the starting index position and ending index position of the head entity according to the set threshold. Then, use the nearest matching principle to pair the identified starting and ending index positions to obtain a set of candidate head entities. The prediction of entity type is considered as a multi-classification problem. Two new linear layers are added to the entity relationship extraction model to predict the entity type label. The five-tuple consisting of the head entity, the head entity type, the semantic relationship, the tail entity, and the tail entity type is extracted. The head entity type prediction is implemented by calling the softmax classification model. Step S34: Construct a tail entity decoding layer. Each semantic relationship corresponds to a tail entity annotation. For the identified head entity, all relationships are traversed. For a semantic relationship r in the semantic relationship set, the semantic enhancement word vector and the head entity feature are concatenated and passed through an RDC module and a linear layer to decode the start and end index positions of the tail entity for the relationship r. At the same time, the tail entity type is predicted, and the softmax classification model is called to output the tail entity type prediction result. Step S35: A binary cross entropy loss function is used to predict the start and end positions of the head entity / tail entity; a multi-class cross entropy loss function is used to predict the entity type; Step S36: During the training of the entity relationship extraction model, the entity relationship extraction loss value of each epoch is monitored in real time and dynamically compared with the current optimal loss value. The model with the minimum loss value is obtained and saved. An early stopping mechanism is added to the model training. As the epoch increases, when the entity relationship extraction loss value increases for five consecutive times, the model training is terminated early and the model before the loss value increases is retained. Step S37: Select the three indicators of precision, recall and F1 to evaluate the performance of the entity relationship extraction model finally saved.

4. The method for identifying key common technical entities according to claim 1, characterized in that: Step S5 includes: Step S51: In combination with social network analysis methods, based on the technology theme co-occurrence matrix, the technology entity types are used as nodes, and the co-occurrence times higher than the average co-occurrence times are used as edge weights to establish a technology theme co-occurrence network; Step S52: quantify the importance of technology topics using degree centrality, betweenness centrality, and closeness centrality, thereby identifying technology nodes with significant status in the network. The calculation formula is: DC i =k i / (N-1) Where, DC i , BC i 、CC i are degree centrality, betweenness centrality, and closeness centrality; k i represents the number of edges connected to node i, represents the number of paths that pass through node i and are the shortest paths, g st represents the number of shortest paths connecting s and t, d ij represents the distance from node i to node j, and N represents the number of technology nodes; Step S53: The leadership of common technologies is measured by the average citation frequency of the topic. The calculation formula is: Where, represents the average citation frequency of topic i, M i represents the number of technical texts belonging to topic i, and F(t) represents the frequency of technical text t belonging to topic i being cited by other technical texts; Step S54: Calculate the technical topic criticality score using the entropy method: Among them, S K (i) is the criticality score of technical topic i, norm represents normalization, and w5-w8 are the criticality indicator weights obtained by the entropy method; Step S55: If the critical score S of the technical topic i K (i)≥avg[S K ], avg represents the average value, and the common technical theme i is determined to be the identified key common technical theme.

5. A system for identifying key common technical entities, characterized in that: Includes the following modules: Data preprocessing module: obtains technical text datasets in the required fields and performs data cleaning on the technical text datasets; Corpus formation module: defines technical entity types and semantic relationships, and annotates text summary content to form a corpus; Model training module: Use the corpus to train the entity relationship extraction model introduced into the neural network; Screening common technology entity module: Screen common technology entities through commonality measurement, efficiency and relevance; Identifying key common technology entities: This module uses social network analysis to measure the importance of technology entities and combines leading indicators to measure technology criticality, accurately and efficiently identifying key common technology entities. The implementation of screening common technical entity modules includes the following steps: Step S41: Using an entity relationship extraction model to obtain a five-tuple of semantic relationships between technical entities to construct a technical topic co-occurrence matrix; the co-occurrence relationship between technical topics means that if the same technical document contains two or more technical topic entities at the same time, then there is a co-occurrence relationship between the two or more topics; Step S42: The versatility of the technical theme is measured by the technical theme co-occurrence rate, and the average value of the co-occurrence times is set as the threshold for the co-occurrence partner statistics. The calculation formula for the technical co-occurrence rate is: R i =coop i / (n-1), where R i represents the technical co-occurrence rate of technical topic i, coop i represents the number of co-occurring partners of technical topic i, and n is the number of technical topics; Step S43: measuring the effectiveness of the technical theme by the average number of homologous systems and the average number of technical contribution points of the theme; Step S44: Measure the relevance of technical topics using the technical topic co-occurrence intensity index, and the calculation formula is: I ij =coo(i,j) / (occ(i)+occ(j)-coo(i,j)), , where I ij represents the co-occurrence intensity of technical topic i and technical topic j, I i represents the co-occurrence intensity of technical topic i, coo(i,j) represents the co-occurrence frequency of technical topic i and technical topic j in the technical text; occ(i) and occ(j) represent the frequency of technical topic i and technical topic j in the technical text respectively; Step S45: Calculate the technical theme commonality score according to the entropy method: Among them, S G (i) is the commonality score of technical topic i, norm represents standardization, w1-w4 are the commonality index weights obtained by entropy method, R i represents the technical co-occurrence rate of topic i, represents the average number of homologous systems for topic i, represents the average number of technical contribution points of topic i, I i represents the co-occurrence strength of topic i; Step S46: If the commonality score S of the technical topic i G (i)≥avg[S G ], avg represents the average value, and the technical topic i is determined to be the identified common technical topic.

6. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for identifying key common technical entities according to any one of claims 1 to 5 are implemented.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the method for identifying the key common technical entity according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Semantic analysis and entity relationship joint extraction method for service resources

    CN117217219A

  • Identifying interesting commonalities between entities

    US9116982B1