Document composite label generation method and device, equipment and storage medium

By generating and combining multiple similarity matrices and extracting TagDC models using composite tags, the problem that traditional single tag classification method is difficult to cover multiple topics of power documents is solved, and higher label coverage and accuracy are achieved.

CN120106013APending Publication Date: 2025-06-06GUANGDONG POWER GRID CO LTD CUSTOMER SERVICE CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510175761.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The traditional single-label classification method is difficult to ensure the coverage and accuracy of power document labels, resulting in the neglect or omission of important information.

Method used

By obtaining the cosine similarity between the content vector of the target document and the label vector in the preset label knowledge base, the topic tag similarity matrix and user tag similarity matrix are generated, combining the user collaborative similarity matrix and topic collaborative similarity matrix, the TagDC model is used to extract the tags to generate a multi-label confidence probability list, and finally select several tag combinations to generate composite tags.

Benefits of technology

The multi-label confidence probability list generation for power documents is realized, which improves the coverage and accuracy of tags, and can more comprehensively cover multiple topics or concepts in power document text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106013A_ABST
    Figure CN120106013A_ABST
Patent Text Reader

Abstract

The invention discloses a document composite label generation method and device, equipment and a storage medium. The method comprises the steps that a theme label similarity matrix corresponding to a target document theme is generated; generating a corresponding user label similarity matrix; generating a corresponding user cooperation similar matrix and a theme cooperation similar matrix; inputting the user collaborative similar matrix and the theme collaborative similar matrix into a preset composite label extraction TagDC model, so that the composite label extraction TagDC model generates a multi-label confidence probability list of each target document according to the user collaborative similar matrix and the theme collaborative similar matrix, and outputs the multi-label confidence probability list of each target document; and according to the multi-label confidence probability list, selecting a plurality of label combinations from the multi-label confidence probability list to generate a composite label of the target document. According to the invention, the composite label of the target document can be generated, and the coverage and accuracy of the label of the power document are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method, device, equipment and storage medium for generating document composite tags. Background Art

[0002] With the popularization of the Internet and the explosion of information, people have an increasing demand for fast and accurate information acquisition. At present, the traditional single-label classification method can no longer meet people's needs, especially in the power industry, because power documents often have multiple attributes, and the power document label generation technology needs to cover multiple topics or concepts in the power document text as comprehensively as possible. However, in practical applications, due to the diversity and complexity of the power document corpus, the traditional single-label classification method is difficult to guarantee the coverage and accuracy of the label, which may cause some important information to be ignored or omitted. Summary of the invention

[0003] The present invention provides a method, device, equipment and storage medium for generating composite tags of documents, so as to solve the technical problem that the traditional single tag classification method can no longer meet the requirements of power document tag classification.

[0004] In order to solve the above technical problems, an embodiment of the present invention provides a method for generating a document composite tag, comprising:

[0005] Obtain the target document, generate the document content vector corresponding to each sentence in the target document, and calculate the cosine similarity between each document content vector and the label vector corresponding to each label in the preset label knowledge base, and use the label with the largest cosine similarity as the label of the corresponding sentence, and then generate the topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document;

[0006] Obtaining a user behavior preference matrix that records the user's historical behavior, and generating a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix;

[0007] Calculate the cosine similarity between the matrix vectors corresponding to the pairwise user tag similarity matrices, obtain the corresponding user collaborative similarity matrix according to the cosine similarity, calculate the weighted average of the first similarity and the second similarity between the target document topics, and obtain the corresponding topic collaborative similarity matrix according to the weighted average;

[0008] Input the user collaborative similarity matrix and the topic collaborative similarity matrix into a preset composite tag extraction TagDC model, so that the composite tag extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs the multi-label confidence probability list for each target document;

[0009] According to the multi-label confidence probability list, several label combinations are selected from the multi-label confidence probability list to generate a composite label for the target document.

[0010] As a preferred solution, the generation of the tag knowledge base includes:

[0011] According to a preset document subject database, new word discovery is performed on the text content of all document subjects in the document subject database to extract a preset number of power industry keywords;

[0012] For each electric power industry keyword, the document topics containing the electric power industry keyword are screened out, and the text content of the screened document topics is clustered and analyzed to find similar tag words containing the electric power industry keyword, and then a corresponding tag knowledge base is generated based on the electric power industry keyword and the similar tag words.

[0013] As a preferred solution, generating a topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document includes:

[0014] According to the labels corresponding to each sentence in the target document and the preset TF-IDF algorithm, the TF-IDF weight between the target document topic and all labels is calculated;

[0015] According to the TF-IDF weights between the target document topic and all tags, the corresponding document topic tag similarities are obtained, and then the corresponding topic tag similarity matrix is ​​constructed according to the document topic tag similarities.

[0016] As a preferred solution, the step of selecting a plurality of label combinations from the multi-label confidence probability list to generate a composite label of the target document according to the multi-label confidence probability list includes:

[0017] According to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted from high to low, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate a composite label for the target document.

[0018] Based on the above embodiment, another embodiment of the present invention provides a device for generating document composite tags, including: a topic tag similarity matrix generation module, a user tag similarity matrix generation module, a collaborative similarity matrix generation module, a multi-tag confidence probability list generation module and a composite tag generation module;

[0019] The topic label similarity matrix generation module is used to obtain the target document, generate the document content vector corresponding to each sentence in the target document, and calculate the cosine similarity between each document content vector and the label vector corresponding to each label in the preset label knowledge base, and use the label with the largest cosine similarity as the label of the corresponding sentence, and then generate the topic label similarity matrix corresponding to the subject of the target document according to the label corresponding to each sentence in the target document;

[0020] The user tag similarity matrix generation module is used to obtain a user behavior preference matrix that records the user's historical behavior, and generate a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix;

[0021] The collaborative similarity matrix generation module is used to calculate the cosine similarity between matrix vectors corresponding to the pairwise user tag similarity matrices, obtain the corresponding user collaborative similarity matrix according to the cosine similarity, calculate the weighted average of the first similarity and the second similarity between the target document topics, and obtain the corresponding topic collaborative similarity matrix according to the weighted average;

[0022] The multi-label confidence probability list generation module is used to input the user collaborative similarity matrix and the topic collaborative similarity matrix into a preset composite label extraction TagDC model, so that the composite label extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs the multi-label confidence probability list for each target document;

[0023] The composite label generation module is used to select several label combinations from the multi-label confidence probability list to generate a composite label for the target document according to the multi-label confidence probability list.

[0024] As a preferred solution, the generation of the tag knowledge base includes:

[0025] According to a preset document subject database, new word discovery is performed on the text content of all document subjects in the document subject database to extract a preset number of power industry keywords;

[0026] For each electric power industry keyword, the document topics containing the electric power industry keyword are screened out, and the text content of the screened document topics is clustered and analyzed to find similar tag words containing the electric power industry keyword, and then a corresponding tag knowledge base is generated based on the electric power industry keyword and the similar tag words.

[0027] As a preferred solution, generating a topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document includes:

[0028] According to the labels corresponding to each sentence in the target document and the preset TF-IDF algorithm, the TF-IDF weight between the target document topic and all labels is calculated;

[0029] According to the TF-IDF weights between the target document topic and all tags, the corresponding document topic tag similarities are obtained, and then the corresponding topic tag similarity matrix is ​​constructed according to the document topic tag similarities.

[0030] As a preferred solution, the step of selecting a plurality of label combinations from the multi-label confidence probability list to generate a composite label of the target document according to the multi-label confidence probability list includes:

[0031] According to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted from high to low, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate a composite label for the target document.

[0032] Based on the above embodiments, another embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, and when the processor executes the computer program, the method for generating document composite tags described in the above embodiments of the invention is implemented.

[0033] Based on the above embodiments, another embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the method for generating document composite tags described in the above invention embodiment.

[0034] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0035] The present invention provides a method for generating composite tags of documents, which comprises the following steps: obtaining a target document, generating a document content vector corresponding to each sentence in the target document, and calculating the cosine similarity between each document content vector and a tag vector corresponding to each tag in a preset tag knowledge base, taking the tag with the largest cosine similarity as the tag of the corresponding sentence, and then generating a topic tag similarity matrix corresponding to the topic of the target document according to the tags corresponding to each sentence in the target document; obtaining a user behavior preference matrix recording user historical behaviors, and generating a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix; calculating the cosine similarity between matrix vectors corresponding to each pair of user tag similarity matrices, and generating a topic tag similarity matrix according to the cosine similarity matrix; The similarity is used to obtain the corresponding user collaborative similarity matrix, and the weighted average of the first similarity and the second similarity between the target document topics is calculated, and the corresponding topic collaborative similarity matrix is ​​obtained according to the weighted average; the user collaborative similarity matrix and the topic collaborative similarity matrix are input into a preset composite tag extraction TagDC model, so that the composite tag extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs a multi-label confidence probability list for each target document; according to the multi-label confidence probability list, a number of label combinations are selected from the multi-label confidence probability list to generate a composite label of the target document. Through the present invention, a composite label of a target document can be generated, and the generated composite label can more comprehensively cover multiple topics or concepts in the power document text than a traditional single label, thereby improving the coverage and accuracy of the power document label. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a flowchart of a method for generating a document composite tag provided by an embodiment of the present invention;

[0037] Figure 2 It is a schematic diagram of the structure of the TagDC model;

[0038] Figure 3 It is a structural schematic diagram of a device for generating document composite tags provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by technicians in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions.

[0041] In the description of the embodiments of the present application, the technical terms "first", "second", etc. are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.

[0042] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0043] In the description of the embodiments of the present application, the term "and / or" is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0044] In the description of the embodiments of the present application, the term "multiple" refers to more than two (including two). Similarly, "multiple groups" refers to more than two groups (including two groups), and "multiple pieces" refers to more than two pieces (including two pieces).

[0045] In the description of the embodiments of the present application, unless otherwise clearly specified and limited, technical terms such as "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to the specific circumstances.

[0046] Embodiment 1

[0047] Please refer to Figure 1, is a flow chart of a method for generating a document composite tag provided by an embodiment of the present invention, comprising the following specific steps:

[0048] S1. Obtain the target document, generate the document content vector corresponding to each sentence in the target document, and calculate the cosine similarity between each document content vector and the label vector corresponding to each label in the preset label knowledge base, and use the label with the largest cosine similarity as the label of the corresponding sentence, and then generate the topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document;

[0049] Preferably, the generation of the label knowledge base includes: based on a preset document subject database, performing new word discovery on the text content of all document subjects in the document subject database, and extracting a preset number of electric power industry keywords; for each electric power industry keyword, screening out document subjects containing the electric power industry keyword, and performing cluster analysis on the text content of the screened document subjects to find similar label words containing the electric power industry keyword, and then generating a corresponding label knowledge base based on the electric power industry keyword and the similar label words.

[0050] Preferably, generating a topic label similarity matrix corresponding to the target document subject according to the label corresponding to each sentence in the target document includes: calculating the TF-IDF weight between the target document subject and all labels according to the label corresponding to each sentence in the target document and a preset TF-IDF algorithm; obtaining the corresponding document topic label similarities according to the TF-IDF weight between the target document subject and all labels, and then constructing a corresponding topic label similarity matrix according to the document topic label similarities.

[0051] The present invention adopts a composite tag extraction method for power documents combining position tags and parts of speech. On the basis of automatic word segmentation, a position tag is added to each entry to form a position mark tag set. Combined with part of speech analysis, the robustness of multi-class composite tag extraction is improved. The specific steps are as follows:

[0052] Step 1: Generate a tag knowledge base

[0053] New word discovery is performed on the text content of all power document topics in the database, and a certain number of power industry keywords are extracted. These power industry keywords are business terms used to describe the key attributes of power document topics. For each power industry keyword, the topics containing the power industry keyword are screened out, and the text content of the screened power document topics is clustered. Through cluster analysis, similar label words containing the power industry keyword are found, and a power industry label knowledge base is formed. This knowledge base contains several categories of power industry label knowledge, and the structure of each knowledge is [standard label, similar label 1, ..., similar label k]. Among them, the power industry standard label is the word with the highest frequency of occurrence among all similar labels, and similar labels are other similar label words in this category except the standard label.

[0054] Among them, a method based on statistics and semantic information is used to discover new words in the text content of the power document topic: First, the Chinese word segmentation tool is used to segment the power document topic. Then, statistical methods and semantic information (such as synonyms, antonyms, etc.) are used to discover new keywords. These keywords may be words or phrases that appear frequently in a specific field but are not common in the general corpus. The steps for word segmentation and new keyword discovery of power documents are as follows:

[0055] Choose a suitable word segmentation tool: Select jieba as the word segmentation tool for Chinese processing.

[0056] Preprocess data: Clean the power document content to remove irrelevant symbols, words and punctuation.

[0057] Word segmentation operation: Use the selected segmentation to segment the cleaned text and generate a word list.

[0058] Construct word frequency statistics: count the frequency of occurrence of each word and generate a word frequency table.

[0059] New word discovery: Frequency-based discovery: Set a threshold to identify words that appear frequently in power texts (such as words with a frequency higher than a certain value). Use semantic information: Combine synonym and antonym word libraries to identify domain-specific terms, generate vectors with the BERT model, calculate the similarity between words, and discover potential new words.

[0060] The trained classifier is used to determine whether a topic contains the specified power industry keywords, thereby filtering out topics containing the power industry keywords; wherein the training and input steps of the classifier are:

[0061] Classification training steps: Data: Build a well-annotated data set, including the topics of power documents and their corresponding power industry keywords. Feature extraction: Extract the text content into feature vectors through Word2Vec; Select model: Select the classification algorithm LSTM; Train model: Use the prepared data set to train the model and adjust the hyperparameters to optimize the classification effect; Evaluate model: Use methods such as cross-validation to evaluate the accuracy and recall of the model.

[0062] Input and output: Input: feature vector of the power document topic to be classified; Output: a binary classification result indicating whether the topic contains specified power industry keywords (such as "yes" or "no").

[0063] The topics of power documents are clustered and analyzed in the following way to find similar label words: the hierarchical value clustering algorithm is used to classify topics with high similarity into the same category, and representative labels are extracted from each cluster as similar label words. In this way, topics containing similar label words and associated with keywords in the power industry can be found. The steps of hierarchical value clustering algorithm for topic classification and label word extraction are as follows:

[0064] Cluster analysis steps: Feature representation: Use the feature extraction method mentioned above to convert each power document topic into a feature vector. Calculate the similarity matrix: Use cosine similarity based on the feature vector to calculate the similarity between topics. Apply a hierarchical clustering algorithm: Select a split hierarchical clustering algorithm. Set the distance threshold for clustering to decide when to merge two topics. Generate a cluster tree: Generate a tree structure (dendrogram) through the clustering process to facilitate observation and selection of cluster results. Extract similar tag words: Select representative topics from each cluster. Common methods include selecting the cluster center or the most frequently occurring topic. Extract representative tags to form a set of similar tag words.

[0065] Step 2: Calculate the similarity matrix

[0066] Using the power industry label knowledge base, a deep learning model is trained to generate sentence vectors. The sentences in each text content are input into the model to generate the corresponding power document content vector. By calculating the cosine similarity between these content vectors and each label vector in the label knowledge base, the standard label corresponding to the label with the highest cosine similarity to the content vector is selected as the topic label of the text. In this way, each topic is associated with at least one label. By using the TF-IDF algorithm, the TF-IDF weights of the power document topic and all the power industry standard labels are calculated, the power document topic label similarity is obtained, and the topic label similarity matrix is ​​constructed.

[0067] The deep learning model is trained by using text data from the power industry label knowledge base. The pre-trained language model BERT is used to capture the semantic and contextual relationships between words, and a model that can convert input sentences into fixed-length vector representations is constructed in combination with contrastive learning training methods. In this way, each sentence input into the model will obtain the corresponding sentence vector as the text content vector. The specific implementation is as follows:

[0068] 1. Data augmentation: Use existing labels to construct sentences. For example, combine words into phrases or sentences to provide richer context for the model. These sentences are descriptions, usages, or related natural language sentences based on industry knowledge.

[0069] 2. Negative samples: During training, some negative samples are introduced, for example, some irrelevant words from the label to help the model learn the similarity and dissimilarity between words.

[0070] 3. Use the pre-trained BERT model: When using the BERT model, a large amount of text data has been learned in the pre-training phase, which enables the model to understand the contextual relationship of words. It can be fine-tuned on the labeled knowledge base. Although the original data may be wordy, BERT can enhance understanding through up and down embedding.

[0071] 4. Build synthetic training data: Generate sentences using text data from the power industry (such as technical documents, papers, reports, etc.) so that your training data will be more contextual.

[0072] The text content vector is a fixed-length vector representation generated by a deep learning model, which captures the semantics and information contained in the sentence. Cosine similarity measures the similarity between two vectors by calculating the cosine value of the angle between them. The specific calculation method is to divide the dot product of the two vectors by the product of their respective norms (lengths).

[0073] When applying the TF-IDF algorithm, we first need to calculate the TF-IDF weights of all words in the power documents, and obtain the ratio of the number of times each word appears in different power documents to the number of times it appears in the entire corpus, as well as information such as the inverse document frequency. Then, based on this weight information, we can calculate the TF-IDF weight value between each power document and all standard tag keywords, thereby obtaining the topic tag similarity matrix. The steps for calculating the TF-IDF weight between each power document and the standard tag keywords are as follows:

[0074] 1. Calculate TF (word frequency): For each power document, calculate the number of occurrences of each word. The TF value is usually defined as the ratio of the number of times a word appears in a document to the total number of words in the document.

[0075] 2. Calculate IDF inverse document frequency: Calculate the frequency of a word in all documents. The IDF value is defined as the total number of documents divided by the number of documents containing the word, and take the logarithm.

[0076] 3. Calculate TF-IDF: For each power document and each standard tag keyword, calculate the TF-IDF value.

[0077] 4. Calculate similarity: After obtaining the TF-IDF weights of each standard tag keyword, these weights can be compared with the TF-IDF weights of the power documents to calculate the similarity of the topic tags.

[0078] Through the above steps, the TF-IDF weights between the power document and all standard tag keywords can be obtained, and a similarity matrix can be constructed. In this way, the similarity of the topic tags of each power document can be quantified and analyzed.

[0079] S2, obtaining a user behavior preference matrix that records the user's historical behavior, and generating a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix;

[0080] The user behavior preference matrix is ​​constructed based on the historical behavior records of users in the database. The behavior score is multiplied by the topic tag similarity matrix to obtain the i-th value in the user tag similarity matrix, which represents the similarity between the user and a single standard tag i, and the user tag similarity matrix is ​​constructed.

[0081] Among them, constructing a user behavior preference matrix usually requires analyzing the historical behavior records of users in the database. These records may include user clicks, consultations, payments and other activities related to the power industry. First, these behaviors need to be classified and sorted, and then the content recommendation algorithm can be used to calculate the user's preference for each category or topic label, and the results can be filled into the matrix. Finally, a user behavior preference matrix reflecting the user's preference for each topic label in the power industry is obtained.

[0082] The behavior score is based on the user's participation or preference for a particular activity or hashtag. Activities related to a particular hashtag are scored based on the frequency and intensity of the user's participation in the activity or expression of preference. These scores can be used as weights when calculating the user's hashtag similarity matrix, thereby more accurately reflecting the similarity between the user and a single standard tag.

[0083] S3, calculating the cosine similarity between the matrix vectors corresponding to the pairwise user tag similarity matrices, obtaining the corresponding user collaborative similarity matrix according to the cosine similarity, calculating the weighted average of the first similarity and the second similarity between the target document topics, and obtaining the corresponding topic collaborative similarity matrix according to the weighted average;

[0084] Step 3: Calculate the collaborative similarity matrix

[0085] For all power document topics, the weighted average of the first similarity and the second similarity between each other is calculated to obtain the topic collaborative similarity matrix; the cosine similarity of the user tag similarity matrix vectors between each other is calculated to obtain the user collaborative similarity matrix.

[0086] The steps of calculating the weighted average and the cosine similarity are as follows:

[0087] First, for the similarity between two topic tags, assuming that the first similarity is similarity_1 and the second similarity is similarity_2, the weighted average can be calculated using the following formula: weighted_average = (similarity_1*weight_1+similarity_2*weight_2) / (weight_1+weight_2), where weight_1 and weight_2 are the weights of the two similarities.

[0088] For the cosine similarity of the user label similarity matrix vector, assuming that there are two vectors A and B, the cosine similarity between them can be calculated by the following formula: cosine_similarity = A·B / (||A||*||B||), where “·” represents the vector inner product, and “||A||” represents the modulus of vector A.

[0089] The first similarity and the second similarity are obtained in the following manner:

[0090] 1. Similarity_1: Content-based similarity: It can be calculated by analyzing the TF-IDF feature of the document content. The cosine similarity is used to calculate the content similarity of two topic documents.

[0091] 2. Second similarity (Similarity_2): Based on tag or topic relationship: By analyzing the tags of the document, the similarity between tags is calculated. The Jaccard similarity coefficient is used to calculate the similarity between two topic tags.

[0092] S4. Input the user collaborative similarity matrix and the topic collaborative similarity matrix into a preset composite tag extraction TagDC model, so that the composite tag extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs the multi-label confidence probability list for each target document;

[0093] Step 4: Establish a composite tag extraction TagDC model

[0094] Please refer to Figure 2 , which is a schematic diagram of the structure of the TagDC model. Figure 2 As shown in the figure, the composite tag extraction TagDC model is a word learning enhanced CNN-capsule module, which includes two parts TagDC-DL and TagDC-CF. TagDC-DL word representation learning is to enhance the semantic expression of the original word vector by combining the word vector with the surrounding context information; TagDC-CF description representation learning is to use the kernel in the convolution layer to slide each description to extract local features and generate feature maps; label probability calculation is to calculate the length of each label category based on the capsule network to obtain a multi-label confidence probability list for each object. Then several labels with the highest confidence probability are assigned to the current object.

[0095] S5. According to the multi-label confidence probability list, select several label combinations from the multi-label confidence probability list to generate a composite label for the target document.

[0096] Preferably, according to the multi-label confidence probability list, a plurality of label combinations are selected from the multi-label confidence probability list to generate a composite label for the target document, including: according to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted in descending order, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate a composite label for the target document.

[0097] Step 5: Calculate the final confidence probability list of labels

[0098] According to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted from high to low, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate the composite label of the target document. By sorting the probability values ​​in the final confidence probability list and using semantic similarity and deep semantic features, the ability to extract composite labels of power documents can be achieved more accurately and efficiently.

[0099] Embodiment 2

[0100] Please refer to Figure 3 , is a schematic diagram of the structure of a device for generating a composite label of a document provided by an embodiment of the present invention, the device comprising: a topic label similarity matrix generating module, a user label similarity matrix generating module, a collaborative similarity matrix generating module, a multi-label confidence probability list generating module and a composite label generating module;

[0101] The topic label similarity matrix generation module is used to obtain the target document, generate the document content vector corresponding to each sentence in the target document, and calculate the cosine similarity between each document content vector and the label vector corresponding to each label in the preset label knowledge base, and use the label with the largest cosine similarity as the label of the corresponding sentence, and then generate the topic label similarity matrix corresponding to the subject of the target document according to the label corresponding to each sentence in the target document;

[0102] The user tag similarity matrix generation module is used to obtain a user behavior preference matrix that records the user's historical behavior, and generate a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix;

[0103] The collaborative similarity matrix generation module is used to calculate the cosine similarity between matrix vectors corresponding to the pairwise user tag similarity matrices, obtain the corresponding user collaborative similarity matrix according to the cosine similarity, calculate the weighted average of the first similarity and the second similarity between the target document topics, and obtain the corresponding topic collaborative similarity matrix according to the weighted average;

[0104] The multi-label confidence probability list generation module is used to input the user collaborative similarity matrix and the topic collaborative similarity matrix into a preset composite label extraction TagDC model, so that the composite label extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs the multi-label confidence probability list for each target document;

[0105] The composite label generation module is used to select several label combinations from the multi-label confidence probability list to generate a composite label for the target document according to the multi-label confidence probability list.

[0106] Preferably, the generation of the label knowledge base includes: based on a preset document subject database, performing new word discovery on the text content of all document subjects in the document subject database, and extracting a preset number of electric power industry keywords; for each electric power industry keyword, screening out document subjects containing the electric power industry keyword, and performing cluster analysis on the text content of the screened document subjects to find similar label words containing the electric power industry keyword, and then generating a corresponding label knowledge base based on the electric power industry keyword and the similar label words.

[0107] Preferably, generating a topic label similarity matrix corresponding to the target document subject according to the label corresponding to each sentence in the target document includes: calculating the TF-IDF weight between the target document subject and all labels according to the label corresponding to each sentence in the target document and a preset TF-IDF algorithm; obtaining the corresponding document topic label similarities according to the TF-IDF weight between the target document subject and all labels, and then constructing a corresponding topic label similarity matrix according to the document topic label similarities.

[0108] Preferably, according to the multi-label confidence probability list, a plurality of label combinations are selected from the multi-label confidence probability list to generate a composite label for the target document, including: according to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted in descending order, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate a composite label for the target document.

[0109] It should be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art may understand and implement it without paying any creative effort.

[0110] Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0111] Embodiment 3

[0112] Accordingly, an embodiment of the present invention provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the method for generating document composite tags as described in the above-mentioned embodiment of the invention when executing the computer program.

[0113] The electronic device may be a computing device such as a desktop computer, a notebook, a palm computer, a cloud server, etc. The device may include, but is not limited to, a processor and a memory.

[0114] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the device, and various interfaces and lines are used to connect various parts of the entire device.

[0115] Embodiment 4

[0116] Accordingly, an embodiment of the present invention provides a storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the method for generating document composite tags described in the above-mentioned embodiment of the invention.

[0117] The memory can be used to store the computer program, and the processor realizes various functions of the device by running or executing the computer program stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0118] The storage medium is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, the steps of each method embodiment described above can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.

[0119] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.

Claims

1. A method for generating a composite tag of a document, characterized in that: include: Obtain the target document, generate the document content vector corresponding to each sentence in the target document, and calculate the cosine similarity between each document content vector and the label vector corresponding to each label in the preset label knowledge base, and use the label with the largest cosine similarity as the label of the corresponding sentence, and then generate the topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document; Obtaining a user behavior preference matrix that records the user's historical behavior, and generating a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix; Calculate the cosine similarity between the matrix vectors corresponding to the pairwise user tag similarity matrices, obtain the corresponding user collaborative similarity matrix according to the cosine similarity, calculate the weighted average of the first similarity and the second similarity between the target document topics, and obtain the corresponding topic collaborative similarity matrix according to the weighted average; Input the user collaborative similarity matrix and the topic collaborative similarity matrix into a preset composite tag extraction TagDC model, so that the composite tag extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs the multi-label confidence probability list for each target document; According to the multi-label confidence probability list, several label combinations are selected from the multi-label confidence probability list to generate a composite label for the target document.

2. The method for generating a document composite tag according to claim 1, characterized in that: The generation of the label knowledge base includes: According to a preset document subject database, new word discovery is performed on the text content of all document subjects in the document subject database to extract a preset number of power industry keywords; For each electric power industry keyword, the document topics containing the electric power industry keyword are screened out, and the text content of the screened document topics is clustered and analyzed to find similar tag words containing the electric power industry keyword, and then a corresponding tag knowledge base is generated based on the electric power industry keyword and the similar tag words.

3. The method for generating a composite document tag according to claim 1, wherein: The step of generating a topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document includes: According to the labels corresponding to each sentence in the target document and the preset TF-IDF algorithm, the TF-IDF weight between the target document topic and all labels is calculated; According to the TF-IDF weights between the target document topic and all tags, the corresponding document topic tag similarities are obtained, and then the corresponding topic tag similarity matrix is ​​constructed according to the document topic tag similarities.

4. The method for generating a composite document tag according to claim 1, wherein: The step of selecting a plurality of label combinations from the multi-label confidence probability list to generate a composite label of the target document according to the multi-label confidence probability list includes: According to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted from high to low, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate a composite label for the target document.

5. A device for generating a composite tag of a document, characterized in that: include: Topic tag similarity matrix generation module, user tag similarity matrix generation module, collaborative similarity matrix generation module, multi-tag confidence probability list generation module and composite tag generation module; The topic label similarity matrix generation module is used to obtain the target document, generate the document content vector corresponding to each sentence in the target document, and calculate the cosine similarity between each document content vector and the label vector corresponding to each label in the preset label knowledge base, and use the label with the largest cosine similarity as the label of the corresponding sentence, and then generate the topic label similarity matrix corresponding to the subject of the target document according to the label corresponding to each sentence in the target document; The user tag similarity matrix generation module is used to obtain a user behavior preference matrix that records the user's historical behavior, and generate a corresponding user tag similarity matrix according to the user behavior preference matrix and the topic tag similarity matrix; The collaborative similarity matrix generation module is used to calculate the cosine similarity between matrix vectors corresponding to the pairwise user tag similarity matrices, obtain the corresponding user collaborative similarity matrix according to the cosine similarity, calculate the weighted average of the first similarity and the second similarity between the target document topics, and obtain the corresponding topic collaborative similarity matrix according to the weighted average; The multi-label confidence probability list generation module is used to input the user collaborative similarity matrix and the topic collaborative similarity matrix into a preset composite label extraction TagDC model, so that the composite label extraction TagDC model generates a multi-label confidence probability list for each target document according to the user collaborative similarity matrix and the topic collaborative similarity matrix, and outputs the multi-label confidence probability list for each target document; The composite label generation module is used to select several label combinations from the multi-label confidence probability list to generate a composite label for the target document according to the multi-label confidence probability list.

6. The device for generating a composite document tag according to claim 5, characterized in that: The generation of the label knowledge base includes: According to a preset document subject database, new word discovery is performed on the text content of all document subjects in the document subject database to extract a preset number of power industry keywords; For each electric power industry keyword, the document topics containing the electric power industry keyword are screened out, and the text content of the screened document topics is clustered and analyzed to find similar tag words containing the electric power industry keyword, and then a corresponding tag knowledge base is generated based on the electric power industry keyword and the similar tag words.

7. The device for generating a composite document tag according to claim 5, characterized in that: The step of generating a topic label similarity matrix corresponding to the topic of the target document according to the labels corresponding to each sentence in the target document includes: According to the labels corresponding to each sentence in the target document and the preset TF-IDF algorithm, the TF-IDF weight between the target document topic and all labels is calculated; According to the TF-IDF weights between the target document topic and all tags, the corresponding document topic tag similarities are obtained, and then the corresponding topic tag similarity matrix is ​​constructed according to the document topic tag similarities.

8. The device for generating a composite document tag according to claim 5, characterized in that: The step of selecting a plurality of label combinations from the multi-label confidence probability list to generate a composite label of the target document according to the multi-label confidence probability list includes: According to the multi-label confidence probability list, the confidence probability of each label in the multi-label confidence probability list is sorted from high to low, and then according to the sorting of the confidence probabilities, a preset number of labels are selected from the multi-label confidence probability list, and the selected labels are combined to generate a composite label for the target document.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the method for generating a document composite tag as claimed in any one of claims 1 to 4 when executing the computer program.

10. A storage medium, characterized in that: The storage medium includes a stored computer program, wherein when the computer program is executed, the device where the storage medium is located is controlled to execute the method for generating a document composite tag according to any one of claims 1 to 4.