Industrial chain construction method, equipment and storage medium

By training the word association model and expanding the associated vocabulary of industrial chain recognition words, the problem of limited identification words manually constructing industrial chain recognition words is solved, the accuracy and comprehensiveness of industrial chain construction are improved, and a more accurate correlation score between industrial chain recognition words and enterprise data is achieved.

CN114595330BActive Publication Date: 2025-05-13JIANGSU FENGYUN TECH SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210254413.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-05-13
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The limited number of recognition words for industrial chain artificially constructed, resulting in incomplete coverage of recognition words for industrial chain, incomplete and inaccurate data results of fuzzy matching, and problems with the degree of matching and correlation.

Method used

By obtaining manually defined industrial chain identification words and enterprise sample data, using existing data sets to train word association models, generate the associated vocabulary of each vocabulary, determine the correlation matrix and correlation degree between the industrial chain vocabulary and enterprise vocabulary, and expand the correlation vocabulary of the industrial chain identification words.

Benefits of technology

The problem of incomplete coverage of identification words in the industrial chain is solved, and the accuracy and comprehensiveness of industrial chain construction is improved. The numerical correlation score shows the correlation between the data results and the industrial chain marker words, which improves the accuracy of the correlation between the identification words and enterprise data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595330B_ABST
    Figure CN114595330B_ABST
Patent Text Reader

Abstract

The present application relates to an industrial chain construction method, device and storage medium, belonging to the field of computer technology, which includes inputting industrial chain identification words into a pre-trained word association model to obtain a vocabulary set; determining the correlation matrix between each industrial chain word and enterprise word in the vocabulary set; determining the degree of correlation between the industrial chain word and the enterprise word based on the matrix information of the correlation matrix; for a group of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, setting the label of the enterprise sample data to which the enterprise word belongs to the industrial chain identification word corresponding to the industrial chain word; using the trained word association model to construct the industrial chain of the enterprise data to be processed; the accuracy and comprehensiveness of the industrial chain construction can be improved; and the problems of incompleteness and inaccuracy of fuzzy matching data results, high and low matching degree of matching data results and correlation between each data result can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to an industrial chain construction method, equipment and storage medium, and belongs to the field of computer technology. [Background technology]

[0002] The industrial chain of an enterprise can often reflect the enterprise's operating capabilities. Based on this, users often need to identify the corresponding industrial chain based on the enterprise data.

[0003] The traditional industrial chain construction method includes: the user labels the sample enterprise data obtained to obtain the corresponding industrial chain identification words; uses the sample enterprise data and the corresponding industrial chain identification words to train the industrial chain construction model to identify the industrial chain corresponding to the enterprise data.

[0004] However, manual labeling may result in incomplete data labeling, which will lead to poor recognition results for industrial chain construction models. [Summary of the invention]

[0005] This application provides an industrial chain construction method, device and storage medium, which can solve the problem of limited artificial industrial chain identification words, resulting in incomplete coverage of industrial chain identification words, and can also solve the problem of incomplete and inaccurate fuzzy matching data results, the problem of matching degree of matching data results and the correlation between data results. This application provides the following technical solutions:

[0006] On the one hand, a method for constructing an industrial chain is provided, the method comprising:

[0007] Obtain manually defined industry chain identification words;

[0008] Obtain enterprise sample data;

[0009] Using an existing data set to train a word association model, the word association model is used to generate associated words for each word in the existing data set;

[0010] Processing the enterprise sample data to obtain enterprise vocabulary;

[0011] Inputting the industrial chain identification word into the word association model to obtain a vocabulary set of the industrial chain identification word;

[0012] Determine a correlation matrix between each industrial chain vocabulary in the vocabulary set and the enterprise vocabulary;

[0013] Determining the degree of association between the industrial chain vocabulary and the enterprise vocabulary based on the matrix information of the association matrix;

[0014] For a group of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, setting the label of the enterprise sample data to which the enterprise word belongs to the industrial chain identification word corresponding to the industrial chain word;

[0015] Determine whether the identified industrial chain identification words meet the preset training requirements;

[0016] In the case where the identified industrial chain identification words meet the preset training requirements, the trained word association model is used to construct the industrial chain of the enterprise data to be processed.

[0017] Optionally, the using an existing data set to train a word association model includes:

[0018] Input the existing data set into a pre-built word vector generation model to obtain a word vector space;

[0019] Inputting the existing data set into the pre-trained BERT model to obtain the word expression distribution of the existing data set;

[0020] Obtaining a semantic modification operation based on the word expression distribution input;

[0021] Correcting the contextual semantics of the word vector space based on the semantic correction operation to obtain a corrected word vector space;

[0022] The word vector generation model is corrected using the corrected word vector space to obtain the word association model.

[0023] Optionally, the matrix information includes at least two types; the determining the degree of association between the industrial chain vocabulary and the enterprise vocabulary based on the matrix information of the association matrix includes:

[0024] Get the information weight corresponding to each matrix information;

[0025] The sum of the products of the quantized values ​​of each matrix information and the corresponding information weight is calculated to obtain the correlation degree.

[0026] Optionally, the matrix information includes at least two of the following information:

[0027] The degree of association in the degree of association matrix;

[0028] The variance of the association matrix;

[0029] The mean of the correlation matrix;

[0030] the dimension of the affinity matrix; and

[0031] The part of speech corresponding to each correlation degree in the correlation matrix.

[0032] Optionally, the method further comprises:

[0033] When there are at least two groups of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, obtaining the segmented search words of the industrial chain identification words and the correlation weights of the segmented search words, wherein the correlation weights are used to indicate the correlation degree between the segmented search words and different enterprise words;

[0034] A group of industrial chain words and enterprise words is selected from at least two groups of industrial chain words and enterprise words based on the association weights.

[0035] Optionally, selecting a group of industrial chain words and enterprise words from at least two groups of industrial chain words and enterprise words based on the association weights includes:

[0036] Get the correlation between each enterprise term and the segmented search term;

[0037] Input each enterprise vocabulary and segmented search term into a preset regular expression to obtain a matching result; the matching result includes a match and a mismatch;

[0038] Calculate the association score of each group of industrial chain words and enterprise words by combining the association weight, the association degree and the matching result;

[0039] A group of industrial chain vocabulary and enterprise vocabulary with the highest correlation score is selected from at least two groups of industrial chain vocabulary and enterprise vocabulary.

[0040] Optionally, the processing of the enterprise sample data to obtain enterprise vocabulary includes:

[0041] Performing word segmentation processing on the enterprise sample data to obtain the enterprise vocabulary;

[0042] or,

[0043] Performing word segmentation and grammatical analysis on the enterprise sample data; removing grammatically incorrect enterprise sample data to obtain the enterprise vocabulary;

[0044] or,

[0045] The enterprise sample data is segmented and negative expressions in the enterprise sample data are removed to obtain the enterprise vocabulary.

[0046] Optionally, the method further comprises:

[0047] In the case that the identified industrial chain identification word does not meet the preset training requirement, the word association model is adjusted, and the step of training the word association model using the existing data set is triggered.

[0048] On the other hand, an electronic device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement the above-mentioned industrial chain construction method.

[0049] On the other hand, a computer-readable storage medium stores a program, and when the program is executed by a processor, it is used to implement the above-mentioned industrial chain construction method.

[0050] The beneficial effects of the present application include at least: obtaining manually defined industrial chain identification words; obtaining enterprise sample data; using an existing data set to train a word association model, the word association model is used to generate associated words for each word in the existing data set; processing the enterprise sample data to obtain an enterprise vocabulary; inputting the industrial chain identification words into the word association model to obtain a vocabulary set of industrial chain identification words; determining the correlation matrix between each industrial chain word and the enterprise vocabulary in the vocabulary set; determining the degree of correlation between the industrial chain vocabulary and the enterprise vocabulary based on the matrix information of the correlation matrix; for a group of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, setting the label of the enterprise sample data to which the enterprise vocabulary belongs to the industrial chain identification word corresponding to the industrial chain vocabulary; determining whether the identified industrial chain identification words meet the preset training requirements; when the identified industrial chain identification words meet the preset training requirements, using the trained word association model to construct the industrial chain of the enterprise data to be processed; the problem of limited artificially constructed industrial chain identification words resulting in incomplete coverage of industrial chain identification words can be solved; since the associated words of the industrial chain identification words can be expanded through the word association model, the industrial chain identification words can cover more enterprise data, thereby improving the accuracy and comprehensiveness of the industrial chain construction.

[0051] At the same time, the correlation score can be quantified to show the correlation between the data results and the corresponding industrial chain marker words, and to improve the accuracy of determining the correlation between the industrial chain identification words and the enterprise data.

[0052] In addition, when there are at least two groups of industrial chain vocabulary and enterprise vocabulary whose correlation degree is greater than or equal to a preset degree threshold, by obtaining the segmented search terms of the industrial chain identification terms and the correlation weights of the segmented search terms; and selecting a group of industrial chain vocabulary and enterprise vocabulary from at least two groups of industrial chain vocabulary and enterprise vocabulary based on the correlation weights, a more accurate correlation relationship between the industrial chain vocabulary and the enterprise vocabulary can be further obtained, thereby improving the accuracy of industrial chain construction.

[0053] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0054] Figure 1 is a flow chart of an industrial chain construction method provided by an embodiment of the present application;

[0055] Figure 2 is a block diagram of an industrial chain construction device provided by an embodiment of the present application;

[0056] Figure 3 It is a block diagram of an electronic device provided by an embodiment of the present application. [Specific implementation method]

[0057] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application but are not intended to limit the scope of the present application.

[0058] First, several terms involved in this application are introduced.

[0059] Natural Language Processing (NLP): It is a method to achieve effective communication between people and computers using natural language.

[0060] Word vector: also known as word embedded NLP, a general term for a set of language modeling and feature learning techniques in which words or phrases from the vocabulary are mapped to vectors of real numbers. Conceptually, it involves mathematical embedding from a one-dimensional space for each word to a continuous vector space with lower dimensions. Methods for generating such mappings include neural networks, dimensionality reduction of word co-occurrence matrices, probabilistic models, interpretable knowledge base methods, and explicit representation of terms in the context in which the words appear.

[0061] Chinese word segmentation: word segmentation is the process of recombining continuous character sequences into word sequences according to certain specifications. Chinese word segmentation refers to the word segmentation that exists because of the particularity of Chinese in basic grammar.

[0062] Bidirectional Encoder Representations from Transformers (BERT): is an autoencoder language model that can extract relational features at multiple levels to more fully reflect the semantics of a sentence, and can also obtain word meanings based on the context of the sentence to avoid ambiguity.

[0063] Optionally, the present application uses the industrial chain construction method provided in each embodiment as an example for explanation in an electronic device, where the electronic device is a terminal or a server. The terminal may be a mobile phone, a computer, a tablet computer, etc. This embodiment does not limit the type of electronic device.

[0064] Figure 1This is a flow chart of an industrial chain construction method provided by an embodiment of the present application, which method includes at least the following steps:

[0065] Step 101, obtaining manually defined industrial chain identification words.

[0066] The electronic device displays a vocabulary input box for the user to input an industrial chain identification word. The industrial chain identification word may be one or at least two, and this embodiment does not limit the number of industrial chain identification words.

[0067] Illustratively, the industrial chain identification words are set based on the industrial chain construction requirements. For example, the industrial chain identification words can be artificial intelligence, new energy, big data, etc. This embodiment does not limit the implementation method of the industrial chain identification words.

[0068] Step 102, obtaining enterprise sample data.

[0069] The enterprise sample data may be read from an industrial and commercial registration agency, or may be input by a user. This embodiment does not limit the source of the enterprise sample data.

[0070] Optionally, each piece of enterprise sample data includes but is not limited to at least one of the following: business scope, product type, enterprise name and registered address, etc. This embodiment does not limit the content included in the enterprise sample data.

[0071] Step 103: Use the existing data set to train a word association model, where the word association model is used to generate associated words for each word in the existing data set.

[0072] The existing data set may be a data set in a field related to the industrial chain identification term, or may also include data sets in other fields. This embodiment does not limit the type of the existing data set.

[0073] In one example, a word association model is trained using an existing data set, including: inputting the existing data set into a pre-built word vector generation model to obtain a word vector space; inputting the existing data set into a pre-trained BERT model to obtain a word expression distribution of the existing data set; obtaining a semantic correction operation based on the word expression distribution input; correcting the contextual semantics of the word vector space based on the semantic correction operation to obtain a corrected word vector space; and using the corrected word vector space to correct the word vector generation model to obtain a word association model.

[0074] The word vector generation model is different from the BERT model. The word vector generation model includes but is not limited to: word2vec, or co-occurrence matrix, or GloVe (Global Vectors for Word Representation) model. This embodiment does not limit the implementation method of the word vector generation model.

[0075] In this embodiment, by using the word representation distribution generated by the BERT model to correct the word vector generation model, the model performance of the word vector generation model can be improved, thereby improving the model performance of the word association model.

[0076] In other examples, an existing data set can also be input into a pre-built word vector generation model, the word vectors generated by the word vector generation model can be manually modified, and the word vector generation model can be corrected using the corrected word vector space to obtain a word association model. This embodiment does not limit the generation method of the word association model.

[0077] Step 104, processing the enterprise sample data to obtain enterprise vocabulary.

[0078] Optionally, the enterprise sample data is processed to obtain the enterprise vocabulary, including: performing word segmentation processing on the enterprise sample data to obtain the enterprise vocabulary; or, performing word segmentation processing and grammatical analysis on the enterprise sample data; removing grammatically incorrect enterprise sample data to obtain the enterprise vocabulary; or, performing word segmentation processing on the enterprise sample data and removing negative expressions in the enterprise sample data to obtain the enterprise vocabulary.

[0079] The word segmentation processing methods include but are not limited to: a method based on string matching, a method based on semantic understanding, or a method based on frequency statistics of combinations of various characters. This embodiment does not limit the word segmentation processing method.

[0080] The syntax analysis method includes, but is not limited to, a recursive descent analysis method or an operator precedence analysis method. This embodiment does not limit the syntax analysis method.

[0081] The method of removing negative expressions includes: identifying blacklist words in the enterprise sample data and removing the blacklist words. The blacklist words include negative expressions, and the blacklist words are pre-stored in the electronic device. The negative expressions include but are not limited to: not, not including, or not, etc. This embodiment does not limit the implementation method of the negative expression.

[0082] In this embodiment, the accuracy of the coverage data results can be improved by performing word segmentation processing and removing negative expressions.

[0083] Step 105 , input the industrial chain identification words into the word association model to obtain a vocabulary set of industrial chain identification words.

[0084] Due to the limitations of personal knowledge and experience, manual data processing is prone to incomplete data labeling. Based on this, in this embodiment, by associating the industrial chain identification words, the comprehensiveness of the coverage data results can be improved.

[0085] Step 106: determine the correlation matrix between each industrial chain vocabulary and enterprise vocabulary in the vocabulary set.

[0086] In one example, based on the word co-occurrence of industrial chain vocabulary and enterprise vocabulary, the correlation degree between industrial chain vocabulary and enterprise vocabulary is determined, so as to obtain the correlation matrix between each industrial chain vocabulary and enterprise vocabulary.

[0087] Among them, word co-occurrence is obtained by statistically analyzing the co-occurrence distribution characteristics of words in a document collection.

[0088] Specifically, each correlation degree in the correlation matrix is ​​calculated by the following formula:

[0089] C il =n il / (n i +n l -n il )

[0090] Among them, n i Represents the industry chain vocabulary k i The document frequency, n l Indicates corporate vocabulary k l The document frequency, n il Indicates that it also contains industry chain vocabulary k i With corporate vocabulary l The document frequency.

[0091] Step 107: Determine the degree of association between the industrial chain vocabulary and the enterprise vocabulary based on the matrix information of the association matrix.

[0092] Optionally, the matrix information includes at least two types; accordingly, based on the matrix information of the correlation matrix, the correlation degree between the industrial chain vocabulary and the enterprise vocabulary is determined, including: obtaining the information weight corresponding to each matrix information; calculating the sum of the products of the quantized values ​​of each matrix information and the corresponding information weight to obtain the correlation degree.

[0093] Illustratively, the matrix information includes at least two of the following information: the degree of association in the association matrix; the variance of the association matrix; the mean of the association matrix; the dimension of the association matrix; and the part of speech corresponding to each degree of association in the association matrix.

[0094] Among them, the information weights corresponding to different matrix information are pre-stored in the electronic device, and the values ​​of the information weights belong to [0, 1]. Optionally, the information weight of the association is greater than the information weight of the variance, greater than the information weight of the mean, greater than the information weight of the dimension, and greater than the information weight of the part of speech. This embodiment does not limit the setting method of the information weights.

[0095] In the case where the matrix information includes only one type, the degree of association can be directly handled as the degree of association.

[0096] Step 108 , for a group of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, the label of the enterprise sample data to which the enterprise word belongs is set to the industrial chain identification word corresponding to the industrial chain word.

[0097] When industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold are grouped together, the label of the enterprise sample data to which the enterprise word belongs is directly set to the industrial chain identification word corresponding to the industrial chain word.

[0098] When there are at least two groups of industrial chain vocabulary and enterprise vocabulary whose correlation degree is greater than or equal to a preset degree threshold, before setting the label of the enterprise sample data to which the enterprise vocabulary belongs to the industrial chain identification word corresponding to the industrial chain vocabulary, it also includes: obtaining the segmented search words of the industrial chain identification word and the correlation weight of the segmented search words; and selecting a group of industrial chain vocabulary and enterprise vocabulary from at least two groups of industrial chain vocabulary and enterprise vocabulary based on the correlation weight.

[0099] The correlation weight is used to indicate the correlation degree between the segmented search terms and different enterprise terms.

[0100] The segmented search terms belong to the subcategories of the industry chain identification terms. For example, if the industry chain identification term is artificial intelligence, then the segmented search terms can be subcategories of artificial intelligence such as smart home or autonomous driving.

[0101] Specifically, the electronic device provides an input box for segmented search terms and associated weights. After the electronic device obtains the segmented search terms and the associated weights, it selects a group of industrial chain terms and enterprise terms from at least two groups of industrial chain terms and enterprise terms based on the associated weights, including: obtaining the correlation between each enterprise term and the segmented search term; inputting each enterprise term and segmented search term into a preset regular expression to obtain a matching result; the matching result includes a match and a mismatch; combining the associated weights, the correlation degrees and the matching results to calculate the association scores of each group of industrial chain terms and enterprise terms; and selecting a group of industrial chain terms and enterprise terms with the highest association scores from at least two groups of industrial chain terms and enterprise terms.

[0102] The calculation method of the correlation between the enterprise vocabulary and the segmented search terms is referred to in step 106 , and will not be described in detail in this embodiment.

[0103] The regular expression is pre-stored in the electronic device and is used to indicate whether the enterprise vocabulary and the segmented search term match the preset language model. If the enterprise vocabulary and the segmented search term indicated by the regular expression do not match the preset language model, the output result is 0; if the enterprise vocabulary and the segmented search term indicated by the regular expression match the preset language model, the output result is 1.

[0104] The association score of each group of industrial chain vocabulary and enterprise vocabulary is calculated by combining the association weight, association degree and matching result, including: calculating the weighted value between the association weight and the association degree, and taking the weighted value and the quantitative value of the matching result as the association score.

[0105] Step 109 , determining whether the identified industrial chain identification words meet the preset training requirements.

[0106] After the electronic device obtains a set of industrial chain vocabulary and enterprise vocabulary, it outputs the enterprise sample data and the corresponding industrial chain identification words to which the enterprise vocabulary belongs, so as to prompt the user to input an operation instruction whether it meets the preset training requirements, such as: outputting a selection box for the user to select whether it meets the preset training requirements or does not meet the preset training requirements; if the operation instruction is that it meets the preset training requirements, it is determined that the identified industrial chain identification words meet the preset training requirements, and step 110 is executed; if the operation instruction is that it does not meet the preset training requirements, it is determined that the identified industrial chain identification words do not meet the preset training requirements. At this time, the word association model is adjusted, and the step of using the existing data set to train the word association model is triggered, that is, step 103 is executed again.

[0107] Step 110 , when the identified industrial chain identification words meet the preset training requirements, the trained word association model is used to construct the industrial chain of the enterprise data to be processed.

[0108] Specifically, after the enterprise data to be processed is processed, the correlation between the obtained enterprise vocabulary and each industrial chain vocabulary in the vocabulary set corresponding to the industrial chain identification word is calculated one by one to obtain a correlation matrix; then, based on the matrix information of the correlation matrix, the correlation degree between the industrial chain vocabulary and the enterprise vocabulary is determined; the industrial chain identification word corresponding to the industrial chain vocabulary with the highest correlation degree with the enterprise vocabulary is used as the constructed industrial chain.

[0109] In summary, the industrial chain construction method provided in this embodiment obtains manually defined industrial chain identification words; obtains enterprise sample data; uses an existing data set to train a word association model, and the word association model is used to generate associated words for each word in the existing data set; processes the enterprise sample data to obtain enterprise vocabulary; inputs the industrial chain identification words into the word association model to obtain a vocabulary set of industrial chain identification words; determines the correlation matrix between each industrial chain word and the enterprise vocabulary in the vocabulary set; determines the degree of correlation between the industrial chain vocabulary and the enterprise vocabulary based on the matrix information of the correlation matrix; for a group of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, sets the label of the enterprise sample data to which the enterprise vocabulary belongs to the industrial chain identification word corresponding to the industrial chain vocabulary; determines whether the identified industrial chain identification words meet the preset training requirements; when the identified industrial chain identification words meet the preset training requirements, uses the trained word association model to construct the industrial chain of the enterprise data to be processed; can solve the problem that the industrial chain identification words are limited in artificial construction, resulting in incomplete coverage of the industrial chain identification words; because the associated vocabulary of the industrial chain identification words can be expanded through the word association model, the industrial chain identification words can cover more enterprise data, thereby improving the accuracy and comprehensiveness of the industrial chain construction.

[0110] At the same time, the correlation score can be quantified to show the correlation between the data results and the corresponding industrial chain marker words, and to improve the accuracy of determining the correlation between the industrial chain identification words and the enterprise data.

[0111] In addition, when there are at least two groups of industrial chain vocabulary and enterprise vocabulary whose correlation degree is greater than or equal to a preset degree threshold, by obtaining the segmented search terms of the industrial chain identification terms and the correlation weights of the segmented search terms; and selecting a group of industrial chain vocabulary and enterprise vocabulary from at least two groups of industrial chain vocabulary and enterprise vocabulary based on the correlation weights, a more accurate correlation relationship between the industrial chain vocabulary and the enterprise vocabulary can be further obtained, thereby improving the accuracy of industrial chain construction.

[0112] Figure 2 1 is a block diagram of an industrial chain construction device provided by an embodiment of the present application. The device includes at least the following modules: a recognition word acquisition module 210, a sample acquisition module 220, a model training module 230, a sample processing module 240, a recognition word association module 250, a vocabulary association module 260, a degree calculation module 270, a label setting module 280, a result determination module 290 and an industrial chain construction module 291.

[0113] An identification word acquisition module 210 is used to acquire manually defined industry chain identification words;

[0114] The sample acquisition module 220 is used to acquire enterprise sample data;

[0115] A model training module 230, for training a word association model using an existing data set, wherein the word association model is used to generate associated words for each word in the existing data set;

[0116] A sample processing module 240 is used to process the enterprise sample data to obtain enterprise vocabulary;

[0117] The identification word association module 250 is used to input the industrial chain identification word into the word association model to obtain a vocabulary set of the industrial chain identification word;

[0118] A vocabulary association module 260, used to determine a correlation matrix between each industrial chain vocabulary in the vocabulary set and the enterprise vocabulary;

[0119] A degree calculation module 270, for determining the degree of association between the industrial chain vocabulary and the enterprise vocabulary based on matrix information of the association matrix;

[0120] The label setting module 280 is used to set the label of the enterprise sample data to which the enterprise vocabulary belongs to the industrial chain identification word corresponding to the industrial chain vocabulary for a group of industrial chain vocabulary and enterprise vocabulary whose correlation degree is greater than or equal to a preset degree threshold;

[0121] A result determination module 290 is used to determine whether the identified industrial chain identification words meet the preset training requirements;

[0122] The industrial chain construction module 291 is used to construct the industrial chain of the enterprise data to be processed using the trained word association model when the identified industrial chain identification words meet the preset training requirements.

[0123] For relevant details, refer to the above method embodiment.

[0124] It should be noted that: the industrial chain construction device provided in the above embodiment only uses the division of the above functional modules as an example when constructing the industrial chain. In actual applications, the above functional distribution can be completed by different functional modules as needed, that is, the internal structure of the industrial chain construction device is divided into different functional modules to complete all or part of the functions described above. In addition, the industrial chain construction device provided in the above embodiment and the industrial chain construction method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0125] Figure 3 3 is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 301 and a memory 302.

[0126] The processor 301 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 301 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 301 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0127] The memory 302 may include one or more computer-readable storage media, which may be non-transitory. The memory 302 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 302 is used to store at least one instruction, which is used to be executed by the processor 301 to implement the industrial chain construction method provided in the method embodiment of the present application.

[0128] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 301, the memory 302 and the peripheral device interface may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface via a bus, a signal line or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply.

[0129] Of course, the electronic device may also include fewer or more components, which is not limited in this embodiment.

[0130] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the industrial chain construction method of the above method embodiment.

[0131] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the industrial chain construction method of the above method embodiment.

[0132] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0133] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for constructing an industrial chain, characterized in that: The method comprises: Obtain manually defined industry chain identification words; Obtain enterprise sample data; Using an existing data set to train a word association model, the word association model is used to generate associated words for each word in the existing data set; Processing the enterprise sample data to obtain enterprise vocabulary; Inputting the industrial chain identification word into the word association model to obtain a vocabulary set of the industrial chain identification word; Determine a correlation matrix between each industrial chain vocabulary in the vocabulary set and the enterprise vocabulary; Determining the degree of association between the industrial chain vocabulary and the enterprise vocabulary based on the matrix information of the association matrix; For a group of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, setting the label of the enterprise sample data to which the enterprise word belongs to the industrial chain identification word corresponding to the industrial chain word; Determine whether the identified industrial chain identification words meet the preset training requirements; In the case where the identified industrial chain identification words meet the preset training requirements, the trained word association model is used to construct the industrial chain of the enterprise data to be processed.

2. The method according to claim 1, characterized in that The method of using an existing data set to train a word association model includes: Input the existing data set into a pre-built word vector generation model to obtain a word vector space; Inputting the existing data set into the pre-trained BERT model to obtain the word expression distribution of the existing data set; Obtaining a semantic modification operation based on the word expression distribution input; Correcting the contextual semantics of the word vector space based on the semantic correction operation to obtain a corrected word vector space; The word vector generation model is corrected using the corrected word vector space to obtain the word association model.

3. The method according to claim 1, characterized in that The matrix information includes at least two types; the matrix information based on the correlation matrix determines the correlation degree between the industrial chain vocabulary and the enterprise vocabulary, including: Get the information weight corresponding to each matrix information; The sum of the products of the quantized values ​​of each matrix information and the corresponding information weight is calculated to obtain the correlation degree.

4. The method according to claim 3, characterized in that The matrix information includes at least two of the following information: The degree of association in the degree of association matrix; The variance of the association matrix; The mean of the correlation matrix; the dimension of the affinity matrix; and The part of speech corresponding to each correlation degree in the correlation matrix.

5. The method according to claim 1, characterized in that The method further comprises: When there are at least two groups of industrial chain words and enterprise words whose correlation degree is greater than or equal to a preset degree threshold, obtaining the segmented search words of the industrial chain identification words and the correlation weights of the segmented search words, wherein the correlation weights are used to indicate the correlation degree between the segmented search words and different enterprise words; A group of industrial chain words and enterprise words is selected from at least two groups of industrial chain words and enterprise words based on the association weights.

6. The method according to claim 5, characterized in that The step of selecting a group of industrial chain words and enterprise words from at least two groups of industrial chain words and enterprise words based on the association weights includes: Get the correlation between each enterprise term and the segmented search term; Input each enterprise vocabulary and segmented search term into a preset regular expression to obtain a matching result; the matching result includes a match and a mismatch; Calculate the association score of each group of industrial chain words and enterprise words by combining the association weight, the association degree and the matching result; A group of industrial chain vocabulary and enterprise vocabulary with the highest correlation score is selected from at least two groups of industrial chain vocabulary and enterprise vocabulary.

7. The method according to claim 1, characterized in that The enterprise sample data is processed to obtain enterprise vocabulary, including: Performing word segmentation processing on the enterprise sample data to obtain the enterprise vocabulary; or, Performing word segmentation and grammatical analysis on the enterprise sample data; removing grammatically incorrect enterprise sample data to obtain the enterprise vocabulary; or, The enterprise sample data is segmented and negative expressions in the enterprise sample data are removed to obtain the enterprise vocabulary.

8. The method according to claim 1, characterized in that The method further comprises: In the case that the identified industrial chain identification word does not meet the preset training requirement, the word association model is adjusted, and the step of training the word association model using the existing data set is triggered.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement the industrial chain construction method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, which, when executed by a processor, is used to implement the industrial chain construction method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data tagging classification method and device, terminal and storage medium

    CN110413775A

  • Policy data processing method and device and storage medium

    CN112102137A