Multi-element market information processing method, system and equipment based on natural language processing
Through the multi-market information processing method based on natural language processing, valuable entities in multi-market information are identified and marked, the relationship between entities is determined, knowledge graphs are constructed and structured data is generated, and the problem of low efficiency of unstructured multi-market information processing is solved, and the high accuracy and standardized processing of data is achieved.
Patent Information
- Application Number
- CN202510324344.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
Smart Images

Figure CN120144610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies, and particularly to a method, system and device for processing multi-source market information based on natural language processing. Background Art
[0002] Currently, the power resource scheduling system generally adopts a regional regulation mode. To analyze the main factors affecting power consumption and electricity volume, it is necessary to access various types of market information data.
[0003] Multi-source market information includes: market information of industry associations and market information of information platforms. Thus, diversified market information is obtained from various markets, and the diversified market information is standardized to obtain a multi-source market information dataset for analyzing the influencing factors of electricity consumption. Among them, the standardization process of multi-source market information includes data cleaning, conversion, merging and other processes.
[0004] Taking the acquisition of market information of industry associations as an example, the information processing system connects to various industry associations such as the China National Machinery Industry Corporation, and regularly accesses the market information data of industry associations through methods such as offline data file import or online system connection. Among them, the market information of industry associations includes: a list of representative enterprises in key industries, which supports the dynamic update of the list of sample users across the network; information such as the output and price of key products monitored by industry associations; industry economic statistical data such as the revenue, profit, and added value of key industries; industry dynamic information such as industry development trends, market demands, and product characteristics.
[0005] Taking the acquisition of market information of information platforms as an example, the information processing system connects to external data information platforms and regularly accesses various types of professional market information data. Among them, the market information of information platforms includes various text-based information, including but not limited to industry news, domestic macro information, market highlights, etc.; macroeconomic data announced by the National Bureau of Statistics, Department of Finance, General Administration of Customs, etc.; economic data of key industries such as energy and manufacturing.
[0006] In reality, a lot of market information is taken from information or reports, etc. Natural language texts and data are mixed together, thus forming unstructured data, and it is difficult to extract valuable data information. Currently, it usually relies on manual reading of these information or reports to extract data. In addition, the data is scattered, which is not conducive to subsequent cleaning, conversion, merging and other processes. Therefore, there is an urgent need for a technology that can process batch and massive unstructured data on a computer system. Summary of the Invention
[0007] The present invention processes multi-source market information through an information processing system, avoiding the situation of manually extracting multi-source market information and improving the accuracy of determining structured data. The specific technical solutions are as follows:
[0008] First aspect, a multi - market information processing method based on natural language processing, includes the following processes:
[0009] S100: The first entity recognition model roughly analyzes the multi - market information to determine whether each sentence in the multi - market information is product - related or industry - related, and regards the product - related or industry - related sentences as valuable entities;
[0010] S200: The second entity recognition model finely analyzes the valuable entities to determine the entities corresponding to the corresponding word segments in the sentences corresponding to the valuable entities, and determines such entities as secondary entities; Determine the types of each secondary entity including: industry, product, economic term, numerical value, unit, and modifier;
[0011] S300: The relationship extraction model determines the relationships between each secondary entity according to the secondary entity and its corresponding type;
[0012] S400: Construct a knowledge graph according to the relationships between each secondary entity;
[0013] S500: Construct structured data according to the knowledge graph.
[0014] Preferably, the first entity recognition model includes: a word embedding model, BERT, Bi_LSTM, and CRF; The recognition process of the first entity recognition model includes the following steps:
[0015] S110: The first entity recognition model receives multi - market information;
[0016] S120: Input the multi - market information into the word embedding model, and process the multi - market information through the word embedding model to generate corresponding word vectors;
[0017] S130: Input the word vectors into BERT, and extract features from the word vectors through BERT to generate corresponding feature vectors;
[0018] S140: Input the feature information output by BERT into Bi_LSTM, and process the feature vectors through Bi_LSTM to output corresponding feature vectors;
[0019] S150: Input the feature vectors output by Bi_LSTM into CRF, and perform entity recognition on the feature vectors according to the preset recognition rules through CRF, so as to recognize entities that conform to the recognition rules from the multi - market information;
[0020] S160: The first entity recognition model eliminates worthless entities and outputs valuable entities. Taking commas or full stops as boundaries, it determines whether each sentence in the multi-source market information is related to the industry or the product, and regards the sentences related to the industry or the product as valuable entities. Preferably, in the S200, the second entity recognition model includes: a word embedding model, an attention module, a Bi_LSTM, and a CRF; the recognition process of the second entity recognition model includes the following steps:
[0021] S210: Input the valuable entities into the second entity recognition model, and the second entity recognition model processes the valuable entities through the word embedding model, thereby outputting first feature information corresponding to each valuable entity;
[0022] S220: Input the first feature information into the attention module, and process the valuable entities through the attention module to generate second feature information;
[0023] S230: Input the second feature information into the Bi_LSTM, and perform semantic analysis on the second feature information through the Bi_LSTM to generate third feature information;
[0024] S240: Input the third feature information into the CRF, and process the third feature information through the CRF to label the types of each word segmentation corresponding to each valuable entity, and determine the secondary entities; where the types include: industry, product, economic term, numerical value, unit, and modifier.
[0025] Preferably, in the S200, the attention module includes:
[0026] Forward local attention model: used to perform attention processing on the target feature information and a predetermined number (for example, L) of first feature information before this;
[0027] Global attention model: used to perform attention processing on the target feature information and all the first feature information;
[0028] Backward local attention model: used to perform attention processing on the target feature information and a predetermined number of first feature information after this.
[0029] Preferably, the forward local attention model respectively generates key vectors K 1 ~C n corresponding to multiple word vectors C 1 ~K n , query vectors Q 1 ~Q n and value vectors V 1 ~V n , and then the multiple first feature information C 1 ~Cn are sequentially used as target feature information; calculate the semantic feature vector F corresponding to the target feature information C through the following formula i i ,
[0030]
[0031] w i,j = softmax(h i,j );
[0032]
[0033] where i = 1 to n, j = i - L to i; w represents the weight value of the target feature information C relative to the target feature information C; h represents the correlation coefficient between the key vector K of the first feature information C and the query vector Q of the target feature information C; d represents the dimension of the key vector; i,j i j i,j j j i j k
[0034] To determine the semantic feature vector F of the target feature information C, calculate the correlation coefficient h of each first feature information C relative to the target feature information C using the formula; i i j i i,j ;
[0035] Perform a softmax function calculation on h to obtain the corresponding weight w; i,j i,j i,j ;
[0036] Calculate the semantic feature vector F of the target feature information C. i i .
[0037] Preferably, the backward local attention model respectively generates key vectors K~K, query vectors Q~Q, and value vectors V~V corresponding to multiple word vectors C~C, and then uses multiple first feature information C~C 1 ~ n 1 ~ n 1 ~ n 1 ~ n 1 ~ n in turn as the target feature information; calculate the semantic feature vector F corresponding to the target feature information C through the following formula i corresponding semantic feature vector F i ,
[0038] w i,j = softmax(h i,j );
[0039] where i = 1 to n, j = i to i + L; w i,j represents the weight value of the target feature information C relative to the target feature information C i ; h j represents the correlation coefficient between the key vector K of the first feature information C i,j and the query vector Q of the target feature information C j ; d j represents the dimension of the key vector; w i j k represents the weight corresponding to the correlation coefficient h i,j i,j 1 n 1 .
[0040] Preferably, the global attention model respectively generates the key vectors K 1 ~ K n , the query vectors Q 1 ~ Q n and the value vectors V 1 ~ V n corresponding to a plurality of word vectors C 1 ~ C n , and then the plurality of first feature information C 1 ~ C n are taken as the target feature information in turn; calculate the semantic feature vector F corresponding to the target feature information C i through the following formula i ,
[0041] w i,j = softmax(h i,j );
[0042] where i = 1 to n, where j = 1 to n; w i,j represents the weight value of the target feature information C i relative to the target feature information C j ; h i,j represents the key vector K of the first feature information C j and the query vector Q of the target feature information C j i j The correlation coefficient between; d k Denote the dimension of the key vector; w i,j Denote the correlation coefficient h i,j The corresponding weight.
[0043] Preferably, in the S300, the relation extraction model includes: BERT, Bi_LSTM, an attention module, a fully connected layer, and a classifier; where the relation extraction model is used to determine the relations between each secondary entity; the relation types include: subordinate relation, attribute relation, numerical relation, and unit relation; the process of the relation extraction model extracting the relations between secondary entities includes the following steps:
[0044] S310: Input the secondary entity into the relation extraction model;
[0045] S320: After receiving the secondary entity, the relation extraction model inputs the secondary entity into BERT, and BERT extracts features from the secondary entity to generate fourth feature information;
[0046] S330: The relation extraction model inputs the fourth feature information into Bi_LSTM, and Bi_LSTM processes the fourth feature information to generate fifth feature information;
[0047] S340: The relation extraction model inputs the fifth feature information into the attention model, and the attention model processes the fifth feature information to generate sixth feature information;
[0048] S350: The relation extraction model inputs the sixth feature information into the fully connected layer, and the fully connected layer processes the sixth feature information to generate seventh feature information;
[0049] S360: The relation extraction model inputs the seventh feature information into the classifier, and the classifier classifies the first feature information to output the relations between each secondary entity.
[0050] In the second aspect, a multi-source market information processing system based on natural language processing includes:
[0051] A first entity recognition module, used to recognize valuable entities in the text; the first entity recognition module is sequentially connected with a word embedding model module, a BERT module, a Bi_LSTM module, and a CRF model module inside;
[0052] A second entity recognition module, used to re-label the valuable entities to determine the secondary entities included in the sentences corresponding to each valuable entity; the second entity recognition module is sequentially connected with a word embedding model module, an attention module, a Bi_LSTM module, and a CRF module inside;
[0053] A relation extraction model module for determining the relationships between various secondary entities; the relation extraction model module is sequentially connected with a BERT module, a Bi_LSTM module, an attention module, a fully connected layer module, and a classifier module;
[0054] A knowledge graph module for constructing a knowledge graph based on the relationships between secondary entities;
[0055] A structured data module for multi-source market information for constructing structured data of multi-source market information based on the knowledge graph to complete the processing of multi-source market information.
[0056] In a third aspect, an electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, any of the above-mentioned multi-source market information processing methods based on natural language processing is implemented.
[0057] The advantages of the present invention over the prior art are as follows:
[0058] (1) By processing multi-source market information through a first entity recognition model and a second entity recognition model based on natural language processing, valuable entities in multi-source market information are screened from coarse to fine, and the relationships between various entities are determined through a relation extraction model, and a knowledge graph is constructed based on the relationships between entities and structured data is generated, so as to perform refined processing on each entity, thereby achieving a high accuracy of the constructed structured data; thus, the standardization of data is initially realized, which is conducive to subsequent data cleaning, transformation, and integration.
[0059] (2) Through a forward local attention model, a global attention model, and a backward local attention model, correlation calculations of different dimensions are respectively performed on the feature information corresponding to entities, that is, through the forward local attention model and the backward local attention model, the correlation between each entity in a sentence or a small paragraph is accurately extracted. Through the global attention model, the correlation between each entity corresponding to multi-source market information can be extracted. Thus, the features of each entity are finely extracted, and multi-dimensional analysis of multi-source market information is realized. Description of the Drawings
[0060] Figure 1 It is a schematic diagram of the process of the second entity recognition model in an embodiment of the present invention.
[0061] Figure 2 It is one of the examples of the knowledge graph constructed by the information processing system according to the relationships between secondary entities in an embodiment of the present invention.
[0062] Figure 3 It is one of the examples of the knowledge graph constructed by the information processing system according to the relationships between secondary entities in an embodiment of the present invention. Specific implementation manner
[0063] A multi - market information processing method based on natural language processing includes the following processes:
[0064] S110: The first entity recognition model receives multi - market information;
[0065] S120: Input the multi - market information into a word embedding model, and process the multi - market information through the word embedding model to generate corresponding word vectors;
[0066] S130: Input the word vectors into BERT, and extract features from the word vectors through BERT to generate corresponding feature vectors;
[0067] S140: Input the feature information output by BERT into Bi_LSTM, and process the feature vectors through Bi_LSTM to output corresponding feature vectors;
[0068] S150: Input the feature vectors output by Bi_LSTM into CRF, and perform entity recognition on the feature vectors according to a preset recognition rule through CRF, so as to identify entities that conform to the recognition rule from the multi - market information; where the recognition rule is used to identify entity content related to products and industries. Entities that conform to this recognition rule are called valuable entities.
[0069] For example, the multi - market information is: "To strengthen statistical law enforcement, relevant bases have been revised according to regulations. The mining industry achieved a total profit of 717.92 billion yuan, a year - on - year decrease of 9.5%; the electricity, heat, gas and water production and supply industry achieved a total profit of 476.71 billion yuan, a growth of 20.1%."
[0070] Thus, the CRF of the first entity recognition model annotates each sentence in the multi - market information bounded by commas or full stops to determine whether it is a valuable entity or a valueless entity. When annotating each sentence, mark the start word of the sentence with a start identifier. For example, the start identifier for a valuable entity is VS, and the start identifier for a valueless entity is US.
[0071] And mark the end word of the sentence with an end identifier. For example, the end identifier for a valuable entity is VE, and the start identifier for a valueless entity is UE. The middle part between the start identifier and the end identifier is the middle content of the corresponding type (valuable entity or valueless), marked with o.
[0072] For example, "To strengthen statistical law enforcement, relevant bases have been revised according to regulations" is a valueless entity, and the marked content is:
[0073]
[0074] For example, if "The total profit realized by the mining industry was 717.92 billion yuan" is a valuable entity, the content after marking is as follows:
[0075]
[0076] The specific annotation information for each sentence is shown in the following table:
[0077]
[0078] S160: The first entity recognition model eliminates worthless entities and outputs valuable entities. Taking commas or full stops as boundaries, it determines whether each sentence in the multi-source market information is related to the industry or the product, and regards the sentences related to the industry or the product as valuable entities; among them, the valuable entity is the entity content related to the product and the industry.
[0079] S200: The second entity recognition model conducts a refined analysis on the valuable entities, determines the entities corresponding to the corresponding word segmentations in the sentences corresponding to the valuable entities, and determines such entities as secondary entities; it determines that the types of each secondary entity include: industry, product, economic term, numerical value, unit, and modifier; the second entity recognition model in S200 includes: a word embedding model, an attention module, a Bi_LSTM, and a CRF; the recognition process of the second entity recognition model includes the following steps:
[0080] S210: Input the valuable entities into the second entity recognition model, and the second entity recognition model processes the valuable entities through the word embedding model, thereby outputting the first feature information corresponding to each valuable entity;
[0081] S220: Input the first feature information into the attention module, and process the valuable entities through the attention module to generate the second feature information; the attention module includes:
[0082] Forward local attention model: used to perform attention processing on the target feature information and a predetermined number (for example, L) of the first feature information before this;
[0083] The forward local attention model respectively generates key vectors K 1 ~C n corresponding to multiple word vectors C 1 ~K n 、query vectors Q 1 ~Q n and value vectors V 1 ~V n ,and then takes the multiple first feature information C 1 ~C n as the target feature information in sequence; calculate the target feature information C through the following formulai The corresponding semantic feature vector F i ,
[0084] w i,j = softmax(h i,j );
[0085] where i = 1 to n, j = i - L to i; w i,j represents the weight value of the target feature information C i relative to the target feature information C j ; h i,j represents the key vector K j of the first feature information C j and the query vector Q i of the target feature information C j ; d k represents the dimension of the key vector;
[0086] To determine the semantic feature vector F i of the target feature information C i , use the formula to calculate the correlation coefficient h j of each first feature information C i relative to the target feature information C i,j ;
[0087] Perform the softmax function calculation on h i,j to obtain the weight w i,j corresponding to the correlation coefficient h i,j ;
[0088] Calculate the semantic feature vector F i of the target feature information C i .
[0089] It should be noted that in the case where the number of first feature information before the target feature information is less than the predetermined number, the semantic feature vector F 1 ~C i can be calculated using only the key vectors and value vectors of the first feature information C i to the target feature information C i .
[0090] Global attention model: used to perform attention processing on the target feature information and all first feature information; the global attention model generates the key vectors K 1 ~K n corresponding to multiple word vectors C 1 ~K n , query vectors Q 1 ~Qn and the value vector V 1 ~V n , and then use multiple first feature information C 1 ~C n as the target feature information in sequence; calculate the semantic feature vector F i corresponding to the target feature information C i ,
[0091] w i,j = softmax(h i,j );
[0092] where i = 1~n, and j = 1~n; w i,j represents the weight value of the target feature information C i relative to the target feature information C j ; h i,j represents the correlation coefficient between the key vector K j of the first feature information C j and the query vector Q i of the target feature information C j ; d k represents the dimension of the key vector; w i,j represents the weight corresponding to the correlation coefficient h i,j .
[0093] Backward local attention model: used to perform attention processing on the target feature information and a predetermined number of first feature information after this; the backward local attention model respectively generates key vectors K 1 ~K n , query vectors Q 1 ~Q n and value vectors V 1 ~V n corresponding to multiple word vectors C 1 ~C n , and then use multiple first feature information C 1 ~C n as the target feature information in sequence; calculate the semantic feature vector F i corresponding to the target feature information C i ,
[0094] w i,j = softmax(h i,j );
[0095] where i = 1~n, j = i~i + L; w i,j represents the target feature information Ci The weight value relative to the target feature information C j ; h i,j Indicates the key vector K j of the first feature information C j and the query vector Q i of the target feature information C j The correlation coefficient between them; d k Indicates the dimension of the key vector; w i,j Indicates the weight corresponding to the correlation coefficient h i,j .
[0096] Fuse the semantic feature vectors respectively output by the forward local attention model, the global attention model, and the backward local attention model, for example, using the splicing method, to generate the second feature information.
[0097] S230: Input the second feature information into the Bi_LSTM, and perform semantic analysis on the second feature information through the Bi_LSTM to generate the third feature information;
[0098] S240: Input the third feature information into the CRF, and process the third feature information through the CRF, so as to label the types of each word segment corresponding to each valuable entity, and determine the secondary entities; where the types include: industry, product, economic term, numerical value, unit, and modifier.
[0099] For example, the sentence corresponding to the valuable entity is: "The industrial electricity cost in the power, heat production and supply industry has increased to 890 billion yuan." Then the secondary entities of the valuable entity are respectively: "power, heat production and supply industry", "industrial electricity", "cost", "increase", "89000", and "billion yuan".
[0100] Then the second entity recognition model determines that the type of "power, heat production and supply industry" is industry, the type of "industrial electricity" is product, the type of "cost" is economic term, the type of "89000" is numerical value, the type of "billion yuan" is unit, and the type of "increase" is modifier.
[0101] S300: The relationship extraction model determines the relationships between each secondary entity according to the secondary entity and the corresponding type; the relationship extraction model includes: BERT, Bi_LSTM, attention module, fully connected layer, and classifier; where the relationship extraction model is used to determine the relationships between each secondary entity; the relationship types include: subordinate relationship, attribute relationship, numerical relationship, and unit relationship; the process of the relationship extraction model extracting the relationships between secondary entities includes the following steps:
[0102] S310: Input the secondary entity into the relationship extraction model;
[0103] S320: After receiving the secondary entity, the relationship extraction model inputs the secondary entity into BERT, and BERT extracts features from the secondary entity to generate fourth feature information;
[0104] S330: The relationship extraction model inputs the fourth feature information into Bi_LSTM, and Bi_LSTM processes the fourth feature information to generate fifth feature information;
[0105] S340: The relationship extraction model inputs the fifth feature information into the attention model, and the attention model processes the fifth feature information to generate sixth feature information;
[0106] S350: The relationship extraction model inputs the sixth feature information into the fully connected layer, and the fully connected layer processes the sixth feature information to generate seventh feature information;
[0107] S360: The relationship extraction model inputs the seventh feature information into the classifier, and the classifier classifies the first feature information to output the relationships between each secondary entity.
[0108] For example, the sentence corresponding to the valuable entity is: "The industrial electricity cost of the power, heat production and supply industry has increased to 890 billion yuan." Then the secondary entities of the valuable entity are: "power, heat production and supply industry", "industrial electricity", "cost", "increase", "89000", and "billion yuan".
[0109] Then the relationship extraction model determines that there is a subordinate relationship between "power, heat production and supply industry" and "industrial electricity", and "industrial electricity" belongs to "power, heat production and supply industry"; there is an attribute relationship between "industrial electricity" and "cost"; there is an attribute relationship between "cost" and "increase"; there is a numerical relationship between "cost" and "89000"; there is a unit relationship between "89000" and "billion yuan".
[0110] S400: Construct a knowledge graph based on the relationships between each secondary entity; as shown in Appendix Figure 2 Appendix Figure 3 shown.
[0111] S500: Construct structured data based on this knowledge graph.
[0112] For example, part of the structured data is referred to the following table.
[0113]
[0114] Among them, in the above table, "Electricity, Heat Production and Supply Industry" is in a subordinate relationship with "Industrial Electricity", "Wind Power Generation", and "Solar Power Generation"; "Industrial Electricity" is in an attribute relationship with "Price", "Cost", "Sales Volume", and "Profit"; "Wind Power Generation" is in an attribute relationship with "Price"; "Price" is in an attribute relationship with "Decrease"; "Cost" is in an attribute relationship with "Increase"; "Sales Volume" is in an attribute relationship with "Increase"; "Profit" is in an attribute relationship with "Increase"; "Price" is in a numerical relationship with "0.8" and "0.3"; "Cost" is in a numerical relationship with "89000" and "3500"; "0.8" is in a unit relationship with "Yuan per Kilowatt-hour"; "0.3" is in a unit relationship with "Yuan per Kilowatt-hour"; "89000" is in a unit relationship with "100 million Yuan"; "3500" is in a unit relationship with "100 million Yuan".
Claims
1. A multi-market information processing method based on natural language processing, characterized in that: The process includes the following: S100: The first entity recognition model roughly analyzes the multi-market information to determine whether each sentence in the multi-market information is product-related or industry-related, and regards the product-related or industry-related sentences as valuable entities; S200: The second entity recognition model performs a detailed analysis on the valuable entity, determines the entity corresponding to the corresponding participle in the sentence corresponding to the valuable entity, and determines such entity as a secondary entity; Identify the types of each secondary entity including: industry, product, economic term, value, unit, and modifier; S300: The relationship extraction model determines the relationship between each secondary entity according to the secondary entities and the corresponding types; S400: construct a knowledge graph based on the relationship between each secondary entity; S500: Construct structured data according to the knowledge graph.
2. The method for processing multi-market information based on natural language processing according to claim 1, characterized in that: The first entity recognition model includes: a word embedding model, BERT, Bi_LSTM and CRF; the first entity recognition model recognition process includes the following steps: S110: The first entity recognition model receives multi-market information; S120: inputting the multi-market information into a word embedding model, processing the multi-market information through the word embedding model, and generating corresponding word vectors; S130: Input the word vector into BERT, perform feature extraction on the word vector through BERT, and generate a corresponding feature vector; S140: Input the feature information output by BERT into Bi_LSTM, process the feature vector through Bi_LSTM, and output the corresponding feature vector; S150: Input the feature vector output by Bi_LSTM into CRF, and perform entity recognition on the feature vector according to a preset recognition rule through CRF, so as to identify entities that meet the recognition rule from the multi-market information; S160: The first entity recognition model removes worthless entities and outputs valuable entities, using commas or periods as delimiters to determine whether each sentence in the multi-dimensional market information is industry-related or product-related, and regards sentences that are industry-related or product-related as valuable entities.
3. The method for processing multi-market information based on natural language processing according to claim 1, characterized in that: The second entity recognition model in S200 includes: a word embedding model, an attention module, a Bi_LSTM and a CRF; the second entity recognition model recognition process includes the following steps: S210: inputting the valuable entities into the second entity recognition model, and the second entity recognition model processes the valuable entities through the word embedding model, thereby outputting first feature information corresponding to each valuable entity; S220: inputting the first feature information into an attention module, processing the valuable entity through the attention module, and generating second feature information; S230: inputting the second feature information into the Bi_LSTM, and performing semantic analysis on the second feature information through the Bi_LSTM to generate third feature information; S240: Input the third feature information into the CRF, and process the third feature information through the CRF, so as to mark the type of each word segment corresponding to each valuable entity and determine the secondary entity; wherein the type includes: Industries, products, economic terms, values, units and modifiers.
4. The method for processing multi-market information based on natural language processing according to claim 3, characterized in that: In S200, the attention module includes: Forward local attention model: used to perform attention processing on the target feature information and the first feature information of a predetermined number (e.g., L) before it; Global attention model: used to process the target feature information and all the first feature information; Backward local attention model: used to pay attention to the target feature information and the first feature information of a predetermined number after it.
5. The method for processing multi-market information based on natural language processing according to claim 4, characterized in that: The forward local attention model generates multiple word vectors C1~C n The corresponding key vectors K1~K n 、Query vector Q1~Q n And the value vectors V1~V n , and then multiple first feature information C1~C n In turn, they are used as target feature information; The target feature information C is calculated by the following formula i The corresponding semantic feature vector F i , w i,j =softmax(h i,j ); Among them, i=1~n, j=iL~i; w i,j Represents target feature information C i Relative to the target feature information C j The weight value of h i,j Indicates the first characteristic information C j The key vector K j and target feature information C i The query vector Q j The correlation coefficient between k represents the dimension of the key vector; In order to determine the target feature information C i The semantic feature vector F i , use the formula to calculate each first feature information C j Relative to the target feature information C i The correlation coefficient h i,j ; For h i,j Calculate the softmax function to get the correlation coefficient h i,j The corresponding weight w i,j ; Calculate target feature information C i The semantic feature vector F i .
6. The method for processing multi-market information based on natural language processing according to claim 4, characterized in that: The backward local attention model generates multiple word vectors C1~C n The corresponding key vectors K1~K n 、Query vector Q1~Q n And the value vectors V1~V n , and then multiple first feature information C1~C n In turn, they are used as target feature information; The target feature information C is calculated by the following formula i The corresponding semantic feature vector F i , w i,j =softmax(h i,j ); Among them, i=1~n, j=i~i+L; w i,j Represents target feature information C i Relative to the target feature information C j The weight value of h i,j Indicates the first characteristic information C j The key vector K j and target feature information C i The query vector Q j The correlation coefficient between k represents the dimension of the key vector; w i,j Represents the correlation coefficient h i,j The corresponding weight.
7. The method for processing multi-market information based on natural language processing according to claim 4, characterized in that: The global attention model generates multiple word vectors C1~C n The corresponding key vectors K1~K n 、Query vector Q1~Q n And the value vectors V1~V n , and then multiple first feature information C1~C n In turn, they are used as target feature information; The target feature information C is calculated by the following formula i The corresponding semantic feature vector F i , w i,j =softmax(h i,j ); Where i = 1 to n, where j = 1 to n; w i,j Represents target feature information C i Relative to the target feature information C j The weight value of h i,j Indicates the first characteristic information C j The key vector K j and target feature information C i The query vector Q j The correlation coefficient between k represents the dimension of the key vector; w i,j Represents the correlation coefficient h i,j The corresponding weight.
8. The method for processing multi-market information based on natural language processing according to claim 1, characterized in that: In S300, the relationship extraction model includes: BERT, Bi_LSTM, attention module, fully connected layer and classifier; wherein the relationship extraction model is used to determine the relationship between each secondary entity; the relationship type includes: subordinate relationship, attribute relationship, numerical relationship and unit relationship; the process of extracting the relationship between the secondary entities by the relationship extraction model includes the following steps: S310: Inputting the secondary entity into the relationship extraction model; S320: After receiving the second-level entity, the relationship extraction model inputs the second-level entity into BERT, and extracts features of the second-level entity through BERT to generate fourth feature information; S330: The relationship extraction model inputs the fourth feature information into the Bi_LSTM, and processes the fourth feature information through the Bi_LSTM to generate fifth feature information; S340: the relationship extraction model inputs the fifth feature information into the attention model, and processes the fifth feature information through the attention model to generate sixth feature information; S350: The relationship extraction model inputs the sixth feature information into the fully connected layer, and processes the sixth feature information through the fully connected layer to generate seventh feature information; S360: The relationship extraction model inputs the seventh feature information into the classifier, classifies the first feature information through the classifier, and outputs the relationship between each secondary entity.
9. A multi-market information processing system based on natural language processing, characterized in that: include: The first entity recognition module is used to identify valuable entities in the text; the first entity recognition module is connected in series with the word embedding model module, the BERT module, the Bi_LSTM module and the CRF model module; The second entity recognition module is used to re-label the valuable entities and determine the secondary entities contained in the sentences corresponding to each valuable entity; the second entity recognition module is sequentially connected with the word embedding model module, the attention module, the Bi_LSTM module, and the CRF module; The relation extraction model module is used to determine the relationship between each secondary entity. The relation extraction model module is connected in series with the BERT module, Bi_LSTM module, attention module, fully connected layer module and classifier module. The knowledge graph module is used to construct a knowledge graph based on the relationships between secondary entities; The structured data module of diversified market information is used to construct structured data of diversified market information based on the knowledge graph and complete the processing of diversified market information.
10. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the multi-market information processing method based on natural language processing as described in any one of claims 1 to 8 when executing the computer program.