Industrial Map Construction Method, Device, Equipment and Storage Medium

By encoding and entity recognition of industrial training texts and generating entity information recognition models, the problem of difficult to accurately and quickly generate industrial map construction in the existing technology is solved, and efficient and accurate industrial map construction is achieved.

CN114528981BActive Publication Date: 2025-06-27PINGAN INT SMART CITY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210255942.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2025-06-27
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

The existing industrial map construction method is difficult to accurately and quickly extract the logical relationship and spatial-temporal layout relationship of the industry, resulting in the low quality of the generated industrial map.

Method used

By obtaining industrial training text and preset networks, the text is encoded and entity recognition using the word vector layer, the two-way long and short-term memory network layer and the entity recognition layer, and the network is adjusted to generate an entity information recognition model, and then parse the text to be parsed and build an industrial map.

Benefits of technology

It improves the accuracy and identification efficiency of industrial entity information and source entity information, and enhances the accuracy and efficiency of generating industrial maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114528981B_ABST
    Figure CN114528981B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence, and provides a method, apparatus, device and storage medium for constructing an industrial map. This method encodes each text character in the industrial training text based on a word vector layer and a bidirectional long short-term memory network layer to obtain a target time series vector, and identifies the target time series vector based on an entity recognition layer to obtain predicted entities and predicted labels. The preset network is adjusted according to the text entity information, predicted entities and predicted labels to obtain an entity information recognition model. The text to be parsed is input into the entity information recognition model to obtain industrial entity information and origin entity information. The industrial entity information and origin entity information are subjected to syntactic dependency matching processing to obtain entity information pairs, and the entity information pairs are spliced according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map, improving the generation efficiency and accuracy of the industrial map. In addition, the present invention also relates to blockchain technology, and the industrial map can be stored in the blockchain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and storage medium for constructing an industrial map. Background Art

[0002] In the current method of constructing an industrial map, mainly by a large number of business personnel reading text information under each link of the industry, and then extracting logical relationships and spatio-temporal layout relationships from the text information, and constructing an industrial map based on the industrial department in the specific logical relationships and spatio-temporal layout relationships. However, due to the complexity of the current industrial information volume, it is impossible to accurately and quickly extract logical relationships and spatio-temporal layout relationships, resulting in a low generation method of the industrial map. Summary of the Invention

[0003] In view of the above, it is necessary to provide a method, apparatus, device, and storage medium for constructing an industrial map, which can accurately and quickly generate an industrial map.

[0004] On the one hand, the present invention proposes a method for constructing an industrial map, and the method for constructing an industrial map includes:

[0005] Obtain industrial training texts, and obtain text entity information of the industrial training texts;

[0006] Obtain a preset network, where the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer;

[0007] Encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain a target time series vector for each text character;

[0008] Identify the target time series vector based on the entity recognition layer to obtain predicted entities and predicted labels of the industrial training text;

[0009] Adjust the preset network according to the text entity information, the predicted entities, and the predicted labels to obtain an entity information recognition model;

[0010] Obtain a text to be parsed, and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information;

[0011] Perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain entity information pairs;

[0012] Splice the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

[0013] According to a preferred embodiment of the present invention, the bidirectional long short-term memory network layer includes a forward long short-term memory network layer and a backward long short-term memory network layer. Encoding each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain the target time series vector of each text character includes:

[0014] Encoding the industrial training text based on the word vector layer to obtain the character vector of each text character;

[0015] Locate the character order of each text character in the industrial training text;

[0016] Input the character vectors into the forward long short-term memory network layer in ascending order according to the character order to obtain the forward time series vector of each text character, and input the character vectors into the backward long short-term memory network layer in descending order according to the character order to obtain the backward time series vector of each text character;

[0017] Concatenate the forward time series vector and the backward time series vector to obtain the target time series vector.

[0018] According to a preferred embodiment of the present invention, the step of inputting the character vectors into the forward long short-term memory network layer in ascending order according to the character order to obtain the forward time series vector of each text character includes:

[0019] For any text character, obtain the neighboring characters whose character order is less than that of the any text character as the target characters;

[0020] Obtain the state vector of the target character;

[0021] Concatenate the state vector and the character vector of the any text character to obtain the input vector;

[0022] Calculate the input vector based on the preset network matrix and preset bias value of the forward long short-term memory network layer to obtain the forward time series vector of the any text character.

[0023] According to a preferred embodiment of the present invention, the step of identifying the target time series vector based on the entity recognition layer to obtain the predicted entity and predicted label of the industrial training text includes:

[0024] Calculate the sum of each vector element in the target time series vector to obtain the character score of each text character;

[0025] Obtain the score threshold and preset weight matrix from the entity recognition layer;

[0026] Determine the text characters whose character scores are greater than the score threshold as the predicted entities;

[0027] Calculate the product of the target time series vector corresponding to the predicted entity and the preset weight matrix to obtain the entity probability of the predicted entity on each preset label;

[0028] Determine the preset label with the maximum entity probability as the predicted label of the predicted entity.

[0029] According to a preferred embodiment of the present invention, the text entity information includes training entities and entity labels. Adjusting the preset network according to the text entity information, the predicted entity and the predicted label to obtain the entity information recognition model includes:

[0030] Count the total number of entities of the training entities;

[0031] Calculate the number of predicted entities identical to the training entities as the first prediction number, and calculate the number of predicted labels identical to the entity labels as the second prediction number;

[0032] Calculate the network loss value of the preset network according to the total number of entities, the first prediction number and the second prediction number;

[0033] Based on the network loss value, adjust the network parameters of the bidirectional long short-term memory network layer and the entity recognition layer until the network loss value is less than a preset threshold to obtain the entity information recognition model.

[0034] According to a preferred embodiment of the present invention, the syntactic dependency matching process for the industrial entity information and the origin entity information to obtain the entity information pair includes:

[0035] Filter text sentences from the text to be parsed based on the industrial entity information and the origin entity information;

[0036] Perform word segmentation on the text sentences to obtain a plurality of sentence word segments;

[0037] Identify the part-of-speech of each sentence word segment in the text sentences;

[0038] Based on the sentence word segments with the part-of-speech being the preset part-of-speech, perform dependency recognition on the industrial entity information and the origin entity information to obtain the dependency relationship between the industrial entity information and the origin entity information;

[0039] Determine the industrial entity information and the origin entity information with the same set of dependency relationships as the entity information pair.

[0040] According to a preferred embodiment of the present invention, the obtaining of the text to be parsed includes:

[0041] Obtain any text to be processed from the text library to be processed, and obtain the text identifier of the any text to be processed;

[0042] Obtain the target identifier corresponding to the text identifier from the text association table of the text library to be processed;

[0043] Based on the target identifier, obtain the text corresponding to the target identifier from the text library to be processed as the associated text of the any text to be processed;

[0044] Splice the any text to be processed and the associated text according to the association order in the text association table to obtain the text to be parsed.

[0045] On the other hand, the present invention also provides an industrial map construction device, which includes:

[0046] An acquisition unit, configured to acquire industrial training text and acquire the text entity information of the industrial training text;

[0047] The acquisition unit is further configured to acquire a preset network, where the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer;

[0048] An encoding unit, configured to encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain the target time series vector of each text character;

[0049] An identification unit, configured to identify the target time series vector based on the entity recognition layer to obtain the predicted entity and predicted label of the industrial training text;

[0050] An adjustment unit, configured to adjust the preset network according to the text entity information, the predicted entity, and the predicted label to obtain an entity information recognition model;

[0051] An input unit, configured to acquire the text to be parsed and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information;

[0052] A matching unit, configured to perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain an entity information pair;

[0053] A splicing unit, configured to splice the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

[0054] On the other hand, the present invention also provides an electronic device, which includes:

[0055] A memory, storing computer-readable instructions; and

[0056] A processor executes computer-readable instructions stored in the memory to implement the industry map construction method.

[0057] On the other hand, the present invention also proposes a computer-readable storage medium, in which computer-readable instructions are stored. The computer-readable instructions are executed by a processor in an electronic device to implement the industrial map construction method.

[0058] It can be seen from the above technical solutions that the present invention trains the preset network based on the industry training text and the text entity information, which can improve the recognition ability of the entity information recognition model for entity information, and then parse the text to be parsed through the entity information recognition model, improve the accuracy of the industry entity information and the place of origin entity information, perform syntactic dependency matching on the industry entity information and the place of origin entity information, improve the accuracy of the entity information pair, and thus improve the generation accuracy of the industry map. In addition, the present invention can quickly extract the industry entity information and the place of origin entity information from the text to be parsed through the entity information recognition model, thereby improving the generation efficiency of the industry map. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a flow chart of a preferred embodiment of the method for constructing an industrial map of the present invention.

[0060] Figure 2 It is a functional module diagram of a preferred embodiment of the industrial map construction device of the present invention.

[0061] Figure 3 It is a schematic diagram of the structure of an electronic device of a preferred embodiment of the method for constructing an industrial map according to the present invention. DETAILED DESCRIPTION

[0062] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0063] like Figure 1 FIG. 1 is a flowchart of a preferred embodiment of the method for constructing an industrial map of the present invention. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.

[0064] The industrial map construction method can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0065] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0066] The industrial map construction method is applied to one or more electronic devices. The electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored computer-readable instructions. Its hardware includes, but is not limited to, microprocessors, Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), Digital Signal Processors (DSPs), embedded devices, etc.

[0067] The electronic device can be any electronic product that can perform human-computer interaction with users. For example, personal computers, tablet computers, smartphones, Personal Digital Assistants (PDAs), game consoles, Internet Protocol Televisions (IPTVs), smart wearable devices, etc.

[0068] The electronic device can include network devices and / or user devices. Among them, the network device includes, but is not limited to, a single network electronic device, a group of electronic devices composed of multiple network electronic devices, or a cloud composed of a large number of hosts or network electronic devices based on Cloud Computing.

[0069] The network where the electronic device is located includes, but is not limited to: the Internet, wide area network, metropolitan area network, local area network, Virtual Private Network (VPN), etc.

[0070] S10. Acquire industrial training texts and acquire the text entity information of the industrial training texts.

[0071] In at least one embodiment of the present invention, the industrial training text refers to text information related to a specific industry, and the industrial training text includes the specific industry and the industrial execution location of the specific industry.

[0072] The text entity information includes a training entity and an entity label. The training entity refers to a specific industry or a specific execution location, and the entity label is used to identify that the training entity is an industry name, an industrial location, etc.

[0073] S11. Obtain a preset network, where the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer.

[0074] In at least one embodiment of the present invention, the word vector layer is used to encode text characters.

[0075] The bidirectional long short-term memory network layer includes a forward long short-term memory network layer and a reverse long short-term memory network layer. The forward long short-term memory network layer is used to perform forward time series prediction on text characters, and the reverse long short-term memory network layer performs reverse time series prediction on text characters.

[0076] The entity recognition layer is used to perform named entity recognition on text characters.

[0077] S12. Encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain a target time series vector for each text character.

[0078] In at least one embodiment of the present invention, the target time series vector includes forward time series information and reverse time series information of each text character.

[0079] In at least one embodiment of the present invention, the electronic device encodes each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer, and obtaining the target time series vector for each text character includes:

[0080] Encoding the industrial training text based on the word vector layer to obtain a character vector for each text character;

[0081] Locate the character order of each text character in the industrial training text;

[0082] Input the character vectors into the forward long short-term memory network layer in ascending order of the character order to obtain a forward time series vector for each text character, and input the character vectors into the reverse long short-term memory network layer in descending order of the character order to obtain a reverse time series vector for each text character;

[0083] Concatenate the forward time series vector and the reverse time series vector to obtain the target time series vector.

[0084] By inputting the character vector into the bidirectional long short-term memory network layer according to the character order, it can assist the bidirectional long short-term memory network layer to quickly recognize the character order of the character vector in the industrial training text, and improve the generation efficiency of the target time series vector.

[0085] Specifically, the electronic device inputs the character vector into the forward long short-term memory network layer in ascending order according to the character order, and the forward time series vector of each text character obtained includes:

[0086] For any text character, obtain the neighboring characters whose character order is less than this text character as the target characters;

[0087] Obtain the state vector of the target character;

[0088] Concatenate the state vector and the character vector of this text character to obtain an input vector;

[0089] Calculate the input vector based on the preset network matrix and preset bias value of the forward long short-term memory network layer to obtain the forward time series vector of this text character.

[0090] Among them, the state vector refers to the forward time series vector of the target character. It should be noted that the state vector of the first text character in the industrial training text is the character vector.

[0091] The preset network matrix and the preset bias value are both network parameters of the forward long short-term memory network layer.

[0092] Through the above implementation manner, the forward time series vector can be quickly generated.

[0093] Specifically, the way that the electronic device inputs the character vector into the reverse long short-term memory network layer in descending order according to the character order to obtain the reverse time series vector of each text character is similar to the way that the electronic device inputs the character vector into the forward long short-term memory network layer in ascending order according to the character order to obtain the forward time series vector of each text character, and the present invention will not elaborate on this.

[0094] S13. Based on the entity recognition layer, recognize the target time series vector to obtain the predicted entity and predicted label of the industrial training text.

[0095] In at least one embodiment of the present invention, the predicted entity refers to the entity information obtained after the preset network identifies the industrial training text, and the predicted label refers to the label information obtained after the preset network identifies and predicts the entity information. Among them, the predicted entity corresponds to the training entity, and the predicted label corresponds to the entity label.

[0096] In at least one embodiment of the present invention, the electronic device identifies the target time series vector based on the entity recognition layer, and obtaining the predicted entity and predicted label of the industrial training text includes:

[0097] Calculate the sum of each vector element in the target time series vector to obtain the character score of each text character;

[0098] Obtain the score threshold and the preset weight matrix from the entity recognition layer;

[0099] Determine the text characters with the character score greater than the score threshold as the predicted entity;

[0100] Calculate the product of the target time series vector corresponding to the predicted entity and the preset weight matrix to obtain the entity probability of the predicted entity on each preset label;

[0101] Determine the preset label with the maximum entity probability as the predicted label of the predicted entity.

[0102] Among them, both the score threshold and the preset weight matrix are network parameters of the entity recognition layer.

[0103] Through the above implementation manner, it is possible to avoid performing label prediction on all text characters, thereby improving the generation efficiency of the predicted label.

[0104] S14. Adjust the preset network according to the text entity information, the predicted entity and the predicted label to obtain an entity information recognition model.

[0105] In at least one embodiment of the present invention, the entity information recognition model refers to the preset network with the smallest network loss value.

[0106] In at least one embodiment of the present invention, the electronic device adjusts the preset network according to the text entity information, the predicted entity and the predicted label to obtain an entity information recognition model, including:

[0107] Count the total number of entities of the training entity;

[0108] Calculate the number of predicted entities that are the same as the training entity as the first prediction number, and calculate the number of predicted labels that are the same as the entity label as the second prediction number;

[0109] Calculate the network loss value of the preset network according to the total amount of entities, the first predicted quantity, and the second predicted quantity. The calculation formula of the network loss value is as follows:

[0110] where y represents the network loss value, n represents the total amount of entities, x1 represents the first predicted quantity, and x2 represents the second predicted quantity;

[0111] Adjust the network parameters of the bidirectional long short-term memory network layer and the entity recognition layer based on the network loss value until the network loss value is less than a preset threshold to obtain the entity information recognition model.

[0112] Through the total amount of entities, the first predicted quantity, and the second predicted quantity, the network loss value can be accurately calculated, thereby improving the training accuracy of the entity information recognition model.

[0113] S15. Obtain the text to be parsed, and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information.

[0114] In at least one embodiment of the present invention, the text to be parsed refers to text information that needs to be analyzed for an industrial map. The text to be parsed can be extracted according to a request sent by a user.

[0115] The industrial entity information refers to a specific industry, for example, Industry A, and the origin entity information refers to a specific place, for example, Place C.

[0116] In at least one embodiment of the present invention, the electronic device obtaining the text to be parsed includes:

[0117] Obtain any text to be processed from a text library to be processed, and obtain the text identifier of the any text to be processed;

[0118] Obtain a target identifier corresponding to the text identifier from a text association table of the text library to be processed;

[0119] Based on the target identifier, obtain the text corresponding to the target identifier from the text library to be processed as the associated text of the any text to be processed;

[0120] Concatenate the any text to be processed and the associated text according to the association order in the text association table to obtain the text to be parsed.

[0121] In at least one embodiment of the present invention, the manner in which the electronic device inputs the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information is similar to the manner in which the electronic device analyzes the industrial training text based on the preset network to obtain the predicted entity and the predicted label, and the present invention will not elaborate on this again.

[0122] S16. Perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain entity information pairs.

[0123] In at least one embodiment of the present invention, the entity information pair refers to the pairing relationship between the industrial entity information and the origin entity information.

[0124] In at least one embodiment of the present invention, the electronic device performs syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain entity information pairs, including:

[0125] Screen text sentences from the text to be parsed based on the industrial entity information and the origin entity information;

[0126] Perform word segmentation processing on the text sentences to obtain multiple sentence word segments;

[0127] Identify the part-of-speech of each sentence word segment in the text sentence;

[0128] Perform dependency recognition on the industrial entity information and the origin entity information based on the sentence word segments whose part-of-speech is the preset part-of-speech to obtain the dependency relationship between the industrial entity information and the origin entity information;

[0129] Determine the industrial entity information and the origin entity information with the same set of dependency relationships as the entity information pair.

[0130] Among them, the preset part-of-speech is usually set as a verb.

[0131] For example, the text sentence is: Industry A and Industry B will focus on implementing in Origin C. After identifying the part-of-speech, it is obtained that the part-of-speech of the sentence word segment "implement" is a verb. After performing dependency recognition on the text sentence, it is obtained that Industry A depends on Origin C and Industry B depends on Origin C. Then the entity information pairs can be obtained as: Industry A - Origin C, Industry B - Origin C.

[0132] Through the above implementation manner, the dependency relationship can be accurately identified, and then based on the dependency relationship, the entity information pair can be quickly generated.

[0133] S17. Concatenate the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

[0134] It should be emphasized that, to further ensure the privacy and security of the above industrial map, the above industrial map can also be stored in a node of a blockchain.

[0135] In at least one embodiment of the present invention, the text order refers to the specific sorting of the industrial entity information in the text to be parsed.

[0136] The industrial map includes links formed by multiple industries and the origin corresponding to each industry. For example, the industrial map can be Industry G (Origin C) → Industry K (Origin F) → Industry L (Origin E).

[0137] It can be seen from the above technical solutions that the present invention trains a preset network based on the industrial training text and the text entity information, which can improve the recognition ability of the entity information recognition model for entity information. Then, the text to be parsed is parsed through the entity information recognition model to improve the accuracy of the industrial entity information and the origin entity information, and syntactic dependency matching processing is performed on the industrial entity information and the origin entity information to improve the accuracy of the entity information pair, thereby improving the generation accuracy of the industrial map. In addition, the present invention can quickly extract the industrial entity information and the origin entity information from the text to be parsed through the entity information recognition model, thereby improving the generation efficiency of the industrial map.

[0138] As Figure 2 shown, it is a functional module diagram of a preferred embodiment of the industrial map construction device of the present invention. The industrial map construction device 11 includes an acquisition unit 110, an encoding unit 111, an identification unit 112, an adjustment unit 113, an input unit 114, a matching unit 115, and a splicing unit 116. The module / unit referred to in the present invention means a series of computer-readable instruction segments that can be acquired by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0139] The acquisition unit 110 acquires an industrial training text and acquires the text entity information of the industrial training text.

[0140] In at least one embodiment of the present invention, the industrial training text refers to text information related to a specific industry, and the industrial training text includes a specific industry and the industrial execution location of the specific industry.

[0141] The text entity information includes a training entity and an entity label. The training entity refers to a specific industry or a specific execution location, and the entity label is used to identify that the training entity is an industry name, an industry location, etc.

[0142] The obtaining unit 110 obtains a preset network, and the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer.

[0143] In at least one embodiment of the present invention, the word vector layer is used to encode text characters.

[0144] The bidirectional long short-term memory network layer includes a forward long short-term memory network layer and a backward long short-term memory network layer. The forward long short-term memory network layer is used to perform forward time series prediction on text characters, and the backward long short-term memory network layer performs backward time series prediction on text characters.

[0145] The entity recognition layer is used to perform named entity recognition on text characters.

[0146] The encoding unit 111 encodes each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain a target time series vector for each text character.

[0147] In at least one embodiment of the present invention, the target time series vector includes forward time series information and backward time series information of each text character.

[0148] In at least one embodiment of the present invention, the encoding unit 111 encodes each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer, and obtaining the target time series vector for each text character includes:

[0149] Encoding the industrial training text based on the word vector layer to obtain a character vector for each text character;

[0150] Locate the character order of each text character in the industrial training text;

[0151] Input the character vectors into the forward long short-term memory network layer in ascending order of the character order to obtain a forward time series vector for each text character, and input the character vectors into the backward long short-term memory network layer in descending order of the character order to obtain a backward time series vector for each text character;

[0152] Concatenate the forward time series vector and the backward time series vector to obtain the target time series vector.

[0153] By inputting the character vectors into the bidirectional long short-term memory network layer according to the character order, it can assist the bidirectional long short-term memory network layer to quickly identify the character order of the character vectors in the industrial training text and improve the generation efficiency of the target time series vector.

[0154] Specifically, the encoding unit 111 inputs the character vectors into the forward long short-term memory network layer in ascending order of the character sequence, and the forward temporal vectors of each text character obtained include:

[0155] For any text character, obtain the neighboring characters whose character sequence is less than that of the text character as target characters;

[0156] Obtain the state vector of the target character;

[0157] Concatenate the state vector and the character vector of the text character to obtain an input vector;

[0158] Calculate the input vector based on the preset network matrix and preset bias value of the forward long short-term memory network layer to obtain the forward temporal vector of the text character.

[0159] Wherein, the state vector refers to the forward temporal vector of the target character. It should be noted that the state vector of the first text character in the industrial training text is the character vector.

[0160] Both the preset network matrix and the preset bias value are network parameters of the forward long short-term memory network layer.

[0161] Through the above implementation manner, the forward temporal vector can be quickly generated.

[0162] Specifically, the encoding unit 111 inputs the character vectors into the backward long short-term memory network layer in descending order of the character sequence, and the manner of obtaining the backward temporal vector of each text character is similar to the manner of the encoding unit 111 inputting the character vectors into the forward long short-term memory network layer in ascending order of the character sequence to obtain the forward temporal vector of each text character, and the present invention will not elaborate on this.

[0163] The recognition unit 112 recognizes the target temporal vector based on the entity recognition layer to obtain the predicted entity and predicted label of the industrial training text.

[0164] In at least one embodiment of the present invention, the predicted entity refers to the entity information obtained after the preset network recognizes the industrial training text, and the predicted label refers to the label information obtained after the preset network recognizes and predicts the entity information. Wherein, the predicted entity corresponds to the training entity, and the predicted label corresponds to the entity label.

[0165] In at least one embodiment of the present invention, the recognition unit 112 recognizes the target time series vector based on the entity recognition layer, and the predicted entities and predicted labels of the industrial training text obtained include:

[0166] Calculate the sum of each vector element in the target time series vector to obtain the character score of each text character;

[0167] Obtain the score threshold and the preset weight matrix from the entity recognition layer;

[0168] Determine the text characters with the character scores greater than the score threshold as the predicted entities;

[0169] Calculate the product of the target time series vector corresponding to the predicted entity and the preset weight matrix to obtain the entity probability of the predicted entity on each preset label;

[0170] Determine the preset label with the maximum entity probability as the predicted label of the predicted entity.

[0171] Wherein, both the score threshold and the preset weight matrix are network parameters of the entity recognition layer.

[0172] Through the above implementation manner, it is possible to avoid predicting labels for all text characters, thereby improving the generation efficiency of the predicted labels.

[0173] The adjustment unit 113 adjusts the preset network according to the text entity information, the predicted entity and the predicted label to obtain an entity information recognition model.

[0174] In at least one embodiment of the present invention, the entity information recognition model refers to the preset network with the minimum network loss value.

[0175] In at least one embodiment of the present invention, the adjustment unit 113 adjusts the preset network according to the text entity information, the predicted entity and the predicted label to obtain an entity information recognition model, including:

[0176] Count the total number of entities of the training entities;

[0177] Calculate the number of predicted entities that are the same as the training entities as the first prediction quantity, and calculate the number of predicted labels that are the same as the entity labels as the second prediction quantity;

[0178] Calculate the network loss value of the preset network according to the total number of entities, the first prediction quantity and the second prediction quantity. The calculation formula of the network loss value is:

[0179] Wherein, y refers to the network loss value, n refers to the total number of entities, x1 refers to the first predicted quantity, and x2 refers to the second predicted quantity;

[0180] Adjust the network parameters of the bidirectional long short-term memory network layer and the entity recognition layer based on the network loss value until the network loss value is less than a preset threshold to obtain the entity information recognition model.

[0181] Through the total number of entities, the first predicted quantity, and the second predicted quantity, the network loss value can be accurately calculated, thereby improving the training accuracy of the entity information recognition model.

[0182] The input unit 114 obtains the text to be parsed and inputs the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information.

[0183] In at least one embodiment of the present invention, the text to be parsed refers to text information that needs to be analyzed for an industrial map. The text to be parsed can be extracted according to a request sent by a user.

[0184] The industrial entity information refers to a specific industry, for example, Industry A, and the origin entity information refers to a specific place, for example, Place C.

[0185] In at least one embodiment of the present invention, the input unit 114 obtaining the text to be parsed includes:

[0186] Obtain any text to be processed from the text library to be processed and obtain the text identifier of the any text to be processed;

[0187] Obtain a target identifier corresponding to the text identifier from the text association table of the text library to be processed;

[0188] Based on the target identifier, obtain the text corresponding to the target identifier from the text library to be processed as the associated text of the any text to be processed;

[0189] Concatenate the any text to be processed and the associated text according to the association order in the text association table to obtain the text to be parsed.

[0190] In at least one embodiment of the present invention, the input unit 114 inputs the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information in a manner similar to analyzing the industrial training text based on the preset network to obtain the predicted entity and the predicted label, and the present invention will not elaborate on this.

[0191] The matching unit 115 performs a syntactic dependency matching process on the industrial entity information and the origin entity information to obtain entity information pairs.

[0192] In at least one embodiment of the present invention, the entity information pair refers to the pairing relationship between the industrial entity information and the origin entity information.

[0193] In at least one embodiment of the present invention, the matching unit 115 performs a syntactic dependency matching process on the industrial entity information and the origin entity information to obtain entity information pairs, including:

[0194] Filter text sentences from the text to be parsed based on the industrial entity information and the origin entity information;

[0195] Perform word segmentation on the text sentences to obtain multiple sentence word segments;

[0196] Identify the part-of-speech of each sentence word segment in the text sentences;

[0197] Based on the sentence word segments whose part-of-speech is the preset part-of-speech, perform dependency recognition on the industrial entity information and the origin entity information to obtain the dependency relationship between the industrial entity information and the origin entity information;

[0198] Determine the industrial entity information and the origin entity information with the same set of dependency relationships as the entity information pair.

[0199] Among them, the preset part-of-speech is usually set to a verb.

[0200] For example, the text sentence is: Industry A and Industry B will focus on implementing in Origin C. After identifying the part-of-speech, it is obtained that the part-of-speech of the sentence word segment "implement" is a verb. After performing dependency recognition on the text sentence, it is obtained that Industry A depends on Origin C and Industry B depends on Origin C. Then the entity information pairs that can be obtained are: Industry A - Origin C, Industry B - Origin C.

[0201] Through the above implementation manner, the dependency relationship can be accurately identified, and then based on the dependency relationship, the entity information pair can be quickly generated.

[0202] The splicing unit 116 splices the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

[0203] It should be emphasized that to further ensure the privacy and security of the above industrial map, the above industrial map can also be stored in a node of a blockchain.

[0204] In at least one embodiment of the present invention, the text order refers to the specific sorting of the industrial entity information in the text to be parsed.

[0205] The industry map includes links formed by multiple industries and the production areas corresponding to each industry. For example, the industry map may be G industry (production area C) → K industry (production area F) → L industry (production area E).

[0206] It can be seen from the above technical solutions that the present invention trains the preset network based on the industry training text and the text entity information, which can improve the recognition ability of the entity information recognition model for entity information, and then parse the text to be parsed through the entity information recognition model, improve the accuracy of the industry entity information and the place of origin entity information, perform syntactic dependency matching on the industry entity information and the place of origin entity information, improve the accuracy of the entity information pair, and thus improve the generation accuracy of the industry map. In addition, the present invention can quickly extract the industry entity information and the place of origin entity information from the text to be parsed through the entity information recognition model, thereby improving the generation efficiency of the industry map.

[0207] like Figure 3 The figure is a schematic diagram of the structure of an electronic device of a preferred embodiment of the method for constructing an industrial map according to the present invention.

[0208] In one embodiment of the present invention, the electronic device 1 includes, but is not limited to, a memory 12, a processor 13, and computer-readable instructions stored in the memory 12 and executable on the processor 13, such as an industry map building program.

[0209] Those skilled in the art will understand that the schematic diagram is merely an example of the electronic device 1 and does not constitute a limitation on the electronic device 1. The electronic device 1 may include more or fewer components than shown in the diagram, or a combination of certain components, or different components. For example, the electronic device 1 may also include input and output devices, network access devices, buses, etc.

[0210] The processor 13 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 13 is the operation core and control center of the electronic device 1, connecting various parts of the entire electronic device 1 through various interfaces and circuits, and executing the operating system of the electronic device 1 and various installed application programs, program codes, etc.

[0211] Exemplarily, the computer-readable instructions may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and the computer-readable instruction segments are used to describe the execution process of the computer-readable instructions in the electronic device 1. For example, the computer-readable instructions may be divided into an acquisition unit 110, an encoding unit 111, an identification unit 112, an adjustment unit 113, an input unit 114, a matching unit 115, and a splicing unit 116.

[0212] The memory 12 can be used to store the computer-readable instructions and / or modules. The processor 13 realizes various functions of the electronic device 1 by running or executing the computer-readable instructions and / or modules stored in the memory 12, and by calling the data stored in the memory 12. The memory 12 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the electronic device. The memory 12 may include non-volatile and volatile memories, such as: hard disks, memories, plug-in hard disks, SmartMedia Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other storage devices.

[0213] The memory 12 can be an external memory and / or an internal memory of the electronic device 1. Further, the memory 12 can be a memory in physical form, such as a memory stick, a TF card (Trans-flash Card), and so on.

[0214] If the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by a processor, the steps of the above method embodiments can be implemented.

[0215] Among them, the computer-readable instructions include computer-readable instruction codes, and the computer-readable instruction codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer-readable instruction codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory).

[0216] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed industrial map construction, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0217] Combined with Figure 1 , the memory 12 in the electronic device 1 stores computer-readable instructions to implement an industrial map construction method, and the processor 13 can execute the computer-readable instructions to thereby implement:

[0218] Obtain industrial training texts and obtain the text entity information of the industrial training texts;

[0219] Obtain a preset network, where the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer;

[0220] Encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain the target time series vector of each text character;

[0221] Based on the entity recognition layer, identify the target time series vector to obtain the predicted entity and predicted label of the industrial training text;

[0222] Adjust the preset network according to the text entity information, the predicted entity and the predicted label to obtain an entity information recognition model;

[0223] Obtain the text to be parsed, and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information;

[0224] Perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain an entity information pair;

[0225] Stitch the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

[0226] Specifically, the specific implementation method of the above computer-readable instructions by the processor 13 can refer to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be repeated here.

[0227] In several embodiments provided by the present invention, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0228] Computer-readable instructions are stored on the computer-readable storage medium, wherein the computer-readable instructions, when executed by the processor 13, are used to implement the following steps:

[0229] Obtain an industrial training text, and obtain the text entity information of the industrial training text;

[0230] Obtain a preset network, the preset network includes a word vector layer, a bidirectional long short-term memory network layer and an entity recognition layer;

[0231] Encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain the target time series vector of each text character;

[0232] Based on the entity recognition layer, identify the target time series vector to obtain the predicted entity and predicted label of the industrial training text;

[0233] Adjust the preset network according to the text entity information, the predicted entity and the predicted label to obtain an entity information recognition model;

[0234] Obtain the text to be parsed, and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information;

[0235] Perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain entity information pairs;

[0236] Splice the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

[0237] The module described as a separate component may or may not be physically separated, and the component shown as a module may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0238] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of hardware plus software functional modules.

[0239] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any appended drawing reference signs in the claims should not be regarded as limiting the claimed rights.

[0240] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The described multiple units or devices can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not represent any specific order.

[0241] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for constructing an industrial map, characterized in that The industrial map construction method includes: Obtain industrial training texts, and obtain the text entity information of the industrial training texts, where the text entity information includes training entities and entity labels; Obtain a preset network, where the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer; Encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain the target time series vector of each text character; Identify the predicted entities and predicted labels of the industrial training text based on the entity recognition layer for the target time series vector; Adjust the preset network according to the text entity information, the predicted entity, and the predicted label to obtain an entity information recognition model, including: counting the total number of entities of the training entities; calculating the number of predicted entities that are the same as the training entities as the first prediction number, and calculating the number of predicted labels that are the same as the entity labels as the second prediction number; calculating the network loss value of the preset network according to the total number of entities, the first prediction number, and the second prediction number, and the calculation formula of the network loss value is: , where refers to the network loss value, refers to the total number of entities, refers to the first prediction number, refers to the second prediction number; adjust the network parameters of the bidirectional long short-term memory network layer and the entity recognition layer based on the network loss value until the network loss value is less than a preset threshold to obtain the entity information recognition model; Obtain the text to be parsed, and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information; Perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain entity information pairs; Splice the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial map.

2. The industrial map construction method according to claim 1, wherein The bidirectional long short-term memory network layer includes a forward long short-term memory network layer and a backward long short-term memory network layer. The encoding of each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain the target time series vector of each text character includes: Encode the industrial training text based on the word vector layer to obtain the character vector of each text character; Locate the character order of each text character in the industrial training text; Input the character vectors into the forward long short-term memory network layer in ascending order of the character order to obtain the forward time series vector of each text character, and input the character vectors into the backward long short-term memory network layer in descending order of the character order to obtain the backward time series vector of each text character; Splice the forward time series vector and the backward time series vector to obtain the target time series vector.

3. The industrial map construction method according to claim 2, wherein The inputting of the character vectors into the forward long short-term memory network layer in ascending order of the character order to obtain the forward time series vector of each text character includes: For any text character, obtain the neighboring characters whose character order is less than that of the any text character as target characters; Obtain the state vector of the target character; Splice the state vector and the character vector of the any text character to obtain an input vector; Calculate the input vector based on the preset network matrix and preset bias value of the forward long short-term memory network layer to obtain the forward time series vector of the any text character.

4. The industrial map construction method according to claim 1, wherein The identifying of the predicted entities and predicted labels of the industrial training text based on the entity recognition layer for the target time series vector includes: Calculate the sum of each vector element in the target time series vector to obtain the character score of each text character; Obtain the score threshold and the preset weight matrix from the entity recognition layer; Determine the text characters with character scores greater than the score threshold as the predicted entities; Calculate the product of the target time series vector corresponding to the predicted entity and the preset weight matrix to obtain the entity probability of the predicted entity on each preset label; Determine the preset label with the maximum entity probability as the predicted label of the predicted entity.

5. The industrial map construction method according to claim 1, characterized in that The syntactic dependency matching process for the industrial entity information and the origin entity information to obtain the entity information pair includes: Filter text sentences from the text to be parsed based on the industrial entity information and the origin entity information; Perform word segmentation on the text sentences to obtain multiple sentence word segments; Identify the word segmentation part-of-speech of each sentence word segment in the text sentence; Based on the sentence word segments with the preset part-of-speech, perform dependency recognition on the industrial entity information and the origin entity information to obtain the dependency relationship between the industrial entity information and the origin entity information; Determine the industrial entity information and the origin entity information with the same set of dependency relationships as the entity information pair.

6. The industrial map construction method according to claim 1, characterized in that, The obtaining of the text to be parsed includes: Obtain any text to be processed from the text library to be processed, and obtain the text identifier of the any text to be processed; Obtain the target identifier corresponding to the text identifier from the text association table of the text library to be processed; Based on the target identifier, obtain the text corresponding to the target identifier from the text library to be processed as the associated text of the any text to be processed; Splice the any text to be processed and the associated text according to the association order in the text association table to obtain the text to be parsed.

7. An industrial map construction device, characterized in that The industrial graph construction device includes: An obtaining unit, configured to obtain industrial training texts, and obtain text entity information of the industrial training texts, where the text entity information includes training entities and entity labels; The obtaining unit is further configured to obtain a preset network, where the preset network includes a word vector layer, a bidirectional long short-term memory network layer, and an entity recognition layer; An encoding unit, configured to encode each text character in the industrial training text based on the word vector layer and the bidirectional long short-term memory network layer to obtain a target time series vector for each text character; An identifying unit, configured to identify the target time series vector based on the entity recognition layer to obtain the predicted entity and predicted label of the industrial training text; An adjustment unit, configured to adjust the preset network according to the text entity information, the predicted entity, and the predicted label to obtain an entity information recognition model, including: counting the total number of entities of the training entities; calculating the number of predicted entities that are the same as the training entities as the first prediction quantity, and calculating the number of predicted labels that are the same as the entity labels as the second prediction quantity; calculating a network loss value of the preset network according to the total number of entities, the first prediction quantity, and the second prediction quantity, and the calculation formula of the network loss value is: , where refers to the network loss value, refers to the total number of entities, refers to the first prediction quantity, refers to the second prediction quantity; adjusting network parameters of the bidirectional long short-term memory network layer and the entity recognition layer based on the network loss value until the network loss value is less than a preset threshold to obtain the entity information recognition model; An input unit, configured to obtain the text to be parsed, and input the text to be parsed into the entity information recognition model to obtain industrial entity information and origin entity information; A matching unit, configured to perform syntactic dependency matching processing on the industrial entity information and the origin entity information to obtain an entity information pair; A splicing unit, configured to splice the entity information pairs according to the text order of the industrial entity information in the text to be parsed to obtain an industrial graph.

8. An electronic device, characterized in that, The electronic device includes: A memory storing computer-readable instructions; and A processor that executes the computer-readable instructions stored in the memory to implement the industrial graph construction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are executed by a processor in an electronic device to implement the industrial map construction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Enterprise information map construction method and device, computer device and storage medium

    CN109376273A

  • Named entity identification method and device, equipment and computer readable storage medium

    CN110298019A