Natural gas customer gas consumption unstructured data extraction method and device
By obtaining the text of gas data used by natural gas customers, converting it into strings and processing it in a classified manner, the complexity and universality of unstructured data extraction are solved, and efficient structured data acquisition is achieved, supporting subsequent analysis and knowledge graph creation.
Patent Information
- Application Number
- CN202311856880.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has problems of complexity and poor versatility in the process of unstructured gas data extraction for natural gas customers, which affects the effect of subsequent customer characteristic analysis and knowledge graph creation.
By obtaining the text of customer gas data, converting it into a string, coarse-grained data is extracted, and divided into description text, table text and reading comprehension text according to the data type. Different processing methods are used to convert it into structured data, including cleaning, target field correspondence, BERT model annotation and efficient global pointer model relationship extraction.
It realizes fast, efficient and accurate gas data extraction for natural gas customers, providing high-quality data for subsequent feature analysis and knowledge graph creation.
Smart Images

Figure CN120234359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of radiation technology, and particularly to a method and device for extracting unstructured data of natural gas customers' gas consumption. Background Art
[0002] To maximize the value of natural gas, it is crucial to deeply understand customers' gas consumption characteristics and create a knowledge graph, which not only helps sales companies optimize the supply chain, but also enhances service quality, improves user satisfaction, and strengthens the competitive advantage of natural gas sales companies. Information collection and extraction is the first and key technical basis for customer characteristic analysis and knowledge graph creation. Structured data collection and cleaning technologies are already very mature. Unstructured data extraction and conversion are highly related to document content and target data formats. The extraction process is somewhat complex and has poor generality. If the extraction quality is poor, it will greatly reduce the effect of subsequent customer characteristic analysis and knowledge graph creation. Therefore, how to automatically extract and convert unstructured data of natural gas customers into structured data has become an urgent problem to be solved. Summary of the Invention
[0003] In view of the above problems, the present invention is proposed to provide a method and device for extracting unstructured data of natural gas customers' gas consumption that overcome or at least partially solve the above problems.
[0004] In a first aspect, an embodiment of the present invention provides a method for extracting unstructured data of natural gas customers' gas consumption, including:
[0005] Obtain the text of customers' gas consumption data and convert the text of the customers' gas consumption data into a string;
[0006] Extract the rough-grained data of customers' gas consumption from the obtained string;
[0007] According to the type of the extracted content, divide the rough-grained data of customers' gas consumption into multiple categories, and process the rough-grained data of customers' gas consumption under each category respectively to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text.
[0008] In one embodiment, the extracting the rough-grained data of customers' gas consumption from the obtained string includes:
[0009] Split the obtained string;
[0010] Filter the split string according to preset conditions to obtain the rough-grained data of customers' gas consumption.
[0011] In one embodiment, the category is descriptive text. The coarse-grained customer gas consumption data under the descriptive text category is processed to obtain structured data under the descriptive text category, including:
[0012] Clean the coarse-grained customer gas consumption data to remove preset special characters and characters that do not conform to the preset rules, so as to obtain structured data under the descriptive text category.
[0013] In one embodiment, the category is tabular text. The coarse-grained customer gas consumption data under the tabular text category is processed to obtain structured data under the tabular text category, including:
[0014] Correspond the coarse-grained customer gas consumption data with preset target fields to obtain the correspondence between the tabular data and the preset target fields;
[0015] Obtain the position of the tabular data in the coarse-grained customer gas consumption data;
[0016] Obtain the coarse-grained customer gas consumption data except for the preset special rows;
[0017] According to the correspondence between the coarse-grained customer gas consumption data and the preset target fields, splice the coarse-grained customer gas consumption data into a complete tabular data, and the complete tabular data is used as the structured data under the tabular text category.
[0018] In one embodiment, the category is reading comprehension text. The coarse-grained customer gas consumption data under the reading comprehension text is processed to obtain structured data under the reading comprehension text category, including:
[0019] Obtain coarse-grained text data of the same type;
[0020] Divide the coarse-grained text data into training set data and validation set data according to a preset ratio;
[0021] For the training set data, use the BIESO annotation method to annotate various entities in the coarse-grained text data;
[0022] Input the annotated data into the pre-trained BERT Chinese pre-training model to obtain the corresponding sequence of word vectors;
[0023] Input the obtained sequence of word vectors into the efficient global pointer model to extract the relationship between the head entity and the tail entity;
[0024] The training set is divided into multiple batches to train the efficient global pointer model, and the model is optimized by minimizing the loss function. The global efficient pointer model is then predicted using the validation set data to obtain the corresponding head entity, tail entity, head relationship of the entity, and tail relationship of the entity. The prediction results are evaluated using precision, recall, and F1 values until the preset prediction results are achieved;
[0025] Using the trained efficient global pointer model, the structural data of the head entity, tail entity, entity head, and entity tail relationship corresponding to the reading comprehension type text data is output.
[0026] In a second aspect, an embodiment of the present invention provides an unstructured data extraction device for natural gas customers' gas usage, including:
[0027] An acquisition module for acquiring the text of the customer gas usage data and converting the text of the customer gas usage data into a string;
[0028] A rough extraction module for extracting the customer gas usage coarse-grained data from the converted string;
[0029] A classification processing module for dividing the customer gas usage coarse-grained data into multiple categories according to the type of the extracted content, and processing the customer gas usage coarse-grained data under each category respectively to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text.
[0030] In one embodiment, the device further includes:
[0031] A splitting module for splitting the converted string;
[0032] A screening module for screening the split string according to preset conditions to obtain the coarse-grained data.
[0033] In one embodiment, the device further includes:
[0034] A descriptive text module for processing the coarse-grained data under the descriptive text category to obtain structured data under the descriptive text category;
[0035] A tabular text module for processing the coarse-grained data under the tabular text category to obtain structured data under the tabular text category;
[0036] A reading comprehension module for processing the coarse-grained data under the reading comprehension text to obtain structured data under the reading comprehension text category.
[0037] In a third aspect, an embodiment of the present invention provides a computing device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for extracting unstructured data of natural gas customer gas consumption as described above is implemented.
[0038] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for extracting unstructured data of natural gas customer gas consumption as described above is implemented.
[0039] The beneficial effects of the above technical solutions provided by the embodiments of the present invention at least include:
[0040] An embodiment of the present invention provides a method for extracting unstructured data of natural gas customer gas consumption, including: obtaining the text of customer gas consumption data and converting the text of customer gas consumption data into a string; extracting the rough-grained data of customer gas consumption from the obtained string; dividing the rough-grained data of customer gas consumption into multiple categories according to the type of the extracted content, and processing the rough-grained data of customer gas consumption under each category respectively to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text. By first obtaining the text of customer gas consumption data and then adopting different data structuring methods according to the target field type, the embodiment of the present invention realizes determining the extraction method according to the characteristics of the natural gas customer data document, and quickly, efficiently, and accurately obtains the natural gas customer gas consumption data, providing high-quality data for subsequent analysis of natural gas customer characteristics and creation of a knowledge graph.
[0041] Other features and advantages of the present invention will be described in the following description, and part of them will be obvious from the description or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written description, claims, and drawings.
[0042] The technical solutions of the present invention will be further described in detail below with reference to the drawings and embodiments. Description of the Drawings
[0043] The drawings are used to provide a further understanding of the present invention, and constitute a part of the description. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0044] Figure 1 is a flowchart of a method for extracting unstructured data of natural gas customer gas consumption provided by an embodiment of the present invention;
[0045] Figure 2 is an example diagram of descriptive text data provided by an embodiment of the present invention;
[0046] Figure 3 It is a structural block diagram of an unstructured data extraction device for natural gas customers' gas consumption;
[0047] Figure 4 A method for reading unstructured data of natural gas customers' gas consumption provided by an embodiment of the present invention. Detailed implementation manners
[0048] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0049] To automatically extract unstructured data of natural gas customers and convert it into structured data, an embodiment of the present invention provides a method for extracting unstructured data of natural gas customers' gas consumption, and its flowchart is as Figure 1 shown, including:
[0050] Step S11: Obtain the text of the customer gas consumption data and convert the text of the customer gas consumption data into a string;
[0051] Step S12: Extract the customer gas consumption coarse-grained data from the obtained string;
[0052] Step S13: According to the type of the extracted content, divide the customer gas consumption coarse-grained data into multiple categories, and process the customer gas consumption coarse-grained data under each category respectively to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text.
[0053] In the above step S11, to obtain user data, for example, the content can be read in the way of Apache Tika. Tika is a top-level project of the Apache organization. By using existing parsing libraries, metadata and structured content can be detected and extracted from documents in different formats (such as HTML, PDF, Doc); the specific steps are as follows:
[0054] Upload the natural gas customer gas consumption data document to the server;
[0055] The system reads and parses the document content by calling the Apache Tika library and converts the content into a string;
[0056] Save the string to the database.
[0057] In the above step S12, the customer gas consumption coarse-grained data is extracted from the converted string. For example, the following method can be used:
[0058] Split the converted string.
[0059] According to the preset conditions, filter the split string to obtain the customer gas consumption coarse-grained data.
[0060] In the above step S13, if the category of the customer gas consumption coarse-grained data is descriptive text, process the customer gas consumption coarse-grained data under the descriptive text category to obtain the structured data under the descriptive text category, including:
[0061] Clean the customer gas consumption coarse-grained data to remove the preset special characters and characters that do not conform to the preset rules, and obtain the structured data under the descriptive text category.
[0062] An example of descriptive text data is as Figure 2 shown.
[0063] If the category of the customer gas consumption coarse-grained data is tabular text, process the customer gas consumption coarse-grained data under the tabular text category to obtain the structured data under the tabular text category, including:
[0064] Correspond the customer gas consumption coarse-grained data with the preset target fields to obtain the correspondence between the tabular data and the preset target fields;
[0065] Obtain the position of the tabular data in the customer gas consumption coarse-grained data;
[0066] Obtain the customer gas consumption coarse-grained data except for the preset special rows;
[0067] According to the correspondence between the customer gas consumption coarse-grained data and the preset target fields, splice the customer gas consumption coarse-grained data into a complete tabular data, and the complete tabular data is used as the structured data under the tabular text category.
[0068] An example of tabular text data is shown in Table 1, and its structured example is shown in Table 2.
[0069] Table 1
[0070]
[0071] Table 2
[0072]
[0073] Among text processing methods for reading comprehension, the most complex one is to extract the relationships between entities from the entire text. In the embodiments of the present invention, a relatively advanced GPLinker (GlobalPointer-based Linking) model is adopted for relationship extraction. This model is based on Efficient GlobalPointer for joint event extraction and expands the traditional relationship triple (i.e., subject, predicate, object) into a five-tuple (i.e., S h , S t , p, O h , O t , where S h , S t are the start and end positions of s respectively, while O h , O t are the start and end positions of o respectively). Based on this, relationship extraction is carried out, and the specific steps are as follows:
[0074] If the category of the customer gas consumption coarse-grained data is reading comprehension text, the customer gas consumption coarse-grained data under the reading comprehension text is processed to obtain structured data under the reading comprehension text category, including:
[0075] Obtain coarse-grained text data of the same category;
[0076] Divide the coarse-grained text data into training set data and validation set data according to a preset ratio;
[0077] For the training set data, use the BIESO annotation method to annotate various entities in the coarse-grained text data; where B represents the start position of the entity, I represents the inside of the entity, E represents the end position of the entity, S represents that the entity is a single character, and 0 represents a non-entity;
[0078] Input the annotated data into the pre-trained MacBERT Chinese pre-trained model to obtain the corresponding text vector sequences h1, h2,... h n ;
[0079] Input the obtained text vector sequences into the efficient global pointer model to extract the relationships between the head entity and the tail entity:
[0080] The first step: Perform a full connection layer conversion on the BERT output vector h n to obtain the sequence vectors [q 1,α , q 2,α ,..., p n,α and [k 1,α , k 2,α ,..., k n,α ;
[0081] Step 2: The sequence vectors are brought into the scoring function for calculation to obtain the head entity, tail entity, the head relation of the entity, and the tail relation of the entity respectively;
[0082] The formula of the scoring function is as follows:
[0083]
[0084] where S α (i,j) is the scoring function of the α-type span S[i:j] from i to j, q i,α is the start position, k j,α is the end position, and R j-i is the relative position.
[0085] When S(S h ,S t )>0, the head entity can be obtained. When S(O h ,O t )>0, the tail entity can be obtained. When S(S h ,O h |P)>0, the head relation of the entity can be obtained. When S(S t ,S t |P)>0, the tail relation of the entity can be obtained.
[0086] Step 3: The training set is divided into multiple batches, the efficient global pointer model is trained, and the model is optimized by minimizing the loss function. The formula of the loss function is as follows:
[0087]
[0088] where N is the set of negative categories of training samples.
[0089] Use the data in the validation set to predict the global efficient pointer model to obtain the corresponding head entity, tail entity, the head relation of the entity, and the tail relation of the entity. Evaluate the prediction results through precision, recall, and F1 value until the preset prediction results are achieved;
[0090] Use the trained efficient global pointer model to output the structural data of the head entity, tail entity, the head of the entity, and the tail of the entity corresponding to the reading comprehension text data.
[0091] The ratio of the aforementioned preset training set to the validation set can be, for example, 8:2.
[0092] After testing, the precision of this model is 93.33%, the recall rate is 92.35%, and the F1 value is 89.56%. It proves that the model adopted in the embodiments of the present invention has high effectiveness.
[0093] The structured data finally obtained in the embodiments of the present invention can be optionally stored in a database. In the embodiments of the present invention, customer gas consumption data is read from the database, and the structured data finally obtained in the embodiments of the present invention is stored only in the database, which is a method for reading unstructured data of natural gas customer gas consumption. The flowchart is as Figure 4 shown. In Figure 4 , the "start" above refers to the start of reading data by the method for reading unstructured data of natural gas customer gas consumption, and the "start" below refers to the successful reading of unstructured data of natural gas customer gas consumption and the start of subsequent customer characteristic analysis and knowledge graph creation.
[0094] Since the principle of the problem solved by the method for reading unstructured data of natural gas customer gas consumption is similar to that of the aforementioned method for extracting unstructured data of natural gas customer gas consumption, the implementation of this method can refer to the implementation of the aforementioned method, and the repeated parts will not be elaborated.
[0095] Based on the same inventive concept, the embodiments of the present invention also provide a device for extracting unstructured data of natural gas customer gas consumption. The structural block diagram is as Figure 3 shown, including:
[0096] An acquisition module 31, configured to acquire the text of customer gas consumption data and convert the text of customer gas consumption data into a string;
[0097] A rough extraction module 32, configured to extract customer gas consumption coarse-grained data from the converted string;
[0098] A classification processing module 33, configured to divide the customer gas consumption coarse-grained data into multiple categories according to the type of the extracted content, and process the customer gas consumption coarse-grained data under each category respectively to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text.
[0099] The above device further includes:
[0100] A splitting module, configured to split the converted string;
[0101] A screening module, configured to screen the split string according to preset conditions to obtain coarse-grained data.
[0102] The above device further includes:
[0103] A descriptive text module, configured to process the coarse-grained data under the descriptive text category to obtain structured data under the descriptive text category;
[0104] A tabular text module, configured to process the coarse-grained data under the tabular text category to obtain structured data under the tabular text category;
[0105] A reading comprehension module is used to process the coarse-grained data under the reading comprehension text to obtain structured data under the reading comprehension text category.
[0106] Based on the same inventive concept, an embodiment of the present invention further provides a computing device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for extracting unstructured data of natural gas customers' gas usage.
[0107] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements a method for extracting unstructured data of natural gas customers' gas usage.
[0108] Since the principles of the problems solved by these devices are similar to those of the aforementioned method for extracting unstructured data of natural gas customers' gas usage, the implementation of these devices can refer to the implementation of the aforementioned method, and the repeated parts will not be elaborated.
[0109] An embodiment of the present invention provides a method for extracting unstructured data of natural gas customers' gas usage, including: obtaining the text of the customers' gas usage data and converting the text of the customers' gas usage data into a string; extracting the coarse-grained data of the customers' gas usage from the obtained string; dividing the coarse-grained data of the customers' gas usage into multiple categories according to the type of the extracted content, and respectively processing the coarse-grained data of the customers' gas usage under each category to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text. By first obtaining the text of the customers' gas usage data and then adopting different data structuring methods according to the target field type, the embodiment of the present invention realizes determining the extraction method according to the characteristics of the natural gas customer data document, quickly, efficiently, and accurately obtains the natural gas customers' gas usage data, and provides high-quality data for subsequent analysis of natural gas customer characteristics and creation of a knowledge graph.
[0110] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0111] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0112] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks. Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for extracting unstructured data of natural gas customers' gas consumption, characterized in that, including: obtaining the text of the customer's gas usage data and converting the text of the customer's gas usage data into a string; extracting the customer's coarse-grained gas usage data from the converted string; dividing the customer's coarse-grained gas usage data into multiple categories according to the type of the extracted content, and processing the customer's coarse-grained gas usage data under each category respectively to obtain structured data under each category; the multiple categories include: descriptive text, tabular text, and reading comprehension text.
2. The method according to claim 1, characterized in that, The extracting the customer's coarse-grained gas usage data from the converted string includes: splitting the converted string; filtering the split string according to preset conditions to obtain the customer's coarse-grained gas usage data.
3. The method according to claim 1, characterized in that When the category is descriptive text, processing the customer's coarse-grained gas usage data under the descriptive text category to obtain structured data under the descriptive text category includes: cleaning the customer's coarse-grained gas usage data to remove preset special characters and characters that do not conform to preset rules to obtain structured data under the descriptive text category.
4. The method according to claim 1, wherein When the category is tabular text, processing the customer's coarse-grained gas usage data under the tabular text category to obtain structured data under the tabular text category includes: corresponding the customer's coarse-grained gas usage data with preset target fields to obtain the corresponding relationship between the tabular data and the preset target fields; obtaining the position of the tabular data in the customer's coarse-grained gas usage data; obtaining the customer's coarse-grained gas usage data except for preset special rows; according to the corresponding relationship between the customer's coarse-grained gas usage data and the preset target fields, splicing the customer's coarse-grained gas usage data into a complete tabular data, and the complete tabular data is used as the structured data under the tabular text category.
5. The method according to claim 1, characterized in that, When the category is reading comprehension text, processing the customer's coarse-grained gas usage data under the reading comprehension text to obtain structured data under the reading comprehension text category includes: obtaining coarse-grained text data of the same type; dividing the coarse-grained text data into training set data and validation set data according to a preset ratio; for the training set data, using the BIESO annotation method to annotate various entities in the coarse-grained text data; inputting the annotated data into a pre-trained BERT Chinese pre-training model to obtain a corresponding sequence of word vectors; inputting the obtained sequence of word vectors into an efficient global pointer model to extract the relationship between the head entity and the tail entity; dividing the training set into multiple batches, training the efficient global pointer model, optimizing the model by reducing the loss function, and using the validation set data to predict the global efficient pointer model to obtain the corresponding head entity, tail entity, head relationship of the entity, and tail relationship of the entity, and evaluating the prediction result through precision, recall, and F1 value until the preset prediction result is achieved; using the trained efficient global pointer model to output the structured data of the corresponding head entity, tail entity, entity head, and entity tail relationship of the reading comprehension text data.
6. An unstructured data extraction device for natural gas customers' gas usage, characterized in that, including: An acquisition module, configured to acquire the text of customer gas consumption data and convert the text of the customer gas consumption data into a string; A rough extraction module, configured to extract the rough-grained customer gas consumption data from the string obtained after conversion; A classification processing module, configured to divide the rough-grained customer gas consumption data into multiple categories according to the type of the extracted content, and process the rough-grained customer gas consumption data under each category respectively to obtain the structured data under each category; The multiple categories include: descriptive text, tabular text, and reading comprehension text.
7. The device according to claim 6, characterized in that, The device further includes: A splitting module, configured to split the string obtained after conversion; A screening module, configured to screen the split string according to preset conditions to obtain the rough-grained data.
8. The device according to claim 6, characterized in that, The device further includes: A descriptive text module, configured to process the rough-grained data under the descriptive text category to obtain the structured data under the descriptive text category; A tabular text module, configured to process the rough-grained data under the tabular text category to obtain the structured data under the tabular text category; A reading comprehension module, configured to process the rough-grained data under the reading comprehension text to obtain the structured data under the reading comprehension text category.
9. A computing device, characterized in that, Including: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the method for extracting unstructured data of natural gas customer gas consumption according to any one of claims 1-5.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method for extracting unstructured data of natural gas customer gas consumption according to any one of claims 1-5.