Method for rapidly extracting data and filling text
By defining and annotating data formats, establishing data relationships and generating data reading template files, biopharmaceutical companies can quickly extract and fill in the required data from large-scale production data, solving the problem of time-consuming and error-prone existing methods, and improving work efficiency and the accuracy of data filling.
Patent Information
- Application Number
- CN202510131577.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-06
AI Technical Summary
When using a production and manufacturing management system, biopharmaceutical companies need to extract and fill in data from a large number of production and testing data. The existing methods need to search, form reports and enter them manually through databases, which is time-consuming and error-prone.
By defining the space data format of the fill-in text and marking the data structure in the database, establishing the relationship between the data set, data points and data application types, generating data reading template files, calling the template file to scan the placeholders in the fill-in file line by line, loading and saving data according to the data format, realizing automatic filling of data.
It realizes the rapid retrieval and accurate reporting of required data from large amounts of production data, greatly improving work efficiency, reducing human errors, and quickly generating and filling documents.
Smart Images

Figure CN120068831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a method for quickly extracting filled text of a data set. Background Art
[0002] The Manufacturing Execution System (MES) is a management system that integrates production processes such as planning, production, quality control, inventory management, and material requisition; it is an indispensable part of modern manufacturing enterprises. It can not only help enterprises improve production efficiency and product quality, but also promote the digital transformation of enterprises. When using the Manufacturing Execution System, especially for biopharmaceutical enterprises, it is necessary to fill in data in a format specified by relevant departments to form a number of batch release text files, which requires obtaining the data to be filled in from a large amount of production and test data; currently, corresponding reports are formed through database queries, and then data is extracted from the reports and entered into the specified format text, which takes a lot of manpower and time and is prone to errors. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for quickly extracting filled text of a data set, which can quickly retrieve and obtain the required data from a large amount of production data, fill it into the corresponding filled text file, and quickly generate a filled document, greatly improving work efficiency.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] A method for quickly extracting data and filled text includes the following steps:
[0006] Step 1: Define the blank data format of the filled text and mark the data structure in the database according to the data format;
[0007] The data format includes data coordinates and data application types. The data coordinates include data sets and data points, and the data application types include multiple variables;
[0008] Accurately obtain the position of the required data according to the assignment of the data coordinates, and accurately obtain the required data according to the data application type;
[0009] Step 2: Establish the relationship between data sets, data points, and data application types;
[0010] Obtain the data attributes of the blanks in the filled template file, generate placeholders for the blanks according to the data format; establish the relationship between data sets, data points, and data application types by configuring and generating placeholders for data format encoding;
[0011] Step 3: Establish a data reading template file;
[0012] Locate the blanks to be filled according to the filled text file, clarify the data attributes of the blanks, configure the corresponding data structure according to the data format and generate all blank placeholders; update the filled text file to form a data reading template file and save it;
[0013] Step 4: Data reading and filling;
[0014] Call the data reading template file, parse and scan the placeholders in the filled file line by line according to the rules, load the data according to the data format, and save the loaded data to the buffer area and the filled file to realize data filling.
[0015] Furthermore, the data application type includes 4 variables, and the data format is expressed as $(#ds.dp#)[a][b][c][d]$, where #ds.dp# is the data coordinate, and [a][b][c][d] are the variables of the data application type.
[0016] Furthermore, the rule for generating placeholders to encode the data format is:
[0017] The encoding rule for the data set is: the label number, date and serial number of the data set;
[0018] The encoding rule for the data point is: the data field number, date and serial number;
[0019] [a][b][c][d] represents different recorded values for different batches, processes and operations, where [0] represents the initial data and [] represents the latest data.
[0020] Furthermore, the rule for configuring the corresponding data structure according to the data format:
[0021] According to the attributes of the filled blanks, select the corresponding data set and data point, generate the #ds.dp# number, and perform data address binding; assign values to the [a][b][c][d] variables according to the requirements of the data application type of the filled blanks.
[0022] Furthermore, parsing and scanning the placeholders in the filled file line by line according to the rules specifically means:
[0023] Parse the data coordinate value, query the corresponding data set and data point according to the internal relationship of binding data by the placeholder; parse the values of the variables [a][b][c][d] of the data application type in sequence.
[0024] Furthermore, after the loaded data is saved to the buffer area, data verification and data cleaning are performed, the validity and legality of the loaded data are judged, and the invalid data is cleaned; the data after cleaning and verification is written from the buffer area into the corresponding blanks in the document to complete data filling.
[0025] Advantages of the present invention: It can quickly retrieve and obtain from a large amount of production data, accurately write it into the filled text spaces, and save the corresponding templates for ready-to-use at any time, greatly improving work efficiency and reducing errors. It is applied to the data filling of text files in a specified format, such as the batch release text of biopharmaceutical enterprises and the data acquisition in archived technical text files, which can improve the production process and quality monitoring of biopharmaceutical enterprises. Description of the Drawings
[0026] Figure 1 It is a flow schematic diagram of the present invention.
[0027] Figure 2 It is a diagram of an application embodiment of the present invention. Detailed Embodiments
[0028] As Figure 1 shown, in this embodiment, taking a biopharmaceutical enterprise as an example, a method for quickly extracting data set filling text is disclosed, including the following steps:
[0029] Step 1: Define the data format of the blank positions in the file to be filled. The data format includes data coordinates and data application types. The data coordinates include data sets and data points. Use ds to represent data sets and dp to represent data points. The data coordinates are (#ds.dp#); the data application type consists of 4 variables, represented by abcd for the 4 variables respectively, and the variables can also be increased or decreased according to the data generated by the enterprise. The data application type is [a][b][c][d].
[0030] [a][b][c][d] have the following values:
[0031] [a] represents obtaining the data of the nth batch, and the value range is 0, 1, 2......n;
[0032] [b] represents the data of the mth process shift, and the value range is 0, 1, 2......n;
[0033] [c] represents the data of the pth operation shift, and the value range is 0, 1, 2......n;
[0034] [d] represents the data of the qth record copy, and the value range is 0, 1, 2......n;
[0035] The final data format is formed as: $(#ds.dp#)[a][b][c][d]$.
[0036] Step 2: Establish the relationship between data sets, data points, and data application types;
[0037] In the database of the MES system, corresponding data sets are established according to the generated products, batches, and data sources, such as production raw material data sets, production process data sets, finished product inspection data sets, etc. In the MES system, these data sets have information such as names and storage addresses. The data sets store fields, types, values, etc. of several data, which are presented in data tables.
[0038] Obtain the space nature in the filling template file. Given the attributes of the data to be filled, the attributes represent the data set to which it belongs and the data names and data fields involved. Generate a placeholder for this empty space according to the data format $(#ds.dp#)[a][b][c][d]$ constructed in step 1. Generate a placeholder code through configuration; the #ds.dp# code, and assign values to [a][b][c][d] to establish the association relationship between the data set, data point, and data application type.
[0039] In specific implementation, the #ds.dp# code points to a certain record in a data table; #ds.dp# represents a group of numbers, and the numbers are automatically completed by the coding program. The rules of the coding program are: ds coding rule: the data set number to which it belongs + date + serial number; dp coding rule: the field number to which it belongs + date + serial number; forming the coordinates of the obtained data, so that the data to be filled can be accurately searched and obtained.
[0040] Since in actual production, especially in the production process with loops, under the same process flow, there are multiple production records in different batches and different processes. When filling in data, selection is required, and by assigning values to the suffix [a][b][c][d] variables of the placeholder, it can be locked according to the filling requirements. The values in [a][b][c][d] represent different record values of different batches, processes, and procedures; [0] represents the initial data, and [] represents the latest data.
[0041] Step 3: Establish a data reading template file;
[0042] Call the text file to be filled, locate and select the spaces to be filled, and clarify the attributes of the data filled in the spaces; that is, the data set and the data names and data fields involved.
[0043] Call the data component program and complete the configuration according to the data format in step 1 to generate the placeholder for the space.
[0044] Configuration means: first, select the corresponding data set and data point according to the attributes of the filled space; then bind the data address according to the #ds.dp# number; finally, assign values to the [a][b][c][d] variables according to the requirements of the data application type of the filled space.
[0045] Traverse the text file to be filled until all spaces are generated as placeholders, update the text file to be filled, form a data reading template file and save it for data reading and filling calls.
[0046] Step 4: Data reading and filling;
[0047] Enter the text filling system and call the data reading template file; Parse and scan the placeholders in the data reading template file line by line according to the specifications and obtain the parsing results. Load the parsing results in the data format of $(#ds.dp#)[a][b][c][d]$ and save them in the buffer. Stop scanning until the last placeholder in the text file. Save the current template file for subsequent reuse.
[0048] Among them, the rule parsing: Parse the data coordinate (#ds.dp#) value, query the corresponding data set and data point according to the internal relationship between the placeholder and the bound data. Parse the values of the placeholder suffixes [a][b][c][d] in turn. If there is no value in [], it is defaulted to 1. [a] represents the number of batches. If a is 2, it is parsed that the data filled this time is the data of the second batch. [b] represents the shift data under the nth process in this batch. If it is 3, it is parsed that the data filled this time is the shift data under the second process........ until [d] is parsed. If there is no data, replace it with a null value.
[0049] Step 5: Judge the validity of the data in the data format recorded in the buffer, clear the invalid data, and prompt to re-judge; Write the cleaned and verified data from the buffer into the corresponding spaces in the document to complete the data filling.
[0050] Application Example 1
[0051] Locate and select the space to be filled, and clarify the attributes of the data filled in the space. The space attributes are: stock solution, corresponding data format: the data set is the tetravalent batch release test (055), the data point is the stock solution data (682), the batch is 1, the process is empty, the operation is 2, and the record is IVR-238. As Figure 2 shown, the data format is ${(055.682)[1][0][2][1]}$. Obtain the data according to the data format, and the name of the stock solution is IVR-238.
[0052] The above are only the preferred embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications and substitutions based on the technical solutions and inventive concepts provided by the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for quickly extracting data and filling in text, characterized in that: The steps include: Step 1: Define the data format of the blanks in the filled text, and mark the data structure in the database according to the data format; The data format includes data coordinates and data application types, the data coordinates include data sets and data points, and the data application types include multiple variables; Accurately obtain the location of required data according to the assignment of data coordinates, and accurately obtain required data according to the data application type; Step 2: Establish the relationship between data sets, data points, and data application types; Get the data attributes of the blanks in the filling template file, and generate placeholders for the blanks according to the data format; Generate placeholders through configuration to encode data formats and establish relationships between data sets, data points, and data application types; Step 3: Create a data reading template file; Locate the blanks to be filled in according to the filled text file, clarify the data attributes of the blanks, configure the corresponding data structure according to the data format and generate all blank placeholders; update the filled text file to form a data reading template file and save it; Step 4: Read and fill in data; Call the data reading template file, parse the placeholders in the reporting file line by line according to the rules, load the data according to the data format, and save the loaded data to the buffer area and the reporting file to realize data reporting.
2. A method for quickly extracting data and filling in text according to claim 1, characterized in that: The data application type includes 4 variables, and the data format is expressed as $(#ds.dp#)[a][b][c][d]$, where #ds.dp# is the data coordinate, and [a][b][c][d] is the variable of the data application type.
3. A method for quickly extracting data and filling in text according to claim 2, characterized in that: The rules for encoding data formats by generating placeholders through configuration are: The coding rules of the dataset are: dataset label, date and serial number; The coding rules for data points are: data field number, date and serial number; The replication of [a][b][c][d] represents different record values of different batches, processes, and procedures, where [0] represents the initial data and [] represents the latest data.
4. A method for quickly extracting data and filling in text according to claim 3, characterized in that: Rules for configuring the corresponding data structure according to the data format: According to the attributes of the filled-in blanks, select the corresponding data set and data point, generate the #ds.dp# number, and bind the data address; assign values to the [a][b][c][d] variables according to the application type requirements of the filled-in blank data.
5. A method for quickly extracting data and filling in text according to claim 4, characterized in that: According to the rules, the placeholders in the line-by-line scan report file are as follows: Parse the data coordinate values, and query the corresponding data set and data point according to the intrinsic relationship of the placeholder binding data; parse the values of the variables [a][b][c][d] of the data application type in turn.
6. A method for rapidly extracting data and filling in text according to claim 5, characterized in that: After the loaded data is saved in the cache area, data verification and data cleaning are performed to determine the validity and legality of the loaded data and to clean up invalid data; the cleaned and verified data is written from the cache area into the corresponding spaces in the document to complete the data reporting.
Citation Information
Patent Citations
Document processing method and device based on configuration template
CN115759024A
Digital management and control system, device and method for whole cycle flow of biopharmaceutical production
CN116341889A
Manufacturing management method and manufacturing management device for sealing measurement process
CN118658798A