A method and system for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration.
By employing a dynamic text format configuration method using LSTM, MLP, Transformer, and DNN neural networks, the heterogeneity problem of radioactive effluent detection equipment in nuclear power plants was solved, enabling efficient and accurate data parsing and maintenance, and meeting nuclear power safety compliance requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAT NUCLEAR INFORMATION TECH CO LTD
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-31
AI Technical Summary
The inconsistent text formats of radioactive effluent detection equipment in nuclear power plants lead to high costs for customized development, poor scalability, and difficult operation and maintenance. Furthermore, the lack of AI adaptive capabilities makes it impossible to intelligently identify nuclide aliases and text variation formats, resulting in a high data processing error rate.
It adopts a dynamic text format configuration method based on LSTM and MLP, combined with Transformer and DNN neural networks, to achieve text feature extraction, rule matching and data standardization. It supports multi-format parsing, and allows rule configuration through a visual interface, and has AI self-learning capabilities.
It significantly improves system compatibility and accuracy, reduces development costs, enhances operation and maintenance efficiency, meets nuclear power safety compliance requirements, achieves automatic data cleaning and intelligent verification, with an error rate approaching 0, and adapts to complex text format changes.
Smart Images

Figure CN122494038A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical fields of radioactive effluent detection in nuclear power plants, industrial data analysis, and intelligent text recognition. Specifically, it relates to a method and system for parsing data of various chemical analysis devices for nuclear power plant emissions based on dynamic text format configuration. Background Art
[0002] In order to protect the surrounding good living environment, nuclear power plants regularly detect the chemical element emissions of radioactive effluents, and statistically analyze the detection results through the group plant and group reactor management platform. Corresponding management measures are taken if the standards are exceeded.
[0003] Nuclear power plants need to regularly monitor radioactive nuclides (such as Co-60, Cs-137, I-131, Ag-110m, etc.) in gaseous and liquid effluents, and the detection data needs to be connected to the group plant and group reactor management platform for unified supervision.
[0004] However, the existing technologies have the following serious defects: 1. Severe device heterogeneity: The output text formats of high purity germanium spectrometers and radiochemical analysis devices from different manufacturers and different models are not unified, including various formats such as fixed columns, delimiters, free text, chaotic headers, and mixed units (Bq / g, Bq / L).
[0005] 2. High cost of customized development: For each newly added / replaced device, a separate parsing module needs to be developed. The adaptation of a single device takes about 20 person-days, and the labor cost per unit is about 50,000 yuan. As the number of devices increases, the cost increases linearly.
[0006] 3. Poor scalability and operation and maintenance: Format changes and device replacements require code modification and recompilation for online deployment, which is not conducive to the promotion of group plants and group reactors, and the system maintenance is difficult.
[0007] 4. Weak data processing ability: It cannot automatically process scientific notation, <MDA detection limit, abnormal markings, and unit conversions. Data verification depends on manual work, and the error rate is high.
[0008] 5. Insufficient security and compliance: Lack of nuclear power plant-level encryption, fine-grained permissions, and full-process auditing, making it difficult to meet the regulatory requirements of the nuclear industry.
[0009] 6. Lack of AI adaptability: Existing technologies only rely on fixed rules for parsing, do not have artificial intelligence self-learning ability, cannot intelligently identify nuclide aliases, symbol differences, and text variant formats, and cannot automatically distinguish abnormal data, resulting in cumbersome configuration, limited accuracy, and still high operation and maintenance costs. [[ID=3〕2]
[0010] Therefore, there is an urgent need for a code-free, dynamically adaptable, AI-enhanced, automatically parsed, and nuclear power plant-specific heterogeneous data parsing solution. Summary of the Invention
[0011] The purpose of this application is to overcome the shortcomings of existing nuclear power plant testing equipment, such as reliance on customized development for data analysis, poor compatibility, high cost, difficult operation and maintenance, unintelligent data processing, and lack of AI adaptive capabilities.
[0012] To achieve the above objectives, this application proposes a method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration, including: Step S1: Extract the specified line of text from the detection device data, input it into the trained LSTM model, and output the text features; Step S2: Input the text features into the trained MLP rule matching model and output the rule ID and matching confidence. When the matching confidence is greater than the upper limit of the set confirmation threshold range, the rule is automatically applied. When the matching confidence is within the set confirmation threshold range, the rule is manually confirmed. When the matching confidence is lower than the lower limit of the set confirmation threshold range, the rule is manually configured. Step S3: Parse the data from the testing equipment according to the rules and extract the testing results; Step S4: Input the extracted detection results into the trained Transformer encoder to normalize the nuclide names; input the normalized detection results into the trained DNN feedforward neural network to output the data status; when the data status is abnormal, it is handled manually. Step S5: Verify the detection results according to the set nuclide threshold; Step S6: Standardize the test results.
[0013] As an improvement to the above method, the text features include: delimiter features, nuclide keyword features, unit features, numerical distribution features, and marker features.
[0014] As an improvement to the above method, the range of the threshold to be confirmed is set to be 0.6 to 0.85.
[0015] As an improvement to the above method, the data state output by the DNN feedforward neural network is normal or abnormal; when the data state is abnormal, the abnormal state type is also output, including: abnormally high activity, abnormally low activity, error range exceeding the limit, MDA inconsistency, or cross-device data deviation.
[0016] As an improvement to the above method, it also includes manual configuration of rules, including: header row numbers, data start row, end row identifier, delimiter type, column mapping relationship, unit conversion rules, MDA rules, scientific notation format, and filtering conditions.
[0017] This application also provides a data parsing system for multiple detection devices of chemical elements emitted by nuclear power plants based on dynamic text format configuration, implemented based on the above method, the system comprising: The rule configuration layer provides a visual web interface that supports configuring text formats for different power plants, units, and equipment models; it supports parameter definitions for header rows, data start rows, delimiters, column mapping relationships, units, MDA rules, and scientific notation formats; and it supports adding, modifying, deleting, previewing, exporting, and batch reusing parsing rules. The dynamic parsing engine layer reads data from the detection equipment and performs format parsing; it uses an LSTM model to extract text features and identify text structure; it utilizes an MLP rule matching model to match the optimal parsing rules, achieving adaptive parsing of heterogeneous text; and it parses the detection equipment data according to the rules to extract the detection results; and... The data standardization layer is used to verify the legality of data based on the nuclear power standard nuclide library and the MDA threshold library; it uses a Transformer encoder to normalize the nuclide names; it uses a DNN feedforward neural network to determine the data status; it verifies the detection results and outputs a standardized data structure.
[0018] Compared with existing technologies, the advantages of this application are: This application can significantly improve system compatibility, robustness, stability, security, and accuracy, facilitate subsequent system maintenance, reduce development and implementation costs, and improve work efficiency, as detailed below: 1. Significant cost reduction: Based on the requirement of 20 man-days for adapting one type of equipment, each piece of equipment can save 50,000 yuan in labor costs for implementation, with significant savings in the scenario of multiple factories and multiple stacks.
[0019] 2. Extremely strong compatibility: compatible with multiple formats such as TXT, CSV, Excel, XML, and PDF, and supports mainstream nuclear power testing equipment at home and abroad (such as high-purity germanium spectrometers from brands like ORTEC, CANBERRA, and PerkinElmer).
[0020] 3. No-code operation and maintenance: Equipment replacement / format upgrade only requires modifying the configuration through a visual interface, without development, downtime, or version iteration, improving operation and maintenance efficiency by more than 90%.
[0021] 4. Improved data accuracy: Automatic cleaning and intelligent verification reduce manual intervention by 80%, the error rate approaches 0, and the data import accuracy can reach 100%.
[0022] 5. Safety and Compliance: Meets the requirements for nuclear power information security level protection, radioactive effluent supervision, and audit traceability, and complies with the Nuclear Safety Law and related laws and standards.
[0023] 6. Support for multiple plants and clusters: Unified access to data from multiple bases, multiple units, and multiple devices, eliminating data silos and achieving unified group-level supervision.
[0024] 7. AI self-learning capability: No need for repeated manual configuration, AI automatically learns new device formats, becoming more accurate the more it is used, and the time for new device access is reduced by more than 90%.
[0025] 8. Intelligent nuclide normalization: Automatically corrects nuclide symbols, capitalization, and format differences with an accuracy of ≥99%, avoiding human input errors.
[0026] 9. AI Anomaly Warning: Automatically identifies abnormal, erroneous, and missed data, reducing the workload of manual review and improving data reliability.
[0027] 10. Strong anti-interference ability: It can handle complex text such as garbled characters, blank lines, misaligned columns, and mixed units, adapt to the original output on site, and has strong robustness.
[0028] 11. Fast new device integration: Transfer learning + rule self-generation reduces the new device integration time from 20 person-days to less than 2 hours.
[0029] 12. Safe and reliable operation: Offline inference + rule-based fallback eliminates the risk of AI misjudgment, meets the high reliability requirements of nuclear power, and achieves system availability of over 99.9%. Attached Figure Description
[0030] Figure 1 The diagram shows the system architecture for analyzing data from multiple chemical element testing devices in nuclear power plant emissions, based on a dynamically configured text format. Figure 2 The diagram shows the overall implementation flowchart of a data parsing system for multiple detection devices for chemical elements emitted by nuclear power plants, based on a dynamically configured text format. Figure 3 The diagram shows a flowchart of a data parsing method for multiple detection devices for chemical elements emitted from nuclear power plants, based on a dynamically configured text format. Detailed Implementation
[0031] The technical solution of this application will be described in detail below with reference to the accompanying drawings.
[0032] Example 1 This application adopts a three-layer decoupled architecture: a rule configuration layer, a dynamic parsing engine layer, and a data standardization layer, to achieve complete separation of format, parsing, and business logic.
[0033] I. Overall Architecture like Figure 1 As shown, the system in this application consists of three parts: user terminals, application services, and a database, which are deployed independently but work collaboratively. 1. User terminal Responsible for visual interactive operations, including: visual rule configuration, data import, preview verification, and log querying. Utilizes a B / S architecture and supports web browser access.
[0034] 2. Application Services The system's core processing units include: a dynamic parsing engine, feature matching, an AI feature extraction engine, an AI nuclide normalization model, an AI anomaly data detection model, data cleaning, verification algorithms, and encryption services. A microservice architecture is adopted, with each module deployed independently.
[0035] 3. Database Responsible for persistent data storage, including: rule base, nuclide base, device base, standard database, log base, and permission base. It employs a hybrid storage approach combining relational and document-oriented databases.
[0036] Two- or three-layer decoupled architecture design 1. Rule Configuration Layer Positioning: Responsible for defining, maintaining, and managing parsing formats, thereby decoupling formats from business logic.
[0037] Features: Provides a visual web interface, supporting configuration of text formats for different power plants, generating units, and equipment models; supports parameter definitions for header rows, data start rows, delimiters, column mapping relationships (nuclide name column, activity column, error column, MDA column), units, MDA rules, scientific notation format, etc.; supports adding, modifying, deleting, previewing, exporting, and batch reusing parsing rules; enables hot-swappable loading of rules, rule updates do not require service restarts, and adding new equipment requires no code development. Data Structure: Rules are stored in JSON format, including equipment identification fields (power plant ID, generating unit ID, equipment model), format definition fields (header row number, data start row, delimiter type, column mapping dictionary), data processing fields (unit conversion factor, MDA identifier, scientific notation regular expression), and metadata fields (creation time, version number, effective status).
[0038] 2. Dynamic parsing engine layer Positioning: Responsible for text feature recognition, adaptive parsing, and AI intelligent processing.
[0039] Functions: Reads uploaded device output text and parses multiple formats such as TXT, CSV, Excel, XML, and PDF; uses an LSTM (Long Short-Term Memory) network to extract temporal features of the text and automatically identifies the text structure (fixed columns / delimiters / free text / mixed tables); performs adaptive parsing of heterogeneous text based on feature matching optimal parsing rules; automatically handles scientific notation, detection limit symbols <, anomaly markers # / *, blank lines, garbled characters, and other noise; performs preliminary extraction, cleaning, and structure transformation of raw data.
[0040] AI model implementation: Text structure classification model: Employs LSTM (Long Short-Term Memory) recurrent neural network. Inputs include text lines, symbols, units, numerical distributions, and keywords. Outputs include fixed column / delimiter / free text / mixed table classifications. LSTM (Long Short-Term Memory) recurrent neural networks are mature models for time-series data processing. Standard applications include: time-series text feature extraction and classification, abnormal time-series data detection, sequence data prediction and trend analysis, natural language text structuring, robust processing of noisy text, adaptive recognition of multi-format text, few-shot transfer learning, key entity recognition and normalization, offline lightweight inference, rule-assisted decision enhancement, etc.
[0041] This application innovatively applies LSTM to heterogeneous text parsing scenarios in nuclear power plant radioactive effluent testing equipment for the first time, making full use of its temporal feature extraction, anti-interference, and adaptive classification capabilities to solve long-standing technical pain points in the industry, such as inconsistent formats, difficult parsing, complex configuration, and insufficient reliability.
[0042] Rule matching model: Employs MLP (Multilayer Perceptron), with text feature vectors as input and optimal rule ID and data state as output.
[0043] LSTM execution process: (1) Initial state input: Output the hidden state h from the previous time step. t-1 1. Cellular memory state C at the previous moment t-1 and the text feature input x at the current moment. t Simultaneously fed into the LSTM unit; (2) Forget gate calculation: h t-1 With x t After weighted summation, the forget gate output f is obtained through the sigmoid activation function. t f t The value is between 0 and 1, and is used to control the information that needs to be discarded in the cell state; (3) Preliminary update of cell state: update the cell state C from the previous moment. t-1 With the Gate of Oblivion t Multiplying elements one by one completes the process of forgetting useless historical information. (4) Input gate and candidate state calculation: h t-1 With x t Simultaneously, the input is fed into the input gate and passed through the sigmoid function to obtain i. t Control the degree to which new information is written; at the same time, h t-1 With x t Candidate cell state C is generated by tanh activation. t ; (5) New cell state generation: Input gate i t With candidate state C t Multiply, then add to the updated C after forgetting. t-1 Add them together to get the new cell state C at the current moment. t This completes the preservation and updating of memories; (6) Output gate calculation: h t-1 With x t The output gate o is obtained through sigmoid. t This is used to control the information strength of the final output; (7) Output the current hidden state: Set the new cell state C t After activation by tanh, and then with the output gate o t Element-wise multiplication yields the LSTM output h at the current time. t ; (8) State transfer: h t With C t As input for the next time step, the subsequent text sequence is processed to complete the extraction of temporal features.
[0044] The technological innovations of the dynamic parsing engine layer include: (1) Innovative segmented timing input of nuclear power laboratory test text: The output text of the equipment is automatically divided into header segment, nuclide data segment and remarks information segment, and only the valid nuclide data segment is input into LSTM; (2) Nuclear power field-specific feature engineering innovation: Construct a nuclear power laboratory-specific feature set, including nuclide keywords, activity units (Bq / g / Bq / L), detection limit symbol <, scientific notation E, and anomaly markers # / *; (3) AI prediction enhancement innovation guided by rule base: LSTM prediction results are jointly verified with the rule base, and rule parsing is automatically switched when the confidence is low, forming a highly reliable parsing mode with AI first and rule as a fallback; (4) Lightweight offline single-batch inference innovation: Single-file single-time offline inference is adopted, without online training, caching of historical data, or cross-file association, to ensure that nuclear power data does not leave the area and is compliant and controllable; (5) Multi-device format transfer learning adaptation innovation: The known device format features are transferred to new devices, and a small number of samples are sufficient to complete the recognition of new formats, enabling rapid access to new devices with zero or minimal configuration; (6) Innovation to enhance the robustness of text noise: For noise such as garbled characters, blank lines, misaligned columns, and redundant notes, local window temporal feature extraction is adopted to identify only valid data rows; (7) Innovation in multi-unit parallel time series feature recognition: Simultaneously recognize multiple unit features such as Bq / g, Bq / L, mL, and g in the same text segment, and automatically distinguish the measurement type; (8) Dynamic rule self-generation assists innovation: After LSTM recognizes the text structure, it automatically generates preliminary parsing rules, and the user only needs to confirm to complete the configuration; (9) Innovation in time-series anomaly jump detection: Automatically identify logical errors such as data jumps, missing rows, missing columns, MDA inversion, and abnormally high / low activity based on time-series relationships; (10) Innovation in cross-file format consistency verification: compare the time sequence characteristics of multiple historical files on the same device, automatically identify device upgrades and output format changes, and provide early warnings to avoid parsing failures.
[0045] 3. Data Standardization Layer Position: Responsible for data verification, unit conversion, normalization, and unified data entry.
[0046] Functions: Performs data validity verification based on nuclear power standard nuclide libraries and MDA threshold libraries; uses a Transformer encoder to unify variants such as Ag-110m / AG-110M and I-131 / i-131 into standard names; employs a DNN feedforward neural network to automatically verify activity rationality, error range, MDA consistency, and cross-device data consistency; automatically performs unit conversions such as Bq / g and Bq / L and mass / volume normalization; outputs standardized data structures, supporting unified data entry, preview, editing, and reporting.
[0047] AI model implementation: Nuclide name normalization model: using a Transformer encoder, the input is the original nuclide string, and the output is the standard nuclide name; Anomaly data discrimination model: Employs a DNN feedforward neural network. Inputs include activity, error, MDA, units, and historical data. Outputs are normal / abnormal / abnormal type. Standardized field structure: Nuclide name, activity, error, MDA, units, sampling time, equipment number, power plant, and generating unit.
[0048] III. Detailed Implementation Steps (e.g.) Figure 2 (As shown) Step 1: Visual configuration of dynamic text rules; Administrators configure rules via a web interface: (1) Select power plant → unit → equipment model and establish a unique equipment identifier; (2) Configure the header row number, data start row, and end row identifier; (3) Configure the delimiter type (comma, tab, fixed width, regular expression); (4) Configure column mapping relationship: Specify the column index or column name of the nuclide name column, activity value column, error column, MDA column, and unit column; (5)Configure unit conversion rules: Set the conversion coefficient between Bq / g and Bq / L, and configure volume / mass normalization parameters; (6)Configure MDA rules: Set the processing logic for detection limit identifiers (such as "<", "ND", "<MDA"); (7)Configure scientific notation format: Define regular expression matching patterns (such as "1.23E+04", "1.23×10^4"); (8)Configure filtering conditions: Set the line prefixes to be skipped (such as lines starting with "Remarks", "Attention", "*"); (9)Click the preview button, and the system will display the parsing effect in real time; (10)After debugging is passed, save it to the rule library, and the rule will take effect immediately without restarting the service.
[0049] Step 2: Text feature self-learning and device automatic adaptation; After the operator uploads the device output text file: (1)The system reads the first N lines (default 50 lines) of the file as samples; (2)The LSTM model extracts text features, including: Separator features: Count the occurrence frequency and distribution of commas, tab characters, and spaces; Radionuclide keyword features: Match radionuclide symbols in the standard radionuclide library (such as Co-60, Cs-137); Unit features: Identify unit strings such as Bq / g, Bq / L, Bq / m 3 and so on; Numerical distribution features: Identify scientific notation, decimals, and integers; Detection limit symbol features: Identify symbols such as <, >, ND, etc.; Marker features: Identify abnormal markers such as #, *; (3)Generate a text feature vector and input it into the MLP rule matching model; (4)The model outputs the optimal rule ID and matching confidence; (5)If the confidence > 0.85, automatically apply this rule; if 0.6 < confidence ≤ 0.85, prompt the user to confirm; if the confidence ≤ 0.6, enter the manual configuration process; (6)The user can confirm the automatic recognition result with one click without manually selecting the rule.
[0050] Step 3: Heterogeneous format unified parsing engine; (1)The system performs parsing according to the matching rules: Format recognition: Select parsing strategies according to the format types in the rules: Fixed column format: Intercept strings according to column positions; Separator format: Split the string by the separator; Free text format: Extract using regular expressions; Mixed table format: First identify the table boundaries and then parse by rows and columns; (2)Data cleaning: Remove blank lines and leading / trailing spaces; Process scientific notation: Convert "1.23E+04" to 12300.0; Process detection limit: Identify "<0.5" as the MDA value, mark the activity value as <MDA or process by half MDA; Process abnormal symbols: Remove markers such as # and *, and retain the original numerical value; Process duplicate rows: Deduplicate based on nuclide name + sampling time; (3)Data extraction: Extract nuclide name, activity value, error, MDA, unit, and sampling time. Mark the rows where extraction fails and enter the exception handling process; (4)Unit conversion: If the original unit is Bq / g and the target unit is Bq / L, perform conversion based on the sample volume / mass parameter; if the original unit is mL or g, normalize the volume / mass to the standard unit; record the values before and after conversion and the conversion factor.
[0051] Step 4: AI nuclide name normalization and intelligent verification; (1)Nuclide name normalization: Input the extracted original nuclide string into the Transformer encoder; the model outputs the standard nuclide name (e.g., normalize "AG-110M" to "Ag-110m", "i-131" to "I-131"); for unrecognized nuclide names, mark them for manual confirmation; (2)Data intelligent verification: Input the activity, error, MDA, unit, and historical statistical data into the DNN feedforward neural network; the model outputs the data status: normal / abnormal; if it is abnormal, further output the abnormal type: activity abnormally high, activity abnormally low, error range exceeded, MDA inconsistent, cross-device data deviation; abnormal data is highlighted (red background) at the front end and the reason for the abnormality is displayed; (3)Threshold verification: Compare with the thresholds in the nuclear power standard nuclide library; Verify the rationality of the activity: Judge whether the activity value is within the physically possible range (e.g., the activity of Co-60 does not exceed the total activity of the sample); Verify the error range: Judge whether the relative error is within a reasonable interval (e.g., usually not exceeding 50%); Verify MDA consistency: Determine whether the relationship between MDA and activity values is logical (e.g., when the activity < MDA, the activity value should be marked as < MDA).
[0052] Step 5: Standardize data for storage and preview verification (1) Data standardization: Unify field formats: Nuclide name (string), activity (float), error (float), MDA (float), unit (enumeration value), sampling time (timestamp), equipment number (string), power plant (string), unit (string); Unify unit conversion to standard units (Bq / L or Bq / g); Unify the time format to ISO 8601 format.
[0053] (2) Preview verification: The front end displays the parsing results in a table form, with each row corresponding to a nuclide data; abnormal data rows are highlighted in red, and the mouse hover shows the details of the abnormality; Support single-line editing: Double-click on the cell to modify the value; Support batch editing: Select multiple rows for batch modification of units, sampling time, etc.; Provide "Confirm for storage" and "Return for modification" buttons. After confirmation, the data is written into the standard database.
[0054] Step 6: Nuclear power level safety and compliance control (1) Transmission encryption: Use the national cipher SM4 algorithm to encrypt data transmission; use the HTTPS protocol, and the SSL certificate uses the national cipher algorithm suite; sensitive fields (such as nuclide activity) are encrypted at the field level before transmission; (2) Storage encryption: Sensitive data in the database is encrypted and stored using the SM2 / SM3 algorithm; the key is stored in a dedicated key management system (KMS) and rotated regularly; (3) Hierarchical permission control: Adopt the RBAC (Role-Based Access Control) model; Role definition: Administrator: Has all permissions, including rule configuration, user management, and system settings; Configurator: Has the permissions of rule configuration and rule modification; Operator: Has the permissions of data import, preview, and confirm for storage; Viewer: Only has the permissions of data query and report viewing; Data isolation: Isolated at four levels: power plant → unit → equipment → data type. Users can only access data within the authorized scope; (4) Full-process log auditing: Record all operation logs, including: user ID, username, operation type (login, configure rules, import data, modify data, delete data), operation object (rule ID, data ID), operation time (accurate to milliseconds), IP address, operation result (success / failure), and changed content (comparison of before and after values); log storage adopts append-only mode, prohibiting modification and deletion; provide a log query interface, supporting multi-dimensional queries by time, user, operation type, device, etc.; logs are archived regularly, and the retention period complies with nuclear industry regulatory requirements (usually no less than 10 years).
[0055] The LSTM, Transformer, and DNN used in this application are all existing standard neural network models. No improvements have been made to the internal structure such as gating structure, activation function, cell state transfer method, and attention mechanism. They are simply used as mature tools to innovatively apply to the parsing scenario of output text from nuclear power plant detection equipment, so as to achieve intelligent recognition and automatic matching of parsing rules for heterogeneous equipment text.
[0056] Example 2 like Figure 3 As shown, this application also provides a method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration. Implemented based on the aforementioned system, the method includes: Step S1: Extract the specified line of text from the detection device data, input it into the trained LSTM model, and output the text features; Step S2: Input the text features into the trained MLP rule matching model and output the rule ID and matching confidence. When the matching confidence is greater than the upper limit of the set confirmation threshold range, the rule is automatically applied. When the matching confidence is within the set confirmation threshold range, the rule is manually confirmed. When the matching confidence is lower than the lower limit of the set confirmation threshold range, the rule is manually configured. Step S3: Parse the data from the testing equipment according to the rules and extract the testing results; Step S4: Input the extracted detection results into the trained Transformer encoder to normalize the nuclide names; input the normalized detection results into the trained DNN feedforward neural network to output the data status; when the data status is abnormal, it is handled manually. Step S5: Verify the detection results according to the set nuclide threshold; Step S6: Standardize the test results.
[0057] Example 3: The entire process of rule configuration and data import; A chemical administrator at a nuclear power plant needs to connect a newly purchased high-purity germanium spectrometer (e.g., ORTEC brand, model GEM-MX7080). The device outputs text in CSV format, but the header structure is different from that of existing devices.
[0058] Step 1: Rule configuration; (1) The administrator logs into the system and enters the "Text Rule Configuration" interface; (2) Select the power plant "XX Nuclear Power", the unit "Unit 1", and the equipment model "ORTEC-GEM-MX7080"; (3) Configure the header row as the 3rd row (the first 2 rows are device information) and the data start row as the 4th row; (4) Configure the separator as a comma (,); (5) Configure column mapping: Nuclide Name Column: Column B Activity value column: Column D Error column: Column E (5th column) MDA column: Column F Unit column: Column G (7th column) (6) Configure unit: The default is Bq / L. If the unit column is empty, the default is Bq / L. (7) Configure MDA rules: The detection limit identifier is "<", and the processing method is "take half of the MDA as the activity value"; (8) Configure scientific notation format: The regular expression is [\d\.]+[Ee][+-]?\d+; (9) Configure filter conditions: skip lines that start with "Remarks", "Note", or "#"; (10) Upload the sample file and click Preview. The system will display the parsing results table. After confirming that there are no errors, save the rules.
[0059] Step 2: Data Import; (1) The operator logs into the system and enters the "Data Import" interface; (2) Upload the test result file generated by the device (e.g., "2025-04-13-sample.csv"); (3) The system automatically identifies the device: The LSTM model extracts file features, matches the rule base, and identifies it as "ORTEC-GEM-MX7080" device with a confidence level of 0.92; (4) The system automatically applies the corresponding rules for parsing; (5) The front-end displays a preview table containing 10 nuclide data, in which the activity value of Cs-137 in the third row is marked as abnormal (highlighted in red) because the value exceeds the historical average by 3 times; (6) The operator checks the original record, confirms that the data is correct, and clicks "Confirm Normal"; (7) The operator clicks "Confirm Inbound" and the data is written to the standard database; (8) System operation log: User "Zhang San", Operation "Import Data", Equipment "ORTEC-GEM-MX7080-1 Unit", Time "2025-04-13 14:30:25", Result "Success", Data volume "10 records".
[0060] Example 4: AI-powered automatic adaptation and anomaly detection; A nuclear power plant needs to connect an older model of radiochemical analyzer. The output of this equipment is in fixed-width TXT format, and there are some garbled characters and non-standard format issues.
[0061] Step 1: AI Automatic Adaptation (1) The operator uploaded the equipment output file without specifying the equipment model; (2) The LSTM model reads the file and extracts features: No explicit delimiter, but column positions are fixed; It contains the nuclide keywords "Co-60" and "Cs-137"; Includes the unit "Bq / g"; There are garbled characters ""; (3) The model is identified as "fixed column format", but there is no rule in the rule base that matches it exactly; (4) Confidence level 0.75, the system prompts "Suspected fixed column format, recommended rule: Generic-FixedWidth-001, confidence level 75%, please confirm or select manually"; (5) After operator confirmation, the system applies and parses the rule; (6) The system prompts "Garbled characters detected, automatically filtered", and displays the cleaned data.
[0062] Step 2: Anomaly detection; (1) The analysis results showed that the activity of Co-60 was 1.23E+05 Bq / g, with an error of ±0.5%; (2) The DNN model detected an anomaly: the error of 0.5% is far lower than the device’s usual error range of 5%-15%; (3) The system marks this row of data as "abnormal: error range is suspicious"; (4) The operator found that the error value of the original file should be ±5.0%, but it was mistakenly identified as 0.5% due to formatting issues; (5) The operator manually corrected the value, and after re-uploading, the verification passed.
[0063] Example 5: Application of multiple plants and multiple reactors; A nuclear power group is implementing a group plant and group reactor management platform, which requires the connection of 12 different brands and models of testing equipment from its three subordinate nuclear power bases.
[0064] (1) Pre-implementation assessment: Traditional method: 12 units × 20 person-days = 240 person-days, costing approximately 600,000 yuan; This application method: all configurations are expected to be completed within 6 hours.
[0065] (2) Implementation process: Day 1 morning: Deploy the system at each base and configure basic data (power plant and unit information); Afternoon of Day 1: Each base's configuration personnel configure the local equipment rules, with an average configuration time of 30 minutes per type of equipment; Day 2: Data import test was conducted, and all 12 devices were successfully connected with 100% data accuracy. Day 3: Officially launched and operational, the group headquarters can view the radioactive effluent detection data of each base in real time.
[0066] (3) Implementation results: Total man-days reduced from 240 man-days to 6 hours; Labor cost savings: approximately 600,000 yuan; Supports unified display, statistics, and alarm functions on a multi-plant / multi-reactor platform; data import accuracy is 100%, and maintenance requires zero code changes; meets the group-level environmental protection and nuclear safety supervision requirements.
[0067] This application may also provide a computer device, including: at least one processor, memory, at least one network interface, and a user interface. The various components in this device are coupled together via a bus system. It is understood that the bus system is used to implement communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0068] The user interface can include a display, keyboard, or clicking device. Examples include a mouse, trackball, touchpad, or touchscreen.
[0069] It is understood that the memory in the embodiments disclosed in this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0070] In some implementations, the memory stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof: operating systems and applications.
[0071] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application programs include various applications, such as media players and browsers, used to implement various application functions. Programs implementing the methods of the embodiments of this disclosure can be included in the application programs.
[0072] In the above embodiments, the processor can also invoke programs or instructions stored in memory, specifically programs or instructions stored in an application program, for the following purposes: Follow the steps described above.
[0073] The above methods can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by software instructions. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic diagrams disclosed above. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the disclosed methods can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.
[0074] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof.
[0075] For software implementation, the technology of this application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of this application. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0076] This application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, it can implement the steps in the above method embodiments.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit it. Although this application has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application, and should all be covered within the scope of the claims of this application.
Claims
1. A method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration, comprising: Step S1: Extract the specified line of text from the detection device data, input it into the trained LSTM model, and output the text features; Step S2: Input the text features into the trained MLP rule matching model and output the rule ID and matching confidence. When the matching confidence is greater than the upper limit of the set confirmation threshold range, the rule is automatically applied. When the matching confidence is within the set confirmation threshold range, the rule is manually confirmed. When the matching confidence is lower than the lower limit of the set confirmation threshold range, the rule is manually configured. Step S3: Parse the data from the testing equipment according to the rules and extract the testing results; Step S4: Input the extracted detection results into the trained Transformer encoder to normalize the nuclide names; input the normalized detection results into the trained DNN feedforward neural network to output the data status. When the data status is abnormal, it should be handled manually; Step S5: Verify the detection results according to the set nuclide threshold; Step S6: Standardize the test results.
2. The method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration according to claim 1, characterized in that, The text features include: delimiter features, nuclide keyword features, unit features, numerical distribution features, and marker features.
3. The method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration according to claim 1, characterized in that, The set threshold range for confirmation is 0.6 to 0.
85.
4. The method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration according to claim 1, characterized in that, The data output by the DNN feedforward neural network is either normal or abnormal. When the data status is abnormal, the abnormal status type is also output, including: abnormally high activity, abnormally low activity, error range exceeded, MDA inconsistency, or cross-device data deviation.
5. The method for parsing data from multiple detection devices for chemical elements emitted by nuclear power plants based on dynamic text format configuration according to claim 1, characterized in that, It also includes manual configuration of rules, including: header row numbers, data start row, end row identifier, delimiter type, column mapping relationship, unit conversion rules, MDA rules, scientific notation format, and filtering conditions.
6. A data parsing system for multiple detection devices of chemical elements emitted by nuclear power plants based on dynamic text format configuration, implemented based on the method described in any one of claims 1-5, characterized in that, The system includes: The rule configuration layer provides a visual web interface that supports configuring text formats for different power plants, units, and equipment models; it supports parameter definitions for header rows, data start rows, delimiters, column mapping relationships, units, MDA rules, and scientific notation formats; and it supports adding, modifying, deleting, previewing, exporting, and batch reusing parsing rules. The dynamic parsing engine layer reads data from the detection equipment and performs format parsing; it uses an LSTM model to extract text features and identify text structure; it utilizes an MLP rule matching model to match the optimal parsing rules, achieving adaptive parsing of heterogeneous text; and it parses the detection equipment data according to the rules to extract the detection results; and... The data standardization layer is used to verify the legality of data based on the nuclear power standard nuclide library and the MDA threshold library; it uses a Transformer encoder to normalize the nuclide names; it uses a DNN feedforward neural network to determine the data status; it verifies the detection results and outputs a standardized data structure.