Food Testing Big Data Analysis System and Methods
By constructing a batch-level contamination indicator path set and trend segment identification, combined with risk level adjustment and map establishment, the problem of insufficient contamination trend identification in food testing big data analysis was solved, realizing dynamic adjustment and panoramic grouping of risks, and improving the reliability of food safety analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNNAN XIANYANG BIOTECHNOLOGY CO LTD
- Filing Date
- 2026-03-19
- Publication Date
- 2026-05-26
AI Technical Summary
Existing food testing big data analysis systems cannot accurately reflect the continuous changes of contaminant indicators over time, lack dynamic identification mechanisms, make it difficult to detect potential risks in a timely manner, and the data classification and screening results are unstable, affecting the reliability of food safety analysis and traceability capabilities.
By constructing a batch-level pollution indicator path set, identifying trend segments with consistent change directions, and combining the critical expression range of risk level for interval discrimination, the risk level is adjusted. Furthermore, by switching trend directions, the breakpoint is identified, and a comprehensive map is established by integrating detection source and project category information.
It enables dynamic extraction of pollution trends, real-time adjustment of risk levels, and precise location of trend reversals, solving the problems of insufficient identification of continuous pollution changes and ambiguous risk classification, and improving the panoramic grouping capability of detection data.
Smart Images

Figure CN121880835B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data analytics, and in particular to a big data analytics system and method for food testing. Background Technology
[0002] Big data analytics technology involves the collection, storage, management, processing, and analysis of large-scale structured or unstructured datasets to discover potential correlations and extract valuable information. Its overall technical framework typically consists of a data acquisition terminal, data transmission channels, a data processing platform, and a data analysis engine, relying on a distributed computing platform for efficient computation and management. It is applicable to data-driven application scenarios in multiple industries such as finance, healthcare, transportation, security, and environmental protection. Traditional food testing big data analytics systems refer to technical systems that collect and organize historical testing data, experimental data, and real-time monitoring data related to food testing to achieve food safety assessment and problem tracking. Traditional food testing big data analytics typically involves setting up data acquisition terminals to obtain food physicochemical indicators and microbiological test results, transmitting them to a local server, storing the data in a database, summarizing the data using statistical methods, and using manually constructed classification models to perform risk rating on food samples, achieving preliminary classification and screening of the test data.
[0003] In current food testing big data analysis processes, the main reliance is on data acquisition terminals to obtain physicochemical indicators and microbiological results, which are then stored and aggregated in local databases. Risk identification methods typically rely on pre-set models and manually constructed classification standards for static judgment, resulting in an inability to accurately reflect the continuous changes in contaminant indicators over time. This lack of a dynamic identification mechanism for contamination trends makes it difficult to promptly identify potential risks when testing data accumulates rapidly. Data classification and screening results are easily affected by noise and lack stability. Furthermore, there is a lack of structured analytical methods for analyzing changes in the behavior of consecutive batches of samples, making it difficult to form a trend tracking and risk identification system covering the entire testing process. Ultimately, this affects the reliability and traceability of food safety analysis results. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a food testing big data analysis system and method.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a food testing big data analysis system comprising:
[0006] The batch information extraction module obtains the pollutant index values corresponding to each batch of food samples in the food physicochemical testing sampling equipment, extracts the sample number and detection time, arranges them in ascending order according to the detection time, and constructs a continuous information chain of pollutant indicators in the time dimension to generate a batch-level pollutant indicator path set.
[0007] The time-series feature recognition module extracts interval segments according to the detection time sequence based on the pollutant indicators in the batch-level pollution indicator path set, compares the changing direction of continuous data within the interval, filters indicator segments with the same changing direction, extracts pollutant detection information at the end of the segment, and generates an indicator trend distribution segment set.
[0008] The risk boundary adjustment module performs interval discrimination based on the positional relationship between the pollutant detection information at the end of the segment and the preset risk level critical expression interval according to the index trend distribution segment set. At the same time, it extracts the positive and negative relationship of the change values of pollutants in adjacent batches and adds directional markers to generate a pollution index level classification structure set.
[0009] The trend break segmentation module performs adjacency comparison of directional marker values between consecutive batches based on the pollution index level classification structure set. When the positive and negative directional switching behavior is identified, the switching position is marked as a breakpoint and the sequence after the breakpoint is re-divided to generate a pollution index trend segmentation set.
[0010] The risk assessment output module, based on the pollution index trend segment division set, combined with food sampling source information, test item category information, and test time record information, performs comprehensive numbering and grouping processing to establish an overall correlation map and output the food testing big data analysis results.
[0011] As a further aspect of the present invention, when performing interval discrimination on the positional relationship, if the pollutant detection information at the end of the paragraph falls within the adjacent judgment zone of the preset risk level critical expression interval, then the corresponding segment as a whole is classified into a higher-level risk level expression interval.
[0012] The batch-level pollution indicator path set includes a pollutant indicator time series chain, sample number sequence, detection time axis, pollutant value fluctuation trajectory, and indicator path mapping relationship. The indicator trend distribution segment set includes change direction markers, continuous trend segments, trend change sequence numbers, segment end node identifiers, and trend segment coverage. The pollution indicator level classification structure set includes risk level attribution labels, risk interval coverage information, directional change markers, level upgrade records, and pollutant change magnitude classification. The pollution indicator trend segment division set includes trend segment boundary points, trend direction switching points, segment attribution adjustment sequences, trend segment grouping identifiers, and trend continuity mapping diagrams. The food testing big data analysis results include a comprehensive numbering structure, grouping and classification labels, sampling source information, detection category information, and detection time series distribution map.
[0013] As a further aspect of the present invention, the batch information extraction module includes:
[0014] The data stream receiving submodule acquires the sample dataset from the food physicochemical testing sampling equipment, parses the original sampling information and corresponding test output records, extracts the data of food sample number, test time and contaminant item value, and outputs the test time data stream in a unified timestamp format.
[0015] The pollution index analysis submodule extracts the corresponding food sample number sequence and pollutant item value group based on the detection time data stream, constructs a pollution data set with the sample number as the primary key, performs an ordered comparison of the pollutant item values in the set according to the preset pollutant item list, performs a sorting operation on the pollutant index in ascending order of time, and generates a pollutant time sequence arrangement matrix.
[0016] The pollution path construction submodule analyzes the direction of numerical change of pollutant items in continuous samples based on the pollutant numerical information at each time node in the pollutant time sequence arrangement matrix, performs order difference judgment on each pollutant item, and aggregates the indicator change trend of each pollutant on the sample time line to generate a batch-level pollution indicator path set.
[0017] As a further aspect of the present invention, the time-series feature recognition module includes:
[0018] The interval extraction submodule obtains the batch-level pollution index path set, divides the pollutant index sequence into equal-step segments according to the detection time order, forms continuous time interval segments for adjacent detection time nodes, and records the corresponding pollutant value sequence in each interval segment to generate a pollution index time interval set.
[0019] The direction consistency filtering submodule compares the changing direction of adjacent pollutant values within the time interval based on the pollution index time interval set, filters the intervals based on the consistency of the changing direction sign, retains the intervals with consistent changing direction, and outputs a sequence of continuously changing index segments to generate consistent changing interval identification data.
[0020] The trend segment generation submodule extracts the pollutant detection values and detection time information at the end of the corresponding interval segment based on the consistent change interval identifier data, marks the continuous change behavior, performs segment aggregation processing, and generates an index trend distribution segment set.
[0021] As a further aspect of the present invention, the risk boundary adjustment module includes:
[0022] The segment discrimination submodule obtains the index trend distribution segment set, reads the pollutant detection values at the end of the segment, performs position matching operation according to the preset risk level critical expression interval boundary value, identifies the interval number corresponding to each detection value, and determines whether it is at the risk judgment zone boundary by comparing the difference between each number and the adjacent interval band number, and generates a pollutant level critical proximity marker set.
[0023] The risk enhancement classification submodule extracts all sequence information of the corresponding detection value segment in the indicator trend distribution segment set based on the pollutant level critical proximity marker set. The segments that meet the boundary conditions are classified into a classification interval one level higher than the current risk level. Combined with the corresponding detection time identifier, a pollutant level mapping sequence is established in chronological order. The pollutant level adjustment judgment value is calculated and obtained. The interval number is updated by combining the level weight threshold of the current classification interval, and a pollutant risk classification interval set is generated.
[0024] The direction marker generation submodule calculates the difference between the pollutant detection values of adjacent batches based on the pollutant risk classification interval set and the detection time sequence, determines the positive or negative relationship and marks the upward or downward trend, and attaches the trend marker to the corresponding pollutant and classification interval number to generate a pollution index level classification structure set.
[0025] As a further aspect of the present invention, the calculation formula for the pollutant level adjustment judgment value is as follows:
[0026] ;
[0027] in, This represents the pollutant level adjustment judgment value for the i-th detection value. This represents the i-th detection value. This represents the boundary value of the risk interval corresponding to the current detection value. Let ik be the ith detection value, n be the calculation span, and R be the normalization coefficient.
[0028] As a further aspect of the present invention, the trend fracture segmentation module includes:
[0029] The directional adjacency comparison submodule obtains the directional marker values in the continuous batch sequence based on the pollution index level classification structure set, performs sign consistency comparison on the directional marker values of adjacent batches, judges whether the positive and negative signs have switched, and records the batch index number of each sign change position to generate a directional switching judgment sequence.
[0030] The breakpoint identifier generation submodule filters the batch index positions where the symbols switch between positive and negative based on the direction switching determination sequence, establishes a breakpoint identifier record for each index position, and associates the breakpoint position with the corresponding detection time to generate a trend breakpoint position identifier set.
[0031] The paragraph attribution reconstruction submodule segments the continuous batch sequence in the pollution index level classification structure set according to the trend break position identifier set. It re-numbers the starting position of the paragraph after the break point and adjusts the paragraph attribution mark of the corresponding batch to generate a pollution index trend paragraph division set.
[0032] As a further aspect of the present invention, the risk assessment output module includes:
[0033] The information merging submodule obtains food sampling source information, test item category information, and test time record information based on the pollution index trend segment division set. It performs joint numbering processing on the segment number and the three types of information, completes batch-level grouping based on the numbering, and generates a joint numbering set of test information.
[0034] The association graph construction submodule reads the paragraph identifiers from the pollution index trend paragraph division set based on the joint number set of the detection information, constructs a relationship mapping structure with paragraph numbers as nodes, records the connection relationships between nodes in the dimensions of sampling source, detection item category and detection time, and establishes an overall association graph.
[0035] The analysis results output submodule, based on the overall correlation map, organizes the correlation density and distribution status of each node, summarizes the node information according to a unified numbering order, and outputs the big data analysis results of food testing.
[0036] The big data analysis method for food testing includes the following steps:
[0037] S1: Obtain the pollutant index values corresponding to each batch of food samples in the food physicochemical testing sampling equipment, extract the sample number and detection time, arrange them in ascending order according to the detection time, construct a continuous information chain of pollutant indicators in the time dimension, and generate a batch-level pollutant indicator path set.
[0038] S2: Based on the pollutant indicators in the batch-level pollution indicator path set, extract intervals in the order of detection time, compare the changing direction of continuous data within the interval, filter indicator segments with the same changing direction, extract pollutant detection information at the end of the segment, and generate an indicator trend distribution segment set.
[0039] S3: Based on the distribution segment set of the index trend, the positional relationship between the pollutant detection information at the end of the segment and the preset risk level critical expression interval is determined by interval discrimination. At the same time, the positive and negative relationship of the change values of pollutants in adjacent batches is extracted and directional markers are added to generate a pollution index level classification structure set.
[0040] S4: Based on the pollution index level classification structure set, perform adjacency comparison of the directional marker values between consecutive batches. When the positive and negative directional switching behavior is identified, mark the switching position as a breakpoint and re-divide the sequence after the breakpoint to generate a pollution index trend segment division set.
[0041] S5: Based on the pollution index trend segment division set, combined with food sampling source information, test item category information and test time record information, perform comprehensive numbering and grouping processing, establish an overall correlation map, and output the food testing big data analysis results.
[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0043] In this invention, by extracting pollutant indicators corresponding to each batch of samples and constructing a time-series information chain, and dividing trend segments by combining a consistency identification method for change direction, the relationship between tail data and risk intervals is classified and judged. Then, the risk level is adjusted according to the pollutant change direction, and the break position is identified by switching the trend direction to reconstruct the data sequence attribution. By integrating multi-dimensional information such as detection source and project category to establish a comprehensive map, this processing method can realize the dynamic extraction of pollution trends, real-time adjustment of risk levels, accurate positioning of trend turning points, and panoramic grouping of detection data. This effectively solves the problems of insufficient identification of the continuity of pollution changes, ambiguous risk classification, and isolated and scattered detection data in existing technologies. Attached Figure Description
[0044] Figure 1 This is a system flowchart of the present invention;
[0045] Figure 2 This is a flowchart of the batch information extraction module of the present invention;
[0046] Figure 3 This is a flowchart of the temporal feature recognition module of the present invention;
[0047] Figure 4 This is a flowchart of the risk boundary adjustment module of the present invention;
[0048] Figure 5 This is a flowchart of the trend fracture segmentation module of the present invention;
[0049] Figure 6 This is a flowchart of the risk assessment output module of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0051] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0052] Please see Figure 1 The food testing big data analysis system includes:
[0053] The batch information extraction module obtains the pollutant index values corresponding to each batch of food samples in the food physicochemical testing sampling equipment, extracts the sample number and detection time, arranges them in ascending order according to the detection time, and constructs a continuous information chain of pollutant indicators in the time dimension to generate a batch-level pollutant indicator path set.
[0054] The temporal feature recognition module extracts intervals based on pollutant indicators in the batch-level pollution indicator path set according to the detection time sequence, compares the changing direction of continuous data within the interval, filters indicator segments with consistent changing directions, extracts pollutant detection information at the end of the segment and marks continuous changing behavior, and generates a set of indicator trend distribution segments.
[0055] The risk boundary adjustment module determines the positional relationship between the pollutant detection information at the end of the segment and the preset risk level critical expression interval based on the distribution segment set of indicator trends. When the segment falls within the adjacent judgment zone of the preset risk level critical expression interval, the entire segment is classified into a higher-level risk level expression interval. At the same time, based on the directional characteristics of the continuous increase or decrease of detection values, the positive and negative relationships of the changes in pollutant values of adjacent batches are extracted and directional markers are added to generate a pollution indicator level classification structure set.
[0056] The trend break segmentation module classifies the structure set according to the pollution index level, performs adjacency comparison of the directional marker values between consecutive batches, and marks the switching position as a breakpoint after identifying the positive and negative switching behavior, and re-divides the sequence after the breakpoint to generate a pollution index trend segmentation set.
[0057] The risk assessment output module is based on the pollution index trend segment division set, combined with food sampling source information, test item category information and test time record information, to perform comprehensive numbering and grouping processing, establish an overall correlation map, and output the big data analysis results of food testing.
[0058] The batch-level pollution indicator path set includes the pollutant indicator time series chain, sample number sequence, detection time axis, pollutant value fluctuation trajectory, and indicator path mapping relationship. The indicator trend distribution segment set includes change direction markers, continuous trend segments, trend change sequence numbers, segment end node identifiers, and trend segment coverage. The pollution indicator level classification structure set includes risk level attribution labels, risk interval coverage information, directional change markers, level upgrade records, and pollutant change magnitude classification. The pollution indicator trend segment division set includes trend segment boundary points, trend direction switching points, segment attribution adjustment sequences, trend segment grouping identifiers, and trend continuity mapping diagrams. The food testing big data analysis results include a comprehensive numbering structure, grouping and classification labels, sampling source information, detection category information, and detection time series distribution map.
[0059] Please see Figure 2 The batch information extraction module includes:
[0060] The data stream receiving submodule acquires the sample dataset from the food physicochemical testing sampling equipment, parses the original sampling information and corresponding test output records, extracts the data of food sample number, test time and contaminant item value, and outputs the test time data stream in a unified timestamp format.
[0061] The data stream receiving submodule establishes a full-duplex communication connection with the physicochemical testing and sampling equipment deployed at the end of the food processing line, listening to and capturing the raw testing messages uploaded by the equipment in real time. This submodule is equipped with a message parsing engine. When it receives a raw sampling information stream in hexadecimal format, the parsing engine, according to a predefined communication protocol frame structure, locates and extracts byte segments containing the unique identifier of the food sample (RFID tag or batch barcode), the testing timestamp, and simultaneously extracts the physicochemical contaminant values (such as lead, cadmium, and aflatoxin content) recorded by the testing equipment. To address the time format differences between different equipment manufacturers, the submodule has a built-in Network Time Protocol (NTP) calibration unit that uniformly converts the parsed testing time into a Unix timestamp format accurate to milliseconds (e.g., 1741021200000), and strongly associates the sample number and contaminant value with this unified timestamp, outputting a standardized testing time data stream.
[0062] The pollution index analysis submodule extracts the corresponding food sample number sequence and pollutant item value group based on the detection time data stream, constructs a pollution data set with the sample number as the primary key, performs an ordered comparison of the pollutant item values in the set according to the preset pollutant item list, performs a sorting operation on the pollutant index in ascending order of time, and generates a pollutant time series arrangement matrix.
[0063] The pollution index analysis submodule is equipped with a high-throughput data cleaning and reassembly unit, receiving the detection time data stream as input. This submodule first allocates a hash map in memory, using the food sample number as the key and all corresponding pollutant items and their numerical values as values, to group the scattered data stream. Subsequently, the submodule calls a quicksort algorithm to sort the pollutant data records under each sample number in ascending order based on the associated timestamp, eliminating data disorder caused by network latency. Based on this, the submodule constructs a two-dimensional data structure as a pollutant time-series arrangement matrix, where the row indices correspond to consecutive time nodes, the column indices correspond to a preset list of pollutant items (e.g., [total arsenic, lead, cadmium, mercury]), and the matrix elements store specific detection values. If a specific pollutant's detection value is missing at a certain time point, the submodule uses linear interpolation to fill the gap based on data from adjacent time points, ensuring the density of the matrix and computational continuity.
[0064] The pollution path construction submodule analyzes the direction of numerical change of pollutant items in continuous samples based on the pollutant numerical information at each time node in the pollutant time sequence arrangement matrix, performs order difference judgment on each pollutant item, and aggregates the indicator change trend of each pollutant on the sample time line to generate a batch-level pollution indicator path set.
[0065] The pollution pathway construction submodule receives the pollutant time-series arrangement matrix and is equipped with a difference operation unit for numerical fluctuation analysis at the microscopic level. The submodule iterates through each column of the matrix (i.e., each pollutant) and calculates the values of adjacent time nodes sequentially according to the row index order of the time dimension. and The difference between pollutant values ( ),in, and Representing time nodes and The submodule calculates the pollutant values. Based on this difference, instead of directly outputting numerical values, it aggregates and generates a vector sequence containing the direction of numerical change (increase, decrease, or no change) and the magnitude of change (rate of change). This sequence intuitively describes the migration and accumulation behavior of a single pollutant between samples during continuous production or sampling. The submodule further merges the vector sequences generated from all pollutant columns to form a batch-level pollutant indicator path set. This path set is essentially a multidimensional tensor that fully records the dynamic evolution trajectory of various physicochemical indicators of the batch of food over time throughout the entire testing period, providing a structured data foundation for subsequent time-series feature analysis.
[0066] Please see Figure 3 The time-series feature recognition module includes:
[0067] The interval extraction submodule obtains the batch-level pollution index path set, divides the pollutant index sequence into equal step sizes according to the detection time order, forms continuous time interval segments for adjacent detection time nodes, and records the corresponding pollutant value sequence in each interval segment to generate a pollution index time interval set.
[0068] The interval extraction submodule loads the batch-level pollution indicator path set and integrates a sliding window segmenter. Based on a preset time step parameter (e.g., a step size of 15 minutes), the segmenter performs equally spaced gridding of the pollutant indicator sequence along the detection time axis. For each segmented time window, the submodule not only extracts all pollutant values within the start and end time of the window but also simultaneously extracts environmental parameters (such as temperature and humidity, if present in the data stream) within that time period, forming an independent data capsule. For continuous detection behavior crossing window boundaries, the submodule employs an overlapping slicing strategy with an overlap rate of 20%, ensuring that feature information located at the interval edges is not lost due to forced truncation. In this way, the continuous path set is discretized into a series of short time-series segments with time labels, generating a pollution indicator time interval set. Each interval set contains a snapshot of the local fluctuations of all pollutant values within that time period.
[0069] The direction consistency filtering submodule is based on the time interval set of pollution indicators. It compares the changing direction of adjacent pollutant values within the interval, and filters the interval based on the consistency of the changing direction sign. It retains the intervals with consistent changing direction and outputs a sequence of continuously changing indicator segments to generate consistent changing interval identification data.
[0070] The direction consistency filtering submodule incorporates a monotonicity verification logic unit to perform feature scanning on each segment of the pollution indicator time interval set. The submodule calculates the sign of the first derivative of pollutant values at adjacent sampling points within the interval. If the signs of the first derivatives between all adjacent points within the interval are consistent (i.e., all positive indicates a continuous increase, and all negative indicates a continuous decrease), then the interval is determined to have a strong consistent direction of change. The submodule filters out cluttered intervals with frequently alternating signs (i.e., exhibiting an oscillating state), retaining only segments reflecting a clear trend of pollution accumulation or decline. For example, in an interval containing 10 sampling points, if the value of each subsequent sampling point is greater than or equal to the value of the previous sampling point, the interval is marked as a "positive consistency interval." After filtering, the submodule links these retained intervals according to the original time series and outputs change consistency interval identification data consisting of the start time, end time, and consistency direction identifier (+1 or -1).
[0071] The trend segment generation submodule extracts the pollutant detection values and detection time information at the end of the corresponding interval segment based on the consistent change interval identifier data, marks the continuous change behavior and performs segment aggregation processing to generate an index trend distribution segment set.
[0072] The trend segment generation submodule receives consistent interval identification data and initiates the segment aggregation engine. This engine identifies adjacent intervals with the same directional identifier and physically merges them on the timeline, thereby constructing a longer-spanning continuous change segment. The submodule focuses on extracting the tail data of each merged segment, namely the pollutant detection value and corresponding timestamp at the moment the trend ends, as this value represents the final cumulative state of the continuous change process. Simultaneously, the submodule calculates the overall slope characteristic of the merged segment to describe the intensity of the trend. After aggregation, the scattered consistent intervals are reconstructed into several macroscopic trend blocks, each carrying core attributes such as "start time, end time, initial value, peak / trough value (tail value), and average rate of change." These structured objects collectively constitute the indicator trend distribution segment set, accurately depicting the phased behavioral characteristics of pollutants over a long time span.
[0073] Please see Figure 4 The risk boundary adjustment module includes:
[0074] The segment discrimination submodule obtains the segment set of indicator trend distribution, reads the pollutant detection values at the end of the segment, performs position matching operation according to the preset risk level critical expression interval boundary value, identifies the interval number corresponding to each detection value, and determines whether it is at the boundary of the risk judgment zone by comparing the difference between each number and the adjacent interval band number, and generates a pollutant level critical proximity marker set.
[0075] The segment discrimination submodule operates based on a set of indicator trend distribution segments, loading a food safety risk level threshold table specified by national or industry standards. This submodule reads the contaminant detection values at the end of each trend segment, using them as probe data and projecting them into a preset risk level critical expression range. For example, for the detection of the heavy metal "lead," the preset safety threshold is 0.2 mg / kg, and the warning zone is set to [0.18, 0.22]. The submodule calculates the Euclidean distance between the tail detection value and the risk boundary values at each level, identifying whether the value falls within the warning zone (i.e., the risk judgment zone boundary). If the detection value... satisfy ,in, Indicates the lower bound of the boundary. This indicates the upper limit of the boundary. The submodule immediately labels the segment as "critically approaching" and records the difference between it and the specific boundary line, generating a set of pollutant level critically approaching markers. This process effectively identifies those "near-exceeding" sample segments that, although not yet exceeding the standard, are already on the verge of high risk.
[0076] The risk enhancement classification submodule extracts all sequence information of the corresponding detection value segment from the pollutant level critical proximity marker set. Segments meeting the boundary conditions are then classified into a classification interval one level higher than the current risk level. Combined with the corresponding detection time identifier, a pollutant level mapping sequence is established in chronological order using the formula:
[0077] ;
[0078] The pollutant level adjustment judgment value is obtained through calculation, and the interval number is updated by combining the level weight threshold of the current classification interval, thereby generating a set of pollutant risk classification intervals; among which, This represents the pollutant level adjustment judgment value for the i-th detection value. This represents the i-th detection value. This represents the boundary value of the risk interval corresponding to the current detection value. Let ik be the ith detection value, n be the calculation span, and R be the normalization coefficient;
[0079] The risk enhancement classification submodule implements a strict risk reassessment strategy for segments marked as "critically close." This submodule extracts all historical data for these segments, aiming to determine whether an artificial risk upgrade is necessary through volatility analysis. Here, dynamic risk assessment logic is introduced; the submodule substitutes the segment data into the risk level adjustment calculation formula.
[0080] ;
[0081] In this operational logic, the submodule first determines the current detection value. and their corresponding risk boundary values Calculate the span Set to the total number of sample points contained in the segment (e.g.) This is used to trace historical fluctuations within this segment. Normalization coefficient. The settings are based on the standard deviation or maximum permissible fluctuation range of this type of pollutant in historical big data (for example, for lead pollutants, the settings are...). ).
[0082] To ensure the reproducibility of the technical solution, the implementation data shown in Table 1 is introduced for illustration:
[0083] Table 1 Calculation Parameters for Pollutant Risk Assessment
[0084] ;
[0085] Referring to Table 1, the submodule performs the following specific operations:
[0086] First, calculate the boundary distance term: .
[0087] Next, the cumulative fluctuation term is calculated: according to the sequence data in Table 1, the sum of the absolute values of adjacent differences is 0.04.
[0088] The total of the molecular parts is: .
[0089] The denominator is the square root of n: .
[0090] The intermediate value is obtained after division: .
[0091] Finally, divide by the normalization coefficient R: .
[0092] The calculation result The submodule compares the result with a preset risk level weight threshold (e.g., 0.6). In this example, 0.447 < 0.6, indicating that although the value is close to the boundary, its internal fluctuations are relatively stable and have not yet reached the point where a mandatory upgrade of the risk level is necessary; therefore, the original risk level classification is maintained. If the calculation result... The submodule will then forcibly reclassify the entire segment into a higher risk category (e.g., from "Attention" to "Warning"). The advantage of this formula is that it considers not only the static distance between the value and the boundary (…). Furthermore, by introducing a summation term, the dynamic instability assessment of the data is enhanced, preventing misjudgments triggered by a single numerical value and upgrading risk assessment from a "point" to a "trend" dimension. After completing the judgment, the submodule updates the classification ID of the segment and generates a set of pollutant risk classification intervals.
[0093] The direction marker generation submodule calculates the difference between the pollutant detection values of adjacent batches based on the pollutant risk classification interval set and the detection time sequence, determines the positive or negative relationship and marks the upward or downward trend, and attaches the trend marker to the corresponding pollutant and classification interval number to generate a pollution index level classification structure set.
[0094] The direction marker generation submodule further refines the trend description based on the established risk classification. The submodule extracts the average pollutant detection values from adjacent batches (i.e., different production shifts or different sampling rounds) according to the detection time sequence and calculates the difference. ),in, This represents the difference between the mean values of pollutant detections. This represents the average value of pollutant tests in the current batch. This represents the average value of pollutant tests in the previous batch. If... Marked as "Upward Risk Trend"; if This is marked as a "risk decline trend." This marker is appended as metadata to the attributes of each risk classification segment. Thus, the generated pollution index level classification structure set not only includes the risk level (high / medium / low) but also carries the dynamic evolution direction (rising / falling), providing complete vector information for subsequent breakpoint analysis.
[0095] Please see Figure 5 The trend break segmentation module includes:
[0096] The directional adjacency comparison submodule classifies the structure set according to the pollution index level, obtains the directional marker value in the continuous batch sequence, compares the sign consistency of the directional marker values of adjacent batches, judges whether the positive and negative signs have switched, and records the batch index number of each sign change position to generate a directional switching judgment sequence.
[0097] The directional adjacency comparison submodule focuses on detecting state abrupt changes during continuous production. This submodule reads the contamination index level classification structure set and extracts the directional marker sequence of consecutive batches (e.g., [up, up, up, down, down]). Internally, the submodule is configured with an XOR logic comparator to compare adjacent batches one by one. and The direction sign is determined by the logic comparator. When the logic comparator detects a sign flip (i.e., from "rising" to "falling", or vice versa), it outputs "1" as a sudden change signal; otherwise, it outputs "0". The submodule records the indexes of all positions where the output is "1", corresponding to time points in the production process where process adjustments, raw material changes, or equipment failures may occur. This process generates a direction switching determination sequence, precisely pinpointing the logical coordinates where a fundamental reversal of the contamination trend occurs.
[0098] The breakpoint identifier generation submodule filters the batch index positions where the symbols switch between positive and negative based on the direction switching judgment sequence, establishes a breakpoint identifier record for each index position, and associates the breakpoint position with the corresponding detection time to generate a trend breakpoint position identifier set.
[0099] The breakpoint identification generation submodule physically instantiates all index positions marked as abrupt changes based on the direction switching determination sequence. For each symbol switching point, the submodule creates a detailed record object containing "breakpoint ID, occurrence time, preceding trend type, subsequent trend type, and pollutant value at the breakpoint." For example, if the trend changes from rising to falling at 14:00, the submodule will generate a breakpoint record and link it to the detection data before and after that moment via a database foreign key. These records constitute a trend breakpoint identification set, which physically divides the continuous time stream into several event windows with distinct characteristics, providing clear time anchors for subsequent attribution analysis.
[0100] The paragraph attribution reconstruction submodule segments the continuous batch sequence in the pollution index level classification structure set according to the trend break position identifier set. It re-numbers the starting position of the paragraph after the break point and adjusts the paragraph attribution mark of the corresponding batch to generate a pollution index trend paragraph division set.
[0101] The segmentation and reorganization submodule performs the final segmentation and reorganization task. This submodule uses the trend breakpoint identifier set as a slicing tool to slice the original pollution index level classification structure set. Each breakpoint is considered the starting position of a new segment. The submodule assigns a unique segment number (Segment_ID) to each independent segment and redefines its attribution label based on the core trend attributes within that segment (such as "lead content - high risk - continuously rising segment"). This process eliminates the original mechanical division based on a fixed time step, replacing it with a semantic division based on "pollution behavior characteristics," generating a pollution index trend segmentation set with actual physical meaning, ensuring that each data segment corresponds to an independent pollution evolution event.
[0102] Please see Figure 6 The risk assessment output module includes:
[0103] The information merging submodule is based on the pollution index trend segment division set, obtains food sampling source information, test item category information and test time record information, performs joint numbering processing on the segment number and the three types of information, completes batch-level grouping based on the numbering, and generates a joint numbering set of test information.
[0104] The information merging submodule serves as the fusion center for multi-dimensional heterogeneous data. This submodule accesses a set of pollution indicator trend segments and simultaneously retrieves food sampling source information (origin, supplier) from the enterprise's ERP system, detection item category information (heavy metals, pesticide residues, microorganisms) from the LIMS system, and on-site time record information. The submodule employs a combined primary key encoding technique, hashing the segment number with the aforementioned three types of metadata to generate a unique combined number (e.g., SEG-2025-02-03-PB-MILK-001). In this way, the otherwise dry detection values are given a complete business context, generating a set of combined detection information numbers, enabling the answering of complex questions such as "What trend does the lead content of milk supplied by a certain supplier show within a specific time period?"
[0105] The association graph construction submodule reads the paragraph identifiers from the pollution index trend paragraph division set based on the joint number set of detection information, constructs a relationship mapping structure with paragraph numbers as nodes, records the connection relationships between nodes in the dimensions of sampling source, detection item category and detection time, and establishes an overall association graph;
[0106] The correlation graph construction submodule reads all entity objects from the joint number set of detection information and instantiates each trend segment as a "node" in the graph. Subsequently, the submodule analyzes the commonalities in metadata between nodes and constructs "edge" relationships: if two nodes come from the same supplier, a "same-source association" edge is established; if two nodes belong to the same type of pollutant, a "homogeneous association" edge is established; if two nodes are sequential in time, a "temporal association" edge is established. Through this fully connected operation, the submodule constructs a multi-dimensional, three-dimensional overall correlation graph. In this graph, isolated detection batches are woven into a network, visually demonstrating the transmission path of pollution risk across the supply chain, timeline, and pollutant categories.
[0107] The analysis results output submodule is based on the overall correlation map. It organizes the correlation density and distribution of each node, summarizes the node information according to a unified numbering order, and outputs the big data analysis results of food testing.
[0108] The analysis results output submodule is equipped with a graph analysis algorithm and a visualization rendering engine. This submodule traverses the overall correlation graph, calculates the degree centrality and clustering coefficient of each node to identify the key nodes with the highest risk (such as a long-term high-risk trend segment associated with a specific supplier). The submodule organizes the analysis results according to a hierarchical structure of "risk level-time-source" to generate a structured JSON report or a visualized heatmap dashboard. The final output of the food testing big data analysis results not only includes specific exceedance alarms but also provides risk tracing paths and future trend predictions, assisting managers in quickly locating the source of contamination and formulating targeted containment measures. This result demonstrates that multidimensional graph correlation analysis can effectively reveal the hidden risk network behind a single test data point.
[0109] The big data analysis method for food testing includes the following steps:
[0110] S1: Obtain the pollutant index values corresponding to each batch of food samples in the food physicochemical testing sampling equipment, extract the sample number and detection time, arrange them in ascending order according to the detection time, construct a continuous information chain of pollutant indicators in the time dimension, and generate a batch-level pollutant indicator path set.
[0111] S2: Based on the pollutant indicators in the batch-level pollution indicator path set, extract intervals in the order of detection time, compare the changing direction of continuous data within the interval, filter the indicator segments with the same changing direction, extract the pollutant detection information at the end of the segment, and generate a set of indicator trend distribution segments.
[0112] S3: Based on the distribution segment set of indicator trends, the positional relationship between the pollutant detection information at the end of the segment and the preset risk level critical expression interval is determined. At the same time, the positive and negative relationships of the change values of pollutants in adjacent batches are extracted and directional markers are added to generate a pollution indicator level classification structure set.
[0113] S4: Based on the pollution index level classification structure set, perform adjacency comparison of directional marker values between consecutive batches. When the positive and negative directional switching behavior is identified, mark the switching position as a breakpoint and re-divide the sequence after the breakpoint to generate a pollution index trend segment division set.
[0114] S5: Based on the pollution index trend segment division set, combined with food sampling source information, test item category information and test time record information, comprehensive numbering and grouping processing is performed to establish an overall correlation map and output the big data analysis results of food testing.
[0115] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A food inspection big data analysis system, characterized by, include: The batch information extraction module obtains the pollutant index values corresponding to each batch of food samples in the food physicochemical testing sampling equipment, extracts the sample number and detection time, arranges them in ascending order according to the detection time, and constructs a continuous information chain of pollutant indicators in the time dimension to generate a batch-level pollutant indicator path set. The time-series feature recognition module extracts interval segments according to the detection time sequence based on the pollutant indicators in the batch-level pollution indicator path set, compares the changing direction of continuous data within the interval, filters indicator segments with the same changing direction, extracts pollutant detection information at the end of the segment, and generates an indicator trend distribution segment set. The risk boundary adjustment module performs interval discrimination based on the positional relationship between the pollutant detection information at the end of the segment and the preset risk level critical expression interval according to the index trend distribution segment set. At the same time, it extracts the positive and negative relationship of the change values of pollutants in adjacent batches and adds directional markers to generate a pollution index level classification structure set. The trend break segmentation module performs adjacency comparison of directional marker values between consecutive batches based on the pollution index level classification structure set. When the positive and negative directional switching behavior is identified, the switching position is marked as a breakpoint and the sequence after the breakpoint is re-divided to generate a pollution index trend segmentation set. The risk assessment output module, based on the pollution index trend segment division set, combined with food sampling source information, test item category information and test time record information, performs comprehensive numbering and grouping processing, establishes an overall correlation map, and outputs food testing big data analysis results; The risk boundary adjustment module includes: The segment discrimination submodule obtains the index trend distribution segment set, reads the pollutant detection values at the end of the segment, performs position matching operation according to the preset risk level critical expression interval boundary value, identifies the interval number corresponding to each detection value, and determines whether it is at the risk judgment zone boundary by comparing the difference between each number and the adjacent interval band number, and generates a pollutant level critical proximity marker set. The risk enhancement classification submodule extracts all sequence information of the corresponding detection value segment in the indicator trend distribution segment set based on the pollutant level critical proximity marker set. The segments that meet the boundary conditions are classified into a classification interval one level higher than the current risk level. Combined with the corresponding detection time identifier, a pollutant level mapping sequence is established in chronological order. The pollutant level adjustment judgment value is calculated and obtained. The interval number is updated by combining the level weight threshold of the current classification interval, and a pollutant risk classification interval set is generated. The direction marker generation submodule calculates the difference between the pollutant detection values of adjacent batches based on the pollutant risk classification interval set and the detection time sequence, determines the positive or negative relationship and marks the upward or downward trend, and attaches the trend marker to the corresponding pollutant and classification interval number to generate a pollution index level classification structure set.
2. The food inspection big data analytics system of claim 1, wherein: When performing interval discrimination on the positional relationship, if the pollutant detection information at the end of the paragraph falls within the adjacent judgment zone of the preset risk level critical expression interval, then the corresponding segment will be classified into a higher-level risk level expression interval. The batch-level pollution indicator path set includes a pollutant indicator time series chain, sample number sequence, detection time axis, pollutant value fluctuation trajectory, and indicator path mapping relationship. The indicator trend distribution segment set includes change direction markers, continuous trend segments, trend change sequence numbers, segment end node identifiers, and trend segment coverage. The pollution indicator level classification structure set includes risk level attribution labels, risk interval coverage information, directional change markers, level upgrade records, and pollutant change magnitude classification. The pollution indicator trend segment division set includes trend segment boundary points, trend direction switching points, segment attribution adjustment sequences, trend segment grouping identifiers, and trend continuity mapping diagrams. The food testing big data analysis results include a comprehensive numbering structure, grouping and classification labels, sampling source information, detection category information, and detection time series distribution map.
3. The food inspection big data analytics system of claim 1, wherein, The batch information extraction module includes: The data stream receiving submodule acquires the sample dataset from the food physicochemical testing sampling equipment, parses the original sampling information and corresponding test output records, extracts the data of food sample number, test time and contaminant item value, and outputs the test time data stream in a unified timestamp format. The pollution index analysis submodule extracts the corresponding food sample number sequence and pollutant item value group based on the detection time data stream, constructs a pollution data set with the sample number as the primary key, performs an ordered comparison of the pollutant item values in the set according to the preset pollutant item list, performs a sorting operation on the pollutant index in ascending order of time, and generates a pollutant time sequence arrangement matrix. The pollution path construction submodule analyzes the direction of numerical change of pollutant items in continuous samples based on the pollutant numerical information at each time node in the pollutant time sequence arrangement matrix, performs order difference judgment on each pollutant item, and aggregates the indicator change trend of each pollutant on the sample time line to generate a batch-level pollution indicator path set.
4. The food inspection big data analytics system of claim 1, wherein, The time-series feature recognition module includes: The interval extraction submodule obtains the batch-level pollution index path set, divides the pollutant index sequence into equal-step segments according to the detection time order, forms continuous time interval segments for adjacent detection time nodes, and records the corresponding pollutant value sequence in each interval segment to generate a pollution index time interval set. The direction consistency filtering submodule compares the changing direction of adjacent pollutant values within the time interval based on the pollution index time interval set, filters the intervals based on the consistency of the changing direction sign, retains the intervals with consistent changing direction, and outputs a sequence of continuously changing index segments to generate consistent changing interval identification data. The trend segment generation submodule extracts the pollutant detection values and detection time information at the end of the corresponding interval segment based on the consistent change interval identifier data, marks the continuous change behavior, performs segment aggregation processing, and generates an index trend distribution segment set.
5. The food testing big data analysis system according to claim 1, characterized in that, The formula for calculating the pollutant level adjustment judgment value is as follows: ; in, This represents the pollutant level adjustment judgment value for the i-th detection value. This represents the i-th detection value. This represents the boundary value of the risk interval corresponding to the current detection value. Let ik be the ith detection value, n be the calculation span, and R be the normalization coefficient.
6. The food testing big data analysis system according to claim 1, characterized in that, The trend fracture segmentation module includes: The directional adjacency comparison submodule obtains the directional marker values in the continuous batch sequence based on the pollution index level classification structure set, performs sign consistency comparison on the directional marker values of adjacent batches, judges whether the positive and negative signs have switched, and records the batch index number of each sign change position to generate a directional switching judgment sequence. The breakpoint identifier generation submodule filters the batch index positions where the symbols switch between positive and negative based on the direction switching determination sequence, establishes a breakpoint identifier record for each index position, and associates the breakpoint position with the corresponding detection time to generate a trend breakpoint position identifier set. The paragraph attribution reconstruction submodule segments the continuous batch sequence in the pollution index level classification structure set according to the trend break position identifier set. It re-numbers the starting position of the paragraph after the break point and adjusts the paragraph attribution mark of the corresponding batch to generate a pollution index trend paragraph division set.
7. The food testing big data analysis system according to claim 1, characterized in that, The risk assessment output module includes: The information merging submodule obtains food sampling source information, test item category information, and test time record information based on the pollution index trend segment division set. It performs joint numbering processing on the segment number and the three types of information, completes batch-level grouping based on the numbering, and generates a joint numbering set of test information. The association graph construction submodule reads the paragraph identifiers from the pollution index trend paragraph division set based on the joint number set of the detection information, constructs a relationship mapping structure with paragraph numbers as nodes, records the connection relationships between nodes in the dimensions of sampling source, detection item category and detection time, and establishes an overall association graph. The analysis results output submodule, based on the overall correlation map, organizes the correlation density and distribution status of each node, summarizes the node information according to a unified numbering order, and outputs the big data analysis results of food testing.
8. A big data analysis method for food testing, characterized in that, The method is used to implement the food testing big data analysis system according to any one of claims 1-7, and includes the following steps: S1: Obtain the pollutant index values corresponding to each batch of food samples in the food physicochemical testing sampling equipment, extract the sample number and detection time, arrange them in ascending order according to the detection time, construct a continuous information chain of pollutant indicators in the time dimension, and generate a batch-level pollutant indicator path set. S2: Based on the pollutant indicators in the batch-level pollution indicator path set, extract intervals in the order of detection time, compare the changing direction of continuous data within the interval, filter indicator segments with the same changing direction, extract pollutant detection information at the end of the segment, and generate an indicator trend distribution segment set. S3: Based on the distribution segment set of the index trend, the positional relationship between the pollutant detection information at the end of the segment and the preset risk level critical expression interval is determined by interval discrimination. At the same time, the positive and negative relationship of the change values of pollutants in adjacent batches is extracted and directional markers are added to generate a pollution index level classification structure set. S4: Based on the pollution index level classification structure set, perform adjacency comparison of the directional marker values between consecutive batches. When the positive and negative directional switching behavior is identified, mark the switching position as a breakpoint and re-divide the sequence after the breakpoint to generate a pollution index trend segment division set. S5: Based on the pollution index trend segment division set, combined with food sampling source information, test item category information and test time record information, perform comprehensive numbering and grouping processing, establish an overall correlation map, and output the food testing big data analysis results.