Food detection sample management method and system based on big data analysis

By acquiring data from the entire process of food testing samples, and using text similarity algorithms and hidden Markov models for semantic clarification and association rule construction, the problem of insufficient semantic association in existing sample management platforms has been solved. This enables intelligent retrieval and knowledge mining, and improves data utilization efficiency and risk identification capabilities.

CN121764977BActive Publication Date: 2026-05-08GUIZHOU SHIKEYUAN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUIZHOU SHIKEYUAN INFORMATION TECH CO LTD
Filing Date
2026-03-04
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing big data platforms for food testing sample management lack multi-dimensional semantic association design, resulting in insufficient intelligent retrieval and knowledge mining capabilities. They are unable to effectively identify the semantic associations and logical relationships between different types of sample information, making it difficult to uncover hidden management patterns and risk characteristics from massive sample data.

Method used

By acquiring data from the entire process of food testing samples, a text similarity algorithm is used to determine whether there is semantic ambiguity in unstructured data. A hidden Markov model is used for sequence feature parsing and semantic completion processing to generate semantic clarification results. Based on this, semantic association rules are established, and a knowledge graph is constructed for intelligent retrieval and knowledge mining.

Benefits of technology

It enables semantic enhancement and deep correlation mining of unstructured text information in the entire process data of food testing samples, improving data quality and usability. It can automatically uncover hidden management patterns and risk characteristics, providing accurate data decision support for sample management process optimization and food safety risk assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121764977B_ABST
    Figure CN121764977B_ABST
Patent Text Reader

Abstract

The application discloses a food detection sample management method and system based on big data analysis, and particularly relates to the technical field of food detection data processing, and is used for solving the problem of insufficient intelligent search and knowledge mining ability caused by unstructured text semantic ambiguity and lack of deep semantic association between data in the existing sample management big data platform; the problem is solved by the following steps: obtaining sample full-process data containing structured and unstructured data, using a text similarity algorithm to determine the semantic ambiguity of unstructured data, using a hidden Markov model combined with space-time correlation characteristics to perform sequence analysis and semantic completion on the ambiguous text to generate a semantic clarification result, then establishing and optimizing semantic association rules based on the semantic clarification result and the structured data, and finally constructing a food detection sample information knowledge graph according to the optimized rules, so as to realize intelligent semantic search and deep knowledge mining of sample information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of food testing data processing technology, and more specifically, to a food testing sample management method and system based on big data analysis. Background Technology

[0002] In food testing sample management, with the advancement of digital transformation, sample management can gradually rely on big data platforms to achieve full lifecycle information control. Currently, big data platforms are mainly used to collect and store relevant information for the entire process of food testing samples, from sampling, receipt, storage, testing and allocation to sample disposal. This information includes structured data such as category, source, and batch, as well as unstructured data such as sampling site descriptions, test anomaly explanations, and disposal notes. The platform's core functions focus on information entry, archiving, and basic query and statistics. Through simple integration of structured data, it provides basic data support for sample management, helping to achieve traceability of sample information and standardized process control.

[0003] In existing big data application solutions for food testing sample management, the organization and management of sample information lack multi-dimensional semantic association design. This results in a significant deficiency in the intelligent retrieval and knowledge mining capabilities of the big data platform. Consequently, the platform can only achieve precise retrieval based on keywords and cannot effectively identify the semantic associations and logical relationships between different types of sample information. It is difficult to extract hidden management patterns and risk characteristics from massive sample data, and the value of massive sample data cannot be fully released. Consequently, it cannot provide accurate and effective data support for optimizing food testing sample management and food safety regulatory decisions. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a food testing sample management method and system based on big data analysis to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] Food testing sample management methods based on big data analysis include:

[0007] S1. Obtain full-process data for food testing samples, including structured and unstructured data;

[0008] S2. Use a text similarity algorithm to determine whether there is semantic ambiguity in unstructured data;

[0009] S3. When semantic ambiguity exists, a hidden Markov model is used from the perspective of the spatiotemporal correlation characteristics of the entire sample process data to perform sequence feature parsing and semantic completion processing on the unstructured data with semantic ambiguity, and generate semantic clarification results.

[0010] S4. Establish semantic association rules based on semantic clarification results and structured data, and define the correspondence between semantic clarification results and structured data;

[0011] S5. Conduct logical self-consistency verification of semantic association rules across sample management stages and rule adaptation for sample risk levels to generate optimized semantic association rules.

[0012] S6. Construct a knowledge graph of food testing sample information based on optimized semantic association rules, and perform intelligent retrieval and knowledge mining of food testing sample information.

[0013] Furthermore, S1 includes:

[0014] Obtain the complete process data of the samples that have been entered from the food testing sample management big data platform;

[0015] The obtained sample process data confirms that it includes structured data such as sample type, source, and batch, as well as unstructured data including sampling site description, test anomaly explanation, and handling notes.

[0016] Furthermore, S2 includes:

[0017] Unstructured data, including sampling site descriptions, abnormal detection explanations, and handling notes, obtained from the entire sample process data, are processed through text segmentation and noise reduction.

[0018] Convert unstructured data that has undergone text segmentation and denoising into text feature representations;

[0019] Calculate the similarity values ​​between text feature representations;

[0020] Based on the comparison results between the similarity value and the preset similarity threshold, it is determined whether there is semantic ambiguity in the unstructured data.

[0021] Furthermore, S3 includes:

[0022] Extract the spatiotemporal correlation features of samples from the entire process data of samples corresponding to unstructured data with semantic ambiguity;

[0023] Based on the extracted spatiotemporal correlation features and unstructured data with semantic ambiguity, a hidden Markov model is constructed, including state sequences and observation sequences.

[0024] By decoding the observed sequence using a hidden Markov model, the implicit sequence features of unstructured data with semantic ambiguity can be extracted.

[0025] Based on the parsed hidden sequence features, semantic completion is performed on unstructured data with semantic ambiguity to generate semantically clarified results.

[0026] Furthermore, the decoding of the observation sequence by the Hidden Markov Model includes: applying the Viterbi algorithm, based on the Hidden Markov Model parameters and the observation sequence, recursively calculating the most likely hidden state sequence path through dynamic programming, and backtracking to obtain the complete hidden state sequence as the parsed hidden sequence features.

[0027] Furthermore, S4 includes:

[0028] Extract key semantic elements from the semantic clarification results;

[0029] Extract data elements associated with key semantic elements from structured data of sample category, origin, and batch;

[0030] Analyze the logical correspondence between key semantic elements and data elements;

[0031] Based on the analyzed logical correspondences, semantic association rules are formed to define the correspondence between semantic clarification results and structured data.

[0032] Furthermore, S5 includes:

[0033] The established semantic association rules are mapped to the sampling, storage, testing and sample retention disposal stages of the entire sample management process;

[0034] Verify the consistency of the mapping logic of the same semantic association rule in different stages, and identify the logical conflicts that arise when different semantic association rules are applied across stages;

[0035] Based on the descriptions and handling notes of detection anomalies associated with semantic association rules in historical sample data, analyze and label the risk characteristics and risk levels implied by each semantic association rule;

[0036] Based on the identification results of logical conflicts and the labeled risk characteristics and risk levels, the semantic association rules are logically reconstructed and weighted.

[0037] The semantic association rules that have undergone logical restructuring and weight calibration are output as optimized semantic association rules.

[0038] Furthermore, the logical reconstruction and weight calibration of semantic association rules include: reconstructing the logic by redefining or combining entities and relational predicates in semantic association rules that have logical conflicts, and assigning differentiated confidence weights to their logical relations based on the risk characteristics and risk levels marked by the semantic association rules to complete the weight calibration.

[0039] Furthermore, S6 includes:

[0040] Using optimized semantic association rules as the rules for graph construction, the definitions of entities and relations in the knowledge graph are determined;

[0041] The semantic clarification results are then converted into entity nodes and relation edges in a knowledge graph based on the correspondence defined by the optimized semantic association rules, according to the structured data of sample category, source, and batch.

[0042] Receive retrieval requests for food testing sample information and parse the natural language queries in the retrieval requests into structured queries based on knowledge graph entities and relationships;

[0043] Structured queries are performed in the completed food testing sample information knowledge graph, and knowledge mining is carried out by traversing and matching paths in the knowledge graph.

[0044] On the other hand, the present invention provides a food testing sample management system based on big data analysis, comprising:

[0045] The data acquisition module is used to acquire data from the entire process of food testing samples, including structured and unstructured data.

[0046] The fuzzy judgment module is used to determine whether unstructured data has semantic ambiguity using a text similarity algorithm;

[0047] The clarification generation module is used to perform sequence feature parsing and semantic completion processing on unstructured data with semantic ambiguity from the perspective of spatiotemporal correlation features of the entire sample process data, and generate semantic clarification results when semantic ambiguity exists.

[0048] The rule-building module is used to build semantic association rules based on semantic clarification results and structured data, and to define the correspondence between semantic clarification results and structured data;

[0049] The rule optimization module is used to perform logical self-consistency verification of semantic association rules across sample management stages and rule adaptation for sample risk levels, and to generate optimized semantic association rules.

[0050] The retrieval and mining module is used to construct a knowledge graph of food testing sample information based on optimized semantic association rules, and to perform intelligent retrieval and knowledge mining of food testing sample information.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] 1. This approach achieves semantic enhancement and deep association mining of unstructured text information in the entire process data of food testing samples. Compared with existing solutions that can only store and perform keyword matching, this approach uses text similarity judgment and sequence parsing technology based on Hidden Markov Models to perform context-aware semantic completion of semantically ambiguous texts such as on-site descriptions and anomaly explanations, generating semantically clear clarification results. This significantly improves the quality and usability of unstructured data, providing a reliable semantic foundation for subsequent analysis. Furthermore, by systematically establishing and optimizing semantic association rules, a multi-dimensional and verifiable logical correspondence is constructed between the semantically clarified results and structured data such as category and origin. This not only achieves semantic connectivity between different types of data but also transforms the data organization method from a simple stacking of fields to an association network containing rich domain logic, creating conditions for deeper information integration and utilization.

[0053] 2. Based on semantic enhancement and association construction, the knowledge graph is built and applied through optimized semantic association rules. The resulting knowledge graph transforms discrete sample information into a structured knowledge network linked by entities and relationships, thereby changing the mode of information retrieval and knowledge discovery. It can perform intelligent retrieval based on semantic graphs, understand the deep intent of user queries, and return related and systematic information, rather than scattered field matching results. More importantly, by performing path traversal and pattern matching in the graph, it can automatically uncover complex association rules and potential risk characteristics hidden in different management links and data dimensions. For example, it can discover implicit coupling relationships between specific sources, specific categories, and specific detection anomalies, so that the management patterns and risk signals contained in massive historical sample data can be fully released, providing data decision support based on deep association analysis for the optimization of sample management processes and the accurate assessment of food safety risks. Attached Figure Description

[0054] Figure 1 This is a flowchart of the food testing sample management method based on big data analysis according to the present invention;

[0055] Figure 2 This is a schematic diagram of the food testing sample management system based on big data analysis according to the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Example 1: Figure 1 This invention presents a food testing sample management method based on big data analysis, comprising:

[0058] S1. Obtain full-process data for food testing samples, including structured and unstructured data;

[0059] S2. Use a text similarity algorithm to determine whether there is semantic ambiguity in unstructured data;

[0060] S3. When semantic ambiguity exists, a hidden Markov model is used from the perspective of the spatiotemporal correlation characteristics of the entire sample process data to perform sequence feature parsing and semantic completion processing on the unstructured data with semantic ambiguity, and generate semantic clarification results.

[0061] S4. Establish semantic association rules based on semantic clarification results and structured data, and define the correspondence between semantic clarification results and structured data;

[0062] S5. Conduct logical self-consistency verification of semantic association rules across sample management stages and rule adaptation for sample risk levels to generate optimized semantic association rules.

[0063] S6. Construct a knowledge graph of food testing sample information based on optimized semantic association rules, and perform intelligent retrieval and knowledge mining of food testing sample information.

[0064] S1. Obtain complete data from the entire food testing process, including structured and unstructured data. The specific implementation is as follows:

[0065] The process involves retrieving complete data from the food testing sample management big data platform. This platform is a software system that centrally stores and manages data generated during food testing activities. Deployed on a server, it provides data entry and access interfaces for testing personnel in different locations. In practice, technicians connect to the platform's database via an application programming interface (API) or a direct data query language and construct a data query request. This request explicitly specifies the conditions that the data to be retrieved must meet. These conditions are set by the technicians based on the objectives of the current analysis task. For example, the time condition could be specified as sample entry between January 1, 2025, and December 31, 2025; the location condition as the sample originating from a testing laboratory in East China; or the category condition as dairy products or meat. Upon receiving this query request, the platform retrieves all matching data records from its relational or non-relational database, based on the conditions specified in the query request. These data records are associated with one or more sample numbers and arranged chronologically, comprehensively covering every data point generated sequentially from the sample collection stage, sample registration stage, internal laboratory storage and transfer stage, testing task allocation stage, to the final sample disposal stage. After retrieval, the food testing sample management big data platform merges and packages these scattered data records according to their inherent sample number and chronological order logic, generating a structured data set, such as a multi-row, multi-column data table file or a data exchange file in a specific format. This data set is then transmitted over the network and returned to the requesting technician or system module, thus completing the acquisition of the entire process data of the entered samples.

[0066] The verification process involves obtaining structured data on sample categories, sources, and batches, as well as unstructured data including sampling site descriptions, anomaly reports, and handling notes. Upon receiving this data set, the system or technical personnel execute a data content parsing and verification process. For the verification of structured data on sample categories, sources, and batches, the specific process is as follows: First, read the metadata or headers of the data set to check for columns whose field names completely match the sample category, sources, and batches. Then, verify the format and validity of the specific data content under these columns. For example, for the sample category column, verify that each data value belongs to a predefined and maintained standard sample category list by the food testing sample management big data platform. This list may include specific names such as pasteurized milk, fermented milk, smoked sausages, and edible vegetable oils; any value not in this list will be marked as abnormal. For the source column, verify that its data values ​​conform to a preset format of enterprise name and address combination or a specific administrative division code format. For batch columns, verify that their data values ​​conform to the encoding rule of production date plus serial number. For example, 20250517001 represents batch number 001 on May 17, 2025. These verifications ensure the standardization and direct computability of the structured data.

[0067] The specific process for confirming unstructured data containing descriptions of sampling sites, anomaly reports, and handling remarks is as follows: Check if the dataset contains text-type columns with field names exactly matching the sampling site description, the anomaly report, and the handling remarks. After confirming the existence of these columns, further check if they contain substantive text information, regardless of their length or specific wording. For example, a record in the sampling site description column might state that the sampling location was a supermarket freezer, the product packaging was intact, and the storage temperature was 4 degrees Celsius; this is considered valid unstructured data. Even if another record states that the site conditions were normal, this is still considered a valid record rather than a null value, as it indicates that there were no special circumstances at that stage. Similarly, the anomaly report column might contain descriptions such as the total bacterial count exceeding the standard limit by 50%, and the handling remarks column might contain descriptions such as the excessive sample being sealed and the manufacturer being notified for a recall. The system verifies that the required unstructured data is fully included in the acquired sample data by checking the existence of column names and the non-emptiness of content in these specified text columns. This verification step is the data foundation guarantee for the entire method, ensuring that all input data necessary for subsequent text similarity calculation, semantic completion, and rule establishment are accurately available.

[0068] S2. Use a text similarity algorithm to determine whether there is semantic ambiguity in unstructured data. The specific implementation is as follows:

[0069] Perform text tokenization and denoising processing on the unstructured data of the sampling site description, detection anomaly description, and disposal remarks obtained from the full-process data of the sample. Text tokenization processing refers to the process of splitting a continuous Chinese text sequence into independent and meaningful basic units, namely words. In implementation, a tokenization algorithm that combines dictionary matching and statistical models is adopted. By loading an extended dictionary containing professional vocabulary in the food detection field, which includes terms such as total number of colonies, pesticide residues, aseptic sampling, and sample storage library, the text in the unstructured data is scanned and split. For the text in the sampling site description, such as the sampling location is a supermarket freezer, the product packaging is intact, and the storage temperature is displayed as 4 degrees Celsius, a possible sequence of words obtained after tokenization processing is sampling, location, is, supermarket, freezer, product, packaging, intact, storage, temperature, display, is, 4, degrees Celsius. Denoising processing is to remove words in the text that contribute little to subsequent semantic analysis or may introduce interference after tokenization. This is usually achieved through a predefined general stop word list and a domain stop word list. The stop word list contains function words lacking independent semantics, such as de, di, de, zai, le, wei, etc., as well as high-frequency but low-information words such as ben, this time, perform, etc. Compare the tokenization results with the stop word list and delete the matching words. For example, remove words such as wei in the above sequence. At the same time, denoising processing also includes removing non-semantic characters such as pure numbers, English letters, and special punctuation marks that may exist in the text, unless the number forms a meaningful expression in combination with a specific measurement unit. For example, 4 in 4 degrees Celsius is retained because it is combined with degrees Celsius. After this step, the unstructured data is transformed into a set composed of a series of clean and meaningful words.

[0070] The process transforms unstructured data, after text segmentation and denoising, into text feature representations. Text feature representation aims to convert the semantic content of text into a numerical vector form that computers can perform mathematical operations on. In this embodiment, a bag-of-words model combined with the term frequency-inverse document frequency (TNF) method is used to generate text feature vectors. First, a global vocabulary is constructed based on all unstructured data samples to be processed. This vocabulary contains all unique words that appear after segmentation and denoising. Then, for each unstructured data sample, such as a specific anomaly detection description, the frequency of each word in the sample is counted based on its word sequence. Simultaneously, the document frequency (VRF) of each word in all samples of the entire dataset is calculated, i.e., the number of samples containing that word. The TRF-inverse document frequency value is obtained by multiplying the term frequency and the VRF, where the VRF is typically calculated by dividing the total number of samples by the number of samples containing that word and then taking the logarithm. Ultimately, each text sample is represented as a high-dimensional numerical vector, with the dimension equal to the size of the global vocabulary. Each position in the vector corresponds to the term frequency-inverse document frequency (IF-VRF) of a specific word in that sample. If a word does not appear in the sample, its corresponding position has a value of 0. In this way, the semantic information of the text is encoded into specific, measurable numerical features.

[0071] The similarity value between text feature representations is calculated. The similarity value is a numerical indicator used to quantify the semantic closeness between two text feature vectors. This embodiment uses the cosine similarity algorithm for calculation. For any two unstructured data sets to be compared, such as a sampled scene description and a historical sampled scene description, they have been converted into two text feature vectors. Cosine similarity calculates the cosine of the angle between these two vectors in vector space. Specifically, the calculation process is as follows: First, the dot product of the two vectors is calculated, that is, the values ​​of the corresponding dimensions of the two vectors are multiplied and then all the product results are added together. Then, the magnitude of each vector is calculated, that is, the square root of the sum of the squares of the values ​​of each dimension of the vectors. Finally, the dot product is divided by the product of the magnitudes of the two vectors, and the result is the cosine similarity value. This value ranges from 0 to 1. The closer the value is to 1, the more consistent the directions of the two vectors are, and the higher the semantic similarity of the corresponding texts; the closer the value is to 0, the greater the semantic difference. This calculation is performed between the unstructured data to be judged and a reference dataset consisting of historically clear and well-defined samples. For each piece of data to be judged, a similarity value is calculated between it and multiple samples in the reference dataset, and the average or maximum value can be taken as its representative similarity value.

[0072] The system compares the similarity values ​​with a preset similarity threshold to determine whether unstructured data exhibits semantic ambiguity. The preset similarity threshold is a pre-defined numerical boundary between semantically clear and semantically ambiguous data. This threshold is obtained and set based on historical data analysis. Specifically, a historical dataset is collected, containing both manually labeled semantically clear and semantically ambiguous text samples. This dataset undergoes the same text segmentation and denoising process, conversion to text feature representation, and calculation of similarity with a clear reference set to obtain the similarity value for each historical sample. Statistical analysis is then performed, such as calculating the distribution range of similarity values ​​for labeled semantically clear samples and the distribution range of similarity values ​​for labeled semantically ambiguous samples, observing the numerical separation between the two. Finally, a value that effectively distinguishes the two groups is selected as the preset similarity threshold. For example, the lower quartile of the distribution of clear sample similarity values ​​can be chosen, or a value that maximizes classification accuracy can be selected through iterative testing. In actual judgment, the similarity value calculated from the unstructured data to be judged is compared with a preset similarity threshold. If the similarity value is lower than the preset similarity threshold, for example, if the calculated similarity value is 0.45 while the preset similarity threshold is 0.65, then the unstructured data is judged to have semantic ambiguity. If the similarity value is equal to or higher than the preset similarity threshold, then it is judged to be semantically clear. For example, if the text feature representation of a certain disposal note has a similarity value of 0.45 with the average of the clear reference set, which is lower than the preset similarity threshold of 0.65, then the disposal note is judged to have semantic ambiguity and needs to proceed to the subsequent semantic completion process.

[0073] S3. When semantic ambiguity exists, a Hidden Markov Model is used from the perspective of the spatiotemporal correlation characteristics of the entire sample data to perform sequence feature parsing and semantic completion processing on the unstructured data with semantic ambiguity, generating semantically clarified results. The specific implementation is as follows:

[0074] This process extracts spatiotemporal correlation features from the entire sample flow data corresponding to unstructured data with semantic ambiguity. Spatiotemporal correlation features refer to quantified or encoded data that characterizes the temporal sequence and spatial location information of a sample within the management process. The specific extraction method for spatiotemporal correlation features is as follows: Standard timestamp fields recorded in each management step are identified and parsed from the entire sample flow data; the time intervals between adjacent steps are calculated; and these time intervals are arranged into a time interval sequence according to the order in which the steps occur. Furthermore, the absolute time of each step can be converted into a cumulative number of hours from the initial sampling time, forming a cumulative time series. For example, for a certain sample, its sampling time, sample receipt time, warehousing time, outbound testing time, and disposal time are recorded respectively. The calculated intervals are: sampling to receipt (2 hours), receipt to warehousing (3 hours), warehousing to outbound (24 hours), and outbound to disposal (720 hours). The time interval sequence is then 2, 3, 24, 720; the corresponding cumulative time series from the sampling time is 0, 2, 5, 29, 749. The specific method for extracting spatial association features is as follows: Geographic information is extracted from structured data fields such as the source of the sample's entire process data. For example, the address of the production enterprise described in the text is parsed into latitude and longitude coordinates using geocoding services. Simultaneously, location information is extracted from unstructured text such as the description of the sampling site by matching a predefined list of keywords. For example, keywords such as "supermarket," "warehouse," and "cold storage" are matched and mapped to predefined location type codes, such as code 1 representing a retail terminal, code 2 representing a warehousing environment, and code 3 representing a cold chain environment. Following the order of management stages, the latitude and longitude coordinates or location type codes corresponding to each stage are arranged to form a spatial location sequence or location type sequence. For example, if the source address of a sample is parsed into latitude and longitude coordinates 121.5 and 31.3, and the keywords "supermarket" and "cold storage" are extracted from the sampling site description and mapped to location type codes 1 and 3, its spatial association features may be represented as location type sequences 1 and 3.

[0075] Based on the extracted spatiotemporal correlation features and semantically ambiguous unstructured data, a state sequence and observation sequence of a Hidden Markov Model (HMM) are constructed. An HMM is a statistical model used to describe observation sequences generated by hidden Markov chains. In this embodiment, the inherent, unobservable stages of the sample's management process are defined as the states of the HMM. The state set can be predefined as including: sampling, transportation, storage awaiting inspection, testing, storage after inspection, and disposal. The observation sequence consists of the actually recordable data generated at each stage, including semantic units extracted from unstructured data and some quantized values ​​from the spatiotemporal correlation features. The method for constructing the state sequence is to map the complete sample management process, based on the stage timestamp and stage type, to a state sequence from the predefined state set, such as: sampling, transportation, storage awaiting inspection, testing, storage after inspection, and disposal. The method for constructing the observation sequence is as follows: For each state in the state sequence, the unstructured data recorded in its corresponding stage is processed by text segmentation and denoising, and the most core keywords are taken as the observation symbol. Simultaneously, the time interval or location type encoding corresponding to that stage can also be included as part of the observation symbol. For example, for the state "transportation," the corresponding observation symbol might be a combination of the keyword "transportation vehicle" and the time interval of 2 hours for that stage. The key parameters of the Hidden Markov Model, namely the state transition probability matrix and the observation probability matrix, need to be learned through historical sample full-process data. The learning process is as follows: collect a large amount of historical sample data, construct multiple historical state sequences and observation sequences according to the above method; count the frequency of transitions from one state to another, and calculate the frequency as an estimate of the state transition probability. For example, the probability of transitioning from state "sampling" to state "transportation" is estimated as the number of transitions divided by the total number of times the state appears in the sampling; count the frequency of a specific observation symbol appearing in a specific state, and calculate the frequency as an estimate of the observation probability. For example, the probability of the observation symbol being "transportation vehicle" in state "transportation" is estimated as the number of times that observation symbol appears divided by the total number of times it appears in state "transportation."

[0076] The process involves decoding the observed sequence using a Hidden Markov Model (HMM) to extract the latent sequence features of semantically ambiguous unstructured data. Specifically, for a piece of unstructured data identified as semantically ambiguous, it corresponds to one or more specific steps in the sample management process, but its textual description is unclear. First, based on complete, unambiguous data from other steps of the sample management process, and extracted spatiotemporal correlation features, an observed sequence is constructed that is partially known and partially missing or uncertain due to semantic ambiguity. The Viterbi algorithm, a dynamic programming-based algorithm, is then applied for decoding. Its purpose is to find the hidden state sequence most likely to have generated the observed sequence, given the HMM parameters and the observed sequence. The decoding process consists of two phases: recursion and backtracking. In the recursion phase, starting from the first observed symbol, the algorithm calculates the maximum probability value among all paths leading to each possible state at each time step and records the previous state that achieved that maximum probability. The calculation requires the state transition probabilities and observation probabilities learned in the previous step. The specific recursive calculation involves multiplying the maximum probability of reaching each state in the previous time step by the probability of transitioning to the current state, then multiplying by the probability of generating the current observation symbol in the current state, and finally taking the maximum value. In the backtracking phase, the algorithm starts from the state with the highest probability at the last time step and, based on the path information recorded in the recursive phase, backtracks to the initial time step to determine the most likely sequence of hidden states. This sequence of states is the parsed hidden sequence feature. For example, if an observation symbol in the detection phase is missing due to text ambiguity in an observation sequence, the Viterbi algorithm can infer, based on the preceding and following observation symbols and model parameters, through the above recursive and backtracking calculations, that the most likely state in the detection phase is "detecting".

[0077] Based on the parsed latent sequence features, semantic completion is performed on unstructured data with semantic ambiguity to generate semantically clarified results. Semantic completion processing supplements, replaces, or reorganizes semantically ambiguous unstructured data text based on the most probable state sequence obtained from decoding and the most frequently occurring observation symbols with clear semantics in that state. Specifically, a state-semantic template mapping library is established, storing a set of clearly expressed text templates or keywords corresponding to each hidden state. These templates and keywords are summarized from historical clear data. The summarization method involves collecting text samples of each state from all historical clear data, extracting their high-frequency keywords, or summarizing their common expression patterns as templates. When the most probable state corresponding to a certain ambiguous text is parsed, the standard semantic template or high-frequency keywords corresponding to that state are retrieved from the mapping library. Then, the ambiguous text is compared and fused with these standard semantic elements, retaining the reasonable parts of the ambiguous text and replacing contradictory, missing, or ambiguous parts with standard semantic elements, thereby generating a new, semantically clear text description, i.e., the semantically clarified result. For example, a semantically ambiguous disposal note might initially state "processed." Parsing using the steps described above reveals a hidden state sequence that strongly suggests a high-risk sample disposal process. The most common keywords in this state include "destruction," "record," and "supervisory personnel present." The semantic completion process combines these keywords with the original text to generate a semantically clarified result: "destructed in accordance with regulations in the high-risk sample disposal area; the disposal process was conducted with the presence of supervisory personnel and recorded."

[0078] S4. Establish semantic association rules based on semantic clarification results and structured data, and define the correspondence between semantic clarification results and structured data. The specific implementation is as follows:

[0079] The process involves extracting key semantic elements from the semantic clarification results. The semantic clarification results are complete and semantically clear text descriptions. Key semantic element extraction is achieved through natural language processing (NLP) techniques, specifically part-of-speech tagging and syntactic analysis of the semantic clarification result text to identify noun phrases and verb phrases as candidate elements. These candidate elements are then matched against a pre-built knowledge dictionary for the food testing domain. This dictionary lists important entity and behavior categories within the domain. Entity categories include, for example, the name of the testing project, the name of the disposal facility, and the name of the personnel role; behavior categories include, for example, disposal actions, testing actions, and recording actions. Successfully matched candidate elements are identified as key semantic elements. For example, for a semantic clarification result stating that samples were disposed of in a high-risk sample disposal area according to regulations, and that regulatory personnel were present and recording the process, part-of-speech tagging identifies the noun phrase "high-risk sample disposal area" and the verb phrase "disposal disposal." After matching with the knowledge dictionary, the key semantic elements extracted include "high-risk sample disposal area," "disposal disposal," and "regulatory personnel present and recording."

[0080] The process extracts data elements associated with key semantic elements from structured data of sample category, origin, and batch. The extraction process involves querying a pre-defined element association mapping table for each extracted key semantic element. This table indicates which categories of structured data fields different types of semantic elements are potentially associated with. For example, the mapping table might specify that semantic elements involving disposal actions or locations need to be associated with the sample category field, semantic elements involving risk descriptions need to be associated with the geographical information portion of the origin field, and semantic elements involving time descriptions need to be associated with the time information portion of the batch field. Based on the mapping table's instructions, corresponding information fragments are extracted as data elements from the specific values ​​of the structured data corresponding to the current sample. For example, if the structured data of the current sample is: sample category: pasteurized milk; origin: a dairy company's factory in Guizhou Province; batch: 20250517001, and the mapping table indicates the association with the category field for the key semantic element "destruction processing," then the data element "pasteurized milk" is extracted; and if the mapping table indicates the association with the origin field for the key semantic element "high-risk sample disposal area," then the geographical data element "Guizhou Province" is extracted through text parsing.

[0081] The analysis examines the logical correspondence between key semantic elements and data elements. The analysis process combines historical data statistics with domain rule reasoning. Historical data statistics involve calculating the co-occurrence frequency (CORF) of a specific key semantic element with a specific data element within a dataset containing numerous historical sample records. CORF can be calculated by counting the number of records in which the key semantic element appears alongside the data element across all historical records, and then dividing by the total number of records in which the data element appears. For example, if the data element "pasteurized milk" appears 1000 times in all historical records, and the key semantic element "destruction processing" appears 800 times, the CORF is 0.8. Domain rule reasoning applies known domain knowledge to determine logical relationships. For instance, a domain knowledge point is that yogurt products typically require cold chain transportation; if the source is far from the laboratory, temperature anomalies are likely to appear in the sampling site description. Through rule reasoning, a conditional relationship can be established between the data elements "yogurt products," "long distance," and the key semantic element "temperature anomalies." By combining statistical frequency and rule-based reasoning results, a relationship strength value is assigned to each pair of key semantic elements and data elements, and the relationship type is labeled, such as conditional relationship, causal relationship, or feature description relationship.

[0082] The process involves executing semantic association rules based on the analyzed logical correspondences, defining the correspondence between semantic clarification results and structured data. Specifically, rule formation transforms key semantic elements and data element association pairs with relationship strength values ​​exceeding a preset relationship strength threshold into explicit logical statement-based rules. The preset relationship strength threshold is set based on statistical analysis of the co-occurrence frequency of all key semantic element and data element association pairs in historical data. By observing the distribution of co-occurrence frequencies, a value that effectively distinguishes between strong and weak associations is selected as the dividing point; for example, the upper quartile of the co-occurrence frequency distribution, 0.75, is chosen as the preset relationship strength threshold. Relationship pairs with relationship strength values ​​not lower than this threshold are used to generate semantic association rules. Rules are expressed in an "if-then" form, where the "if" part consists of data elements and their value conditions, and the "then" part consists of key semantic elements. For example, based on the high co-occurrence frequency of the data element "pasteurized milk" and the key semantic element "destruction treatment," a rule is formed: if the sample category is pasteurized milk, then the semantic clarification result of its disposal process should include the key semantic element "destruction treatment." Each rule can also be appended with a confidence parameter, which directly uses the calculated relation strength value. Ultimately, all generated semantic association rules are collected and stored in a rule base. This rule base systematically defines the range of semantic features that the corresponding unstructured text should present after semantic clarification, given structured data such as sample category, source, and batch, thus completing the definition of the correspondence between the two.

[0083] S5. Conduct logical consistency verification and rule adaptation for sample risk levels across sample management stages for semantic association rules, and generate optimized semantic association rules. The specific implementation is as follows:

[0084] The system maps established semantic association rules to the sampling, storage, testing, and sample disposal stages throughout the entire sample management process. Specifically, it establishes a stage-rule mapping table, defining the management stages primarily involved in the preconditions and conclusions of each semantic association rule. For example, if a rule's "if" part includes the condition that the sample category is fresh meat, then the "if" part includes the key semantic element "cold chain transportation." Analysis shows that fresh meat category information is recorded during the sampling stage, while cold chain transportation requires passage through storage and transportation. Therefore, this rule is mapped to the sampling and storage stages. The mapping process is achieved by automatically parsing the "if" structure of the rules, extracting data elements and key semantic elements, and matching them with a predefined element-stage dictionary. This dictionary indicates that sample category elements are typically associated with sampling and all subsequent stages, testing item elements are associated with the testing stage, and disposal action elements are associated with the sample disposal stage. The system labels each rule with one or more applicable stage tags, thus completing the mapping.

[0085] The system verifies the consistency of the mapping logic of the same semantic association rule across different stages and identifies logical conflicts arising when different semantic association rules are applied across stages. The method for verifying logical consistency is to examine the semantic consistency and reasoning rationality of a rule mapped to multiple stages in different stage contexts. For example, a rule might be interpreted in the sampling stage as "If the sample is aquatic product, then the sampling site should have low-temperature preservation equipment," and in the storage stage as "If the sample is aquatic product, then the storage environment should be a cold storage facility." The system checks whether low-temperature preservation equipment and cold storage facilities conceptually constitute a consistent cold chain requirement. This verification is performed by querying a domain ontology that defines the hierarchy and relationships between concepts. For example, if low-temperature preservation equipment is a prerequisite or component concept of cold storage, then the system is considered consistent. The method for identifying logical conflicts is to construct a cross-stage rule dependency graph. When the conclusion output of rule A becomes the precondition input of rule B in a subsequent stage, the system checks whether the two are logically connected. Logical conflicts mainly manifest as contradictory conditions and contradictory inferences. Conditional contradictions refer to mutually exclusive conditions imposed on the same data element in rules at different stages. For example, a rule in the sampling stage requires close attention if the source location is location A, while another rule in the detection stage implicitly states that the risk is lower if the source location is location A. Inference contradictions refer to a situation where the final conclusion reached through a series of cross-stage rules contradicts the conclusion of an independent rule at a particular stage. The system automatically detects and reports such contradictions by simulating the transmission and reasoning of data within the rule chain.

[0086] This process involves analyzing and labeling the risk characteristics and risk levels implied by each semantic association rule based on the anomaly descriptions and handling notes associated with historical sample data. The specific workflow is as follows: First, based on the semantic association rule, all data records that meet some conditions of the rule are selected from the historical sample data. Then, the content of the anomaly description and handling notes fields in these records is extracted, and keyword extraction and statistical analysis are performed on these texts. Risk characteristics are labeled by identifying frequently occurring negative or warning words. For example, if keywords such as "exceeding the standard," "positive," and "non-compliant" are frequently extracted from the anomaly descriptions of sample records that meet a certain rule, then that rule is labeled as implying a risk characteristic of microbial exceeding the standard. The risk level is determined by combining statistical indicators and expert experience. The statistical indicator is the risk event incidence rate. The risk event incidence rate is calculated by counting the total number of historical records that trigger a certain rule, then counting the number of specific entries in which the handling notes include severe measures such as destruction or recall, dividing the latter by the former, and multiplying by 100% to obtain the percentage form of the risk event incidence rate. For example, if 30 out of 100 historical records triggering a certain rule have "destroy" or "recall" notes, the risk event occurrence rate is 30%. The method for obtaining the risk level classification threshold is to calculate the risk event occurrence rate for all historical rules, sort these values ​​from smallest to largest, and take the value at one-third of the total data volume as the threshold for low-risk and medium-risk, and the value at two-thirds of the total data volume as the threshold for medium-risk and high-risk. For example, if the 33rd quartile is 10% and the 66th quartile is 25%, then a risk event occurrence rate below 10% is considered low-risk, between 10% and 25% is medium-risk, and above 25% is high-risk. Therefore, if a rule has a risk event occurrence rate of 30%, which is above the 25% high-risk threshold, then that rule is marked as high-risk.

[0087] Based on the identification results of logical conflicts and the labeled risk characteristics and risk levels, the semantic association rules are logically reconstructed and weighted. Logical reconstruction addresses the identified logical conflicts, specifically redefining entities and relational predicates within the rules. For example, for two rules with contradictory conditions, the contradiction can be eliminated by introducing more refined entity categories or adding constraints such as time or batch. Assuming the conflict stems from differing risk assessments of source location A in the rules, the rule premise can be refined during reconstruction to specify that if the source location is A and the batch is from the first half of the year, it requires close attention, and if the source location is A and the batch is from the second half of the year, the risk is lower. For inference contradictions, it may be necessary to split an overly broad rule into multiple sub-rules applicable to different scenarios, or adjust the definition of intermediate conclusions in the rule chain to ensure consistency in logical transmission. Weighting involves assigning a confidence weight value to each semantic association rule. The specific method for setting the weight is to first assign a base weight based on the rule's risk level. This base weight is calculated using a preset linear function, whose input is the normalized risk event occurrence rate. For example, the base weight is set to 0.5 + 0.4 × normalized risk value, where the normalized risk value maps the occurrence rate of risk events for a rule to a range of 0 to 1. Low risk corresponds to a normalized risk value of 0 to 0.33, medium risk to 0.33 to 0.66, and high risk to 0.66 to 1.0. If a high-risk rule has a normalized risk value of 0.8, its base weight is 0.5 + 0.4 × 0.8 = 0.82. Next, rules with logical conflicts are penalized by multiplying the base weight by a conflict attenuation coefficient. This coefficient is set according to the severity of the conflict; for example, a coefficient of 0.9 for minor logical inconsistencies and 0.7 for serious contradictions. The final calibrated weight is then calculated.

[0088] The execution process outputs optimized semantic association rules after logical reconstruction and weight calibration. The output process involves re-serializing and storing all rules, after the aforementioned verification, annotation, reconstruction, and calibration, in a structured format. Each optimized semantic association rule includes a complete if premise, a conclusion, a set of labels for the mapped application stages, a description of the labeled risk characteristics and risk level labels, and calibrated weight values. These rules are integrated into a new rule base, replacing the original set of semantic association rules. This optimized semantic association rule base serves as the direct rule foundation for subsequent knowledge graph construction. It ensures the logical consistency of rules when applied across stages and reflects the relative importance differences of different rules in risk assessment and decision support through weights.

[0089] S6. Construct a knowledge graph of food testing sample information based on optimized semantic association rules, and perform intelligent retrieval and knowledge mining of food testing sample information. The specific implementation is as follows:

[0090] The execution of semantic association rules is based on the rules for constructing the knowledge graph, defining entities and relationships within it. The specific method involves parsing and optimizing each rule in the semantic association rule base. For the "if" part of each rule, the data elements used as judgment conditions are extracted, and the concept categories to which these data elements belong are defined as entity types in the knowledge graph. For example, from the rule "If the sample category is pasteurized milk and the origin includes Guizhou Province," the data elements "sample category," "pasteurized milk," "origin," and "Guizhou Province" are extracted, thus defining the entity types "sample category" and "geographical region." Pasteurized milk is defined as a specific entity instance under the "sample category" entity type, and Guizhou Province is defined as a specific entity instance under the "geographical region" entity type. For the "then" part of each rule, the key semantic elements used as conclusions are extracted, and these elements are defined as relational predicates or another entity type in the knowledge graph. For example, from the semantic clarification result of "then its disposal notes should include the key semantic element 'high risk,' 'high risk' is extracted and defined as a relational predicate with a risk level, used to connect the sample entity with an entity representing the concept of risk. Simultaneously, based on the risk level label assigned to each rule in the optimized semantic association rules, a weight attribute is attached to the defined relation predicates. The value of this weight attribute directly references the calibrated weight value recorded in the rule. Finally, by integrating the definitions extracted from all rules, a structured knowledge graph schema layer is formed. This schema layer, in the form of a list or configuration file, explicitly specifies the set of allowed entity types, the set of attributes each entity type can possess, the set of relation predicates, and the combinations of entity types that each relation predicate can connect, along with their allowed attributes such as weights.

[0091] The process involves transforming the semantically clarified results with structured data on sample category, origin, and batch, based on the correspondence defined by optimized semantic association rules, into entity nodes and relational edges in a knowledge graph to complete the construction. This transformation and construction is an automated data processing flow. The process begins for each complete sample record in the historical database. First, the specific values ​​of the sample category, origin, and batch fields in the record are read. Based on the knowledge graph schema layer determined in the previous step, corresponding entity nodes are created or located in the knowledge graph database. For example, if the category is pasteurized milk, an entity node of type "sample category" is created or located in the database, with the identifier "pasteurized milk"; if the origin is a dairy company's Guizhou province plant, an entity node of type "geographic region" (Guizhou province) is created or located through address resolution. Next, a core sample entity node is created for this record, for example, named "sample_batch20250517001," and its attributes are assigned: the category attribute points to the pasteurized milk entity node, the origin attribute points to the Guizhou province entity node, and the batch attribute is assigned the value 20250517001. Then, the semantic clarification result text corresponding to the record is read and pattern matched with the optimized semantic association rule base. The matching process involves segmenting the semantic clarification result text and extracting key information, comparing the similarity of the extracted fragments with the key semantic elements described by the "if" part of various rules in the rule base. When a fragment successfully matches the key semantic element of a rule, that rule is triggered. Based on the logical conditions defined in the "if" part of the triggering rule, after confirming that the conditions are met from the structured data of the current sample record, the entity and relation creation operation defined by the rule is executed. For example, if the semantic clarification result contains the text "detected total bacterial count exceeding the standard," and this fragment matches the key semantic element "total bacterial count exceeding the standard" of a rule, and the "if" part of the rule requires the sample category to be dairy products, since the current sample category, pasteurized milk, meets the condition, then according to the definition of the rule, an entity node of type "detected anomaly, total bacterial count exceeding the standard" is created in the knowledge graph, and a relation edge of type "detected" is created between the sample entity node "sample_batch20250517001" and this new entity node. This relation edge can carry a weight attribute, the value of which is taken from the calibrated weight value recorded in the optimized semantic association rule that triggered it. By traversing all historical sample records and repeating this process, a knowledge graph of food testing sample information containing a large number of interconnected entity nodes and relation edges is constructed.

[0092] The system receives retrieval requests for food testing sample information and parses the natural language queries in the retrieval requests into structured queries based on knowledge graph entities and relationships. The receiving function is implemented through a web application programming interface (API), where users send requests containing natural language questions via client software. The parsing process is handled by a natural language processing component. This component first performs word segmentation and part-of-speech tagging on the query statement, and then performs entity recognition. Entity recognition employs a hybrid approach combining dictionary-based and statistical model-based methods. The dictionary used is derived from all entity type names and known entity instance names defined in the knowledge graph's schema layer. For example, for the query statement "find pasteurized milk samples from Guizhou Province in 2025 that were found to have excessive total bacterial counts," dictionary matching identifies the entities "pasteurized milk" and "Guizhou Province," while the statistical model identifies the time entity "2025" and the testing item entity "excessive total bacterial counts." Based on entity recognition, dependency parsing is performed on the query statement to identify core verbs and the grammatical relationships between verbs and entities. For example, the core verbs "find" and "detect" are identified, and it is determined that "detect" connects the pasteurized milk sample to the entity with excessive total bacterial count. Then, these syntactic relations are mapped to predefined relational predicates in the knowledge graph; for example, "detect" is mapped to the relational predicate "detect". Finally, based on the identified entities, relations, and the logical structure obtained from syntactic analysis, a structured query statement that a graph database query language can understand is assembled. For example, the above query can be transformed into: Find entity nodes of type "sample", requiring that their category attribute points to the entity "pasteurized milk", their origin attribute points to the entity "Guizhou Province", their batch attribute year part equals 2025, and there exists a relational edge of type "detect" connecting to the entity with excessive total bacterial count.

[0093] The system executes structured queries within the constructed food testing sample information knowledge graph and performs knowledge mining by traversing and matching paths within the knowledge graph. The query execution process involves submitting the structured query statement generated in the previous step to the graph database management system. The graph database management system searches its stored graph data according to the query logic. The search process typically begins by matching the entity nodes specified in the query, then traverses along the specified relational edges, checking whether all connected nodes and edges meet all attribute filtering conditions set in the query, and finally returning a list of all subgraphs or entity nodes that meet the conditions as the direct query result. Knowledge mining is a deeper analysis based on this. Path traversal and matching are the core mining methods. The system allows users or programs to specify a starting entity, a relational path pattern, and a depth limit, and then automatically explores all possible paths in the graph that match the pattern. For example, starting from a production company entity that has repeatedly encountered problems, the system explores all sample entities, abnormal detection entities, and treatment measure entities connected to it within a 3-step relation. By collecting all these paths and performing statistical frequency analysis, hidden patterns can be discovered. For example, it might be discovered that although a company's samples are of different categories, over 60% of the paths ultimately point to the same detection anomaly, revealing a potential systemic risk for that company. Another approach is heuristic reasoning based on rule weights. During graph traversal and path search, the system can prioritize or filter relationship edges generated by high-weight optimized semantic association rules for exploration. This results in association patterns with higher confidence, focusing on procedurally validated high-risk logical connections, thereby outputting more valuable risk characteristic rules or management insights.

[0094] Example 2: Figure 2 A schematic diagram of the food testing sample management system based on big data analysis of the present invention is provided. The food testing sample management system based on big data analysis includes:

[0095] The data acquisition module is used to acquire data from the entire process of food testing samples, including structured and unstructured data.

[0096] The fuzzy judgment module is used to determine whether unstructured data has semantic ambiguity using a text similarity algorithm;

[0097] The clarification generation module is used to perform sequence feature parsing and semantic completion processing on unstructured data with semantic ambiguity from the perspective of spatiotemporal correlation features of the entire sample process data, and generate semantic clarification results when semantic ambiguity exists.

[0098] The rule-building module is used to build semantic association rules based on semantic clarification results and structured data, and to define the correspondence between semantic clarification results and structured data;

[0099] The rule optimization module is used to perform logical self-consistency verification of semantic association rules across sample management stages and rule adaptation for sample risk levels, and to generate optimized semantic association rules.

[0100] The retrieval and mining module is used to construct a knowledge graph of food testing sample information based on optimized semantic association rules, and to perform intelligent retrieval and knowledge mining of food testing sample information.

[0101] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.

[0102] It should be noted that this invention can be deployed on the device itself to realize embedded applications, or it can run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.

[0103] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wireless or wired transmission; wired transmission methods include optical fiber, twisted pair, coaxial cable, etc.; wireless transmission includes infrared, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center containing one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0104] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0105] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0106] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0107] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0108] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0109] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0110] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A food testing sample management method based on big data analysis, characterized in that, include: S1. Obtain full-process data for food testing samples, including structured and unstructured data; S2. Use a text similarity algorithm to determine whether there is semantic ambiguity in unstructured data; S3. When semantic ambiguity exists, a Hidden Markov Model is used from the perspective of the spatiotemporal correlation characteristics of the entire sample data to perform sequence feature parsing and semantic completion processing on the unstructured data with semantic ambiguity, generating semantically clarified results, including: Extract the spatiotemporal correlation features of samples from the entire process data of samples corresponding to unstructured data with semantic ambiguity; Based on the extracted spatiotemporal correlation features and unstructured data with semantic ambiguity, a hidden Markov model is constructed, including state sequences and observation sequences. By decoding the observed sequence using a hidden Markov model, the implicit sequence features of unstructured data with semantic ambiguity can be extracted. Based on the parsed implicit sequence features, semantic completion is performed on unstructured data with semantic ambiguity to generate semantically clarified results; S4. Establish semantic association rules based on semantic clarification results and structured data, and define the correspondence between semantic clarification results and structured data; S5. Conduct logical consistency verification and sample risk level rule adaptation for semantic association rules across sample management stages, and generate optimized semantic association rules, including: The established semantic association rules are mapped to the sampling, storage, testing and sample retention disposal stages of the entire sample management process; Verify the consistency of the mapping logic of the same semantic association rule in different stages, and identify the logical conflicts that arise when different semantic association rules are applied across stages; Based on the descriptions and handling notes of detection anomalies associated with semantic association rules in historical sample data, analyze and label the risk characteristics and risk levels implied by each semantic association rule; Based on the identification results of logical conflicts and the labeled risk characteristics and risk levels, the semantic association rules are logically reconstructed and weighted. The semantic association rules that have undergone logical reconstruction and weight calibration are output as optimized semantic association rules; S6. Construct a knowledge graph of food testing sample information based on optimized semantic association rules, and perform intelligent retrieval and knowledge mining of food testing sample information.

2. The food testing sample management method based on big data analysis according to claim 1, characterized in that, S1 includes: Obtain the complete process data of the samples that have been entered from the food testing sample management big data platform; The obtained sample process data confirms that it includes structured data such as sample type, source, and batch, as well as unstructured data including sampling site description, test anomaly explanation, and handling notes.

3. The food testing sample management method based on big data analysis according to claim 1, characterized in that, S2 include: Unstructured data, including sampling site descriptions, abnormal detection explanations, and handling notes, obtained from the entire sample process data, are processed through text segmentation and noise reduction. Convert unstructured data that has undergone text segmentation and denoising into text feature representations; Calculate the similarity values ​​between text feature representations; Based on the comparison results between the similarity value and the preset similarity threshold, it is determined whether there is semantic ambiguity in the unstructured data.

4. The food testing sample management method based on big data analysis according to claim 1, characterized in that, Decoding the observed sequence using a Hidden Markov Model (HMM) involves: applying the Viterbi algorithm, based on the HMM parameters and the observed sequence, recursively calculating the most likely hidden state sequence path through dynamic programming, and backtracking to obtain the complete hidden state sequence as the parsed hidden sequence features.

5. The food testing sample management method based on big data analysis according to claim 1, characterized in that, S4 include: Extract key semantic elements from the semantic clarification results; Extract data elements associated with key semantic elements from structured data of sample category, origin, and batch; Analyze the logical correspondence between key semantic elements and data elements; Based on the analyzed logical correspondences, semantic association rules are formed to define the correspondence between semantic clarification results and structured data.

6. The food testing sample management method based on big data analysis according to claim 1, characterized in that, Logical reconstruction and weight calibration of semantic association rules include: redefining or combining entities and relational predicates in semantic association rules with logical conflicts to complete logical reconstruction, and assigning differentiated confidence weights to their logical relations based on the risk characteristics and risk levels marked by the semantic association rules to complete weight calibration.

7. The food testing sample management method based on big data analysis according to claim 1, characterized in that, S6 include: Using optimized semantic association rules as the rules for graph construction, the definitions of entities and relations in the knowledge graph are determined; The semantic clarification results are then converted into entity nodes and relation edges in a knowledge graph based on the correspondence defined by the optimized semantic association rules, according to the structured data of sample category, source, and batch. Receive retrieval requests for food testing sample information and parse the natural language queries in the retrieval requests into structured queries based on knowledge graph entities and relationships; Structured queries are performed in the completed food testing sample information knowledge graph, and knowledge mining is carried out by traversing and matching paths in the knowledge graph.

8. A food testing sample management system based on big data analysis, used to implement the food testing sample management method based on big data analysis as described in any one of claims 1-7, characterized in that, include: The data acquisition module is used to acquire data from the entire process of food testing samples, including structured and unstructured data. The fuzzy judgment module is used to determine whether unstructured data has semantic ambiguity using a text similarity algorithm; The clarification generation module is used to perform sequence feature parsing and semantic completion processing on unstructured data with semantic ambiguity from the perspective of spatiotemporal correlation features of the entire sample process data, and generate semantic clarification results when semantic ambiguity exists. The rule-building module is used to build semantic association rules based on semantic clarification results and structured data, and to define the correspondence between semantic clarification results and structured data; The rule optimization module is used to perform logical self-consistency verification of semantic association rules across sample management stages and rule adaptation for sample risk levels, and to generate optimized semantic association rules. The retrieval and mining module is used to construct a knowledge graph of food testing sample information based on optimized semantic association rules, and to perform intelligent retrieval and knowledge mining of food testing sample information.

Citation Information

Patent Citations

  • Power grid fault intelligent analysis and disposal method and system based on knowledge graph

    CN117992743A

  • Intelligent exploratory data mining system

    CN120542437A