Food biological pollution source identification model training method, identification method and device

By building a circulation network graph model and graph convolution neural network, the problems of low reliability and high computing difficulty in identifying pollution sources in food safety are solved, and more efficient and accurate pollution sources are achieved.

CN119941268AActive Publication Date: 2025-05-06BEIJING RES CENT FOR INFORMATION TECH & AGRI
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411868728.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

The prior art has low reliability in the identification of food safety pollution sources, the identification results based on effective distance are inaccurate, and the calculation of methods based on petri network and bass inference is difficult, resulting in low recognition efficiency.

Method used

By constructing a circulation network graph model based on the regional information and transaction information of the target agricultural products, multiple food risk propagation simulations were carried out, training data sets were generated, and the graph convolution neural network was used to iteratively train the propagation result snapshot data to obtain the food biological pollution source identification model.

Benefits of technology

It improves the efficiency and accuracy of identifying food safety pollution sources, and reduces the difficulty of computing and the complexity of data observation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941268A_ABST
    Figure CN119941268A_ABST
Patent Text Reader

Abstract

The invention provides a food biological pollution source identification model training method, an identification method and a device. The food biological pollution source identification model training method comprises the following steps: constructing a circulation network graph model based on regional information and transaction information of a target agricultural product; initializing a circulation network graph model, and performing multiple food risk propagation simulation according to the initialized circulation network graph model to obtain a training data set; wherein each group of sample data in the training data set comprises target graph nodes and propagation result snapshot data; the target graph node is at least one of different graph nodes; and for each group of training data in the training data set, taking the propagation result snapshot data as input and taking the graph node corresponding to the propagation result snapshot data as output to carry out iterative training on the graph convolutional neural network, and obtaining a food biological pollution source identification model under the condition that the graph convolutional neural network converges. According to the method, the food safety pollution source identification efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of food safety detection, and in particular to a food biological contamination source identification model training method, an identification method and a device. Background Art

[0002] The high incidence of food safety incidents not only endangers people's health, but also has a serious impact on economic development and social stability. When a food safety incident occurs, if the source of pollution can be accurately identified, the transmission chain can be quickly blocked, the loss of the incident can be minimized, and the scope of risk transmission can be estimated, which is of great significance to food safety supervision.

[0003] In the related technology, the effective distance-based method is usually used to convert the geographical distance between nodes into the effective distance. However, the strong assumption that the pollution source has the centrality characteristic is identified by evaluating the centrality of each node one by one, which affects the reliability of the identification result. The pollution source identification method based on the Petri net uses the data structure of the Petri net to accurately record the circulation batches of each link in the agricultural product supply chain. When the circulation batches are large, it is inefficient to accurately record the supply chain circulation batch information one by one. In addition, the source identification method based on Bayesian inference is achieved by calculating the maximum a posteriori probability. As the number of pollution sources and the amount of observation data increase, the amount of calculation will increase exponentially, the calculation difficulty will increase, and at the same time, the efficiency of food pollution source identification will be low. Summary of the invention

[0004] The present invention provides a food biological contamination source identification model training method, identification method and device, which are used to solve the problems that the prior art sampling method based on effective distance in identifying food safety contamination sources has low reliability, and when using the pollution source identification method based on Petri net and Bayesian inference, the data observation and calculation are difficult, resulting in low efficiency in identifying food safety contamination sources, thereby improving the efficiency and accuracy of food safety contamination source identification.

[0005] The present invention provides a method for training a food biological contamination source identification model, comprising: A circulation network graph model is constructed based on the regional information and transaction information of the target agricultural product, wherein the circulation network graph model includes a plurality of graph nodes, different graph nodes correspond to different regions, and edges between different graph nodes correspond to food circulation information between different regions; the weight of the edge is determined based on at least one of the food transaction quantity, transaction amount, geographical distance and population between the two corresponding regions; Initializing the circulation network graph model, and performing multiple food risk propagation simulations according to the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes; For each set of training data in the training data set, the graph convolutional neural network is iteratively trained with the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output, and when the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

[0006] According to a food biological contamination source identification model training method provided by the present invention, the graph convolutional neural network includes a stack of graph convolutional layers and fully connected layers; wherein the graph convolutional layer calculates neighbor node information from the propagation result snapshot data by using a first-order Chebyshev polynomial, and the graph convolutional layer uses L2 regularization and dropout to optimize network parameters, and uses a ReLU activation function; the fully connected layer is used to convert the matrix output by the graph convolutional layer into a vector; the loss function of the graph convolutional neural network is determined based on the sigmoid activation function and the cross entropy loss.

[0007] According to a method for training a food biological contamination source identification model provided by the present invention, the food risk propagation simulation is performed multiple times according to the initialized circulation network graph model to obtain a training data set including: Based on the target propagation diffusion model, the food risk propagation path information is simulated according to the initialized circulation network graph model to obtain a food risk propagation chain; the food risk propagation chain includes target graph nodes and corresponding edges; Among them, the target propagation diffusion model includes one of the SI infectious disease model, the independent cascade model, the SIR model and the linear threshold model.

[0008] According to a method for training a food biological contamination source identification model provided by the present invention, after obtaining the food biological contamination source identification model, the method further comprises: Processing the test sample based on the food biological contamination source identification model to obtain a test result; Based on the test results, a target evaluation index is obtained, and the food biological contamination source identification model is fine-tuned and optimized according to the target evaluation index to obtain an optimized food biological contamination source identification model; the target evaluation index includes at least one of the F score, precision and recall rate.

[0009] The present invention also provides an identification method, comprising: Acquire information on pollution sources to be detected, wherein the information on pollution sources to be detected includes the occurrence area and type of food safety pollution sources; The information of the pollution source to be detected is processed based on the food biological pollution source identification model to obtain a food safety pollution source identification result; wherein the food biological pollution source identification model is trained based on the food biological pollution source identification model training method.

[0010] The present invention also provides a food biological contamination source identification model training device, comprising: A graph model construction module, for constructing a circulation network graph model based on the regional information and transaction information of the target agricultural product, wherein the circulation network graph model includes a plurality of graph nodes, wherein different graph nodes correspond to different regions, and edges between different graph nodes correspond to food circulation information between different regions; and the weight of the edge is determined based on at least one of the number of food transactions, transaction amount, geographical distance, and population between the two corresponding regions; A sample acquisition module, used to initialize the circulation network graph model, and perform multiple food risk propagation simulations according to the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes; The training module is used to iteratively train the graph convolutional neural network for each set of training data in the training data set, taking the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output, and obtain a food biological contamination source identification model when the graph convolutional neural network converges.

[0011] The present invention also provides an identification device, comprising: A data acquisition module, used to acquire information about pollution sources to be detected, wherein the information about pollution sources to be detected includes the occurrence area and type of food safety pollution sources; The identification module is used to process the information of the pollution source to be detected based on the food biological pollution source identification model to obtain the food safety pollution source identification result; wherein, the food biological pollution source identification model is trained based on the food biological pollution source identification model training method.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements any of the above-mentioned food biological contamination source identification model training methods or identification methods.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned food biological contamination source identification model training methods or identification methods.

[0014] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned food biological contamination source identification model training methods or identification methods.

[0015] The food biological contamination source identification model training method, identification method and device provided by the present invention construct a circulation network graph model through the geographical information and transaction information of the target agricultural product, and then perform multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set. Finally, the graph convolutional neural network is iteratively trained with the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output to obtain a food biological contamination source identification model, thereby improving the efficiency and accuracy of food safety contamination source identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 It is a flow chart of the food biological contamination source identification model training method provided by the present invention.

[0018] Figure 2 This is one of the flow charts of the identification method provided by the present invention.

[0019] Figure 3 This is the second flow chart of the identification method provided by the present invention.

[0020] Figure 4 It is a structural schematic diagram of a food biological contamination source identification model training device provided by the present invention.

[0021] Figure 5 It is a structural schematic diagram of the identification device provided by the present invention.

[0022] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] Combine the following Figure 1-Figure 5 The present invention describes the food biological contamination source identification model training method, identification method and device. Figure 1 FIG. 1 is a flow chart of a method for training a food biological contamination source identification model provided by the present invention. Figure 1 As shown, the food biological contamination source identification model training method includes the following steps: Step 110: construct a circulation network graph model based on the geographical information and transaction information of the target agricultural products. The circulation network graph model includes multiple graph nodes, different graph nodes correspond to different regions, and the edges between different graph nodes correspond to the food circulation information between different regions; the weight of the edge is determined based on at least one of the food transaction quantity, transaction amount, geographical distance and population size between the two corresponding regions.

[0025] In this step, the target agricultural products include, but are not limited to, one or more of grain and oil crops, vegetables, fruits, livestock products and aquatic products.

[0026] Different regional divisions of target agricultural products include different administrative regions in the same country or region, or regional divisions according to different countries or regions around the world.

[0027] In this step, the division of graph nodes can be determined based on the identification accuracy requirements and data acquisition capabilities of food contamination sources; for example, graph nodes can be economically developed areas in the current region, such as the administrative center of a province or an area with a large agricultural product trade volume.

[0028] In this step, the transaction information may be the transaction quantity or transaction amount of agricultural products in the corresponding areas of the two nodes.

[0029] For example, in the circulation network graph model, each graph node is a provincial capital city, and the edges between the graph nodes indicate that there is a food transaction relationship between the corresponding provincial capital cities. The weight of the edge includes the number of food transactions or the transaction amount.

[0030] In this embodiment, the weights between the graph nodes are based on the transaction quantity and average transaction amount of the areas corresponding to the two graph nodes over a period of time, the geographical distance, and the population correlation between the two points.

[0031] Specifically, there are two types of factors that affect the weight of the edge: social factors f and trade factors t ; From the graph node m To graph node n The edge between is a directed edge, and the weight of the directed edge is calculated by the following steps: Step 1: Calculate the absolute value of the impact of social factors according to the following formula: ; in, express m population, express m and n The distance between is the average distance; 0.5.

[0032] Step 2: Convert the absolute value of social factors into a value between 0 and 1 using the following formula: .

[0033] Step 3: Absolute value of trade factors , which represents the total trade volume of this type of agricultural products from node m to node n over a period of time.

[0034] Step 4: Convert the absolute value of the trade factor to between 0 and 1 using the following formula: .

[0035] Step 5: Merge by the following formula and : 1 + 2 ; in, for m and n The edges between are directed edges. 1 and 2 is the weight coefficient, which is less than 1 and can be set according to user needs.

[0036] In this embodiment, each node of the circulation network graph model can also be a sub-region at each level of the same region. For example, there are M provinces in country A, and there are multiple cities in different provinces, and each city corresponds to a graph node. It should be noted that there can be an edge and its edge value between any two nodes (indicating that there are agricultural product transactions between the two nodes within a certain time range), or there can be no edge (there are no agricultural product transactions between the two nodes within a certain time range).

[0037] Step 120, initialize the circulation network graph model, and perform multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes.

[0038] In this step, during the food risk transmission process, food contamination can be caused by biological contamination, such as microbial contamination such as bacteria, viruses and parasites.

[0039] In this step, food safety risk propagation simulation can be performed through a target propagation diffusion model, wherein the target propagation diffusion model includes but is not limited to an SI infectious disease model, an independent cascade model, an SIR model, and a linear threshold model.

[0040] In this embodiment, the circulation network graph model is initialized through the following steps: (1) Create an empty directed graph using graphing software; (2) A specified number of nodes are added to the directed graph through a loop, and some edges are added to the target graph nodes according to the accuracy requirements of pollution source identification and data acquisition capabilities; each graph node can also be connected to its k randomly selected neighbor nodes.

[0041] (3) Initialize the infection state of each node to -1 (uninfected) and set the infection state of the target graph node to 1 (infected).

[0042] In this embodiment, multiple food risk transmission simulations are performed according to the initialized circulation network graph model to obtain a training data set including: simulating the food risk transmission path information according to the initialized circulation network graph model based on the target propagation diffusion model to obtain a food risk transmission chain; the food risk transmission chain includes target graph nodes and corresponding edges; wherein the target propagation diffusion model includes one of the SI infectious disease model, the independent cascade model, the SIR model and the linear threshold model.

[0043] In this embodiment, after the circulation network graph model is initialized, the infected nodes in the circulation network graph model propagate the risk to the susceptible nodes connected thereto with a certain probability through the target propagation diffusion model.

[0044] Specifically, a propagation probability β is set to represent the probability that an infected node will spread the risk to adjacent susceptible nodes in each time step; a number of simulation time steps (or iterations) is set; in each time step, all infected nodes are traversed according to the directed edges between nodes from one node to another, and attempts are made to spread the risk to their adjacent susceptible nodes with a propagation probability β; if a susceptible node is successfully infected, its state is updated (the feature update of a graph node depends on the features of its neighboring nodes and the direction of the edges), and the node state after each time step is recorded to generate snapshot data of risk propagation; after the simulation, the risk propagation results of a specific source node are output, including the final state of each node (uninfected or infected); by repeating the above three steps several times, several sets of training data can be obtained, which can constitute the training data set of the pollution source identification model.

[0045] Step 130: For each set of training data in the training data set, the graph convolutional neural network is iteratively trained with the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output, and when the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

[0046] In this step, the graph convolutional neural network consists of a stack of graph convolutional layers and fully connected layers; the graph convolutional layer calculates the neighbor node information from the propagation result snapshot data by using the first-order Chebyshev polynomial, and the graph convolutional layer uses L2 regularization and dropout to optimize the network parameters, and uses the ReLU activation function; the fully connected layer is used to convert the matrix output by the graph convolutional layer into a vector; the loss function of the graph convolutional neural network is determined based on the sigmoid activation function and the cross entropy loss.

[0047] In this embodiment, the graph convolutional neural network is trained through the following steps to obtain a food biological contamination source identification model: (1) Perform preprocessing operations such as data cleaning, normalization, and data enhancement on each sample in the training data set to obtain the preprocessed training data set.

[0048] (2) The propagation result snapshot data extracted from the preprocessed training dataset. The snapshot data should contain features such as the node status (uninfected -1 or infected 1), node type, transaction volume, etc. after each time step.

[0049] (3) In each snapshot, the infected nodes are marked as contamination sources (positive samples) and the susceptible nodes are marked as non-contamination sources (negative samples) to determine the sample labels.

[0050] (4) In this embodiment, a graph theory library (such as NetworkX) can be used to construct a food circulation network diagram, where nodes represent foods or circulation links, edges represent the connection relationship of food flow, and feature vectors are added to each node, such as features extracted from snapshot data.

[0051] (5) Divide the preprocessed training dataset into a training set and a test set, ensuring that each set contains a sufficient number of infected nodes and sufficient snapshot time steps; when dividing, temporal continuity can be considered to avoid data leakage between the training set and the test set.

[0052] (6) Select a suitable GCN architecture (such as GCN based on spectral graph convolution or GCN based on spatial graph convolution), set the input layer and hidden layer of GCN according to the number and type of input features, and set a binary classifier (such as sigmoid function) for each node in the output layer to predict whether the node is a pollution source.

[0053] (7) Use the snapshot data and labels of the training set to train the GCN model. In each training iteration, the prediction results are calculated through forward propagation, and the model parameters are updated through back propagation. The cross entropy loss function is used to evaluate the difference between the prediction results and the true labels, and optimization (such as Adam) is used to minimize the loss. Finally, after reaching the maximum number of training times, the food biological contamination source identification model is obtained.

[0054] The food biological contamination source identification model training method of the embodiment of the present invention constructs a circulation network graph model through the geographical information and transaction information of the target agricultural product, and then performs multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set. Finally, the propagation result snapshot data is used as input and the graph nodes corresponding to the propagation result snapshot data are used as output to iteratively train the graph convolutional neural network to obtain a food biological contamination source identification model, thereby improving the efficiency and accuracy of food safety contamination source identification.

[0055] In some embodiments, after obtaining the food biological contamination source identification model, the food biological contamination source identification model training method also includes: processing the test sample based on the food biological contamination source identification model to obtain the test result; obtaining the target evaluation index based on the test result, and fine-tuning and optimizing the food biological contamination source identification model according to the target evaluation index to obtain the optimized food biological contamination source identification model; the target evaluation index includes at least one of the F score, precision and recall rate.

[0056] In this embodiment, the performance of the food biological contamination source identification model is evaluated on a test set, and indicators such as accuracy, recall, and F1 score are used to evaluate the model's ability to identify contamination sources.

[0057] In this embodiment, representative food samples are first collected from the actual environment to update the circulation network graph model to obtain test samples, and the test samples are pre-processed such as extraction, purification, and concentration to facilitate subsequent pollution source detection. The pre-processed test samples are then input into the established food biological pollution source identification model to perform pollution source detection to obtain test results. Based on the test results, the F score of the model in identifying pollution sources (used to comprehensively measure the performance of the model) is calculated using the following formula: F-score = 2 × (precision × recall) / (precision + recall).

[0058] In this embodiment, the food biological contamination source identification model is comprehensively evaluated according to evaluation indicators such as F-score, precision and recall rate; if the evaluation result is not ideal, the model needs to be fine-tuned and optimized, for example, adjusting the model structure (including increasing or decreasing the number of layers of the model, changing the number of neurons or the connection method, etc.), optimizing parameter settings (including adjusting the values ​​of hyperparameters such as learning rate and regularization coefficient to improve the performance of the model) or increasing training data, etc.; finally, the model is re-trained using the adjusted model structure and optimized parameter settings to obtain an optimized food biological contamination source identification model.

[0059] The food biological contamination source identification model training method of the embodiment of the present invention processes the test sample through the food biological contamination source identification model to obtain the test result, obtains the target evaluation index based on the test result, and fine-tunes and optimizes the food biological contamination source identification model according to the target evaluation index to obtain the optimized food biological contamination source identification model, thereby improving the robustness and generalization ability of the food biological contamination source identification model.

[0060] The identification method provided by the present invention is described below. The identification method described below and the food biological contamination source identification model training method described above can be referenced to each other.

[0061] Figure 2 It is one of the flow charts of the identification method provided by the present invention, such as Figure 2 As shown, the identification method comprises the following steps: Step 210: Obtain information on pollution sources to be detected, where the information on pollution sources to be detected includes the occurrence area and type of food safety pollution sources; In this step, the information of the pollution source to be detected includes the current status of any given food safety incident within a certain range, that is, which nodes in the latest circulation network graph model are in an infected state. These nodes can be used as the information of the pollution source to be detected.

[0062] Step 220: Process the information of the pollution source to be detected based on the food biological pollution source identification model to obtain a food safety pollution source identification result; wherein the food biological pollution source identification model is trained based on the food biological pollution source identification model training method.

[0063] In this step, the food biological contamination source identification model is trained through the following steps: (1) A circulation network graph model is constructed based on the regional information and transaction information of the target agricultural products. The circulation network graph model includes multiple graph nodes. Different graph nodes correspond to different regions. The edges between different graph nodes correspond to the food circulation information between different regions. Different values ​​of the edges correspond to the number of food transactions between different regions.

[0064] (2) Initializing the circulation network graph model, and performing multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; and the target graph node is at least one of the different graph nodes.

[0065] (3) For each set of training data in the training data set, the graph convolutional neural network is iteratively trained with the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output. When the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

[0066] In this embodiment, the specific implementation methods of the above steps are as described in the corresponding embodiments of the above steps 110 to 130, and will not be repeated in this embodiment.

[0067] In this embodiment, the occurrence status of food safety incidents within a certain range is input into the trained food biological contamination source identification model at the graph node corresponding to the circulation network graph model, and the contamination source node (including the real geographical location of the area corresponding to the graph node, such as longitude and latitude and administrative division identification, etc.) is output, and the calculation result of this contamination source node is used to guide the food safety regulatory department to conduct on-site verification and prevention and control.

[0068] The identification method provided in the embodiment of the present invention processes the information of the pollution source to be detected through the food biological pollution source identification model to obtain the food safety pollution source identification result. It utilizes the feature learning ability of the graph neural network for graph structure data and combines the characteristics of the food circulation network to realize the identification of the source of food safety pollution.

[0069] Figure 3 This is the second flow chart of the identification method provided by the present invention. Figure 3In the illustrated embodiment, an identification method is implemented by the following steps: constructing a food circulation network graph model; constructing a training data set; constructing a pollution source identification model; and identifying pollution sources.

[0070] The food biological contamination source identification model training device provided by the present invention is described below. The food biological contamination source identification model training device described below and the food biological contamination source identification model training method described above can be referenced to each other.

[0071] Figure 4 Schematic diagram of the structure of the food biological contamination source identification model training device provided by the present invention. Figure 4 As shown, the food biological contamination source identification model training device includes: a graph model construction module 410, a sample acquisition module 420 and a training module 430.

[0072] A graph model building module 410 is used to build a circulation network graph model based on the regional information and transaction information of the target agricultural product, wherein the circulation network graph model includes a plurality of graph nodes, wherein different graph nodes correspond to different regions, and edges between different graph nodes correspond to food circulation information between different regions; and the weight of the edge is determined based on at least one of the food transaction quantity, transaction amount, geographical distance, and population between the two corresponding regions; The sample acquisition module 420 is used to initialize the circulation network graph model, and perform multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes; The training module 430 is used to iteratively train the graph convolutional neural network for each set of training data in the training data set, using the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output, and obtain a food biological contamination source identification model when the graph convolutional neural network converges.

[0073] The food biological contamination source identification model training device of the embodiment of the present invention constructs a circulation network graph model through the geographical information and transaction information of the target agricultural product, and then performs multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set. Finally, the propagation result snapshot data is used as input and the graph nodes corresponding to the propagation result snapshot data are used as output to iteratively train the graph convolutional neural network to obtain a food biological contamination source identification model, thereby improving the efficiency and accuracy of food safety contamination source identification.

[0074] The identification device provided by the present invention is described below. The identification device described below and the identification method described above can be referenced to each other.

[0075] Figure 5 is a schematic diagram of the structure of the identification device provided by the present invention, such as Figure 5 As shown, the identification device includes: a data acquisition module 510 and an identification module 520.

[0076] The data acquisition module 510 is used to acquire information about the pollution source to be detected, and the information about the pollution source to be detected includes the occurrence area and type of the food safety pollution source; The identification module 520 is used to process the information of the pollution source to be detected based on the food biological pollution source identification model to obtain the food safety pollution source identification result; wherein the food biological pollution source identification model is trained based on the food biological pollution source identification model training method.

[0077] The identification device provided in the embodiment of the present invention processes the information of the pollution source to be detected through the food biological pollution source identification model to obtain the food safety pollution source identification result. It utilizes the feature learning ability of the graph neural network for graph structure data and combines the characteristics of the food circulation network to realize the identification of the source of food safety pollution.

[0078] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6 As shown, the electronic device may include: a processor (processor) 610 , a communication interface (Communications Interface) 620 , a memory (memory) 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other through the communication bus 640 . The processor 610 can call the logic instructions in the memory 630 to execute the food biological contamination source identification model training method, which includes: constructing a circulation network graph model based on the geographical information and transaction information of the target agricultural product, the circulation network graph model includes multiple graph nodes, different graph nodes correspond to different regions, and the edges between different graph nodes correspond to the food circulation information between different regions; the weight of the edge is determined based on at least one of the number of food transactions, transaction amount, geographical distance and population between the two corresponding regions; initializing the circulation network graph model, and performing multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set; wherein each group of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes; for each group of training data in the training data set, the graph convolutional neural network is iteratively trained with the propagation result snapshot data as input and the graph node corresponding to the propagation result snapshot data as output, and when the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

[0079] Or execute an identification method, the method comprising: obtaining information on pollution sources to be detected, the information on pollution sources to be detected including the occurrence area and type of food safety pollution sources; processing the information on pollution sources to be detected based on a food biological pollution source identification model to obtain a food safety pollution source identification result; wherein the food biological pollution source identification model is trained based on a food biological pollution source identification model training method.

[0080] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0081] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the food biological contamination source identification model training method provided by the above methods, the method including: constructing a circulation network graph model based on the geographical information and transaction information of the target agricultural product, the circulation network graph model including multiple graph nodes, different graph nodes corresponding to different regions, and the edges between different graph nodes corresponding to the food circulation information between different regions; the weight of the edge is determined based on at least one of the number of food transactions, transaction amount, geographical distance and population between the two corresponding regions; initializing the circulation network graph model, and performing multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set; wherein each group of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of different graph nodes; for each group of training data in the training data set, taking the propagation result snapshot data as input and the graph node corresponding to the propagation result snapshot data as output, the graph convolutional neural network is iteratively trained, and when the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

[0082] Or execute an identification method, the method comprising: obtaining information on pollution sources to be detected, the information on pollution sources to be detected including the occurrence area and type of food safety pollution sources; processing the information on pollution sources to be detected based on a food biological pollution source identification model to obtain a food safety pollution source identification result; wherein the food biological pollution source identification model is trained based on a food biological pollution source identification model training method.

[0083] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the food biological contamination source identification model training method provided by the above-mentioned methods, the method comprising: constructing a circulation network graph model based on the geographical information and transaction information of the target agricultural product, the circulation network graph model comprising a plurality of graph nodes, different graph nodes corresponding to different regions, and edges between different graph nodes corresponding to food circulation information between different regions; the weight of the edge is determined based on at least one of the number of food transactions, transaction amount, geographical distance and population between the two corresponding regions; initializing the circulation network graph model, and performing multiple food risk propagation simulations based on the initialized circulation network graph model to obtain a training data set; wherein each group of sample data in the training data set comprises a target graph node and propagation result snapshot data; the target graph node is at least one of different graph nodes; for each group of training data in the training data set, taking the propagation result snapshot data as input and taking the graph node corresponding to the propagation result snapshot data as output, the graph convolutional neural network is iteratively trained, and when the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

[0084] Or execute an identification method, the method comprising: obtaining information on pollution sources to be detected, the information on pollution sources to be detected including the occurrence area and type of food safety pollution sources; processing the information on pollution sources to be detected based on a food biological pollution source identification model to obtain a food safety pollution source identification result; wherein the food biological pollution source identification model is trained based on a food biological pollution source identification model training method.

[0085] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0086] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for training a food biological contamination source identification model, characterized in that: include: A circulation network graph model is constructed based on the regional information and transaction information of the target agricultural product, wherein the circulation network graph model includes a plurality of graph nodes, different graph nodes correspond to different regions, and edges between different graph nodes correspond to food circulation information between different regions; the weight of the edge is determined based on at least one of the food transaction quantity, transaction amount, geographical distance and population between the two corresponding regions; Initializing the circulation network graph model, and performing multiple food risk propagation simulations according to the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes; For each set of training data in the training data set, the graph convolutional neural network is iteratively trained with the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output, and when the graph convolutional neural network converges, a food biological contamination source identification model is obtained.

2. The method for training a food biological contamination source identification model according to claim 1, characterized in that: The graph convolutional neural network includes a stack of graph convolutional layers and fully connected layers; wherein the graph convolutional layer calculates neighbor node information from the propagation result snapshot data by using a first-order Chebyshev polynomial, and the graph convolutional layer uses L2 regularization and dropout to optimize network parameters, and uses a ReLU activation function; the fully connected layer is used to convert the matrix output by the graph convolutional layer into a vector; the loss function of the graph convolutional neural network is determined based on the sigmoid activation function and the cross entropy loss.

3. The method for training a food biological contamination source identification model according to claim 1, characterized in that: The food risk propagation simulation is performed multiple times according to the initialized circulation network graph model to obtain a training data set including: Based on the target propagation diffusion model, the food risk propagation path information is simulated according to the initialized circulation network graph model to obtain a food risk propagation chain; the food risk propagation chain includes target graph nodes and corresponding edges; Among them, the target propagation diffusion model includes one of the SI infectious disease model, the independent cascade model, the SIR model and the linear threshold model.

4. The method for training a food biological contamination source identification model according to claim 1, characterized in that: After obtaining the food biological contamination source identification model, the method further includes: Processing the test sample based on the food biological contamination source identification model to obtain a test result; Based on the test results, a target evaluation index is obtained, and the food biological contamination source identification model is fine-tuned and optimized according to the target evaluation index to obtain an optimized food biological contamination source identification model; the target evaluation index includes at least one of the F score, precision and recall rate.

5. A recognition method, characterized in that: include: Acquire information on pollution sources to be detected, wherein the information on pollution sources to be detected includes the occurrence area and type of food safety pollution sources; The information of the pollution source to be detected is processed based on the food biological pollution source identification model to obtain a food safety pollution source identification result; wherein, the food biological pollution source identification model is trained based on the food biological pollution source identification model training method as described in any one of claims 1-4.

6. A food biological contamination source identification model training device, characterized in that: include: A graph model construction module, for constructing a circulation network graph model based on the regional information and transaction information of the target agricultural product, wherein the circulation network graph model includes a plurality of graph nodes, wherein different graph nodes correspond to different regions, and edges between different graph nodes correspond to food circulation information between different regions; and the weight of the edge is determined based on at least one of the number of food transactions, transaction amount, geographical distance, and population between the two corresponding regions; A sample acquisition module, used to initialize the circulation network graph model, and perform multiple food risk propagation simulations according to the initialized circulation network graph model to obtain a training data set; wherein each set of sample data in the training data set includes a target graph node and propagation result snapshot data; the target graph node is at least one of the different graph nodes; The training module is used to iteratively train the graph convolutional neural network for each set of training data in the training data set, taking the propagation result snapshot data as input and the graph nodes corresponding to the propagation result snapshot data as output, and obtain a food biological contamination source identification model when the graph convolutional neural network converges.

7. An identification device, characterized in that: include: A data acquisition module, used to acquire information about pollution sources to be detected, wherein the information about pollution sources to be detected includes the occurrence area and type of food safety pollution sources; An identification module is used to process the information of the pollution source to be detected based on a food biological pollution source identification model to obtain a food safety pollution source identification result; wherein, the food biological pollution source identification model is trained based on the food biological pollution source identification model training method as described in any one of claims 1-4.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • A pollution source abnormal data identification method based on a deep learning algorithm

    CN109711547A

  • Pollution source type automatic identification method based on machine learning

    CN111985567A

  • Soil heavy metal source analysis method

    CN113470765A

  • Environment pollution behavior recognition method based on image analysis

    CN118094327A

  • System and method for dynamic management and control of air pollution

    US20230252487A1