Biological risk factor link tracing algorithm for African swine fever in market circulation field
Through the fusion framework of multi-nuclear heterogeneous measurement learning and multi-task learning, the problem of data heterogeneity in the pork traceability system is solved, efficient and accurate traceability of African swine fever epidemic is achieved, the traceability accuracy and decision-making efficiency are improved, and data security is ensured.
Patent Information
- Application Number
- CN202510406957.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing pork traceability system is difficult to achieve accurate risk traceability when processing a variety of heterogeneous data, especially when facing epidemics such as African swine fever, traditional methods are difficult to quickly locate pollution sources, resulting in inefficient epidemic control.
The fusion framework of multi-core heterogeneous metric learning and multi-task learning is adopted. By building a multi-core metric learning framework, combining multi-task learning and dynamic traceability algorithms, numerical, type and picture data in slaughtering, transportation, warehousing and market links is integrated to achieve efficient and accurate traceability.
It improves the accuracy of traceability, shortens the traceability time, improves decision-making efficiency, and ensures data security, which has significant practical value.
Smart Images

Figure CN120338480A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of food safety and animal disease prevention and control, and particularly relates to a whole-process traceability method for pork based on a metric learning algorithm, which fuses numerical, categorical, and image data in the slaughtering, transportation, warehousing, and market links through multi-core heterogeneous distance metrics, solves the problems of data heterogeneity and the influence of the previous link on the next link, and realizes the accurate traceability of African swine fever. Background Art
[0002] As one of the most important meat consumer products globally, the pork supply chain involves multiple links, including slaughtering, transportation, warehousing, and market links. In recent years, the outbreak of diseases such as African swine fever has posed a serious threat to the pork supply chain, not only causing huge economic losses but also triggering widespread public concern about food safety. Traditional pork traceability methods mainly rely on manual records and simple electronic tag technologies, which have significant deficiencies in data integrity, real-time performance, and accuracy. Especially when faced with complex disease transmission paths, traditional methods are difficult to quickly locate the pollution source, resulting in low epidemic control efficiency.
[0003] In existing pork traceability systems, data heterogeneity is a major technical challenge. The data types involved in the pork supply chain are diverse, including numerical data (such as temperature, humidity), categorical data (such as disinfection records, quarantine results), and picture data (such as surveillance images, transportation vehicle images). These data have significant differences in dimension, distribution, and semantics, and traditional single distance metric methods (such as Euclidean distance, Hamming distance) are difficult to effectively fuse these heterogeneous data, resulting in inaccurate traceability results. In addition, existing systems usually lack real-time monitoring and analysis of environmental parameters in the transportation and warehousing links and cannot detect potential risk factors in a timely manner.
[0004] In recent years, with the development of Internet of Things (IoT), blockchain, and artificial intelligence technologies, significant progress has been made in data collection, storage, and analysis of pork traceability systems. For example, electronic tag systems based on RFID (Radio Frequency Identification) and two-dimensional code technologies can achieve the whole-process tracking of pork products; blockchain technology ensures the immutability and transparency of data through a distributed ledger; machine learning algorithms are used to analyze complex supply chain data to improve traceability accuracy. However, existing systems still have deficiencies in data fusion and model generalization capabilities, especially when dealing with heterogeneous data and multi-task learning, lacking effective technical means.
[0005] In the field of metric learning, existing research has proposed various distance metric methods for numerical and categorical data. For example, the Gaussian kernel function is widely used for calculating the similarity of numerical data, while the Hamming distance and value difference metric (VDM) are applicable to categorical data. However, these methods are usually designed for a single data type and are difficult to be directly applied to mixed data scenarios. In addition, most existing metric learning algorithms rely on manually defined distance metrics and lack adaptability to data distribution and task objectives, resulting in poor performance in practical applications.
[0006] As an effective machine learning paradigm, multi-task learning can significantly improve the generalization ability and robustness of the model by optimizing multiple related tasks while sharing feature layers. In the pork traceability system, multi-task learning can simultaneously optimize the tasks of pollution source location and environmental anomaly prediction, so as to more comprehensively understand the risk factors in the supply chain. However, there are still deficiencies in the task weight allocation and feature sharing mechanism of existing research on multi-task learning, especially when dealing with heterogeneous data, it is difficult to balance the relationship between different tasks.
[0007] In summary, the existing pork traceability system has significant technical bottlenecks in data heterogeneity, environmental monitoring, and multi-task learning. To solve these problems, a method capable of effectively integrating multi-source data is urgently needed. The present invention proposes a brand-new full-process pork traceability method by introducing metric learning and multi-task learning techniques, aiming to improve the efficiency of epidemic prevention and control and the level of food safety. Summary of the Invention
[0008] For the risk traceability in the circulation field, this study conducts hierarchical traceability for the key links in the circulation chain. For the host pork of African swine fever, its circulation chain mainly consists of four parts: the slaughtering link, the transportation link, the warehousing link, and the market link. Except for the slaughtering link, the risk traceability of each link is affected by the risk level of the previous link and the hidden elements of biological risk factors in the current link.
[0009] In the field of the circulation ring, since the hidden element data comes from multiple different distributors, and the data types provided by distributors are rich and diverse. For example, the disinfection situation provided by the slaughterhouse generally has only two situations: disinfected and not disinfected. This kind of data belongs to categorical data. The temperature and humidity provided in the transportation link belong to numerical data, and the photos of the pork appearance belong to picture data. The traditional method to solve data heterogeneity is to uniformly convert data from different sources and different formats into a standard format and specification. By establishing a data warehouse or a data integration platform, data from different data sources are integrated together. During the integration process, operations such as data cleaning, transformation, and loading (ETL) are performed on the data to eliminate noise and inconsistencies in the data and store the data in a unified storage structure. However, pork traceability data not only contains simple conventional data types such as dates and numerical values, but also involves various complex data types such as pictures (such as environmental photos of the slaughterhouse and appearance photos of pork). Data standardization can only handle some basic data format and numerical specification problems and cannot effectively unify these complex and diverse data types.
[0010] To solve this data heterogeneity problem, a pork full-process traceability system based on multi-core heterogeneous metric learning and multi-task learning is proposed, aiming to solve the deficiencies of existing technologies in data fusion, dynamic weight allocation, model generalization ability, and data security. The system integrates numerical data, categorical data, and picture data in the slaughtering link, transportation link, warehousing link, and market link, constructs a unified multi-core metric learning framework, and combines multi-task learning and dynamic traceability algorithms to achieve efficient and accurate traceability of diseases such as African swine fever. The technical solutions are elaborated in detail as follows from aspects such as data preprocessing, metric learning, multi-task training, dynamic correction, and data security:
[0011] First, data collection and standardization processing are carried out. The data types involved in the pork supply chain are complex, and standardization and encoding need to be carried out according to different data characteristics.
[0012] For numerical data (such as temperature and humidity), the dimensionality difference is eliminated through z-score standardization. The specific formula is as follows:
[0013]
[0014] Among them, μ is the mean of historical data, and σ is the standard deviation. Numerical data is standardized through z-score to eliminate the dimensionality difference and make the features follow a distribution with a mean of 0 and a variance of 1. For bounded feature data such as humidity percentage, normalization processing is carried out to scale the data to the interval [0,1]. If missing values are found, the missing samples are deleted, or filled with the mean, median, or model prediction values. For example, after the temperature data in the transportation link is standardized, it can be uniformly mapped to the interval [-2,2] to avoid the impact of magnitude differences on the model.
[0015] For categorical data (such as "disinfected" or "not disinfected" in the disinfection record), it is mapped to a low-dimensional vector e ∈ R through the EmbeddingLayer d , where the embedding weights are optimized through backpropagation to preserve semantic relationships.
[0016] For picture-type data (such as photos of pork freshness), first, the images are uniformly adjusted to a fixed size, the pixel values are normalized to [0, 1], or the normalization parameters of the pre-trained model are used. Data diversity is increased through rotation, flipping, cropping, etc. to enhance the robustness of the model. The Region Proposal Network (RPN) is used to generate candidate regions (Bounding Box) that cover potential deteriorated regions (such as discolored, oozing, and abnormally textured parts). Then, ResNet-50 is used to extract high-level semantic features f(I) ∈ R 2048 , extract the high-level semantic features of the candidate regions, and perform L2 normalization to ensure the stability of the feature vectors. The class probabilities of the candidate regions are predicted through the fully connected layer, and the coordinates of the candidate boxes are refined to make the boxes fit tightly to the edges of the target regions.
[0017] Feature extraction for picture-type data is performed through a convolutional neural network, and the specific process is as Figure 2 shown. This link consists of multiple convolutional layers, pooling layers, and fully connected layers. During the training process, the network automatically learns various features in the image, from low-level edges and textures to high-level semantic information. Through forward propagation, the input image undergoes multiple convolutional and pooling operations, and the output of the last layer or several layers can be used as the extracted feature vector.
[0018] After inputting the image data, candidate region selection is first performed to find several regions in the image that may contain the target. These candidate regions are the objects of subsequent feature extraction and contain potential target information in the image. Then, normalization operations are performed on the selected candidate regions to adjust them to a unified size and format. The normalized candidate region images are sequentially input into the CNN network. After being processed by multiple convolutional layers and pooling layers, the feature representation corresponding to each candidate region is finally obtained. The extracted features are input into a Support Vector Machine (SVM), and the SVM classifies the candidate regions based on these features to determine whether each region contains the target and what kind of target it contains. At the same time, the features are also used for Bounding Box Regression to adjust the position and size of the candidate regions through calculation, so that the detection box can more accurately surround the target object.
[0019] Secondly, a multi-core heterogeneous metric learning framework is constructed. To fuse multi-source data, the system designs kernel functions for different data types.
[0020] Numerical data uses the Gaussian kernel function, and the specific formula is as follows:
[0021]
[0022] Among them, σ controls the similarity decay rate and is applicable to the similarity calculation of continuous features such as temperature and humidity. Let the numerical data set be (x 11 , x 12 , ···, x 1k ).
[0023] Categorical data uses a mixed kernel function, and the specific formula is as follows:
[0024] K cat (x 2i , x 2j ) = λD Hamming (x 2i , x 2j ) + (1 - λ)D VDM (x 2i , x 2j )
[0025] Among them, D Hamming counts the number of dimensions with different statistical attribute values (for example, the Hamming distance between "disinfected" and "not disinfected" in the disinfection record is 1), D VDM calculates the semantic distance based on the category distribution (for example, the difference in the occurrence frequency of a certain disinfection record in different disease samples), λ is a learnable weight that balances the contributions of the two distances. Let the categorical data set be (x 21 , x 22 , ···, x 2h ).
[0026] After the image data is extracted by pre-trained features, the Gaussian kernel function is used, and the specific formula is as follows:
[0027]
[0028] Captures the similarity of image features. Among them, let the image data set be (x 31 , x 32 , ···, x3n).
[0029] To achieve cross-modal data fusion, the system constructs a unified metric space by jointly optimizing the kernel weight and the Mahalanobis matrix M, and optimizes α1, β1, γ1. The objective function is designed as:
[0030]
[0031] Among them, S is the set of similar sample pairs (such as uncontaminated batch data), D is the set of dissimilar sample pairs (such as contaminated and uncontaminated data), η is the boundary threshold (usually set to 1.0), and μ and γ are regularization coefficients to prevent overfitting. The Adam algorithm is used in the optimization process, and the initial learning rate is set to 10 -3 , by alternately updating the kernel weight λ and the Mahalanobis matrix M until the loss function converges (the change rate is less than 10 -5 or iterate 1000 times). According to the optimized weight values, calculate the K total1 value for each link.
[0032] Based on the calculated K total1 value, design an asymmetric causal kernel, and its specific formula is:
[0033]
[0034] The formula describes the causal relationship between the (i + 1)-th link and the i-th link. Sigmoid is a common activation function, and its expression is The in this formula is the weighted fusion value of all risk factors related to the (i + 1)-th link. The role of Sigmoid(K i+1 ) is to map to the interval (0, 1), so as to normalize the influence of the (i + 1)-th link. This can ensure that the value of the K causal function is within a certain range, which is convenient for analyzing and comparing the causal relationships between different links. The slaughter link is already the source link in this study, so the final result of the asymmetric causal kernel formula for the slaughter link is 0. This formula mainly calculates the influence of the slaughter link on the transportation link, the influence of the transportation link on the storage link, and the influence of the storage link on the final market link.
[0035] Finally, weighted-fuse multiple results to obtain the final K total value, and its formula is:
[0036] K total = αK num + βK cat + γK img + εK causal
[0037] Among them, α + β + γ + ε = 1. Optimize again to generate a unified low-dimensional feature space and achieve data fusion.
[0038] Then perform KNN retrieval to correct the pollution probability, and its objective function is optimized as:
[0039]
[0040] To improve the generalization ability of the model, the system introduces a multi-task learning framework and optimizes both the main task (pollution source location) and the auxiliary task (environmental anomaly prediction) simultaneously. The main task extracts multi-source features h ∈ R 128 and outputs the probability distribution y main ∈ R C (C is the number of link categories), and the loss function uses cross-entropy loss
[0041]
[0042] to minimize the prediction error. The auxiliary task predicts the probability of temperature and humidity anomalies in the transportation link y aux ∈ [0, 1], and the loss function is the mean squared error. The specific formula is as follows:
[0043]
[0044] The combined loss function is obtained by weighted summation, and its specific formula is as follows:
[0045]
[0046] By balancing the task contributions, the weights are dynamically adjusted according to the AUC of the validation set. During training, the batch size is set to 32, and an early stopping strategy (terminate if the validation loss does not decrease for 10 consecutive rounds) is adopted to avoid overfitting of the model.
[0047] Based on the learned Mahalanobis distance, the system uses the K-Nearest Neighbor algorithm (KNN) to retrieve historical similar cases and correct the pollution probability of the current batch. For the sample s test to be measured, calculate its Mahalanobis distance from the historical sample s k Select the Top-K nearest neighbor samples and weight-average their pollution labels. The correction formula is:
[0048]
[0049] where the weight The user interface displays the contribution degree of each link in the form of a heat map. For example, the abnormal disinfection in the slaughterhouse is highlighted in red (contribution degree > 80%), and the abnormal transportation temperature is marked in orange (50% - 80%), assisting the supervision department to quickly locate the risk links.
[0050] The system uses the SHA-256 hashing algorithm to generate data fingerprints to ensure the immutability of data. A hash value Hash = SHA-256(Data||Timestamp) is generated for each record and stored in association with the timestamp to prevent replay attacks. The permission control module divides access permissions based on roles (RBAC): slaughterhouse administrators can only view the data of their own facilities, logistics managers manage transportation data, and regulatory agencies have full-link read and write permissions. Multiple-signature verification is required for data modification (such as joint authorization by the slaughterhouse and regulatory agencies) to ensure the traceability of operations.
[0051] Taking a batch of African swine fever positive samples as an example, the system retrieves its slaughter and disinfection records (categorized), transportation temperature, transportation humidity (numerical), and vehicle monitoring images (pictorial). After multi-core fusion, it outputs an 85% probability of abnormal slaughterhouse disinfection and a 70% probability of abnormal transportation temperature. Five historical similar cases are retrieved through KNN (such as batches with low disinfection scores and high transportation temperatures). After weighted correction, the probability of the slaughterhouse is increased to 88%. The visualization interface highlights the slaughterhouse node and generates a report recommending key verification of the disinfection process. Experimental data shows that the traceability accuracy of this system (AUC = 0.94) is improved by 23.7% compared with the traditional Euclidean distance method (AUC = 0.76), and the average time consumption is shortened from 2 hours to 15 minutes, with an 87.5% efficiency improvement.
[0052] The core innovation of this invention lies in: for the first time, a fusion framework of multi-core heterogeneous metric learning and multi-task learning is proposed to overcome the problem of data heterogeneity in pork traceability; through dynamic optimization of kernel weights and Mahalanobis matrices, environment-adaptive weight allocation is achieved; combined with the KNN algorithm and visualization interface, decision-making efficiency is improved; hash fingerprints and RBAC mechanisms are introduced to ensure data security. This system provides efficient and reliable technical support for food safety management, with significant practical value and promotion prospects.
[0053] The beneficial effects of this invention are:
[0054] 1. Aiming at the problem of data heterogeneity in risk traceability in the market circulation field, a multi-core metric learning algorithm is proposed to fuse numerical data, categorical data, and pictorial data.
[0055] 2. Aiming at the influence of the previous link on the next link in the market circulation field, an asymmetric causal kernel function is introduced to calculate the influence of the previous link on the current link for different links.
[0056] 3. Aiming at the risks of different links, an M matrix is constructed to intuitively calculate the influence of each link on the final link. Description of the Drawings
[0057] Figure 1 : Architecture diagram of the pork whole-process traceability system
[0058] Figure 2 : Flow chart of image feature extraction
[0059] Figure 3 : Multi-core heterogeneous metric learning framework diagram
[0060] Figure 4 : Dynamic traceability flow chart Detailed implementation manners
[0061] Refer to Figures 1 to 4 to further illustrate the present invention
[0062] According to Figure 1 the shown flow chart, and combined with actual data processing and model construction, each step of the method is introduced in detail
[0063] This patent defines and classifies the pork traceability process and the collected data according to the characteristics of the pork traceability mode. The collected data can be roughly divided into structured data and unstructured data. Due to the particularity of pork traceability, the collected data includes numerical data, categorical data, and image data. Traditional traceability models are difficult to solve data heterogeneity and the fusion of the three types of data, which will have a great impact on the final traceability
[0064] Metric learning narrows the distance between similar data. For similar scenarios in pork traceability (such as temperature and humidity, disinfection records, etc. in the compliance transportation link), metric learning adjusts the distance calculation method in the feature space (such as Euclidean distance, cosine similarity, etc.) to make similar data gather in the feature space. For pollution-related data (such as the transportation temperature and humidity of African swine fever positive samples, abnormal disinfection records, etc.) and normal data, metric learning optimizes the metric rule to increase the interval between the two types of data in the feature space. For example, the temperature data in the polluted transportation link and the normal transportation temperature data are clearly distinguished by the learned metric rule, providing a basis for subsequent detection of pollution links. Gaussian kernel function is used for its numerical data, hybrid kernel function is used for categorical data, and Gaussian kernel function is used for image data to capture the similarity of images
[0065] The present invention verifies the risk traceability model for 180 batches of pork samples transported from Sichuan to Shanghai. Among these batches of pork transported from Sichuan to Shanghai, 52 batches of pork are from Chengdu, 40 batches are from Nanchong, 35 batches are from Leshan, and 53 batches are from Yibin. Among these pork, a total of 78 batches of pork from Chengdu and Leshan are transported to Shanghai Jiangyang Agricultural Products Wholesale Market for warehousing and sales, and a total of 63 batches of pork from Nanchong and Leshan are transported to Shanghai Agricultural Products Central Wholesale Market for warehousing and sales. Different production areas and different warehousing environments will affect the microbial contamination risk of pork, thus affecting the risk traceability. In this study, the size of the sampling detection rate of microorganisms (the number of detected microorganisms / the quantity of sampled pork samples) is used as the basis for judging the size of the microbial contamination risk. According to the expert marking, the size of the risk and the originating area of the risk are marked.
[0066] Step 1: Randomly divide the above data according to the ratio of 7:3, where 70% is used as the training set samples to train the model, and 30% is used as the test set samples to test the goodness of the model.
[0067] 1.1 Data cleaning: Check and process the missing values and outliers in the data. For missing values, appropriate methods can be selected to fill according to the data characteristics. For example, for numerical data, if the missing ratio is small, the mean or median can be used for filling; if the missing ratio is large, a prediction model can be considered for filling. For outliers, they can be identified through statistical analysis (such as the 3σ principle) or visualization methods (such as box plots), and then decide whether to delete or correct them.
[0068] 1.2 Data encoding: Encode the categorical data so that it can be processed by the computer. For example, for the categorical variable "pork variety", one-hot encoding can be used to encode different varieties such as "Duroc pig" and "Landrace pig" into different vectors.
[0069] 1.3 Data standardization: For numerical data, in order to eliminate the dimension difference, a standardization method is adopted, and its specific formula is as follows:
[0070]
[0071] where x is the original data, μ is the mean, and σ is the standard deviation.
[0072] Step 2: Calculate the values of various kernel functions through the processed data. For the numerical data after standardization processing, such as the breeding environment temperature and transportation duration, the Gaussian kernel function is used
[0073]
[0074] Calculate the similarity between samples, where ||x 1i -x 1j|| represents the Euclidean distance between the same type of data of different samples. σ is the bandwidth parameter of the Gaussian kernel function, and its optimal value can be determined by methods such as cross-validation. The size of the bandwidth σ affects the smoothness of the kernel function. A smaller σ value makes the kernel function more localized and more sensitive to nearby data points, while a larger σ value makes the kernel function smoother and considers a wider range of data points.
[0075] For the encoded categorical data, a hybrid kernel function is used to calculate the similarity, and the specific formula is as follows:
[0076] K cat (x 2i ,x 2j )=λD Hamming (x 2i ,x 2j )+(1-λ)D VDM (x 2i ,x 2j )
[0077] If there are relevant picture data such as the internal condition of the transport vehicle and the appearance of pork, first use the pre-trained convolutional neural network to extract the features of the pictures to obtain feature vectors, and then use the Gaussian kernel function to calculate the similarity between the pictures. The specific formula is as follows:
[0078]
[0079] Calculate the influence of the previous link on the next link, and use the asymmetric causal kernel function. The specific formula is as follows:
[0080] K causal (x i ,x i+1 )=K i ·Sigmoid(K i+1 )
[0081] Step 3: Perform weighted fusion according to the formula K total =αK num +βK cat +γK img +εK causal where α+β+γ+ε=1. Here, first take α, β, γ, ε as 0.25, and define a loss function (mean square error loss function):
[0082]
[0083] By continuously adjusting the weight coefficients, the loss function is minimized. This method usually requires multiple iterative calculations to gradually approach the optimal solution. Finally, the optimal weight values of the slaughtering process are α = 0.4, β = 0.25, γ = 0.35, ε = 0.0; the optimal weight values of the transportation process are α = 0.2, β = 0.25, γ = 0.35, ε = 0.3; the optimal weight values of the warehousing process are α = 0.2, β = 0.2, γ = 0.3, ε = 0.05; and the optimal weight values of the market process are α = 0.3, β = 0.25, γ = 0.35, ε = 0.05.
[0084] Step 4: Obtain the fused and optimized K total value, unify the similarity measure of multi-modal data, and use it as the basis for optimizing the Mahalanobis matrix M and classification decision-making. K total is the input of the objective function of metric learning. By minimizing the K total distance between similar samples and maximizing the margin between dissimilar samples, the weights of M are optimized. Then, the M matrix is generated after correction using the KNN algorithm as shown below:
[0085]
[0086] Sub-matrix M 屠宰 contains elements:
[0087]
[0088] Quantify the independent risks of internal characteristics in the slaughtering process, such as the weights of the hygiene score (1.2), disinfection frequency (0.9), and operating room image features (1.1). Negative values indicate negative correlations between features (such as the negative correlation between the hygiene score and image features).
[0089] Sub-matrix M 运输 contains elements:
[0090]
[0091] The independent risks of transportation duration (0.8) and temperature control records (1.0). The off-diagonal 0.2 represents their synergistic effect.
[0092] Sub-matrix M 仓储 contains elements:
[0093]
[0094] The independent risks of warehousing temperature (0.7) and ventilation frequency (0.6), and 0.05 represents a weak association.
[0095] Sub-matrix M 市场 contains elements:
[0096] M市场 = [0.5]
[0097] Independent risk weight of the humidity in the market sales environment.
[0098] Among them, the influence matrix of the previous link on the next link includes:
[0099]
[0100] Step 5: Calculate the total risk of each link. Among them, the independent risk refers to the sum of the diagonal elements, and the total risk = independent risk + conduction influence. Combine the M matrix to obtain the total risk of each link. Among them, the total risk of the slaughter link is 1.2 + 0.9 + 1.1 + 0 = 3.2, the total risk of the transportation link is 0.8 + 1 + 0.3 = 2.1, the total risk of the warehousing link is 0.7 + 0.6 + 0.1 + 0.1 + 0.05 = 1.55, and the total risk of the market link is 0.5 + 0.1 + 0.05 = 0.65. By comparison, the total risk of the slaughter link is the largest. Therefore, the slaughter link has the greatest impact on the detection of African swine fever in the final pork samples.
Claims
1. An algorithm for tracing risk factors of African swine fever in the market circulation field, characterized in that The method includes the following steps: Step 1: For the biological risk traceability of the host pork of African swine fever in the market circulation field, which mainly includes 4 circulation links: slaughtering link, transportation link, warehousing link and market link. In each circulation link, there will be different slaughterhouses for slaughtering, different logistics companies for transportation, and different markets for sales. During this process, the data collection includes the following characteristics: 1.1 In each slaughterhouse, there will be data on relevant hidden elements, including the temperature in the slaughter workshop Workshop humidity Workshop disinfection situation Tool disinfection temperature Tool disinfection duration Quick-freezing temperature Quick-freezing duration Slaughterhouse hygiene grade Pig breed Tool disinfection method Pork photos etc.; 1.2 There will be data on relevant hidden elements in each transportation link, including transportation temperature Transportation humidity Transportation duration Packaging type Transportation type (Normal temperature transportation, cold chain transportation), pictures of pork etc.; 1.3 There will be data on relevant hidden elements in each warehousing link, including warehousing temperature Warehousing humidity Warehousing time Warehousing sanitation level Pork pictures etc.; 1.4 There will be data on relevant hidden elements in each market segment, including market temperature Market humidity Hygiene situation etc.; 1.5 Due to the transfer and transportation of meat between different transfer links, affected by the previous link, the probability of African swine fever virus occurrence or the probability of being infected with African swine fever in this link will increase; The data of the hidden elements of biological risk factors in the above-mentioned market circulation field can be defined as follows represents the i-th transfer link. There are 4 transfer links in the market circulation field, namely the slaughter link, the transportation link, the warehousing link, and the market link. The data of each link is divided into three categories: numerical data, categorical data, and picture data. Among them, the 1 in x1k represents the first category of numerical data, and k represents a total of k numerical data. x2h represents there are h categorical data, and x 3n represents there are n picture data. Generally speaking represents that there are z data of the n-th (n = 1, 2, 3, 4) category in the i-th (i = 1, 2, 3, 4) transfer link. Among them, X 1 represents the slaughter link, where X 2 represents the transportation link, where X 3 represents the warehousing link, where X 4 represents the market link. Step 2: There will be different types of data in one link. The data is roughly divided into numerical data, categorical data and picture data. Before using the kernel function to calculate the similarity, each type of data needs to be preprocessed; Step 3: Input the data of each processed link in Step 2 (represented as ) into the kernel function respectively, where the numerical data uses the Gaussian kernel function and outputs K num , the categorical data uses the mixed kernel function and outputs K cat , and the image data uses the Gaussian kernel function and outputs K img ; Step 4: Based on the obtained K num , K cat , K img , weighted fusion of the multi-core results, and gradually optimize each weight, and finally obtain the K total1 value; Step 5: Capture the asymmetric driving effect of the previous link on the next link, and use the asymmetric causal kernel formula to capture the risk transmission of the previous link to the next link; Step 6: Finally, fuse the non-causal values into K total1 and continue to gradually optimize each weight until finally obtaining the K total value.
2. According to the biological risk factor metric learning traceability algorithm of African swine fever in the market circulation field described in claim 1, characterized in that the hidden element data of the biological risk factor defined in step 1 is as follows: According to the form of the hidden element data of the biological risk factor in the market circulation field, this data can be regarded as a 4-dimensional data including the transfer link dimension, the distributor dimension, the variable dimension and the sample dimension. Each distributor contains multiple samples, and the variables in different distributors are different. And there are the following explanations for these data: 2.1 The hidden element data of African swine fever comes from different distributors in each transfer link of the transfer chain, and the African swine fever data provided by different distributors has different formats; 2.2 The hidden element data between different transfer links are independent of each other and do not affect each other. However, the previous link may affect the next link. In the slaughtering link, if the pork of this batch carries the African swine fever virus because the slaughtering tools are not disinfected and reaches the next link, and then due to abnormal temperature or other risk factors, African swine fever breaks out and the pork deteriorates.
3. A whole-process traceability method for pork based on a metric learning algorithm, characterized in that, Including the following steps: Step 1: For the differences in data characteristics of different links, the numerical data, categorical data and picture data are processed by matrices M1, M2, and M3 respectively to better capture the internal structure of the data, improve the accuracy and effectiveness of the traceability data processing, and help with accurate tracing; Step 2: Perform feature encoding on the heterogeneous data of the slaughtering link, transportation link, warehousing link and the final market link. The numerical data is standardized to z-score, and its z-score standard formula is: Where x is the original data value, μ is the mean of the data, and σ is the standard deviation of the data (reflecting the degree of data dispersion). If no standardization is performed, the numerical range of data such as temperature data is small and may be ignored when calculating distances, while the numerical range of data such as humidity is large and may dominate the calculation results. Through z-score standardization, this dimensional difference can be eliminated, making different features equally important in metric learning; categorical data is mapped to low-dimensional vectors through the Embedding Layer, and image data undergoes feature extraction through CNN; Step 3: Generate the K value for each link. Input the feature vector of the current sample (generated through the kernel function), as well as the feature vectors and known contamination labels (0 / 1) or probabilities of historical samples. For the current sample, use the Mahalanobis distance, which is a basic distance metric. total Retrieve the K nearest neighbor samples in the historical data. If the contamination labels of the neighboring samples are known, use weighted voting (with the weight being the reciprocal of the distance). If the neighboring samples have a contamination probability p , use weighted average correction. The specific formula is as follows: j where w is the reciprocal of the Mahalanobis distance; k Step 4: Based on the processed data above, the kernel functions of numerical, categorical, and image data are weighted and fused. The original objective function is: Modify the original objective function, incorporate the corrected pollution probability into supervised learning, and the optimized objective function formula is as follows: New item Force the metric space generated by the Markov matrix to be consistent with the corrected contamination probability. Use the modified objective function to optimize M so that it reflects the inter-link correlations (such as the block matrix in the document), and enhance the cross-link similarity metric; Step 5: Design the asymmetric Mahalanobis matrix M, and represent the independent risks and causal conduction of each link in a block manner, where the diagonal blocks represent the independent risk weight matrix, and the off-diagonal blocks represent the causal conduction weight matrix.
4. The multi-core heterogeneous metric learning algorithm according to claim 2, wherein The metric learning engine includes: Heterogeneous kernel function construction unit: The Gaussian kernel function is used for numerical data, the Hamming-value difference hybrid kernel function is used for categorical data, and the Gaussian kernel function is used for image data; Joint optimization unit: Optimize the kernel weights and the Mahalanobis matrix M through the objective function of maximizing the distance between classes and minimizing the distance within classes; Regarding the influence of the next link of the previous link, the present invention introduces an asymmetric causal kernel function based on the Gaussian kernel, Hamming-value difference hybrid kernel, and Gaussian kernel. The specific formula is as follows: K causal (x i ,x i+1 ) = K i · Sigmoid(K i+1 ).
5. The system according to claim 2, wherein In the dynamic pollution traceability module: Based on the learned Mahalanobis distance, retrieve the Top-K similar historical samples, use the pollution labels of the weighted average similar samples to correct the probability of the current batch. This method utilizes similarity calculation and simultaneously performs probability correction, and can obtain a better traceability effect.
6. The system according to claim 3, characterized in that, The metric learning engine supports multi-task joint training: Main task: Minimize the matching error between market positive samples and the polluted links. Its objective is to accurately obtain the pollution source (slaughter link, transportation link, warehousing link, market link), and output the probability value of the polluted link; Auxiliary task: Predict the probability of abnormal temperature and humidity in the transportation and warehousing links. Its objective is to give early warnings of abnormal transportation environments (such as too high temperature, too low humidity); The main task trains a classification model through multi-modal data (numerical, categorical, image) to predict the probability of the polluted link. The auxiliary task constructs a regression model to warn of abnormal transportation / warehousing environments. The two are combined through a multi-task learning framework: sharing the underlying feature extraction layer to capture cross-link commonalities. The main task branch outputs the pollution probability of the four links, and the auxiliary task branch outputs the abnormal score. They are jointly optimized through a dynamic weight loss function, and the abnormal probability of the auxiliary task is used as the attention weight and injected into the feature cross-layer of the main task, so that the high-abnormal links obtain higher feature weights in pollution determination, and finally realize the closed-loop enhancement of "abnormal warning - pollution traceability".
7. The system according to claim 3, wherein The characteristics of the asymmetric Mahalanobis matrix M include: The present invention maps data to a high-dimensional space through a kernel function and optimizes the Mahalanobis matrix M in combination with metric learning, which can effectively capture non-linear feature relationships, while quantifying the independent risks and causal conduction effects of each link. The finally obtained M matrix not only retains the flexibility of the kernel method, but also provides an accurate risk analysis tool for pork traceability in complex scenarios through interpretable weight allocation; First, standardize, encode, and extract features from numerical, categorical, and image data. Calculate the similarities of each data type through kernel functions (such as Gaussian kernel, hybrid Hamming-VDM kernel), and weighted fusion into a global kernel matrix. Design an objective function (minimize the distance between similar samples, maximize the interval between different samples, and add a regularization term), and use gradient descent to optimize M to ensure its positive semi-definiteness. In a multi-task framework, combine the loss functions of the main task (pollution source classification) and the auxiliary task (temperature and humidity anomaly warning), dynamically adjust the weights, and quantify the one-way causal influence through an asymmetric block structure (such as the conduction weight from slaughter to transportation). Finally, the M matrix combines the independent risks and causal conduction weights to provide an interpretable measurement basis for accurate traceability.
Citation Information
Cited By
Livestock and poultry epidemic disease information tracing method and system
CN121416127A