Single pig weight self-supervised reconstruction method for noisy automatic weighing data
By constructing a self-supervised skeleton through a self-supervised reconstruction method, the problems of uncertainty in multiple readings and multi-cluster structure in noisy automatic weighing data are solved. Robust reconstruction of single-pig daily-scale weight trajectory and anomaly labeling are achieved under unlabeled conditions, improving the reliability and deployability of the data.
Patent Information
- Application Number
- CN202610469656.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-10
- Publication Date
- 2026-06-23
Smart Images

Figure CN122265440A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart animal husbandry, precision farming, and agricultural information technology, and in particular to a method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data. Background Technology
[0002] In fattening pig production, body weight is a key phenotypic indicator reflecting growth status and production efficiency, and is widely used for growth monitoring, group management, slaughter decisions, and breeding evaluation. With the development of precision animal husbandry and large-scale production, automated weighing systems or electronic feeding systems (EFS) can continuously collect individual-level body weight readings, timestamps, ear tag numbers, pen locations, and other information without increasing labor costs, providing a basic data source for data-driven production management.
[0003] Existing technologies for processing automatic weighing data can be mainly categorized into the following three types: 1. Rule-based threshold filtering and empirical rule chains: Common practices include removing zero / duplicate values, limiting abrupt jumps between adjacent readings, setting daily weight gain limits, and using fixed weight thresholds to delete outliers. These methods are simple to implement, highly interpretable, and easy to deploy. However, their thresholds often rely on empirical settings and are sensitive to variations in weight stage, data entry, equipment status, and batch size. When noise mechanisms change with age, data entry position, or equipment calibration status, fixed thresholds are prone to cross-scenario mismatch, potentially leading to excessive deletion or retention of outliers, thus introducing systematic bias.
[0004] 2. Same-Day Compression Aggregation and Smoothing / Interpolation Reconstruction: Another common method involves compressing multiple readings from the same pig-day into a single statistic (such as the mean or median), and then smoothing and interpolating the daily-scale series. This can be achieved using techniques such as Lowes smoothing, Savitzky-Golay multinomial filtering, and moving averages to obtain a continuous curve; or by employing state-space models (such as Kalman filtering and smoothing) to suppress random noise. While this approach can improve series continuity and reduce high-frequency fluctuations, in situations such as "same-day multi-cluster mixing," "weighing equipment drift," "significant deviations of the median and mean from the true weight," and "non-major clusters potentially being closer to the true weight in the early high-noise phase," compression followed by smoothing can irreversibly mix multimodal structures, masking true candidates. Furthermore, smoothing methods tend to "smooth out" systematic anomalies, potentially weakening the ability to identify and audit structural problems such as drift and jumps, and are prone to producing unreliable interpolation extrapolations for cases with long gaps in data.
[0005] 3. Anomaly Detection and Weight Regression Based on Supervised Learning: Another approach involves training classification or regression models to identify abnormal readings and predict true weight. Inputs can be single load signals, short-time window statistical features, or multimodal information (such as images / point clouds). These methods can achieve high accuracy under controlled conditions, but typically rely on a large amount of high-quality labels or reference weight data. However, in pig farms, accurate labeling is costly, definitions differ across farms, and equipment and management conditions change frequently, making it difficult to maintain and stably generalize supervised models. Furthermore, supervised learning often focuses on single-point accuracy or short-term predictions, rarely explicitly addressing issues such as "uncertainty of multiple readings on the same day, multi-cluster structure, drift, and missing data," making it difficult to stably recover auditable daily-scale individual growth trajectories under low-labeling conditions.
[0006] In summary, existing technologies generally suffer from the following shortcomings: (1) It is difficult to explicitly handle the uncertainty and multi-cluster structure of multiple readings of the same pig within a day under conditions of no or weak labeling; (2) It is difficult to simultaneously take into account the error suppression of the high noise stage in the early age and the continuity of the weight of a single pig at the daily scale; (3) There is a lack of auditable identification and labeling mechanism for structural anomalies such as scale drift, local jumps and missing gaps, which makes it difficult to support equipment diagnosis and production audit; (4) There is a lack of a general solution that can recover the daily weight trajectory of a single pig from noisy automatic weighing data and provide stable input for downstream prediction in a low-cost manner with less reliance on additional hardware and manual labeling.
[0007] Therefore, there is an urgent need for a method that can reconstruct the daily weight trajectory of a single pig under noisy automatic weighing data and without labeling. This method should not only preserve the uncertainty of multiple readings on the same day and complete self-supervised screening and repair on this basis, but also output auditable quality labels and abnormal event markers to support production management decisions and equipment status diagnosis. Summary of the Invention
[0008] The purpose of this invention is to provide a self-supervised reconstruction method for single-pig weight from noisy automated weighing data. Under unlabeled conditions, it explicitly handles the uncertainty and multi-cluster structure of multiple readings of the same pig within a day. A high-confidence skeleton is constructed through stable daily gating and short-term feasible region consistency constraints. The skeleton-guided adsorption mechanism is used to repair daily-scale coverage and continuity, and finally outputs a single-pig daily-scale weight trajectory that can be used for production decisions. At the same time, auditable quality labels and anomaly event markers are generated to reduce the dependence on manual labeling and additional hardware, improve robustness and deployability under different pig houses, pens, batches and equipment status changes, and provide stable and reliable data input for downstream growth assessment, predictive modeling and equipment status diagnosis.
[0009] To achieve the above objectives, the present invention provides the following solution: A self-supervised reconstruction method for single-pig weight from noisy automated weighing data includes: Obtain individual-level weighing records for a single pig and perform standardization processing to obtain a set of multiple weight readings for a single pig per day; Based on the set of multiple body weight readings for a single pig per day, a stability index and a set of candidate readings for a single pig per day are obtained, wherein the stability index is used to divide a single pig per day into stable days and unstable days; On the stable day set, candidate readings that satisfy the feasible region consistency constraint are selected from the candidate reading set of single pig-day, and a multi-level branching strategy is used to construct an initial single pig self-supervised skeleton, and a quality label is output for single pig-day; The missing / unstable days of the initial single-pig self-supervised skeleton were repaired to obtain the reconstructed day-scale body weight sequence.
[0010] The present invention also provides a self-supervised reconstruction system for single-pig weight from noisy automatic weighing data, comprising: a stable day gating module, a candidate reading set construction module, a self-supervised skeleton construction module, and a skeleton-guided adsorption repair module.
[0011] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data.
[0012] The beneficial effects of this invention are as follows: Under unlabeled conditions, this invention explicitly handles the uncertainty and multi-cluster structure of multiple readings of the same pig within a day, constructs a high-confidence skeleton through stable daily gating and short-term feasible region consistency constraints, and uses the skeleton-guided adsorption mechanism to repair daily-scale coverage and continuity, ultimately outputting a single-pig daily-scale weight trajectory that can be used for production decisions. At the same time, it generates auditable quality labels and anomaly event markers to reduce the dependence on manual labeling and additional hardware, improve robustness and deployability under different pig houses, pens, batches and equipment status changes, and provide stable and reliable data input for downstream growth assessment, predictive modeling and equipment status diagnosis. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of a method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to an embodiment of the present invention; Figure 2Loss rates at different age stages according to embodiments of the present invention With relative dispersion Relationship diagram; Figure 3 This is a density distribution diagram of the range of weight changes with age within a stable intraday period according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the self-supervised skeleton initialization and 5-day consistency in an embodiment of the present invention, wherein (a) is the initialization starting from age d, and (b) is the identification of lonely days and the 5-day consistency rule with multi-level branches; Figure 5 This is a diagram illustrating the self-supervised skeleton of segmentation and inter-segment alignment and the 2-day consistency requirement in an embodiment of the present invention, wherein (a) represents segmentation and inter-segment alignment, and (b) represents a more stringent 2-day consistency verification. Figure 6 This is a column-level anomaly event detection and auditing diagram for daily weight gain mutations and short-window weight gain variance peak values in an embodiment of the present invention. Figure 7 This is a distribution map and ranking of the day and night patterns and specific time periods of the weighing activities according to an embodiment of the present invention; Figure 8 This is a typical noise pattern in the weighing data of the electronic feeding station in an embodiment of the present invention; Figure 9 This is a comparison of the daily weight wash coverage rate of BioSSR in this embodiment of the invention with six benchmark methods; Figure 10 This is a comparison chart of the average weight trajectory of pigs reconstructed from EFS readings by the various methods of this invention, with the actual weight and the PIC410 reference curve. Figure 11 The weight prediction performance of long-span recursive extrapolation (n=1–80) in an embodiment of the present invention; Figure 12 The image shows the cleaning and marking visualization of the individual pig weight trajectory (four representative pig pens) after BioSSR processing in this embodiment of the invention, where (a) represents pens 35 and 37, and (b) represents pens 69 and 82. Figure 13 This refers to the count of abnormal days detected in this embodiment of the invention and the matching status with the production calibration log; Figure 14 This is a comparison chart of the cleaning effects of BioSSR and three baseline models: MID, IQR-MID, and LOWESS, according to embodiments of the present invention. Figure 15 This is a comparison chart of the cleaning effects of BioSSR and three baseline models: SG, Kalman, and Kalman RTS, in embodiments of the present invention. Detailed Implementation
[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0017] like Figure 1 As shown, this embodiment proposes a self-supervised reconstruction method for single-pig weight from noisy automatic weighing data, including: Obtain individual-level weighing records for a single pig and perform standardization processing to obtain a set of multiple weight readings for a single pig per day; Based on the set of multiple body weight readings for a single pig per day, a stability index and a set of candidate readings for a single pig per day are obtained, wherein the stability index is used to divide a single pig per day into stable days and unstable days; On the stable day set, candidate readings that satisfy the feasible region consistency constraint are selected from the candidate reading set of single pig-day, and a multi-level branching strategy is used to construct an initial single pig self-supervised skeleton, and a quality label is output for single pig-day; The missing / unstable days of the initial single-pig self-supervised skeleton were repaired to obtain the reconstructed day-scale body weight sequence.
[0018] Furthermore, based on the set of multiple body weight readings per pig per day, the stability index and the candidate set of readings per pig per day are obtained, including: Calculate the robust center of body weight readings based on the set of multiple body weight readings for a single pig per day; The median absolute deviation is calculated based on the set of multiple body weight readings for a single pig per day and the robust center of the body weight readings. The stability index is obtained by the ratio of the median absolute deviation to the sum of the robust center of the weight reading and a preset constant. Density clustering is performed on the set of multiple weight readings of a single pig per day to obtain several clusters. The actual readings corresponding to the density peak of each cluster are extracted as cluster representatives to obtain the candidate set of readings of the single pig per day.
[0019] Specifically, data acquisition and standardization preprocessing include: acquiring individual-level weighing records output by automatic weighing systems or electronic feeding systems, wherein the records include at least: ear tag identification, timestamp, age or date, weight reading, and pen / house identification. Standardization processing is performed on the raw records, including but not limited to: removing null and invalid values, standardizing the time format, and grouping weight readings by "same pig, same day / age (same pig-day)" to obtain a set of multiple weight readings for each pig-day. ,in p Representative pigs pig , d Representing the pig's age in days d , i Representative pigs p age d The following weighing i The system uses "multiple readings per pig per day" as the basic processing unit, providing a data structure foundation for subsequent uncertainty preservation, gating, and skeleton construction.
[0020] Furthermore, density clustering of the single-pig-day set of multiple weight readings includes: Calculate the robust center for each age in each pen and multiply it by the scaling factor to obtain the adaptive clustering tolerance for weight stage. Perform density clustering on the set of multiple weight readings of a single pig-day based on the clustering tolerance to obtain several clusters.
[0021] Specifically, stable daily gating based on the stability of body weight readings includes: For each pig-day, the robustness center and dispersion of the multi-readout set are calculated, and a dimensionless stability index is constructed to distinguish between stable and unstable days that can be used as trajectory anchor points. (1) Calculate the weight readings per pig per day / age at the robust center : ,in, The number of times pig p was weighed at age d.
[0022] (2) Calculate the median absolute deviation per pig per day / age. : ; (3) Calculate the relative dispersion (stability index) of each pig per day / age. : ; in, It is a small constant, introduced to prevent numerical instability when the denominator is too small. ϵ = 1e-6. It is a dimensionless index: the smaller the value, the more tightly clustered and therefore more stable the multiple weight readings on the same day; while the larger the value, the more likely the readings on that day may be affected by scale jumping events, mixed (half-body weighing or multiple pigs weighed), unstable posture, or abnormal equipment condition. Therefore, It was used as a priori measure of the stability of daily body weight readings in pigs.
[0023] (4) When ≤ τ A day is determined to be stable if it is stable, otherwise it is determined to be unstable; among which τ Threshold. Stable day: (default τ= 0.05); Unstable days: . The judgment is: loss rate The "unstable days" will not be discarded. They will still be retained in subsequent steps for the construction of the "candidate readout set" and absorption-based correction, but will be excluded from the anchor points during the construction of the "self-supervised skeleton" to prevent noise contamination of long-term trajectories. The selection follows the principle of "suppressing highly dispersed readings while maintaining sufficient daily coverage." Loss rate ( This is used to quantify the retention cost incurred by this threshold, which is defined as... The percentage of pigs classified as unstable by day. (The default...) It was set to 0.05 because it significantly suppressed highly scattered days while still maintaining sufficient age coverage and a sufficient number of effective pig-day observations (e.g., ...). Figure 2 (As shown).
[0024] Among them, loss rate ( ) represents the given relative dispersion. Below, the proportion of pigs per day was filtered as unstable; the curves were stratified by the same 7-day age window (labels indicate the window range and the number of pig houses providing data). The default stable day gating threshold was used. = 0.05 was chosen after a trade-off between suppressing highly discrete pig-days and preserving sufficient stable day coverage for subsequent self-supervised skeleton initialization.
[0025] Furthermore, on the stable day set, candidate readings that satisfy the feasible region consistency constraint are selected from the candidate reading set of single-pig-day data, and a multi-level branching strategy is used to construct the initialization single-pig self-supervised framework, including: S1. Obtain the initial weight value of a single pig for a single day based on the distance between the candidate reading set of a single pig per day and the corresponding robust center; S2. On the stable day set, the short-term relative change of weight readings is quantified by a sliding window of a preset length as the consistency constraint of the feasible region. S3. Scan backward from the new age to the old age, select candidate readings that satisfy the feasible region consistency constraint from the candidate reading set of single pig-day, and attach them to the existing layer; S4. When the candidate reading does not meet the feasible region consistency constraint of the existing branch, a new branch layer is created within a sliding window of a preset number of days. If the number of layers reaches the preset value, return to S3; otherwise, proceed to S5. S5. Align the branches between segments and snap the short segments according to the proximity trend, update the initial weight value of the single pig-day, and obtain the initial single pig self-supervised skeleton, wherein the branch layer containing more than or equal to the preset number of consecutive points is regarded as a long growth segment, and the branch layer containing less than the preset number of consecutive points is regarded as a short segment. S6. The initialized single-pig self-supervised skeleton is verified by a sliding window with a length less than the preset day to obtain the final initialized single-pig self-supervised skeleton.
[0026] Furthermore, segment alignment of the branch layer includes: The longest segment containing the most points is used as the anchor point to fit the local linear trend, and the remaining long segments are regarded as candidate long segments to be matched. Alignment is performed by minimizing the mean absolute residual of each candidate long segment relative to the trend line. After aligning long segments, the residuals of short segments are compared with the trends of adjacent long segments, and the long segment with the best trend match is selected as the alignment target for alignment.
[0027] Specifically, to address the issue that multiple readings from the same pig-day may present a mixed cluster structure, this embodiment does not directly compress the readings of the same day into a single value. Instead, it clusters the readings within each pig-day and forms a candidate set by representing the values of each cluster. To preserve uncertainty: (1) Calculate each column pen Each day d Down( pen-d ) stable center And set clustering tolerance that adapts to weight stage. : ; ; in, It is the scaling factor (default value) = 0.05), which is conceptually independent of , This is the median after removing extreme values. For the age of dayd This is the collection of all weight data recorded in the time column. This design creates a tolerance that is linearly proportional to the weight level, thereby reducing systematic mismatches that occur when using fixed absolute thresholds across different weight ranges.
[0028] (2) Set of multiple weight readings for each pig-day within the same pig-day according to Density clustering (e.g., DBSCAN) is performed to obtain several clusters; (3) Extract the actual reading corresponding to the density peak for each cluster as the cluster representative (e.g., map the KDE peak to the original reading), and form a candidate set. .
[0029] (4) and based on the candidate set for each pig-day With the corresponding column stability center The distance is used to give the initial weight values for each pig-day on a temporary daily scale. : ; This approach retains the multi-peak structure of the same day, incorporating the possibility that "the true weight may be located in a secondary cluster" into the subsequent self-supervised selection space; and avoids irreversible information loss of the median or mean when multiple clusters are mixed. Alternative solutions include: clustering algorithms can be replaced with Gaussian mixture models, mean shift, hierarchical clustering, etc.; cluster representation can be replaced with the cluster median, the mean after removing extreme values, etc.; tolerance... It can adaptively estimate based on weight stage, age segment, or equipment accuracy.
[0030] Self-supervised skeleton construction based on short-run feasible region consistency constraints includes: In scenarios where reliable real labels are unavailable, short-term feasible growth regions are encoded as a self-supervised consistency constraint: real weight should not exhibit abrupt jumps that violate biophysical rationality on short timescales. Therefore, the consistency constraint-driven self-supervised process automatically selects the most self-consistent time branch from the candidate read set as the trusted skeleton and outputs auditable three-state quality labels (T: trusted skeleton point, O: lonely day, F: discard).
[0031] (1) Define Loneliness Day O: For some pigs, the temporary daily body weight at a specific age. Lack of sufficient neighboring observations (e.g., missing) , , and In this situation, the local consistency of the trajectory cannot be reliably assessed. Therefore, for each pig-day, the number of available observations within the 5-day neighborhood (±4 days) and the 2-day neighborhood (±1 day) are counted, denoted as […]. and : ; ; If there is insufficient neighborhood information (i.e., only data for the current day is available), the day is marked as O (lonely day; insufficient evidence) and excluded from the initialization of the self-supervised skeleton, thereby avoiding the forced judgment of making a credible skeleton point T or discarding F in the absence of evidence.
[0032] (2) Define the relative change in body weight (or daily weight gain constraint). : .in, It is a small constant, introduced to prevent numerical instability when the denominator is too small. ϵ = 1e-6.
[0033] (3) Set a window consistency threshold (e.g., a 5-day window threshold). = 0.28 and 2-day window threshold =0.22), to select candidate values that satisfy the constraints from the candidate set, and to allow updating the weight of each temporary pig-day with the candidate set when the constraints are satisfied. .
[0034] Among them, on the stable day ( Under the condition of ≤ 0.05, a sliding window of 2-7 days is used to quantify the short-term relative change in body weight readings. (like Figure 3 As shown in Table 1, the statistics for each window length are highly consistent (the median shifts to the right as the window length increases (e.g., from about 0.02 to 0.08), reflecting a larger cumulative change over a longer age range). However, the upper bound of the high quantile varies little across different window lengths (P90 ≈ 0.17–0.20, P95 ≈ 0.22–0.28, P99 ≈ 0.38–0.41). The upper bounds of extreme changes (P90, P95, P99) are not dominated by short-term window length or physiological growth. Even during stable days, they are "capped" by residual anomalies or structural events. Therefore, the upper bounds of extreme changes can serve as a conservative upper bound for the short-term feasible region. Secondly, the estimation of the broader short-term feasible region threshold is insensitive to intermediate stages such as window length and physiological growth, indicating that the statistical characteristics of broader short-term relative changes have good stability and reproducibility in the stable day set. Therefore, a 5-day window threshold is set. = 0.28 and 2-day window threshold = 0.22.
[0035] Table 1 (4) Set a window consistency threshold (e.g., a 5-day window threshold). = 0.28 and 2-day window threshold =0.22), to select candidate values that satisfy the constraints from the candidate set, and to allow updating the body weight of each pig-day at the temporary daily scale with the candidate set when the constraints are satisfied. .
[0036] During the stabilization period, a consistency threshold for the short-term feasible region is set within a 5-day window. = 0.28, and scans backward from the newer age to the older age (e.g. Figure 4 (as shown in (a)-(b)). For the age of... d Heavenly Pig p For candidate reading sets Perform a "sticky" selection. Specifically, if candidate readings exist... c ∈ Its nearest confirmed point on the left (i.e. , ... It can satisfy the consistency constraint. ≤ Then the candidate reading c It will be attached to the corresponding layer, and Allowed c ∈ Update (i.e., effectively correct the temporary daily weights in the candidate set).
[0037] When there is no c ∈ of c When attaching to any existing layer, a fixed-length waiting window of 5 days is used (i.e., allowing a maximum of 5 consecutive days in an unresolved state), and within this window, a search is performed to find layers that satisfy the same consistency constraints. ≤ Any pair of points. If such a pair of points is found, a new branch layer is created. To control model complexity and reduce overfitting, the maximum number of layers is set to [value missing]. L max = 6 This multi-level branching process explicitly represents the uncertainty caused by multimodal daily readings and provides a structural basis for subsequent segment alignment.
[0038] (5) Align the long segments after segmentation (e.g., fit the local trend with the longest segment and minimize the residual to align other segments), and snap the short segments according to the proximity trend to obtain the initial skeleton 1. Skeleton_1 ): After layering, the Lonely Day O is considered the dividing point, and each layer is divided into a group of continuous segments. Segments with at least 5 consecutive points are considered long segments (long segments). Within a long segment, the layer containing the most points is considered the anchored segment (longest segment, 1). st The remaining long segments are considered as candidate long segments to be matched. Segments with fewer than 5 consecutive points are defined as short long segments (short segments), and all short segments are considered as candidate short segments.
[0039] In pig-day readings, the multi-layered structure can lead to mismatches between different growth segments, causing a trajectory to be divided into multiple reasonable stages with different connectivity patterns. To ensure continuity between long segments, a segment-align strategy is employed. The longest segment (1... st The first candidate long segment is used as an anchor point to fit a local linear trend, and they are aligned by minimizing the mean absolute residual of each candidate long segment relative to that trend line. During the validation phase, the anchor points are iteratively replaced with the second and third longest (2... nd and 3 rd The process iterates through candidate long segments and repeats the alignment process; if the minimum residual does not improve the situation, the current solution is retained; otherwise, the updated solution is accepted. This process progressively corrects mismatches in long segments, such as... Figure 5 As shown in (a).
[0040] For candidate segments, after aligning long segments, their residuals are compared with the trends of adjacent long segments, and the long segment with the best trend match is selected as the alignment target. Then, by analyzing the candidate set... Select the candidate value that is closest to the selected target to update. This yields the initial self-supervised skeleton 1 for each pig, denoted as "". skeleton 1 ".
[0041] (6) Stricter consistency rules and verification: skeleton_1 Global self-consistency checks are performed, but single-point jumps and short gaps require more sensitive local checks. Therefore, a stricter 2-day window consistency threshold is used in local validation: ≤ = 0.22.
[0042] If a point has already been classified as F under the 5-day consistency rule, it cannot be "saved" by the 2-day rule. skeleton_1 For the points in the data, perform a 2-day verification process and mark the points that fail the verification as F to prevent small local anomalies from being masked by the global trend. Figure 5 As shown in (b).
[0043] After this step, each pig will obtain the final initialized self-supervised skeleton 2, denoted as skeleton_2 Each pig is assigned an auditable three-state label every day: T, O, or F.
[0044] Furthermore, the repair of missing / unstable days in the initialized single-pig self-supervised skeleton includes: S1. Fit the curve of the reliable skeleton points of the final initialized single pig self-supervised skeleton, perform bandwidth screening on the candidate set of lonely days and unstable days, select the points that fall within the preset bandwidth range of the curve, and if there are multiple points, select the point that minimizes the residual of the previous and next neighborhoods to update the daily scale weight, and obtain the single pig self-supervised skeleton after adsorption in the skeleton. S2. For each target day outside the final initialized single-pig self-supervised skeleton range, based on the nearest anchor points of the single-pig self-supervised skeleton after adsorption within the skeleton, fit the local linear trend, construct a conservative feasible band to search the candidate reading set of the single-pig-day, select the candidate points falling within the band and use them as new anchor points, repeat S2 until there are no new anchor points, and obtain the single-pig self-supervised skeleton after exoskeleton adsorption, thus reconstructing the daily scale weight sequence.
[0045] Specifically, with the skeleton skeleton_2 The T-point is used as the teacher signal to repair missing / unstable days both within and outside the skeleton, improving coverage and continuity while avoiding cumulative drift due to erroneous extrapolation. (1) Adsorption and repair within the skeletal framework: right skeleton_2 The skeleton T-point is fitted with a smoothed teacher curve (e.g., using smoothing or linear interpolation methods like LOWESS). Bandwidth filtering is then applied to the candidate set for O and unstable days, retaining only those candidates falling within the relative bandwidth of the teacher curve. skeleton_3 (Single-pig self-supervised skeletal skeleton after intraskeletal adsorption); if multiple candidates satisfy the criteria, the candidate that minimizes the residuals of the preceding and following neighborhoods is selected to update the daily-scale body weight. The bandwidth setting remains the same. ≤ = 0.28 and ≤ =0.22.
[0046] (2) Conservative adsorption extension outside the framework: For each target day outside the skeleton2 range, based on the skeleton skeleton_3Several recent anchor points were fitted to local linear trends, and a conservative feasible band P90 was constructed. ≤ = 0.20 and ≤ =0.17, search candidate reading set C p,d The candidate set is expanded and promoted to a new anchor only if there are candidates falling within the band, until no more can be satisfied (i.e., when there are no new anchors). skeleton_4 (Self-supervised skeleton of a single pig after exoskeleton attachment). If multiple candidate values satisfy the constraint, the one that minimizes the local residual (i.e., minimizes the mismatch with the teacher trend given nearby anchor points) will be selected and promoted to the next new anchor point.
[0047] The above steps, while ensuring reliability, fill in gaps and omissions, improve daily-scale coverage and trajectory continuity, and suppress the cumulative contamination of drift noise through extrapolation. Alternative solutions: the teacher curve can be replaced with splines, Kalman smoothing, or robust local regression; the adsorption bandwidth can be segmented by day age or adaptively based on historical fluctuation ranges; the extrapolation range can be limited by a maximum extrapolation length or a minimum evidence density.
[0048] Furthermore, obtaining the single-pig self-supervised skeleton after adsorption within the skeleton also includes: The growth period is divided into several age stages. The weight gain sequence of pigs in each age stage is estimated based on the self-supervised skeleton of a single pig after adsorption in the skeleton. The weight gain sequence is clustered, and the adaptive fluctuation radius is used as the neighborhood radius. The largest cluster is taken as the normal weight gain pattern, and the remaining clusters and noise points are taken as abnormal weight gain / loss events on the daily scale. If the number of valid pigs in the pigpen is greater than the preset value, the daily median of the short window variance term of all pigs in the pigpen is calculated, and the IQR rule is used to analyze the daily median to identify candidate abnormal days.
[0049] Specifically, after repair within the skeleton, the final determined initial self-supervised skeleton is... skeleton_2 It will be optimized skeleton_3 To improve the integrity of the reconstructed growth trajectory, additional attempts were made to repair unresolved points outside the skeleton (mainly concentrated around 70-100 days of age). However, repairing points outside the skeleton is inherently less reliable because: (i) the local temporal neighborhood may be sparse; and (ii) this process may be affected by potential enclosure-level events such as scale drift or collective disturbances, which may introduce systematic biases.
[0050] Therefore, explicitly outputting event labels at both the individual and group (column) scales aims to: prevent abnormal days from being incorrectly "included" in normal trajectories; enable auditing of device-level data quality; and provide actionable signals for downstream management decisions or device diagnostics, such as... Figure 6 The diagram shown is an abnormal event detection and auditing diagram using columns 35 and 37 as examples.
[0051] use skeleton_3 Temporary daily weight readings in state T Two daily-level features were constructed: for pigs p age d Daily weight gain is defined as the difference between two consecutive days. And calculations are only performed when data from both days is available: ; To capture the abrupt change in a two-day weight gain pattern, the variance of daily weight gain within a rolling window of length 2 was calculated. : ; (1) Individual scale: Due to the daily weight gain in the individual's growth trajectory The distribution of growth factors varies across different age stages, and using a fixed threshold may lead to systematic mismatches between different stages. To reduce the confounding effect of age stage on anomaly identification, the growth period is divided into several age stages (70–90, 90–110, 110–130, 130–150, and 150+), and the values for each age stage are estimated. E Inner pig p weight gain sequence The stage-specific fluctuation range.
[0052] Specifically, in the stage E Inside, collect pigs p All available daily weight gain samples are denoted as And calculate the HPD interval (Highest A posteriori density interval: the shortest interval covering a specified proportion of samples). Set the coverage level to HPD_MASS=0.95 to obtain the shortest coverage interval [ l , h ],in: l :stage E The lower limit of the main weight gain distribution (lower endpoint of the HPD interval); h :stage E The upper limit of the main weight gain distribution (the upper end of the HPD interval).
[0053] Then, this stage Adaptive fluctuation radius Defined as half the interval width: ; Next, the weight gain sequence at a specific stage. Perform DBSCAN clustering and use The neighborhood radius (minimum number of samples = 1) was used. The largest cluster was considered the "normal weight gain pattern" for that stage, while all remaining clusters and noise points were marked "individual = red," indicating that the pig... p In the stage E An anomalous weight gain / loss event on a daily scale has occurred.
[0054] (2) Population scale: When the number of effective pigs in a pen on a certain day is ≥3, the short window variance term of all pigs in that pen will be calculated. The daily median. Then, the IQR rule (where = 1.5) Analyze the daily median to identify candidate anomalous days and detect sudden changes.
[0055] If a candidate day contains pigs marked "individual = red", then that day is marked "group = red" (this is stronger evidence of a real event). If no "individual = red" cases are observed on that day, it is marked "group = orange", indicating that there may be minor disturbances at the equipment, environmental, or group level, rather than confirmed individual growth abnormalities.
[0056] This system differentiates between "individual abnormal growth / abnormal readings" and "equipment / pen level anomalies (such as drift, collective disturbances)," providing interpretable signals for data auditing, equipment diagnosis, and management decisions, and preventing abnormal days from being mistakenly included in the normal trajectory. Alternative solutions include: anomaly identification can be replaced by change point detection, CUSUM, robust Z-scores, quantile-based adaptive thresholds, etc.; the group scale can be extended to the pig house / batch scale.
[0057] Furthermore, before inputting the reconstructed daily body weight sequence into the machine learning-based prediction model, the following steps are included: The reliable skeleton points in the reconstructed daily weight sequence are used as high-confidence daily anchor points. The original readings corresponding to the high-confidence daily anchor points are traced back to the anchor cluster, and the original readings and their timestamps in the anchor cluster are used for intraday structured representation and correction.
[0058] Specifically, intraday period normalization: Automatic weighing readings exhibit a clear intra-diurnal rhythm: the diurnal activity differences of pigs and fluctuations in gastrointestinal fullness caused by feed and water intake can lead to systematic shifts in weight readings at different times of the same day. If all daily readings are directly summarized into a single daily weight, inter-diurnal variations will be mixed with intra-diurnal rhythms, thus affecting daily weight gain estimation and attribution of abnormal events. Therefore, this invention, after completing extraskeletal adsorption, uses high-confidence daily anchor points ( skeleton_4 The one marked with T Based on the anchor point, the original readings corresponding to the anchor point are traced back to the anchor cluster of that day, and only the original readings and their timestamps within the cluster are used for intraday structured representation and correction.
[0059] Specifically, the trigger frequency of all original weighing records was first statistically analyzed hourly, and hourly frequency curves were plotted at different age stages. Figure 7 This method was used to verify that weighing behavior has a stable temporal preference and to determine the intraday window structure accordingly. Based on frequency peaks and troughs and typical feeding rhythms, the day was divided into four time windows: 20:00–04:00, 04:00–07:00, 07:00–15:00, and 15:00–20:00. Subsequently, the anchor cluster readings for each pig-day were grouped according to the above time windows, and the median of the readings within each time window was calculated to form a structured weight observation of "pig-day-time window".
[0060] Considering that intraday bias patterns may differ across age groups, this embodiment uses a reference window with sufficient sample size and relatively stable behavior (e.g., 04:00–15:00) as a unified benchmark (Table 2 shows the relative weight changes at different times within the next day for different age groups). Within each age group, the stage-specific ratio of other time windows relative to the reference window is estimated, and weight observations outside the reference window are mapped to the reference window scale accordingly. The final output is the daily-scale weight. The priority and backoff rules are applied: if the pig has sufficient readings within the reference window, the median of the reference window (within the anchor cluster) is directly taken as the reference value. If the reference window is missing or insufficient, the representative values of other time windows are converted to the reference scale according to the stage scaling factor, and the median is taken as the reference value. .
[0061] Table 2 It is important to emphasize that this step is a "time period normalization correction" rather than interpolation point creation: the output daily scale weight can still be traced back to the original readings of the same day's anchor cluster. It only performs a unified scale mapping on the systematic shifts in different time periods, thereby improving cross-day comparability and reducing the interference of intraday rhythms on daily weight gain estimation and subsequent anomaly detection.
[0062] After completing the daily-scale weight output and obtaining the cleaned dataset, the prediction model is then trained. The cleaned data is used as the input to the prediction model, and the output includes: a reliable daily-scale weight trajectory (continuous or maximum coverage sequence) for each pig and its trajectory point quality labels (T / O / F labels); individual and group-scale anomalous event labels (such as data level labels like red / orange); Optional output: Fix source markers (which cluster / time period from the candidate set) to support audit traceability.
[0063] To verify the effectiveness of this invention (BioSSR: A Self-Supervised Reconstruction Method for Single-Pig Weight Trajectory in Noisy Automated Weighing Data) in "daily-scale weight trajectory reconstruction and usability improvement for single pigs" in a real-world production scenario, a systematic evaluation was conducted on data from a commercial finishing pig farm. The data came from multiple batches of finishing pigs from two commercial finishing pig farms and 20 pig houses, totaling 1392 pigs; the age at entry was approximately 70±1 days, with most pigs slaughtered at ≥150 days. An automated weighing / electronic feeding system (EFS, Osborne Industries, Inc.) continuously recorded individual-level raw weight readings, with data fields including timestamp (date-time-minute), age, weight reading, ear tag number, pen number, and pig house number. To obtain the reference true body weight (truth-BW), an independent weighing scale was installed in each pen of the above 20 pig houses, and each pig was weighed daily according to a fixed procedure (e.g., 9:00-11:00). The equipment was calibrated at a fixed cycle (once every 5 days) to obtain a daily true body weight dataset that can be used for alignment assessment.
[0064] To ensure a fair and interpretable comparison, this assessment divides the data into three categories: (1) Truth-BW dataset: Daily-scale true body weight obtained through a standardized process; (2) Raw EFS BWreadings dataset: Raw pig-day multi-reading records output by EFS; (3) Cleaned daily BW datasets: Daily-scale body weight trajectories reconstructed from raw multi-readings using the self-supervised reconstruction method for single-pig body weight trajectories at the daily scale of this invention and six commonly used baseline methods. The six baseline methods include: Same-day median aggregation (MID), IQR trimmed median (IQR-MID), LOWESS smoothing, Savitzky-Golay filtering (SG), Kalman filtering, and Kalman RTS smoothing. The above baselines cover typical paradigms such as regular aggregation, smoothing filtering, and state-space denoising, and can form a directly comparable reference trajectory with this invention.
[0065] The evaluation index system consists of two parts: (1) Cleaning / reconstruction quality (aligned with truth-BW): Align the daily scale body weight output by each method with the truth-BW by individual-day age. If the reconstructed value falls within ±5% of the true value, it is recorded as a valid day. The percentage of valid days is used as the coverage rate to measure the availability and continuity of the daily scale trajectory.
[0066] (2) Downstream learnability and prediction benefits (keeping the prediction framework unchanged, only changing the source of the input sequence): Construct a unified lightweight prediction task, and use historical weight lag features as the fixed method ( , , The input consists of time index features (age in days and months); the input is based on column type. pen Five-fold cross-validation of groups is used to reduce optimistic bias caused by homology correlation, and the test set is drawn from unseen fields, thus more closely resembling production generalization scenarios. In this setting, only the source of the "daily scale body weight sequence" (this invention and each baseline / ablation variant) is replaced, and the prediction results are uniformly aligned with the same truth-BW to calculate indicators such as MAE, RMSE, MAPE, and R². Two scenarios are set up: short-term rolling prediction (n=1, predicting the next day) and long-span recursive extrapolation (rollout, n=1–80, continuously predicting the next 1–80 days) to evaluate the stability of error accumulation over time.
[0067] Based on the aforementioned real data and a unified comparative standard, the present invention can achieve the following technical effects (providing key quantitative evidence): 1. Significantly improves coverage and continuity on a diurnal scale during the early high-noise phase: In the scenario, the 70–90 day age group exhibits high dispersion and a clear multi-cluster structure among multiple readings on the same day (e.g. Figure 8 As shown, baseline methods often suffer from breakpoints and insufficient available days. This invention, through a chain mechanism of "uncertainty retention in candidate reading sets—stable day gating—short-term feasible domain consistency framework—framework-guided adsorption and remediation," can still output high-coverage daily-scale trajectories in the most challenging early stages: coverage reaches 98% in the 70–90 day age range, and a stable coverage level of over 95% can be maintained at approximately 75 days, significantly outperforming the baseline method. Figure 9 This effect stems from: stable day-gating blocking the contamination of the framework by highly discrete pig-days; candidate sets preventing irreversible compression of multiple clusters on the same day; and adsorption repair filling in gaps and breaks, increasing the density of usable days.
[0068] The overall trends of all cleaning methods are similar. Figure 10However, systematic differences exist in the early stages (before 90 days of age). PIC410, a commercial reference curve, deviates significantly from the actual average trajectory of the data in this study (e.g., the time point when the weight trajectory reaches 100 kg in this embodiment is approximately 10 days earlier than the PIC410 timeframe), indicating that a single static template is insufficient to cover the real differences across different batches, seasons, and management conditions; production management requires continuously updated single-pig-level dynamic trajectories as decision input. Furthermore, the BioSSR method of this invention most closely approximates the actual weight.
[0069] 2. Under the condition of a fixed prediction model, it has significant advantages in short-term rolling forecast performance and long-span extrapolation: In downstream forecasting tasks, in the short-term rolling forecasting (n=1, forecasting the next day) of the 5-fold cross-validation framework grouped by column, only the daily-scale body weight sequence source is replaced. When this invention is used as input, its significant advantage over the other six baseline models is concentrated in the 70–90 day age window (Table 3 shows the downstream forecasting performance and data cleaning coverage for different age stages): In the 70–90 day age stage, the error levels corresponding to this invention are MAE=1.11 kg, RMSE=1.48 kg, MAPE=2.65%, and R... 2 =0.951, indicating that even during the period of strongest noise, a daily-scale weight sequence that is closer to the actual growth can still be recovered. This effect is directly caused by the "selectable space + feasible region restriction" brought about by the candidate set and the consistency constraint, which mechanistically suppresses the propagation of structural noise such as drift, half-weighing, and multi-pig interference over time.
[0070] Table 3 In long-span recursive extrapolation scenarios (n=1–80), as the prediction step size n increases, the errors of each method increase, and R... 2 A decrease is a general trend; however, under the same prediction model and training / evaluation process, when only the input weight sequence is replaced, BioSSR shows a more significant advantage in long spans of n>60. Figure 11 Table 4 shows the ranking prediction performance of long-span recursive extrapolation (n=60–80). This "long-span advantage amplification" is directly related to its higher coverage and lower input bias in the early stage of 70–90 days: the extrapolation of a longer span is essentially more dependent on the estimation of the initial state and short-term trend in the early window, and the larger missing and biased baseline in the early stage will be amplified in the recursion (n=1-80); while the BioSSR of this invention can provide more accurate and continuous weight input in the early stage of high noise, thereby providing a more reliable initial state and trend basis for long-span prediction and improving the overall long-term stability.
[0071] Table 4 3. Provide auditable quality labels and anomaly event markers to support equipment drift diagnosis and production auditing: In addition to outputting daily weight trajectories, the BioSSR of this invention also outputs pig-day quality tags (e.g., T / O / F) and dual-scale abnormal event markers (individual level and pen level). Figure 12 (a)-(b)). Compare the detected abnormal events with the scale calibration records in the production log: the log records 88 calibrations. If a calibration for "scale drift" is usually accompanied by abnormal behavior before and after, there are at least 176 corresponding abnormal events. Figure 13 The method in this embodiment identified 181 "red" events (combining "individual-level weight gain anomalies" and "short-term variance surges at the column level"), of which 170 were aligned with production logs, a 93.92% alignment rate. This indicates that the "red" tag can effectively capture recordable and perceptible explicit equipment drift events in the production system. In contrast, only 6 out of 201 "orange" events (short-term variance surges at the column level) were aligned with logs (2.99%), suggesting that "orange" events mostly reflect mild drift, environmental influences, and other data quality fluctuations, and are "warning" signals rather than explicit calibration events. This demonstrates that the event tagging can effectively capture recordable and perceptible explicit equipment drift / calibration-related anomalies in the production system. This capability makes automated weighing data cleaning no longer a "black box outputting a curve," but rather forms traceable, interpretable, and closed-loop maintenance-compatible audit information.
[0072] 4. Reduces cross-day confounding caused by intraday rhythms and gastrointestinal fullness, improving the reliability of daily weight gain estimation and abnormal attribution: Statistics show that weighing triggers exhibit a stable peak-valley structure within 24 hours, with a systematic bias of approximately 1% in weight across different time windows (Table 2). Directly summarizing all readings for daily weight would cause intraday rhythms to confound inter-diurnal variations. This invention employs an intraday time-period normalization strategy of "anchor point backtracking candidate clusters—structuring by time window—estimating ratios by age stage and mapping to a reference window scale" to correct time window bias without interpolation, improve cross-day comparability, and reduce the risk of confounding attributions for abnormal events (environment / management / equipment).
[0073] In summary, this invention achieves daily-scale single-pig weight trajectory reconstruction with "high coverage and low error in the early high-noise stage" using only noisy single-modal weight data from real commercial pig farms, without relying on large-scale manual annotation or additional visual or 3D sensing hardware. It also provides auditable quality labels and equipment anomaly event signals, which can directly support production management decisions, equipment maintenance diagnosis, and subsequent predictive modeling, and has clear technical, economic, and management application value.
[0074] The method of this embodiment is applied to the self-supervised reconstruction of single-pig weight trajectory from automatic weighing data in fattening pig farms, including: 1) Data sources and experimental scenarios: The data in this example comes from multiple batches of fattening pigs from two commercial fattening farms and 20 pig houses, totaling 1392 fattening pigs. The age at entry into the pen was approximately 70±1 days, and most were slaughtered at ≥150 days. The electronic feeding system (EFS) continuously recorded individual-level weighing information, including at least: timestamp, age, weight reading, ear tag number, pen and house number.
[0075] To obtain the true body weight (truth-BW) for alignment assessment, an independent weighing scale was installed in each pen, and each pig was weighed daily according to a fixed procedure to obtain the true body weight.
[0076] 2) Implementation process: Aggregate the raw data into multiple reading sets based on the "pig-day" ratio for each pig, and then perform the following steps sequentially: (1) Stable daily gating: Calculate the relative dispersion index of daily readings for a single pig. And use a threshold (default 0.05) to filter stable days as the source of skeleton anchor points; (2) Candidate reading set: Cluster the multiple readings of the same single pig on the same day and extract the cluster representatives to form a candidate set, retaining the uncertainty of multiple clusters on the same day; (3) Self-supervised skeleton construction: Select the most self-consistent branch on the stable day set based on the short-term feasible region consistency constraint, and output the skeleton and quality label (T / O / F). (4) Skeleton-guided adsorption repair: Using the skeleton as a teacher signal, intra-skeleton repair and extra-skeleton conservative expansion are performed on missing / unstable days to improve coverage and continuity; (5) Abnormal event labeling: Output abnormal event labels (red / orange) at both the individual and field scales for auditing and equipment diagnosis; (6) Intraday time period normalization (optional): Based on the readings in the candidate cluster of structured anchor points according to the preset time window, and the deviation ratio estimated according to the age stage, the readings of different time windows are mapped to a unified reference window scale to obtain the final daily scale body weight.
[0077] 3) Examples of key parameters (not limited): Stable daily threshold: τ = 0.05; Short-term feasible region thresholds: 5-day window P95≈0.28, 2-day window P95≈0.22, 2-day window P90≈0.17; Examples of intraday time windows: 20:00–04:00, 04:00–07:00, 07:00–15:00, 15:00–20:00; Example of a reference window: 04:00–15:00.
[0078] 4) Output results: The output includes: daily weight trajectory of a single pig ( final_weight The system includes pig-day quality tags (T / O / F) and individual / pen-scale anomaly markers (red / orange), and can include candidate clusters and time window sources for each day for traceability auditing.
[0079] 5) Comparative examples and experimental results (examples): To verify the effectiveness of this invention, daily body weight trajectories were generated using six baseline methods and compared with truth-BW alignment. The results are as follows: Figures 14-15 As shown, these include: MID, IQR-MID, LOWESS, SG, Kalman, and Kalman RTS.
[0080] Evaluation metrics include: Coverage: the percentage of effective days where the reconstructed body weight falls within ±5% of the true value; Error metrics: MAE, RMSE, MAPE, R² (calculated in alignment with truth-BW); Audit consistency: the alignment rate between abnormal events and production calibration logs.
[0081] Example of experimental results: During the 70–90 day period, the coverage of this invention was approximately 98%, achieving the predicted results and reaching MAE≈1.11 kg, RMSE≈1.48 kg, MAPE≈2.65%, and R 2 ≈0.951; In terms of device auditing, the alignment rate between identified red events and calibration logs is approximately 93.92% (170 / 181).
[0082] This embodiment also provides a self-supervised reconstruction system for single-pig weight from noisy automatic weighing data, including: The stable day gating module is used to acquire individual-level weighing records of a single pig and perform standardized processing to obtain a set of multiple weight readings of a single pig per day and generate stability indicators. The candidate reading set construction module is used to obtain a candidate reading set for a single pig per day based on the multi-weight reading set for a single pig per day. The self-supervised skeleton construction module is used to select candidate readings that satisfy the feasible region consistency constraint from the candidate reading set of single pig-day on the stable day set, construct and initialize the single pig self-supervised skeleton using a multi-level branching strategy, and output quality labels for single pig-day. The skeleton-guided adsorption repair module is used to repair the missing / unstable days of the initialized single-pig self-supervised skeleton and obtain the reconstructed day-scale body weight sequence. Anomaly identification module is used to output anomaly event markers at both the individual and column scales. The intraday time-period normalization module is used to structure the raw readings within the candidate cluster according to time windows and perform phased bias correction to generate the final diurnal body weight. The output module is used to input the reconstructed daily weight sequence into a machine learning-based prediction model and output the daily weight trajectory of a single pig, quality labels, and abnormal event markers.
[0083] This embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data.
[0084] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data, characterized in that, include: Obtain individual-level weighing records for a single pig and perform standardization processing to obtain a set of multiple weight readings for a single pig per day; Based on the set of multiple body weight readings for a single pig per day, a stability index and a set of candidate readings for a single pig per day are obtained, wherein the stability index is used to divide a single pig per day into stable days and unstable days; On the stable day set, candidate readings that satisfy the feasible region consistency constraint are selected from the candidate reading set of single pig-day, and a multi-level branching strategy is used to construct an initial single pig self-supervised skeleton, and a quality label is output for single pig-day; The missing / unstable days of the initial single-pig self-supervised skeleton were repaired to obtain the reconstructed day-scale body weight sequence.
2. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 1, characterized in that, Based on the set of multiple body weight readings per pig per day, the stability index and the candidate set of readings per pig per day are obtained, including: Calculate the robust center of body weight readings based on the set of multiple body weight readings for a single pig per day; The median absolute deviation is calculated based on the set of multiple body weight readings for a single pig per day and the robust center of the body weight readings. The stability index is obtained by the ratio of the median absolute deviation to the sum of the robust center of the weight reading and a preset constant. Density clustering is performed on the set of multiple weight readings of a single pig per day to obtain several clusters. The actual readings corresponding to the density peak of each cluster are extracted as cluster representatives to obtain the candidate set of readings of the single pig per day.
3. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 2, characterized in that, Density clustering of the single-pig-day set of multiple body weight readings includes: Calculate the robust center for each age in each pen and multiply it by the scaling factor to obtain the adaptive clustering tolerance for weight stage. Perform density clustering on the set of multiple weight readings of a single pig-day based on the clustering tolerance to obtain several clusters.
4. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 3, characterized in that, On the stable day set, candidate readings that satisfy the feasible region consistency constraint are selected from the candidate reading set of single-pig-day readings, and a multi-level branching strategy is used to construct the initialization single-pig self-supervised skeleton, including: S1. Obtain the initial weight value of a single pig for a single day based on the distance between the candidate reading set of a single pig per day and the corresponding robust center; S2. On the stable day set, the short-term relative change of weight readings is quantified by a sliding window of a preset length as the consistency constraint of the feasible region. S3. Scan backward from the new age to the old age, select candidate readings that satisfy the feasible region consistency constraint from the candidate reading set of single pig-day, and attach them to the existing layer; S4. When the candidate reading does not meet the feasible region consistency constraint of the existing branch, a new branch layer is created within a sliding window of a preset number of days. If the number of layers reaches the preset value, return to S3; otherwise, proceed to S5. S5. Align the branches between segments and snap the short segments according to the proximity trend, update the initial weight value of the single pig-day, and obtain the initial single pig self-supervised skeleton, wherein the branch layer containing more than or equal to the preset number of consecutive points is regarded as a long growth segment, and the branch layer containing less than the preset number of consecutive points is regarded as a short segment. S6. The initialized single-pig self-supervised skeleton is verified by a sliding window with a length less than the preset day to obtain the final initialized single-pig self-supervised skeleton.
5. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 4, characterized in that, Inter-segment alignment of branch layers includes: The longest segment containing the most points is used as the anchor point to fit the local linear trend, and the remaining long segments are regarded as candidate long segments to be matched. Alignment is performed by minimizing the mean absolute residual of each candidate long segment relative to the trend line. After aligning long segments, the residuals of short segments are compared with the trends of adjacent long segments, and the long segment with the best trend match is selected as the alignment target for alignment.
6. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 5, characterized in that, Repairing the missing / unstable days of the initialized single-pig self-supervised scaffold includes: S1. Fit the curve of the reliable skeleton points of the final initialized single pig self-supervised skeleton, perform bandwidth screening on the candidate set of lonely days and unstable days, select the points that fall within the preset bandwidth range of the curve, and if there are multiple points, select the point that minimizes the residual of the previous and next neighborhoods to update the daily scale weight, and obtain the single pig self-supervised skeleton after adsorption in the skeleton. S2. For each target day outside the final initialized single-pig self-supervised skeleton range, based on the nearest anchor points of the single-pig self-supervised skeleton after adsorption within the skeleton, fit the local linear trend, construct a conservative feasible band to search the candidate reading set of the single-pig-day, select the candidate points falling within the band and use them as new anchor points, repeat S2 until there are no new anchor points, and obtain the single-pig self-supervised skeleton after exoskeleton adsorption, thus reconstructing the daily scale weight sequence.
7. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 6, characterized in that, After obtaining the skeletal adsorption within the skeleton, the following steps are also included: The growth period is divided into several age stages. The weight gain sequence of pigs in each age stage is estimated based on the self-supervised skeleton of a single pig after adsorption in the skeleton. The weight gain sequence is clustered, and the adaptive fluctuation radius is used as the neighborhood radius. The largest cluster is taken as the normal weight gain pattern, and the remaining clusters and noise points are taken as abnormal weight gain / loss events on the daily scale. If the number of valid pigs in the pigpen is greater than the preset value, the daily median of the short window variance term of all pigs in the pigpen is calculated, and the IQR rule is used to analyze the daily median to identify candidate abnormal days.
8. The method for self-supervised reconstruction of single-pig weight from noisy automatic weighing data according to claim 1, characterized in that, Before inputting the reconstructed diurnal body weight sequence into the machine learning-based prediction model, the following steps are included: The reliable skeleton points in the reconstructed daily weight sequence are used as high-confidence daily anchor points. The original readings corresponding to the high-confidence daily anchor points are traced back to the anchor cluster, and the original readings and their timestamps in the anchor cluster are used for intraday structured representation and correction.
9. A self-supervised reconstruction system for single-pig weight based on noisy automatic weighing data, used to implement the method as described in any one of claims 1-8, characterized in that, include: The stable day gating module is used to acquire individual-level weighing records of a single pig and perform standardized processing to obtain a set of multiple weight readings of a single pig per day and generate stability indicators. The candidate reading set construction module is used to obtain a candidate reading set for a single pig per day based on the multi-weight reading set for a single pig per day. The self-supervised skeleton construction module is used to select candidate readings that satisfy the feasible region consistency constraint from the candidate reading set of single pig-day on the stable day set, construct and initialize the single pig self-supervised skeleton using a multi-level branching strategy, and output quality labels for single pig-day. The skeleton-guided adsorption repair module is used to repair the missing / unstable days of the initialized single-pig self-supervised skeleton and obtain the reconstructed day-scale body weight sequence.
10. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the method described in any one of claims 1-8.