Pork pig weight prediction method and device based on multiple models and storage medium
By extracting audio, video, and text features from multi-source data and optimizing the prediction model in conjunction with environmental parameters, the problem of relying on manual judgment for monitoring the weight of pigs has been solved, achieving high-precision and automated weight prediction.
Patent Information
- Application Number
- CN202511693820.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-02-24
AI Technical Summary
Current technologies for monitoring the weight of pigs rely on manual judgment and lack unified objective standards, resulting in large errors in weight estimation data and poor monitoring effectiveness.
By extracting key audio and video segments and effective text data of stable pig activity from multi-source associated datasets, integrating multi-source effective data, using a small model to extract visual, audio and text features, quantifying and generating multimodal effective feature vectors, and fusing these features in a multimodal large model base, combined with environmental feeding text embedding vectors, the target weight prediction model is optimized, and finally, a description of pig weight estimation is generated through a time-series aggregation strategy.
It improves the automation, scenario adaptability, and prediction accuracy of pig weight prediction, reduces the cost of manual intervention, and enhances prediction accuracy and process efficiency.
Smart Images

Figure CN121561305A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of livestock and poultry breeding technology, and in particular to a method, device and storage medium for predicting the weight of meat pigs based on multiple models. Background Technology
[0002] In large-scale pig farming, monitoring pig weight is crucial for optimizing feed formulation, managing growth cycles, and determining slaughter time. The accuracy of this data directly impacts resource allocation and efficiency control in the farming process. Currently, related technologies rely heavily on subjective judgment by farmers based on pig body shape. The accuracy of this estimation method depends entirely on the farmers' experience, lacking a unified and objective standard. This can easily lead to estimated weights with errors far exceeding practical application requirements, resulting in ineffective pig weight monitoring.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, device and storage medium for predicting the weight of fattening pigs based on multiple models, in order to solve the technical problem of poor monitoring effect of fattening pig weight.
[0005] To achieve the above objectives, this application proposes a multi-model-based method for predicting the weight of fattening pigs, the method comprising: From the cleaned multi-source associated dataset, key audio and video segments and effective text data of stable pig activity are extracted and integrated to obtain multi-source effective data. Visual features, audio features, and text features are extracted from the multi-source effective data using a small model, and then quantized to generate a multimodal effective feature vector. Based on the multimodal large model base, the effective feature vectors of the multimodal model and the environmental feeding text embedding vectors are fused together to collaboratively optimize the output target weight prediction model; The target weight prediction model generates a description of the estimated weight of the pigs by analyzing the continuous periodic prediction results through a time-series aggregation strategy based on the multimodal effective feature vectors.
[0006] In one embodiment, audio and video data showing stable pig activity status are identified and filtered from the cleaned multi-source associated dataset, and multi-source candidate data are generated by combining environmental parameters and feeding log text data. By analyzing the rate of change and audio energy intensity between the audio frames summarized from the multi-source candidate data, the key audio and video segments are extracted. The key audio and video segments are matched with corresponding valid text data according to the data collection timestamp, and integrated to generate the multi-source valid data.
[0007] In one embodiment, the audio and video key segments and effective text data in the multi-source effective data are classified through the processing branch of the small model to obtain the classified audio and video key segments and effective text data; Based on the processing branch of the small model, the visual features, audio features, and text features are extracted from the classified audio and video key segments and effective text data, respectively. The visual features, audio features, and text features are quantized and compressed, and the three types of features are aligned according to the data acquisition timestamp to obtain the multimodal effective feature vector.
[0008] In one embodiment, the multimodal effective feature vector and the environmental feeding text embedding vector are input into the multimodal large model base, and feature encoding and association mapping are performed on the multimodal effective feature vector and the environmental feeding text embedding vector to output a hybrid feature vector; Based on the multimodal large model base, the feature dimensions of the hybrid feature vector are unified and then processed into a matrix to obtain a cross-modal feature matrix. The cross-modal feature matrix and group body size parameters are processed through group adaptation multi-task enhancement to output the target weight prediction model.
[0009] In one embodiment, a multi-task training objective is constructed based on the cross-modal feature matrix and the group body shape parameter annotation, and the weight of extreme scenario samples in training is increased to output an initial weight prediction model; The features obtained from the initial weight prediction model are distilled and then transferred to the parameter template of the adapted pigpen scenario to generate an adapted weight prediction model. The model parameters of the adapted weight prediction model are adjusted based on the environmental feedback from the target pigpen to obtain the target weight prediction model.
[0010] In one embodiment, the target weight prediction model is used to infer and calculate the multimodal effective feature vectors within the collection period, and output the predicted weight of the pig population for each period. A time-series aggregation strategy was used to perform correlation analysis on the predicted weight values of the pig population over consecutive periods, identify and remove outliers that deviate from the normal fluctuation range, and obtain an effective weight prediction sequence. Based on the effective weight prediction sequence, combined with the number of fattening pigs in the pigpen and the characteristics of the growth stage, the sequence mean is calculated and a reasonable error range is determined to generate the estimated weight description of the fattening pigs.
[0011] In one embodiment, by collecting audio and video data, environmental parameters, and feeding logs of the pigs in the pen, the total calibrated weight of the pigs at the final slaughter time is recorded synchronously, and the raw multi-source data is output. Based on the timestamps of the data collection in the original multi-source data, the audio and video data are associated and matched with the environmental parameters, feeding logs and overall calibrated weight at the corresponding time to form a multi-source associated dataset. The small model identifies and removes empty columns, severely blurry audio and video, incomplete text records, and incorrectly labeled weight data from the multi-source associated dataset to obtain the cleaned multi-source associated dataset.
[0012] In one embodiment, based on the estimated weight description of the pigs and the actual slaughter weight, the deviation rate between the estimated weight and the actual weight is calculated, and the error distribution characteristics are analyzed in conjunction with the reliability description to generate an error assessment report. By using the error assessment report, the source of error can be located, the model module that needs to be optimized can be determined, and the direction of model optimization can be obtained. Based on the model optimization direction and the training data supplemented by error samples, adjust the parameters of the corresponding model modules that need to be optimized, verify and output the iteratively optimized target weight prediction model.
[0013] In addition, to achieve the above objectives, this application also proposes a pig weight prediction device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multi-model-based pig weight prediction method described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-model-based method for predicting the weight of fattening pigs as described above.
[0015] This application provides a multi-model-based method for predicting the weight of fattening pigs. The method involves first associating and cleaning collected multi-source data; then extracting key audio-visual segments and effective text data representing stable pig activity from the cleaned multi-source associated dataset and integrating them into multi-source effective data; next, using a small model to extract visual, audio, and text features from this multi-source effective data and quantifying them to generate multi-modal effective feature vectors; subsequently, fusing these multi-modal effective feature vectors with environmental feeding text embedding vectors based on a multi-modal large model base and collaboratively optimizing the output target weight prediction model; finally, using this target weight prediction model combined with the multi-modal effective feature vectors, and analyzing the continuous periodic prediction results through a time-series aggregation strategy to generate a description of the estimated weight of the fattening pigs. This method solves the technical problems of traditional fattening pig weight prediction, such as low efficiency due to reliance on manual measurement, insufficient multi-modal data fusion, poor adaptability to extreme scenarios, and insufficient prediction accuracy. It improves the automation level, scenario adaptability, and prediction accuracy of fattening pig weight prediction, while reducing the cost of manual intervention.
[0016] In summary, this application solves the technical problem of poor monitoring effect of pig weight by collecting and cleaning multi-source data, extracting effective audio, video and text as multi-source effective data, generating multimodal vectors through small models, optimizing the prediction model by integrating environmental text into a large model, and obtaining the estimated weight description through time-series aggregation. This improves the automation level and process efficiency of pig weight prediction, while enhancing prediction accuracy and scenario adaptability. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the first embodiment of the multi-model-based method for predicting the weight of meat pigs according to this application; Figure 2 This is a flowchart illustrating the fifth embodiment of the multi-model-based method for predicting the weight of meat pigs in this application; Figure 3 This is a flowchart illustrating the seventh embodiment of the multi-model-based method for predicting the weight of fattening pigs in this application. Figure 4 This is a flowchart illustrating the eighth embodiment of the multi-model-based method for predicting the weight of meat pigs in this application; Figure 5 This is a schematic diagram of the structure of the pig weight prediction device of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] The relevant technology mainly relies on farmers' subjective judgment based on the size of the pigs. The accuracy of this estimation method depends entirely on the farmers' experience. It lacks a unified and objective judgment standard, which can easily lead to errors in the estimated weight data that far exceed the actual application requirements, resulting in poor monitoring of pig weight.
[0023] This application provides a solution: First, from a cleaned multi-source associated dataset, key audio and video segments and effective text data of stable pig activity are extracted and integrated to obtain multi-source effective data. Then, visual features, audio features, and text features in the multi-source effective data are extracted using a small model and quantified to generate multimodal effective feature vectors. Next, based on a multimodal large model base, the multimodal effective feature vectors and environmental feeding text embedding vectors are fused to collaboratively optimize and output a target weight prediction model. Finally, based on the multimodal effective feature vectors, the target weight prediction model analyzes the continuous periodic prediction results using a time-series aggregation strategy to generate a description of pig weight estimation.
[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a pig weight prediction device. The following description uses a pig weight prediction device as an example to illustrate this embodiment and the subsequent embodiments.
[0025] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0026] This application provides a multi-model-based method for predicting the weight of fattening pigs, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the multi-model-based method for predicting the weight of fattening pigs according to this application.
[0027] In this embodiment, the multi-model-based method for predicting the weight of fattening pigs includes steps S10 to S40: Step S10: Extract key audio and video clips and effective text data of stable pig activity from the cleaned multi-source associated dataset, and integrate them to obtain multi-source effective data.
[0028] In this embodiment, the multi-source associated dataset after cleaning refers to a collection of various data such as audio and video, environmental parameters, and feeding logs, which have been filtered to remove invalid and disordered data and associated and bound according to a unified identifier. Key audio and video clips refer to unobstructed audio and video segments with consistent postures and no drastic fluctuations in activity. Valid text data refers to complete and accurate text records related to the environment and feeding. Multi-source valid data refers to the unified set of integrated, stable key audio and video clips and valid text data.
[0029] As an optional implementation method, the multi-source associated dataset after cleaning is parsed frame by frame and segment by segment. By analyzing the stability of the audio and video footage, the continuity of the main elements, and the logical integrity of the text data, key audio and video segments with stable pig activity and no interference information, as well as valid text data without omissions or contradictions, are selected. Then, based on the time dimension of data collection and the associated attributes, one-to-one matching is performed, and data with invalid matches is removed to complete the integration and obtain valid multi-source data. This method has high data screening accuracy and strong correlation, but the processing flow is cumbersome. It can provide a high-quality data foundation for subsequent feature extraction and reduce errors caused by feature noise.
[0030] As an alternative implementation, a grouped parallel processing mechanism is used to split the multi-source associated dataset for cleaning. Simultaneously, activity stability threshold detection is performed on the audio and video data in each group, and validity rule verification is applied to the text data. This quickly extracts key audio and video segments and valid text data that meet basic standards. A global data index is established to achieve cross-group data association and matching, directly integrating all data that meet the conditions to form multi-source valid data. This method is highly efficient and time-saving, but its filtering granularity is relatively coarse. It can quickly output multi-source valid data to adapt to high-efficiency processing scenarios and improve the overall workflow speed.
[0031] Step S20: Extract visual features, audio features, and text features from the multi-source effective data using a small model, and quantize to generate a multimodal effective feature vector.
[0032] In this embodiment, "small model" refers to a lightweight feature extraction model adapted to multi-source data processing. Visual features refer to representational information related to the body shape and posture of pigs extracted from audio and video. Audio features refer to sound representational information related to pig activities extracted from audio and video. Text features refer to semantic representational information related to environmental feeding extracted from valid text. Multimodal effective feature vector refers to a unified feature carrier formed by fusing visual, audio, and textual quantified features.
[0033] As an optional implementation, multi-source valid data is split by type, with key audio and video segments sequentially sent to the visual extraction branch, audio extraction branch, and valid text data sent to the text extraction branch. Feature details are then mined in depth for each category, and each feature category is individually standardized and quantified. The dimensions and structures of various features are aligned according to the data acquisition sequence, and integration is achieved through feature concatenation to generate a multimodal valid feature vector. This method offers detailed feature extraction and accurate representation, but the processing time is relatively long. It can retain key feature information to the greatest extent and improve the accuracy of model prediction.
[0034] As an alternative implementation, from a cleaned multi-source associated dataset, a dynamic threshold for activity stability is set based on the growth stage of the pigs. Key audio and video segments that meet the threshold are extracted. Simultaneously, core text fields are prioritized according to growth stage, and integrated to obtain effective multi-source data for different stages. Visual features are extracted using a small model, and the feature extraction weights are dynamically adjusted based on the pig's growth stage. Audio features are clustered using the small model and assigned recognition weights for different growth stages. Then, text features are parsed using the small model and associated with growth stage labels. After standardizing and quantizing the three types of features, feature fusion weights are dynamically allocated according to the growth stage, generating a multimodal effective feature vector with stage-adaptive weights.
[0035] Step S30: Based on the multimodal large model base, the effective feature vectors of the multimodal model and the environmental feeding text embedding vectors are fused to collaboratively optimize the output target weight prediction model.
[0036] In this embodiment, the multimodal large model base refers to the infrastructure that supports cross-modal feature fusion and model training optimization. The environmental feeding text embedding vector refers to the semantic representation vector converted from text such as environmental parameters and feeding logs. The target weight prediction model refers to the model that, after fusion optimization, can output the predicted weight of meat pigs.
[0037] As an optional implementation, the multimodal effective feature vectors and environmental feeding text embedding vectors are first subjected to dimensionality encoding and semantic enhancement, respectively. The intrinsic correlation between the two types of vectors is then mined through a cross-modal attention mechanism, and this correlation information is integrated into the fusion module of the multimodal large model base. Multiple rounds of parameter iteration are then performed with weight prediction accuracy as the target. The feature fusion weights and model training strategy are continuously optimized, ultimately outputting the target weight prediction model. This method has sufficient feature fusion depth and strong model adaptability, but a long iteration cycle. It can fully leverage the synergistic effect of multi-source features, significantly improving the model's adaptability to complex scenarios and prediction accuracy.
[0038] As an alternative implementation, the effective feature vectors of the multimodal model and the environmental feeding text embedding vectors are directly processed to unify their dimensions. A feature concatenation method is used to quickly integrate the two types of vectors and input them into a large multimodal model base. Based on a preset basic training framework, the core fusion logic is fixed, and only a few iterative optimizations are performed on key parameters to quickly output a target weight prediction model. This method has a simple processing flow and fast iteration speed, enabling model construction and output in a short time, meeting the needs of application scenarios with high processing time requirements.
[0039] Step S40: The target weight prediction model generates a description of the estimated weight of the pigs by analyzing the continuous periodic prediction results through a time-series aggregation strategy based on the multimodal effective feature vector.
[0040] In this embodiment, the time-series aggregation strategy refers to a processing method that performs time-series correlation analysis on continuous data. Continuous period prediction results refer to the weight prediction values output by the model for multiple consecutive time periods. The estimated weight description for pigs refers to structured information containing the weight prediction results and related explanations.
[0041] As an optional implementation method, multimodal effective feature vectors are input into the target weight prediction model. After obtaining continuous periodic prediction results, the temporal continuity and logical consistency of each periodic prediction value are first checked, and abnormal data with abrupt changes in value or contradictions with the trends of previous and subsequent periods are identified and eliminated. Then, the remaining effective prediction values are hierarchically classified according to the completeness of data collection and the clarity of feature representation. A correlation mapping between each level is established through a temporal aggregation strategy, and a weighted calculation is performed by combining the importance weight of the level and the temporal weight of the period to obtain a comprehensive estimated value. At the same time, the deviation distribution range and temporal fluctuation amplitude of the prediction values at each level are statistically analyzed, and supplementary data confidence rating, error source analysis, and applicable scenario descriptions are provided, ultimately generating a comprehensive and detailed description of the estimated weight of meat pigs. This method thoroughly eliminates anomalies, has rigorous aggregation logic, and fully mines information.
[0042] As an alternative implementation, the multimodal effective feature vectors are input into the target weight prediction model to directly obtain continuous periodic prediction results. A parallelized temporal aggregation strategy is used to synchronously statistically analyze all predicted values, quickly integrating them by capturing the overall trend and concentrated features of the predicted values, and simultaneously extracting the core change node information within the prediction period. This directly generates a weight estimation description of the pigs containing the core estimated weight, overall change trend, and key node descriptions. This method has a simple processing flow, high operating efficiency, requires no complex calculations, and consumes few resources.
[0043] For example, in the scenario of predicting the weight of fattening pigs, the stability of audio and video footage and the integrity of text logic are analyzed from a multi-source associated dataset after cleaning. Key audio and video segments showing stable pig activity and valid environmental feeding text data without contradictions are extracted and integrated by timestamp to obtain multi-source valid data. This data is then processed by a lightweight feature extraction small model to extract visual features of body shape and posture, audio features of activity sounds, and semantic text features. After standardization and quantization, the dimensions are aligned to generate a multimodal valid feature vector. This vector and the environmental feeding text are embedded into the Transformer multimodal large model base. Features are fused through a cross-modal attention mechanism, and multi-task training is carried out in conjunction with the annotation of group body shape parameters to improve the weight of extreme scenario samples. After online distillation and adaptation, the target weight prediction model is output through collaborative optimization. This model outputs continuous periodic prediction results based on the multimodal valid feature vector. A sliding window temporal aggregation strategy is used to filter out outliers and integrate them to generate a fattening pig weight estimation description that includes the estimated weight and reliability description.
[0044] By cleaning and integrating multi-source data, extracting and quantifying features from small models, optimizing cross-modal fusion of large models, and performing time-series aggregation analysis, the problems of low efficiency in traditional manual measurement, insufficient utilization of multimodal data, and inaccurate prediction in extreme scenarios have been solved, thereby reducing labor costs and scenario adaptation errors.
[0045] Based on any of the above embodiments, in Embodiment 2 of this application, step S10 includes steps A11 to A13: Step A11: From the cleaned multi-source associated dataset, identify and filter audio and video data showing stable pig activity status, and combine them with environmental parameters and feeding log text data to generate multi-source candidate data.
[0046] In this embodiment, audio and video data refers to audio and video recordings with no violent shaking, no obstruction of the subject, and no significant fluctuations in the range of motion. Environmental parameters refer to various environmental attribute information recorded related to feeding. Feeding log text data refers to text records detailing feeding operations, the status of the pigs, etc. Multi-source candidate data refers to a data set to be further processed, formed by integrating and filtering stable audio and video data, environmental parameters, and feeding log text data.
[0047] As an optional implementation method, the audio and video data in the multi-source associated dataset after cleaning are first analyzed segment by segment. By analyzing the continuity of the footage, the presence of the main subject, and changes in the amplitude of activity, segments with stable pig activity are identified. Then, corresponding environmental parameters and feeding log text data are extracted from the dataset, and each segment is associated with the audio and video segments according to the collection timestamp. Logical consistency is verified on the associated data. Data with invalid associations or contradictory information is removed, and finally, multi-source candidate data is integrated. This method has strict selection criteria, strong data correlation, and high information consistency.
[0048] As an alternative implementation, the multi-source associated dataset after cleaning is classified and split. Simultaneously, preliminary stability checks on audio and video data and integrity verification of environmental parameters and feeding log text data are initiated to quickly identify audio and video data that meet basic stability conditions and text and parameter data that meet integrity standards. A global time index is established to achieve rapid matching and association of the three types of data. Without additional deep logical verification, all successfully matched data are directly integrated to form multi-source candidate data. This method features a simple processing flow, high parallel operation efficiency, and short processing time, ensuring efficient workflow.
[0049] Step A12: By analyzing the rate of change and audio energy intensity between the audio frames of the multi-source candidate data, the key audio and video segments are extracted.
[0050] In this embodiment, the rate of change between audio frames refers to the degree of difference in signal characteristics between consecutive audio frames. Audio energy intensity refers to the magnitude of energy carried by the audio signal. Key audio-visual segments refer to the key segments extracted from audio-visual data that contain core and effective information and are suitable for subsequent processing requirements.
[0051] As an optional implementation, audio and video data are first separated from multi-source candidate data. The audio portion is extracted separately and split into a continuous frame sequence in chronological order. The signal feature differences between adjacent audio frames are calculated frame by frame to obtain the rate of change between audio frames. Simultaneously, the energy representation of the audio signal in each frame is calculated to obtain the audio energy intensity. Based on the joint judgment criteria of the two, continuous audio frame sequences with stable rates of change and energy intensities within a reasonable range are selected. Then, the corresponding segments in the audio and video data are located according to the timestamp of the audio sequence. The information integrity of the segments is verified, and fragmented parts and parts missing core information are removed. Finally, the key audio and video segments are extracted. This method has rigorous extraction logic, high concentration of core information in key segments, and low redundancy, improving the efficiency and accuracy of subsequent model training.
[0052] As an alternative implementation, audio and video data are separated from multi-source candidate data, and the audio portion is extracted. The audio is segmented according to a fixed duration, and the average inter-frame change rate and average energy intensity of each audio segment are calculated in batches. Audio segments that meet the requirements for average change rate and average energy intensity are quickly selected directly based on preset basic screening conditions. The corresponding segments in the audio and video data are directly matched according to the time range of the audio segments, and the matched segments are directly extracted and output as key audio and video segments. This method has a simple processing flow, does not require frame-by-frame calculation, has high batch processing efficiency, and short processing time, ensuring efficient workflow.
[0053] Step A13: The key audio and video segments are matched with corresponding valid text data according to the data collection timestamp, and integrated to generate the multi-source valid data.
[0054] In this embodiment, the data acquisition timestamp refers to the time identifier that records the specific moment of data acquisition, such as audio, video, and text. Valid text data refers to complete and accurate text records related to animal husbandry, such as environmental parameters and feeding logs.
[0055] As an optional implementation, the start and end range of the timestamps for each key audio / video segment is first extracted. Then, all text records whose timestamps fall within this range are selected from the valid text data. The key audio / video segments and corresponding text records are aligned one by one according to their timestamps. The aligned data undergoes information correlation verification to check whether the scene represented by the audio / video is logically consistent with the content of the text records, eliminating data that matches in time but contains contradictory information or has invalid correlation. The verified key audio / video segments and valid text data are reordered along the time dimension and finally integrated to generate multi-source valid data. This method offers accurate time matching, strong data correlation, and consistent information logic.
[0056] For example, in the scenario of predicting the weight of fattening pigs, from a cleaned multi-source associated dataset (containing 150,000 hours of audio and video, 8 million environmental parameter records, and 300,000 feeding logs), 50,000 hours of key audio and video clips showing stable activities such as feeding, resting, and uniform activity of fattening pigs are selected. 250,000 pieces of effective text data containing core information such as feed amount, disease prevention details, and temperature and humidity compliance records are extracted and integrated by timestamp to obtain multi-source effective data. A YOLO mini-model is used to extract visual features such as the body shape, limb movements, and group density of fattening pigs from the audio and video. An MFCC feature extraction mini-model captures audio features such as activity sounds and environmental noise. A lightweight BERT mini-model parses text features such as feeding operations and growth stage annotations from the text data. The three types of features are standardized and quantized to generate a 128-dimensional multimodal effective feature vector. The effective multimodal feature vector and the 64-dimensional environmental feeding text embedding vector encoded by BERT are input into the Transformer-based multimodal large model base. The weight correlation logic of the bimodal features is mined through the cross-attention mechanism. A training framework with accurate weight prediction as the main task and group body size adaptation judgment as the auxiliary task is constructed. The feature fusion weight and model parameters are iteratively optimized. After 10 rounds of verification, the target weight prediction model is output. The target weight prediction model receives multimodal effective feature vectors collected every 24 hours. After 30 consecutive cycles of inference, it obtains the corresponding weight prediction value for the pig population. The model uses a time-series aggregation strategy to analyze the temporal consistency and fluctuation pattern of the prediction values. Two outliers caused by temporary equipment interference are removed, resulting in 28 effective weight prediction sequences. Combining the number of pigs in the target pigpen (600 heads) and their growth characteristics during the growth period, the model calculates the sequence mean and determines a reasonable error range of ±1.5%. This generates a pig weight estimation description: "The current total weight of the pig population is estimated at 36 tons, the average weight is 60 kg, the error range is ±1.5%, the weight gain is in line with the normal growth rate during the growth period, and it is suitable for the current feeding program."
[0057] By first screening stable audio and video, then extracting audio features to locate key segments, and finally matching text data by timestamp, the process solves the problems of redundant multi-source data, inefficient extraction of key information, and loose data association, thereby improving data processing efficiency and data purity.
[0058] Based on any of the above embodiments, in Embodiment 3 of this application, step S20 includes steps B11 to B13: Step B11: Through the processing branch of the small model, classify the audio and video key segments and effective text data in the multi-source effective data to obtain the classified audio and video key segments and effective text data.
[0059] In this embodiment, the processing branch of the small model refers to an independent functional module within the small model specifically responsible for classifying and processing different types of data. Key audio / video segments refer to key segments containing core, valid information extracted from audio / video data. Classified key audio / video segments and valid text data refer to sets of audio / video segments and text data that, after being split according to data type, are respectively assigned to their corresponding processing categories.
[0060] As an optional implementation, a global analysis of the multi-source valid data is first performed to extract core attributes such as format features and content identifiers for each data point. These attributes are then compared one by one with the adaptation standards of each processing branch of the small model. Based on the comparison results, key audio and video segments are accurately assigned to the audio and video processing branch, and valid text data is assigned to the text processing branch. Subsequently, a type consistency check is performed on the data received by each branch to remove misclassified or cross-type mixed data, ultimately obtaining the classified key audio and video segments and valid text data. This method boasts high classification accuracy, high data type purity, and no confounding interference.
[0061] As an alternative implementation, a pre-defined data source format identifier is used for each processing branch of the small model. This allows for direct surface-level format screening of multi-source valid data, quickly determining the data type based solely on surface features such as storage format and file identifier. Data conforming to audio / video format characteristics is directly assigned to the audio / video processing branch, and data conforming to text format characteristics is directly assigned to the text processing branch. The data received by each branch is then directly output as the classified audio / video key segments and valid text data. This method features a simple processing flow, fast classification speed, short processing time, and low resource consumption.
[0062] Step B12: Based on the processing branch of the small model, extract the visual features, audio features, and text features from the classified audio and video key segments and effective text data, respectively.
[0063] As an optional implementation, the audio / video processing branch of the small model performs hierarchical analysis on the classified key audio / video segments. First, it separates the video frame sequence and audio signal. For each video frame, it extracts visual features by mining related representations such as subject shape, outline, and posture. For the audio signal, it analyzes related representations such as frequency, amplitude, and rhythm to extract audio features. Simultaneously, the text processing branch of the small model performs semantic decomposition on the effective text data, mining related representations such as word associations, information dimensions, and core semantics to extract text features. Throughout the entire process, detailed verification is performed on the extraction process for each type of feature to ensure the completeness and accuracy of the feature representation. This method offers sufficient feature extraction depth, comprehensive representation dimensions, and complete information preservation.
[0064] As an alternative implementation, the visual extraction branch, audio extraction branch, and text extraction branch of the small model are launched in parallel. The visual extraction branch quickly captures the overall shape and dynamic trend-related features of the classified key audio and video segments. The audio extraction branch captures the core frequency and energy change-related features of the audio signal. The text extraction branch extracts the core keywords and key information points related to the effective text data, simplifying the feature extraction process of each branch and retaining only the core representation information to complete feature extraction. This method has fast extraction speed, high parallel efficiency, low resource consumption, and short processing time, ensuring efficient process advancement.
[0065] Step B13: Quantize and compress the visual features, audio features, and text features, and align the three types of features according to the data acquisition timestamp to obtain the multimodal effective feature vector.
[0066] In this embodiment, quantization compression refers to the normalization operation on feature data, reducing data volume and unifying data format while retaining core information. Aligning the three types of features refers to the operation of establishing a one-to-one correspondence between visual, audio, and text features according to the acquisition time.
[0067] As an optional implementation, quantization compression methods adapted to the attributes of visual, audio, and text features are adopted respectively. During compression, core representation information is preserved while redundant data is removed. After compression, the original data acquisition timestamps corresponding to each feature type are extracted, and the three types of features are precisely matched one by one according to the chronological order of the timestamps to establish the correspondence between time and feature. Simultaneously, the integrity of the aligned features is verified, and feature interpolation for missing timestamps is supplemented, ultimately obtaining a multimodal effective feature vector. This method offers highly targeted quantization compression, complete preservation of core information, and high timestamp alignment accuracy, thereby improving the accuracy of model training.
[0068] As an alternative implementation, a unified quantization compression standard is used to batch process visual, audio, and text features, without distinguishing between feature attribute differences, quickly reducing data volume and unifying the format. Subsequently, the acquisition timestamps corresponding to all features are aggregated, and the three types of features are grouped at fixed time intervals. Features within the same time interval are automatically bound together, and each group of bound features is directly integrated into a vector form to obtain a multimodal effective feature vector. This method has simple processing logic, high batch operation efficiency, significantly reduces quantization and alignment time, and consumes few resources.
[0069] For example, in the scenario of predicting the weight of fattening pigs, model training and optimization based on pig farm environmental parameters are crucial. Factors such as light intensity, temperature, humidity, pen structure, and obstructions (e.g., feed troughs, waterers) in the pig farm environment directly affect image quality and feature extraction accuracy. Traditional visual models are mostly trained in standardized laboratory environments, and their fitting ability drops significantly after being transferred to pig farm scenarios, with estimation errors generally exceeding 10%. To address this issue, this paper proposes a three-pronged training scheme integrating "environmental parameters," "image features," and "weight labels": Multi-dimensional environmental data collection: Based on the lighting, pen structure, camera installation angle, and camera installation height of each pig pen in the pig farm, image data of fattening pigs and environmental parameters are collected simultaneously to construct an "environment-image-weight" dataset containing a massive number of samples. The weight labels are obtained through various methods, including manual weighing, slaughter weighing, and instrument-assisted weighing, and through scenario transfer for weight estimation, ensuring sufficient data volume. Adaptive feature enhancement: During model training, environmental parameters are used as additional input features and integrated into the feature extraction layer of a convolutional neural network (CNN). For example, to address the increased image noise in low-light environments (<500 lux), the model automatically activates the noise reduction module by learning the mapping relationship between light intensity and noise distribution. To address the increased lying posture of pigs in high-temperature environments (>30℃), the model strengthens the feature capture of the torso contour to avoid misjudgment of body shape due to posture changes. Cross-scene transfer learning: Based on parameter template technology, a federated learning framework is used to distribute the training of environmental data from different pig farms, achieving cross-scene adaptation of the model through parameter sharing. Experimental data shows that the model optimized with environmental parameters controls the weight estimation error in different pig farm scenarios and improves the model's fitting ability. Environmental parameters can be configured through the interface. Input video preprocessing based on a small model: The computing power and storage resources of edge devices are limited. Directly inputting the original video into a large model for processing would lead to high inference latency, failing to meet real-time weight estimation requirements. Furthermore, the original video contains many non-compliant samples (such as empty pen images, images severely occluded by pigs, and blurry images), which would reduce weight estimation accuracy if used directly for inference. To address this, this paper designs a preprocessing module based on a lightweight small model to achieve integrated processing of "sample screening, dimensionality reduction and compression, and preliminary feature extraction": Non-compliant sample removal: The small model is trained through a binary classification task to identify and remove three types of non-compliant samples: 1) Empty pen samples, determined by detecting the presence of pig outlines in the image; 2) Occlusion samples, determined by whether key parts of the pig are occluded, with excessively large occlusion areas marking them as invalid samples; 3) Blurred samples, removed by calculating image sharpness indices (such as variance gradient); and 4) Posture samples, removed by determining whether the pig is standing, lying down, or curled up, and whether its angle is sideways. Experiments show that this module can increase the effective sample rate from 65% of the original video to 98%, significantly reducing invalid inference overhead.Video frame dimensionality reduction and keyframe extraction: The small model uses an inter-frame difference algorithm to analyze the pixel change rate of adjacent video frames. When the change rate is <5%, it is identified as a redundant frame and removed, reducing the video frame rate from 30fps to 5fps. Simultaneously, it extracts the keyframes with the most stable pig posture (e.g., standing and unobstructed frames), reducing the data volume by 75% while ensuring feature integrity, thus shortening the edge-side inference latency. Preliminary feature extraction and quantization: The small model performs preliminary feature extraction on the selected keyframes, outputting a 256-dimensional feature vector. Quantization and pruning further reduce the model size to meet the storage requirements of edge devices, while providing high-quality feature input for subsequent large-scale model inference.
[0070] By employing a process of small-model branching classification, targeted feature extraction, and time-series aligned integration, the problems of chaotic classification of multi-source data, inefficient feature extraction, and asynchronous cross-modal feature integration are solved, thereby improving the training efficiency of the model.
[0071] Based on any of the above embodiments, in Embodiment 4 of this application, step S30 includes steps C11 to C13: Step C11: Input the multimodal effective feature vector and the environmental feeding text embedding vector into the multimodal large model base, perform feature encoding and association mapping on the multimodal effective feature vector and the environmental feeding text embedding vector, and output the hybrid feature vector.
[0072] In this embodiment, the environmental feeding text embedding vector refers to the semantic representation vector transformed from text data such as environmental parameters and feeding logs. Feature encoding refers to the operation of transforming the original features into a standardized form that the model can recognize and compute. Association mapping refers to the operation of mining the inherent logical relationships between different types of features and establishing corresponding relationships. Hybrid feature vector refers to a unified feature carrier formed by fusing two types of input features after encoding and association mapping.
[0073] As an optional implementation, the multimodal effective feature vectors and environmental feeding text embedding vectors are first input into the dedicated encoding module of the multimodal large model base. A differentiated encoding strategy is adopted for the feature attributes of the two types of vectors, preserving their respective core semantics and representational information. After encoding, a cross-modal association mining mechanism is initiated. By analyzing the semantic dimensions and representational logic of the two types of features layer by layer, a fine-grained association mapping relationship is established. The mapped features are then dimensionally aligned and weighted, eliminating redundant features with extremely low correlation, and finally integrated to form a hybrid feature vector. This method features highly targeted feature encoding, fine-grained association mapping, and sufficient fusion depth, improving the adaptability and accuracy of the target weight prediction model.
[0074] As an alternative implementation, a semantic dynamic adaptation layer is added to the multimodal large model base. This layer captures the dynamic changes of effective multimodal feature vectors through a temporal sliding window and mines the temporal semantic associations of the embedded text vectors through text temporal analysis. A semantic association adaptive model is established, dynamically adjusting the semantic mapping rules according to the growth stage and calculating the semantic matching degree of cross-modal features for each period. If the matching degree is lower than a preset threshold, a semantic completion mechanism is triggered, supplementing missing semantic links based on historical semantic association data from the same period. A feature fusion network integrates the dynamically adapted cross-modal semantic features with the original encoded features to generate a hybrid feature vector that is temporally semantically coherent and adapted to the growth stage, while simultaneously recording the semantic adaptation adjustment log.
[0075] Step C12: Based on the multimodal large model base, the feature dimensions of the hybrid feature vector are unified and matrix-processed to obtain the cross-modal feature matrix.
[0076] In this embodiment, matrix processing refers to the operation of arranging the regularized features in an orderly manner according to preset rules and transforming them into a structured matrix. A cross-modal feature matrix refers to a structured data set that carries various fused features after dimensionality unification and matrix processing.
[0077] As an optional implementation method, this approach first decomposes each sub-feature in the hybrid feature vector, statistically analyzing the number of dimensions, data value range, and distribution characteristics of each sub-feature. Based on the optimal fusion dimension standard preset by the multimodal large model foundation, sub-features with excessive dimensions undergo gradient dimensionality reduction using a core representation preservation strategy, while sub-features with insufficient dimensions undergo intelligent interpolation for dimensionality enhancement based on adjacent semantic associations, ensuring that all sub-feature dimensions are fully matched and the data distribution is coherent. Then, the sub-features are hierarchically sorted according to their association strength and temporal logic. The sorted feature data is then arranged in an orderly manner according to the rules of row-to-sample and column-to-dimension, and the data is standardized and converted one by one. After matrix processing, the consistency and completeness of the matrix data are checked row by row and column by column, correcting numerical outliers and supplementing missing data, ultimately obtaining the cross-modal feature matrix. This method features high accuracy with unified dimensions, a rigorous matrix structure, and reliable data quality, maximizing the preservation of core feature information.
[0078] As an alternative implementation, the fixed-dimensional template built into the multimodal large model base is directly invoked to batch read all sub-features in the mixed feature vectors. A unified linear scaling algorithm is then used to quickly adjust all sub-features to the target dimension set by the template. Subsequently, the features are batch-sorted according to their original acquisition time order, and the sorted feature data is directly mapped to the row and column structure of a preset matrix. Only basic format conversion is performed to complete the matrix processing, and the cross-modal feature matrix is quickly output. This method has a simple operation process, fast batch processing speed, low resource consumption, and can significantly shorten the processing cycle.
[0079] Step C13: The cross-modal feature matrix and the group body size parameters are processed through group adaptation multi-task enhancement to output the target weight prediction model.
[0080] In this embodiment, the group body size parameter refers to reference data reflecting the commonalities and differences in body size among the group of pigs. Group adaptation multi-task reinforcement processing refers to a method that combines group parameters and optimizes model adaptability through multi-task collaborative training.
[0081] As an optional implementation, this method analyzes the row-to-row feature meanings of the cross-modal feature matrix one by one, extracting the core feature dimensions. Then, the group body shape parameters are categorized by attribute and associated with the matrix features, separating the main task of weight prediction with auxiliary tasks such as body shape feature adaptation and group difference calibration. Group adaptation training is conducted hierarchically according to task priority, dynamically adjusting the weights of each task during training to enhance the collaborative representation ability of core features and group parameters. After each round of training, the fit between the prediction results and group parameters is compared, iteratively optimizing the model parameters to reduce bias. After multiple rounds of enhancement, the model's integrity is verified, invalid parameter modules are removed, and the final target weight prediction model is output. This method exhibits strong group adaptability, multi-task collaborative enhancement of model representation ability, and accurate parameter optimization, enabling it to meet the prediction needs of diverse body shape scenarios.
[0082] As an alternative implementation, the cross-modal feature matrix and group body shape parameters are directly surface-level spliced and fused. A fixed multi-task framework is preset, including the core task of weight prediction and the basic group adaptation task. A batch reinforcement training mode is used to input the fused data all at once, and training proceeds according to a preset learning rhythm. Only at the end of training is the model prediction results simply validated, and the core parameters are adjusted to meet the basic adaptation requirements, quickly completing model construction and outputting the target weight prediction model. This method has a simple processing flow, efficient fusion and training operations, significantly shortens the model output cycle, and consumes few resources.
[0083] For example, in the scenario of predicting the weight of fattening pigs, the necessity of using a large model base based on a multimodal large model and augmented training with fused text parameters is highlighted: Choosing large pre-trained models such as ViT and Swin Transformer as the base, rather than training a small model from scratch, is based on their superior general visual representation capabilities. These models, pre-trained on ultra-large-scale datasets, have learned to extract rich and robust features and have a deep understanding of shape, texture, and contextual relationships. This strong foundational capability is a prerequisite for their ability to quickly adapt to downstream tasks (such as pig weight estimation) and effectively handle diverse, time-varying data provided by preprocessing modules. Targeted training based on this foundation yields twice the results with half the effort. The base capabilities of the edge AI vision large model (such as general image feature extraction and 3D reconstruction capabilities) are the foundation for accurate weight estimation. However, for the specific target of fattening pigs, augmented training is also needed to improve the model's targeted recognition ability of fattening pig body shape features (such as trunk length, chest circumference, and hip circumference): Augmented training with fused text parameters: The core objective is to improve the model's fitting power and generalization ability across different pig farms. Environmental variations in pig farms (e.g., open / closed environments, bright / dim lighting, camera model, installation height, and angle) are major factors contributing to model instability. Relying solely on image data, the model struggles to adapt to these changes. We treat these pig farm environmental parameters as textual descriptions, encoding them as embedding vectors. Using multimodal fusion techniques (e.g., FiLM, Cross-Attention), these environmental embedding vectors are injected as conditional signals into the feature extraction process of the large-scale visual model. Body shape feature annotation and supervised training: During the dataset annotation phase, in addition to weight labels, key body shape parameters of the pigs (trunk length, chest circumference, hip circumference) are additionally annotated to construct multi-task training objectives. While learning the weight prediction task, the model optimizes the prediction accuracy of body shape parameters through a regression loss function, achieving a precise mapping between "body shape features and weight." Data Augmentation and Extreme Sample Training: Addressing the significant differences in body size (e.g., piglets, fattening pigs, breeding pigs) and posture (standing, lying down, side-lying) throughout the growth cycle of fattening pigs, data augmentation techniques such as random rotation, scaling, flipping, and occlusion are employed to expand sample diversity. Simultaneously, the training weights of extreme samples (e.g., extra-large fattening pigs, deformed fattening pigs) are increased to enhance the model's adaptability to special situations. Furthermore, the model effectively utilizes data from pigs with surrounding deformities. Video Temporal Optimization Strategy: To overcome the challenge of instantaneous errors in single-frame weight estimation, a temporal optimization strategy is introduced. This strategy comprehensively analyzes the weight estimation results of multiple consecutive frames within a longer time window (e.g., 30 seconds or 1 minute), eliminating outliers, aggregating valid values, and ultimately outputting a more accurate and stable weight estimate based on the comprehensive results over 24 hours. Online Distillation and Real-Time Optimization: Utilizing real-time data collected from edge devices, online distillation technology is used to transfer knowledge from a large model to a smaller model, enabling real-time updates of model parameters.For example, when the system detects that the weight estimation error of a certain pen is consistently high, it automatically collects environmental and image data of that pen, makes local fine adjustments, and quickly brings the error back to the normal level.
[0084] By employing a process of feature encoding association, dimension unification matrixing, and group adaptation multi-task enhancement, the problems of insufficient cross-modal feature fusion, poor model group adaptability, and insufficient prediction accuracy are solved. This process can accurately adapt to the weight prediction needs of different groups of meat pigs, significantly reducing the cost of manual weighing and prediction bias.
[0085] Based on any of the above embodiments, in Embodiment 5 of this application, referring to Figure 2 , Figure 2 This is a flowchart illustrating the fifth embodiment of the multi-model-based method for predicting the weight of fattening pigs in this application. Step C13 includes steps D11 to D13: Step D11: Based on the cross-modal feature matrix, construct a multi-task training objective by combining the group body shape parameter annotation, increase the weight of extreme scenario samples in training, and output the initial weight prediction model.
[0086] In this embodiment, the multi-task training objective refers to the set of training objectives that includes the main task of weight prediction and related auxiliary tasks such as body shape feature adaptation and scene adaptation. Extreme scene samples refer to data samples corresponding to unconventional situations such as special body shapes and complex environments. The initial weight prediction model refers to the basic weight prediction model that has not undergone subsequent optimization after the first multi-task training.
[0087] As an optional implementation, this method analyzes the feature dimensions, data distribution, and correlation strength with weight prediction of the cross-modal feature matrix one by one, decomposes the attributes of the group body shape parameter annotation, and distinguishes between conventional and special body shape annotations. Combining these two approaches, a multi-task training objective is constructed with accurate weight prediction as the primary task and body shape feature classification and scene adaptability judgment as auxiliary tasks. Extreme scene samples are filtered by analyzing the complexity of sample scenes and the degree of body shape deviation. Higher training weights are dynamically allocated according to the sample specificity level. The weighted samples and the multi-task training objective are input into the training framework, and the task collaboration logic and model parameters are adjusted iteratively in rounds. After each round of training, the prediction deviation of the primary task and the adaptability of the auxiliary tasks are verified, invalid training iterations are eliminated, and the initial weight prediction model is output. This method's multi-task objective design fits the data characteristics, the weight allocation of extreme scene samples is accurate, and the initial model representation ability is strong, laying a high-quality foundation for subsequent optimization.
[0088] Step D12: The features obtained from the distillation process of the initial weight prediction model are transferred to the parameter template of the adapted pigpen scenario to generate an adapted weight prediction model.
[0089] In this embodiment, distillation refers to the model compression and optimization operation of extracting core features and removing redundant parameters from the initial weight prediction model. Knowledge transfer refers to the process of transferring the core features and empirical parameters obtained from distillation to the target scenario parameter template. The parameter template refers to a pre-set model parameter framework for the environment, feeding conditions, and characteristics of the fattening pig population in a specific pigpen scenario. The adapted weight prediction model refers to the weight prediction model that can match the target pigpen scenario after knowledge transfer and scenario parameter adaptation.
[0090] As an optional implementation, a hierarchical distillation process is performed on the initial weight prediction model. The model structure is decomposed layer by layer to extract core feature layers and key parameters strongly correlated with weight prediction, while redundant modules irrelevant to scene adaptation are removed. Simultaneously, the parameter template adapted to the pigpen scene is comprehensively analyzed to determine the type, value range, and correlation logic between scene adaptation parameters and features. The core features obtained from distillation are semantically correlated with the hierarchical structure of the parameter template, and the fit between feature representations and template parameters is adjusted point by point. Iterative fine-tuning ensures the core features are fully embedded in the template framework. After each round of adjustment, the fit between features and scene parameters is verified, adaptation deviations are corrected, and a suitable weight prediction model is finally generated. This method yields features with high purity, accurate adaptation to the scene template, and strong model specificity.
[0091] As an alternative implementation, the initial weight prediction model undergoes modular distillation, separating it into a general feature extraction module, a weight mapping calculation module, and a scene adaptation interface module. Only the general feature extraction and weight mapping calculation modules are distilled, retaining the core calculation logic and feature association rules while removing scene-specific redundant parameters. A scene parameter template library is constructed, containing differentiated parameter templates for different feeding modes, pen sizes, and climate regions. Using a scene feature matching algorithm, the optimal reference template is selected, and parameter templates adapted to the pigpen scene are generated. A cross-scene feature adaptation layer is established, transforming the core features obtained from distillation into scene features. Simultaneously, a scene difference compensation factor is introduced to dynamically correct the binding relationship between features and templates, generating an adapted weight prediction model with a scene-adaptive compensation mechanism.
[0092] Step D13: Adjust the model parameters of the adapted weight prediction model based on the environmental feedback of the target pigpen to obtain the target weight prediction model.
[0093] In this embodiment, the target pigpen refers to a specific pigpen where weight prediction is required. Environmental feedback refers to feedback information such as real-time environmental data and changes in feeding status collected from the target pigpen. Model parameters refer to the adjustable configurations and values in the model that affect the prediction results.
[0094] As an optional implementation method, environmental feedback from the target pigpen is continuously collected and analyzed according to environmental type. Key fluctuation indicators related to weight prediction are extracted from the feedback, and these indicators are mapped to the corresponding parameter modules in the adapted weight prediction model. Parameters are adjusted in stages according to the fluctuation range of the indicators. Parameters associated with indicators that fluctuate drastically are iteratively fine-tuned, while parameters associated with stable indicators are slightly calibrated. After each round of adjustment, the prediction deviation is verified using real-time samples from the target pigpen. If the deviation exceeds a threshold, the parameter adjustment node is traced back and corrected. This process is repeated until the deviation stabilizes within a preset range, resulting in the target weight prediction model. This method has a close correlation between parameter adjustment and environmental feedback, a high degree of refinement, and extremely strong model adaptability.
[0095] For example, in the scenario of predicting the weight of fattening pigs, a multi-modal feature matrix is constructed based on the multimodal features of audio / video, environment, and feeding text, which are processed through dimensional unification and matrixing. This matrix is combined with group body shape parameter annotations that include the range of normal body shapes, special body shapes (underweight / overweight), and weight-related information. The main task is to accurately predict the weight of fattening pigs, while the auxiliary tasks are body shape classification and scenario adaptation judgment. 1200 sets of extreme scenario samples, such as high temperature and humidity environment, high density feeding, and extreme body shapes, are selected. Their training weights are increased to 2.5 times that of normal samples. They are then input into the ResNet+Transformer fusion framework for multiple rounds of iterative training. The collaborative logic of the main and auxiliary tasks and the model parameters are dynamically optimized. After integrity verification, the initial weight prediction model is output. The initial weight prediction model underwent hierarchical distillation to extract core feature layers and key parameters strongly correlated with weight prediction, eliminating redundant modules. Simultaneously, parameter templates for a suitable pigpen scenario, including parameters such as target pigpen stocking density, feed type, and ventilation standards, were analyzed. The core features obtained from distillation were semantically correlated with the template hierarchy, and feature embedding and binding were completed through iterative fine-tuning to generate a suitable weight prediction model. Environmental feedback data such as temperature and humidity, ammonia concentration, and ventilation rate in the target pigpen were continuously collected. Key fluctuation indicators were extracted and mapped to the corresponding parameter modules in the suitable weight prediction model. Parameters were adjusted in stages according to the fluctuation range of the indicators (large-scale fine-tuning for parameters with drastic fluctuations and small-scale calibration for parameters with stable indicators). After each round of adjustment, the prediction deviation was verified using real-time samples from the target pigpen until the deviation stabilized within a preset range, thus obtaining the target weight prediction model.
[0096] By enhancing extreme scenario adaptation through multi-task enhancement, accelerating scenario implementation through distillation and transfer, and dynamically calibrating through environmental feedback, the traditional model has solved the problems of inaccurate prediction in extreme scenarios, long scenario adaptation cycle, and poor stability in dynamic environments, thereby reducing the cost of human intervention and prediction bias.
[0097] Based on any of the above embodiments, in Embodiment Six of this application, step S40 includes steps E11 to E13: Step E11: Using the target weight prediction model, infer and calculate the multimodal effective feature vectors within the collection period, and output the predicted weight of the pig population for each period.
[0098] In this embodiment, the multimodal effective feature vector refers to the visual, audio, and text fusion feature carrier formed after quantization compression and timestamp alignment within a set collection time period. The predicted weight of the pig herd for each period refers to the estimated weight of the pig herd output by the model for the feature vector of each collection period.
[0099] As an optional implementation method, the effective feature vectors of multiple modalities within the collection period are sorted out according to the collection time sequence, and the core representation dimensions, key semantic information, and time-series correlation features are extracted from the features one by one in each period. The sorted single-period feature vectors are then input into the target weight prediction model one by one. The model decomposes the mapping relationship between features and the weight of the pig population layer by layer, and obtains the initial prediction value through multi-dimensional correlation calculation. Subsequently, the prediction value is compared and verified with the feature trends of adjacent periods, and values that exceed the reasonable fluctuation range are corrected. This input, calculation, and verification process is repeated until all collection periods are covered. Finally, the predicted weight of the pig population for each period is output in cyclical order. This method has a refined inference process, the predicted values are verified for fluctuations, and the numerical stability is strong, which can avoid the errors caused by the bias of single-period features.
[0100] Step E12: Using a time-series aggregation strategy, the predicted weight values of the pig population over consecutive periods are correlated and analyzed to identify and remove outliers that deviate from the normal fluctuation range, thus obtaining an effective weight prediction sequence.
[0101] In this embodiment, the time-series aggregation strategy refers to a systematic method for associating, integrating, and analyzing predicted values for consecutive periods based on time sequence. Outliers refer to estimated values that exceed the natural growth or variation pattern of pig weight. The effective weight prediction sequence refers to the sequence of predicted weight values that maintains temporal continuity and conforms to normal fluctuation patterns after removing outliers.
[0102] As an optional implementation method, the predicted weight values of a pig herd over consecutive periods are compiled and arranged chronologically to form an original sequence. The overall growth and trend of the original sequence are calculated, and the normal fluctuation level standards for different stages are defined. For each predicted value in the original sequence, the difference is first calculated with the predicted values of the three adjacent periods. Then, it is compared with the fluctuation level standard of its current stage to determine whether it exceeds the limit. At the same time, the integrity of the multimodal features of the corresponding period is traced. If the difference exceeds the limit and there are missing or abnormal features, it is marked as an outlier. Subsequently, the marked outliers are placed in a temporary validation set. By analyzing the trend continuity of multiple periods before and after the outlier, it is confirmed whether it is an isolated outlier. Finally, all outliers that have undergone double validation are removed. The remaining predicted values are reassembled chronologically to obtain the effective weight prediction sequence. This method has comprehensive identification dimensions, combines trend, stage features, and the integrity of the original data, and has high accuracy, supporting refined feeding decisions.
[0103] As an alternative implementation, the predicted weight values of all consecutive-period pig populations are clustered according to their numerical distribution. The core group with the highest sample size is selected as the normal value benchmark, and the remaining groups are marked as suspicious groups. The temporal distribution of the predicted values within each suspicious group is analyzed to check for any continuous patterns. If the predicted values within a suspicious group do not have a continuous temporal distribution and the difference between the predicted values and those of the core group exceeds a set proportion, they are directly identified as outliers. If a continuous temporal distribution exists, environmental feedback data for that period is further correlated. It is determined whether the abnormal weight change is caused by a sudden change in the environment. If there is no environmental anomaly, it is identified as an outlier; otherwise, the values of that group are retained. Finally, all identified outliers are removed, and the values of the core group and the retained suspicious groups are integrated and sorted in chronological order to form an effective weight prediction sequence. This method quickly divides the normal and suspicious ranges through clustering, simplifies the judgment logic by combining temporal distribution and environmental correlation, has high processing efficiency, and can quickly identify batches of outliers.
[0104] Step E13: Based on the effective weight prediction sequence, combined with the number of fattening pigs in the pigpen and the characteristics of the growth stage, calculate the sequence mean and determine a reasonable error range to generate the estimated weight description of the fattening pigs.
[0105] In this embodiment, the number of pigs in a pigpen refers to the total number of pigs actually raised in the target pigpen. Growth stage characteristics refer to the core growth characteristics of the pigs in their growth cycle. The sequence mean refers to the average value of all values in the effective weight prediction sequence. The reasonable error range refers to the acceptable deviation range of the estimated value determined by combining growth characteristics and data fluctuations.
[0106] As an optional implementation method, this approach analyzes and generates stage characteristics to clarify the growth characteristics of fattening pigs. It breaks down the effective weight prediction sequence into growth stages, extracts the predicted values within each stage, calculates the sequence mean for each stage, and statistically analyzes the fluctuation range and dispersion of the mean for each stage. Combined with the number of fattening pigs in the pens, it calculates the estimated total weight of the herd for each growth stage. Then, based on the growth stability, predicted value dispersion, and the scale effect of each stage, the error range for each stage is dynamically adjusted. Subsequently, the average weight, total weight, and error range of each stage are integrated, supplemented with stage growth characteristic adaptation descriptions and a connection analysis of weight gain at each stage. Finally, all information is summarized to form a fattening pig weight estimation description that includes stage-specific data, overall aggregation results, and an interpretation of growth trends. This method calculates in stages to fit the growth pattern, has strong error calibration targeting, and provides comprehensive and refined information dimensions.
[0107] As an alternative implementation, this method integrates effective weight prediction sequences to calculate a global sequence mean, statistically analyzes the overall dispersion and distribution characteristics of the global values, obtains the number of pigs in the pen, and multiplies it by the global mean to obtain a total population weight estimate. Subsequently, it retrieves effective weight prediction sequence data from the same historical period, with similar numbers of pigs and characteristics of similar growth stages, analyzes the error distribution patterns and fluctuation ranges of historical data, and establishes a correlation mapping between historical errors and the current sequence dispersion, number of pigs, and characteristics of the growth stage. Based on this mapping relationship, it determines a reasonable error range for the current scenario and then extracts the global growth trend of the effective weight prediction sequence. It combines characteristics of the growth stage to supplement the prediction of weight growth potential and subsequent changes. Finally, it integrates the global average weight, total population weight, reasonable error range, and trend prediction to generate a pig weight estimation description focusing on core data and overall trends. This method is highly efficient in global aggregation calculation, and the error range based on historical data mapping provides more practical reference.
[0108] As an alternative implementation, based on the effective weight prediction sequence, the sequence data is split by growth stage, and the sequence mean, growth rate, and variance of each stage are calculated. Combined with the number of pigs in the pen, the total weight, average weight, and total growth of each stage are calculated. A growth stage error model is established, and a reasonable error range is calculated based on the variance of each stage, the completeness of feature data, and model fit. The weight data and error range of each stage are integrated, and stage growth assessments and feeding adaptation suggestions are added to generate a structured description of pig weight estimation that includes stage details, an overall summary, growth assessments, and feeding suggestions. It also supports switching between simplified and detailed output formats according to user needs.
[0109] For example, in the scenario of predicting the weight of fattening pigs, the optimized Transformer multimodal target weight prediction model receives multimodal effective feature vectors (including fattening pig visual behavior features, environmental perception features, and feeding text embedding features) within a collection period (each collection period is 24 hours). It then analyzes the mapping relationship between features and weight cycle by cycle and completes inference calculations, outputting the predicted weight values of the fattening pig population for 30 consecutive cycles. A temporal aggregation strategy is used to arrange the predicted values of all consecutive cycles in chronological order. The normal fluctuation range for each stage is set based on the growth patterns of fattening pigs in their juvenile, growth, and maturity stages. By calculating the mean deviation of each predicted value from the predicted values of the adjacent three cycles and analyzing the temporal continuity of the values, three outliers deviating from the fluctuation range (such as abnormally high / low values caused by temporary equipment failures) are identified and removed, resulting in an effective weight prediction sequence for 27 cycles. Based on this effective weight prediction sequence, and combined with the number of 500 fattening pigs in the target pigpen, the mean of each stage sequence is calculated according to the growth stage. Based on the characteristics of larger weight fluctuations in the growth stage and smaller fluctuations in the maturity stage, a reasonable error range is dynamically determined (±2.1% in the growth stage and ±1.3% in the maturity stage). The average weight of each stage, the total weight of the herd (sequence mean × number of pigs), the error range, and the growth trend adaptation description are integrated to generate a fattening pig weight estimation description that includes "The current total weight of the fattening pig herd is estimated to be XX tons, the average weight is XX kilograms, the error range in the growth stage is ±2.1%, and the error range in the maturity stage is ±1.3%, and the weight gain is in line with the normal rate of growth".
[0110] By using a process of cyclical inference to output predicted values, time-series aggregation to remove outliers, and phased accounting to generate descriptions, the problems of chaotic time-series weight predictions, outlier interference, and poor fit between weight estimation and growth stages are solved, providing accurate data for phased feeding management.
[0111] Based on any of the above embodiments, in Embodiment Seven of this application, referring to Figure 3 , Figure 3 This is a flowchart illustrating the seventh embodiment of the multi-model-based method for predicting the weight of fattening pigs according to this application. Before step S10, steps F11-F13 are also included: Step F11 involves collecting audio and video data, environmental parameters, and feeding logs of the pigs in the pen, simultaneously recording the total calibrated weight of the pigs at the final slaughter time, and outputting raw multi-source data.
[0112] In this embodiment, environmental parameters refer to quantitative data on environmental conditions such as temperature, humidity, ventilation, and gas concentration within the pigpen. Feeding logs are text records documenting feeding operations such as feed types, feed amounts, disease prevention measures, and health status. Total calibrated weight refers to the total weight of the herd after the pigs have reached slaughter standards, obtained through precise measurement. Raw multi-source data refers to the initial data set after integrating audio / video data, environmental parameters, feeding logs, and total calibrated weight, without any processing.
[0113] As an optional implementation method, data acquisition devices deployed in different areas of the pigpens are activated at fixed time intervals to capture the activity and audio of the pigs from all angles. Simultaneously, environmental parameters such as temperature, humidity, ventilation rate, and ammonia concentration are collected and continuously recorded in real time. After each feeding, cleaning, and disease prevention operation, the staff immediately and meticulously records the operation content, time points, and related details to form a feeding log. When the pigs reach the preset slaughter standard, all slaughter pigs are transferred to a standard weighing area, where they are weighed batch by batch using precise weighing equipment, and the total calibrated weight is calculated. The collected audio and video data are then segmented and named according to timestamps, environmental parameters are sorted and organized according to the collection time, and feeding logs are categorized and archived according to operation type. Finally, an index is established using timestamps to link the three types of data with the total calibrated weight, integrating all information to form a structurally complete dataset, and outputting the original multi-source data. This method offers comprehensive data collection coverage, accurate time correlation, and high information integrity.
[0114] As an alternative implementation, sensor-based data acquisition devices are deployed at key locations in the pigpens. These devices automatically activate when they detect specific behaviors in pigs, such as concentrated activity, feeding, or resting, capturing audio and video data for the corresponding time periods. Simultaneously, baseline values of environmental parameters are collected periodically. Additional data collection and recording of fluctuations are triggered when parameters exceed normal ranges. Feeders summarize daily feeding operations at fixed times, uniformly recording information such as total feed intake, disease prevention details, and any abnormalities to form a feeding log. Before slaughter, different batches of pigs are sampled and weighed three times. A conversion relationship is established based on the sampled weight, the number of pigs in the pen, and growth uniformity to calculate the total weight. After all pigs are slaughtered, a total weighing is performed for calibration to determine the final overall calibrated weight. Subsequently, the sensor-collected audio and video data are categorized by behavior type, environmental parameters are archived separately according to baseline values and fluctuation data, and the feeding logs are organized chronologically. Finally, the three types of data are integrated with the calibrated overall calibrated weight, linked by event type and time node, to form a dataset focusing on key behaviors and core parameters, outputting raw multi-source data. This method is highly targeted, effectively reduces redundant data, and batch sampling and weighing does not interfere with the normal slaughtering process, resulting in higher data collection efficiency.
[0115] Step F12: According to the timestamp of the data collection in the original multi-source data, the audio and video data are associated and matched with the environmental parameters, feeding logs and overall calibrated weight at the corresponding time to form a multi-source associated dataset.
[0116] In this embodiment, a multi-source associated dataset refers to a structured dataset that is closely associated and time-corresponding, formed by integrating various types of data through timestamp matching.
[0117] As an optional implementation method, this approach analyzes the timestamp formats of all audio / video data, environmental parameters, and feeding logs from the original multi-source dataset to unify the time representation standard. The dataset is divided into continuous time units, and the audio / video data is further segmented into segments of corresponding durations and labeled with unit numbers. Core information such as the mean and fluctuation range of environmental parameters within each time unit is extracted simultaneously, and the corresponding feeding log entries for each time period are associated with the unit number. The integrated data from all time units is then concatenated chronologically, and combined with the slaughter time node of the overall calibrated weight, a full-cycle timeline is established. The associated data of each time unit is bound to the slaughter cycle, and the existence of missing data in each time unit is checked. Missing items are marked with their status and supplemented with associated timestamp descriptions. This ultimately forms a multi-source associated dataset with time units as the core and a coherent temporal sequence. This method offers fine-grained temporal association and strong temporal logic, fully reconstructing the correspondence of data throughout the entire cycle, and providing accurate support for subsequent model mining of the correlation between full-cycle temporal features and weight.
[0118] Step F13: The small model is used to identify and remove empty columns, severely blurry audio and video, incomplete text records, and incorrectly labeled weight data from the multi-source associated dataset to obtain the cleaned multi-source associated dataset.
[0119] In this embodiment, empty pen data refers to various meaningless data collected when there are no pigs moving around in the pen. Severely blurry audio and video refers to multimedia data where the details of the picture are indistinguishable and the sound signal is chaotic. Incomplete text records refer to feeding logs or parameter records that lack core fields or have incomplete information. Incorrect calibrated weight data refers to inaccurate weight values that conflict with information such as growth patterns and the number of pigs in stock.
[0120] As an optional implementation method, the audio and video data in the multi-source association dataset is first analyzed frame by frame to extract core indicators such as image clarity, subject recognition, and sound signal stability. Data is then marked as severely blurry based on indicator thresholds. Next, the completeness of each field in the text records is checked to confirm whether core content such as feed intake and disease prevention information is missing, and incomplete entries are marked. Subsequently, the compatibility of the calibrated weight data with the corresponding growth cycle and number of pigs is compared, and the reasonableness of the values is judged in conjunction with the conventional growth rate, marking erroneous data. Finally, the presence of pig activity trajectories in the audio and video and the presence of pig registration in the text records are analyzed to comprehensively determine and mark empty pens. All unmarked valid data is then extracted and reorganized according to the original association logic to obtain a cleaned multi-source association dataset. This method has a single and clear data verification dimension, intuitive judgment criteria, and high accuracy in identifying invalid data.
[0121] For example, in a scenario for predicting the weight of fattening pigs, high-definition equipment deployed in key areas of the pigpens continuously collects audio and video data of pig activity. Temperature and humidity sensors, ammonia concentration detectors, and other devices record environmental parameters in real time. Farmers record daily feeding amounts, disease prevention measures, and health status logs according to regulations. Once the pigs reach slaughter standards, they are weighed batch by batch using a high-precision electronic weighbridge, and the total calibrated weight is calculated cumulatively. All collected information is simultaneously linked to output raw multi-source data. The collection timestamps of various data types in the raw multi-source data are extracted. The audio and video data are segmented by time nodes, and a one-to-one correspondence is established with the corresponding environmental parameters and farming log entries. The full-cycle time-series data is then bound to the total calibrated weight to form a complete link, constructing a multi-source associated dataset. A lightweight YOLO+LSTM model is used to verify each item in the multi-source associated dataset. YOLO is used to identify whether there is pig activity in the audio and video to determine empty pen data. Severely blurry audio and video were filtered based on image clarity threshold and sound signal-to-noise ratio. The integrity of core fields in the feeding log was verified to remove incomplete text records. The rationality of the calibrated weight was identified by comparing the growth cycle of meat pigs with the number of pigs in stock to identify erroneous data. After removing all invalid data, the valid information was integrated according to the original timestamp association logic to obtain a cleaned multi-source associated dataset.
[0122] By using a process of cyclical inference to output predicted values, time-series aggregation to remove outliers, and phased accounting to generate descriptions, the problems of chaotic time-series weight predictions, outlier interference, and poor fit between weight estimation and growth stages are solved, providing accurate data for phased feeding management.
[0123] Based on any of the above embodiments, in Embodiment Eight of this application, referring to Figure 4 , Figure 4 This is a flowchart illustrating the eighth embodiment of the multi-model-based method for predicting the weight of fattening pigs according to this application. Following step S40, steps G11-G13 are also included: Step G11: Based on the estimated weight description of the pigs and the actual slaughter weight, calculate the deviation rate between the estimated weight and the actual weight, analyze the error distribution characteristics in conjunction with the reliability description, and generate an error assessment report.
[0124] In this embodiment, the actual slaughter weight refers to the total weight of the herd after the pigs have reached the slaughter standard, obtained through precise measurement. The deviation rate is the ratio of the difference between the estimated weight and the actual weight to the actual weight. The reliability description provides supplementary information on factors affecting error during the weight estimation process, such as data quality, model suitability, and data collection conditions. The error distribution characteristics refer to the central tendency, dispersion, and distribution pattern of error values across different intervals. The error assessment report is a structured evaluation document that integrates the deviation rate calculation results, error distribution analysis, and improvement suggestions.
[0125] As an optional implementation method, the overall re-estimation value of the pig population in the weight estimation description is extracted and the difference is calculated with the actual slaughter weight. The deviation rate is obtained by combining the ratio of the difference to the actual slaughter weight. Subsequently, the deviation rate data is broken down by growth stage, and information such as the data collection completeness, model scenario adaptability, and environmental interference of each stage in the reliability specification is matched one by one. The error values corresponding to different influencing factors are classified and statistically analyzed. Multiple error intervals are divided and the error frequency of each interval is counted. The core intervals of error concentration and the distribution range of extreme errors are analyzed to trace the problems in the weight estimation process corresponding to extreme errors. The main causes of errors at each growth stage are summarized. Finally, the deviation rate statistics, stage-by-stage error distribution maps, correlation analysis of influencing factors, and targeted improvement suggestions are integrated to generate a comprehensive error assessment report. This method has fine-grained error analysis and accurate attribution, and can clearly present the error differences at different stages.
[0126] Step G12: Locate the source of error through the error assessment report, determine the model module that needs to be optimized, and obtain the direction of model optimization.
[0127] In this embodiment, the source of error refers to the specific reasons that cause the deviation between the estimated weight and the actual weight, covering aspects such as data acquisition, feature processing, and model training. The model module requiring optimization refers to the functional unit within the model that directly affects the generation of error. The model optimization direction refers to the specific functional improvement strategies and key adjustments formulated for the module requiring optimization.
[0128] As an optional implementation method, the distribution characteristics and influencing factors of different types of errors in the error assessment report are extracted and correlated. These errors are then categorized into three types based on their performance: temporal feature adaptation errors, cross-modal fusion errors, and scene parameter binding errors. The model processing stages corresponding to each type of error are traced one by one: temporal feature adaptation errors correspond to the feature timestamp alignment module, cross-modal fusion errors to the feature encoding and association mapping module, and scene parameter binding errors to the group adaptation multi-task training module. The functional deficiencies of each module are analyzed in depth. Combined with the error attribution conclusions, specific optimization directions are formulated for the deficiencies of each module. This method ensures accurate correspondence between errors and modules, highly targeted optimization directions, and direct access to the core causes of subdivided errors.
[0129] As an alternative implementation method, a multi-dimensional error quantification index system is constructed, including indicators for absolute error deviation, relative deviation, time-series fluctuation coefficient, and scenario adaptation deviation. The values of each index are calculated based on error assessment report data. A quantitative correlation model between error and influencing factors is established, quantifying the error indicators with influencing factors such as data quality, feature performance, model status, and scenario characteristics. Correlation analysis identifies core factors significantly affecting the error indicators, and contribution calculation determines the error impact weight of each core factor. Based on the quantification results, the core sources of error are located. For the model modules corresponding to the core error sources, quantitative optimization goals and specific optimization directions are formulated, resulting in a quantifiable and implementable model optimization scheme.
[0130] Step G13: Based on the model optimization direction and the training data supplemented by error samples, adjust the parameters of the corresponding model module that needs to be optimized, verify and output the iteratively optimized target weight prediction model.
[0131] In this embodiment, the training data supplemented by error samples refers to the set of supplementary training samples selected from the error assessment to enhance the model's error adaptation capability. The iteratively optimized target weight prediction model refers to the final weight prediction model with reduced error and improved performance after parameter adjustment and effect verification.
[0132] As an optional implementation, the training data supplemented with error samples is classified and labeled according to error type. The labeled data is then linked to the model optimization direction one by one. The core parameter dimensions of each model module that needs optimization are broken down, and the adjustment boundaries and optimization logic of each parameter are determined. First, for the highest priority optimization module, the core parameters are fine-tuned step by step according to the optimization direction. After each adjustment, the corresponding category of error samples is input into the module for individual verification, and the correlation between parameter changes and error reduction is recorded. Then, the parameter adjustments of other modules are advanced sequentially according to priority. After each module has completed its individual optimization, all optimization modules are integrated, and mixed test data is input for overall verification. The parameter compatibility between modules is analyzed, and conflicting parameters are fine-tuned to achieve synergistic optimization. This process of individual adjustment, individual verification, overall integration, and overall verification is repeated until the error stabilizes within the expected range, and finally, the iteratively optimized target weight prediction model is output. This method features fine-tuned parameters, strong inter-module synergy, and can maximize the optimization value of each module.
[0133] As an alternative implementation, the training data supplemented by error samples is proportionally mixed with the original training data to construct a comprehensive training dataset. Core improvement requirements in the model optimization direction are extracted, and the parameter adjustment needs of all model modules requiring optimization are integrated. A unified parameter adjustment range and step size are defined, and the parameters of all target modules are updated in a single batch. After the update, the comprehensive training dataset is proportionally split into a validation set and a test set. The validation set is first used to verify the overall error reduction effect of the model, identifying the parameter range corresponding to modules with still high errors. Based on the validation results, the parameters in this range are specifically corrected within a preset range. Then, the test set is used to verify the final effect and determine whether the model performance meets the optimization expectations. This method eliminates the need for separate adjustments and verifications for each module, directly completing parameter optimization and integration, and outputting the iteratively optimized target weight prediction model. This method has high parameter update efficiency, a simple process, and can quickly complete model iteration, significantly shortening the optimization cycle and rapidly meeting the core performance requirements of practical applications.
[0134] For example, in the scenario of predicting the weight of fattening pigs, the deviation rate between the overall re-estimation value of the population and the actual slaughter weight is extracted from the description of the fattening pig weight estimate. Combined with information such as the accuracy of the data collection equipment and the model's adaptability to different scenarios in the reliability description, the central tendency and discrete characteristics of the error at different growth stages are analyzed. Error intervals are divided and the causes of extreme errors are traced, generating an error assessment report that includes deviation rate statistics, error distribution maps, and correlation analysis of influencing factors. Through this report, the source of error is located, and it is identified that unreasonable weight allocation of the cross-modal feature fusion module, insufficient training of extreme scenario samples, and lag in environmental feedback parameter adjustment are the main causes. The corresponding model modules that need to be optimized are determined, and the model optimization direction is obtained as "reconstructing the feature fusion weight rules, strengthening extreme sample training, and optimizing the parameter dynamic adjustment algorithm". The 1200 sets of selected error samples are added to the training dataset, and the core parameters of the corresponding modules are adjusted one by one according to the optimization direction. First, the optimization effect of a single module is tested through the validation set, and then the modules are integrated for overall performance testing to check whether the model error has been reduced to the expected range. Finally, the iteratively optimized target weight prediction model is output.
[0135] By employing a process of error quantification analysis, precise source tracing and location, and sample supplementation and iteration, the problems of ambiguous model error causes, unclear optimization directions, and limited accuracy improvement after iteration have been solved, thereby enhancing the model's predictive stability and accuracy in complex scenarios.
[0136] This application provides a pig weight prediction device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the multi-model-based pig weight prediction method in Embodiment 1 above.
[0137] The following is for reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the pig weight prediction device in the embodiments of this application. The pig weight prediction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, multimodal fusion devices, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), non-contact visual measurement devices, etc., as well as fixed terminals such as portable AI weight estimators, desktop computers, etc. Figure 5 The illustrated pig weight prediction device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0138] like Figure 5 As shown, the pig weight prediction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the pig weight prediction device. The processing unit 1001, the ROM 1002, and the RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the pig weight prediction device to communicate wirelessly or wiredly with other devices to exchange data. Although pig weight prediction devices with various systems are shown in the figure, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.
[0139] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0140] The hog weight prediction device provided in this application employs the multi-model-based hog weight prediction method described in the above embodiments, which can solve the technical problem of poor hog weight monitoring results. Compared with the prior art, the beneficial effects of the hog weight prediction device provided in this application are the same as those of the multi-model-based hog weight prediction method provided in the above embodiments, and other technical features of the hog weight prediction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0141] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0142] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0143] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the multi-model-based method for predicting the weight of fattening pigs in the above embodiments.
[0144] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0145] The aforementioned computer-readable storage medium may be included in the pig weight prediction device; or it may exist independently and not assembled into the pig weight prediction device.
[0146] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the pig weight prediction device, the pig weight prediction device: extracts key audio-visual segments and effective text data showing stable pig activity from a cleaned multi-source associated dataset, and integrates them to obtain multi-source effective data; extracts visual features, audio features, and text features from the multi-source effective data using a small model, and quantifies them to generate multi-modal effective feature vectors; fuses the multi-modal effective feature vectors and environmental feeding text embedding vectors based on a multi-modal large model base, and collaboratively optimizes and outputs a target weight prediction model; the target weight prediction model, based on the multi-modal effective feature vectors, analyzes the continuous periodic prediction results using a time-series aggregation strategy to generate a pig weight estimation description.
[0147] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings.
[0149] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described multi-model-based method for predicting the weight of fattening pigs, thereby solving the technical problem of poor monitoring results for fattening pig weight. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the multi-model-based method for predicting the weight of fattening pigs provided in the above embodiments, and will not be repeated here.
[0150] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for predicting the weight of meat pigs based on multiple models, characterized in that, The method includes: From the cleaned multi-source associated dataset, key audio and video segments and effective text data of stable pig activity are extracted and integrated to obtain multi-source effective data. Visual features, audio features, and text features are extracted from the multi-source effective data using a small model, and then quantized to generate a multimodal effective feature vector. Based on the multimodal large model base, the effective feature vectors of the multimodal model and the environmental feeding text embedding vectors are fused together to collaboratively optimize the output target weight prediction model; The target weight prediction model generates a description of the estimated weight of the pigs by analyzing the continuous periodic prediction results through a time-series aggregation strategy based on the multimodal effective feature vectors.
2. The method for predicting the weight of meat pigs based on multiple models as described in claim 1, characterized in that, The steps of extracting key audio-visual segments and effective text data showing stable pig activity from the cleaned multi-source associated dataset and integrating them to obtain multi-source effective data include: From the cleaned multi-source associated dataset, audio and video data showing stable pig activity status are identified and filtered out, and combined with environmental parameters and feeding log text data to generate multi-source candidate data. By analyzing the rate of change and audio energy intensity between the audio frames summarized from the multi-source candidate data, the key audio and video segments are extracted. The key audio and video segments are matched with corresponding valid text data according to the data collection timestamp, and integrated to generate the multi-source valid data.
3. The method for predicting the weight of meat pigs based on multiple models as described in claim 1, characterized in that, The step of extracting visual features, audio features, and text features from the multi-source effective data using a small model, and quantizing them to generate a multimodal effective feature vector, includes: Through the processing branch of the small model, the audio and video key segments and effective text data in the multi-source effective data are classified to obtain the classified audio and video key segments and effective text data; Based on the processing branch of the small model, the visual features, audio features, and text features are extracted from the classified audio and video key segments and effective text data, respectively. The visual features, audio features, and text features are quantized and compressed, and the three types of features are aligned according to the data acquisition timestamp to obtain the multimodal effective feature vector.
4. The method for predicting the weight of meat pigs based on multiple models as described in claim 1, characterized in that, The steps of fusing the effective multimodal feature vectors and environmental feeding text embedding vectors based on the multimodal large model base to collaboratively optimize the output target weight prediction model include: The multimodal effective feature vector and the environmental feeding text embedding vector are input into the multimodal large model base. Feature encoding and association mapping are performed on the multimodal effective feature vector and the environmental feeding text embedding vector to output a hybrid feature vector. Based on the multimodal large model base, the feature dimensions of the hybrid feature vector are unified and then processed into a matrix to obtain a cross-modal feature matrix. The cross-modal feature matrix and group body size parameters are processed through group adaptation multi-task enhancement to output the target weight prediction model.
5. The method for predicting the weight of meat pigs based on multiple models as described in claim 4, characterized in that, The step of outputting the target weight prediction model by processing the cross-modal feature matrix and group body shape parameters through group adaptation multi-task enhancement includes: Based on the cross-modal feature matrix, a multi-task training objective is constructed by combining the group body shape parameter annotation, and the weight of extreme scenario samples in training is increased to output an initial weight prediction model. The features obtained from the initial weight prediction model are distilled and then transferred to the parameter template of the adapted pigpen scenario to generate an adapted weight prediction model. The model parameters of the adapted weight prediction model are adjusted based on the environmental feedback from the target pigpen to obtain the target weight prediction model.
6. The method for predicting the weight of meat pigs based on multiple models as described in claim 1, characterized in that, The steps of generating a description of pig weight estimation based on the multimodal effective feature vectors and continuous periodic prediction results through a time-series aggregation strategy include: The target weight prediction model is used to infer and calculate the effective feature vectors of the multimodal group within the collection period, and output the predicted weight of the pig population for each period. A time-series aggregation strategy was used to perform correlation analysis on the predicted weight values of the pig population over consecutive periods, identify and remove outliers that deviate from the normal fluctuation range, and obtain an effective weight prediction sequence. Based on the effective weight prediction sequence, combined with the number of fattening pigs in the pigpen and the characteristics of the growth stage, the sequence mean is calculated and a reasonable error range is determined to generate the estimated weight description of the fattening pigs.
7. The method for predicting the weight of meat pigs based on multiple models as described in claim 1, characterized in that, Before the step of extracting key audio-visual segments and effective text data showing stable pig activity from the cleaned multi-source associated dataset and integrating them to obtain multi-source effective data, the multi-model-based pig weight prediction method further includes: By collecting audio and video data, environmental parameters, and feeding logs of pigs in the pigpen, the total calibrated weight of pigs at the final slaughter time is recorded synchronously, and raw multi-source data is output. Based on the timestamps of the data collection in the original multi-source data, the audio and video data are associated and matched with the environmental parameters, feeding logs and overall calibrated weight at the corresponding time to form a multi-source associated dataset. The small model identifies and removes empty columns, severely blurry audio and video, incomplete text records, and incorrectly labeled weight data from the multi-source associated dataset to obtain the cleaned multi-source associated dataset.
8. The method for predicting the weight of meat pigs based on multiple models as described in claim 1, characterized in that, The steps of generating a description of pig weight estimation based on the multimodal effective feature vectors and continuous periodic prediction results through a time-series aggregation strategy include: Based on the estimated weight description of the pigs and the actual slaughter weight, the deviation rate between the estimated weight and the actual weight is calculated. Combined with the reliability description, the error distribution characteristics are analyzed, and an error assessment report is generated. By using the error assessment report, the source of error can be located, the model module that needs to be optimized can be determined, and the direction of model optimization can be obtained. Based on the model optimization direction and the training data supplemented by error samples, adjust the parameters of the corresponding model modules that need to be optimized, verify and output the iteratively optimized target weight prediction model.
9. A device for predicting the weight of pigs, characterized in that, The pig weight prediction device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multi-model-based pig weight prediction method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the multi-model-based method for predicting the weight of fattening pigs as described in any one of claims 1 to 8.