Livestock and poultry meat tissue recognition and intelligent segmentation method and system based on multi-modal vision

By using multimodal vision technology and deep learning models to identify the bone, fat, and muscle fiber orientation of fresh livestock and poultry meat, a high-quality three-dimensional segmentation path is generated. This solves the problems of low standardization and unreasonable multimodal data fusion in existing technologies, and realizes automated and accurate segmentation of fresh livestock and poultry meat.

CN122199966APending Publication Date: 2026-06-12北京二商肉类食品集团有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京二商肉类食品集团有限公司
Filing Date
2026-03-13
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing livestock and poultry fresh meat cutting technologies have shortcomings in terms of standardization, skeletal avoidance capabilities, utilization of muscle fiber orientation, and multimodal data fusion, resulting in unstable cutting quality and limited value enhancement.

Method used

Employing multimodal vision technology, this method acquires multispectral images and 3D point cloud data for spatiotemporal registration and feature fusion. It then uses a deep learning model to identify the orientation of bones, fat, and muscle fibers, constructs a multi-objective optimization function to generate 3D segmentation paths, and combines an adaptive weighted fusion algorithm and confidence assessment to generate high-quality segmentation paths.

Benefits of technology

It enables automated and precise cutting of fresh meat from livestock and poultry, improves the standardization of cutting, and can simultaneously take into account bone avoidance and muscle fiber orientation, thereby improving the cutting quality and automation level.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal vision's livestock and poultry meat tissue recognition and intelligent segmentation method and system, belong to image data processing technical field.For solving the problem that artificial segmentation is low in standardization degree, it is difficult to accurately avoid bone and muscle fiber trend, the multispectral image and three-dimensional point cloud data of the same livestock and poultry fresh meat carcass are obtained, and multi-modal feature map is generated after space-time registration and feature fusion;Bone structure, fat distribution and muscle fiber trend vector field are identified from the feature map using a deep learning model;Based on the identification result, a multi-objective optimization function is constructed, and a three-dimensional segmentation path is obtained from the preset starting point to the end point, avoiding bone, following the muscle fiber trend and satisfying the optimal meat yield. The application is used for livestock and poultry meat intelligent segmentation, and can improve the meat yield and cutting standardization level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image data processing technology, and relates to a method and system for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision. Background Technology

[0002] In the processing of fresh livestock and poultry meat, carcass cutting is a crucial step that determines the value of the meat and the quality of subsequent processing. For a long time, carcass cutting in my country has relied mainly on manual experience or semi-mechanized methods, making it difficult to effectively guarantee cutting accuracy and standardization. Existing cutting methods mainly suffer from the following problems and shortcomings.

[0003] Firstly, my country has a wide variety of livestock and poultry breeds, and the feeding conditions and fattening methods vary significantly across different regions, resulting in inconsistent carcass size, skeletal structure, and fat deposition levels even among the same type of livestock and poultry. In traditional manual butchering operations, workers typically rely on personal experience to determine the cutting location, making it difficult to maintain consistent cutting standards between different operators, and even among the same operator at different times. This subjectivity leads to confusion in the naming of meat parts, and the widespread phenomenon of the same piece having different names or the same piece having different names, seriously affecting the communication efficiency between slaughtering and butchering enterprises and the consumer market. Although some enterprises have introduced foreign butchering equipment, this equipment is usually designed for the standardized body shapes of foreign livestock and poultry breeds and is difficult to adapt to the individual differences in Chinese livestock and poultry carcasses, often requiring significant manual intervention and adjustments in practical applications.

[0004] Secondly, skilled meat cutters need to simultaneously determine the location of bones, the distribution of fat, and the direction of muscle texture, which places high demands on the operator's anatomical knowledge and practical experience. In practice, workers usually can only judge the approximate location of bones by visual observation and touch, making it difficult to accurately identify the boundary between bones and soft tissues. This often results in the blade intruding into the bone area during the cutting process. This increases blade wear and may also create the safety hazard of bone fragments mixed into the meat. At the same time, in order to ensure cutting efficiency, workers often fail to consider the natural direction of muscle fibers, resulting in a large angle between the cut surface and the direction of muscle fibers, which damages the integrity of muscle tissue and affects the cooking quality and taste of the meat.

[0005] Thirdly, from a meat science perspective, cutting perpendicular to the muscle fibers disrupts the integrity of the muscle bundles, leading to increased juice loss during heating and directly impacting the tenderness and flavor of the meat. However, existing cutting methods lack effective means to identify and utilize muscle fiber orientation information during planning and execution. Workers struggle to clearly discern the specific orientation of muscle fibers with the naked eye, especially when the meat surface is covered by fascia or fat. Even experienced workers who can roughly determine the muscle fiber direction find it difficult to maintain consistency between the cutting direction and the muscle fiber orientation throughout continuous operation.

[0006] Fourth, researchers in this field have attempted to introduce visual perception technology to assist in segmentation decisions. However, in practical applications, single-modal visual data often has its own limitations. For example, conventional color images can provide surface texture information, but in areas of fat infiltration or under varying lighting conditions, it is difficult to accurately identify tissue boundaries relying solely on color images. Three-dimensional point cloud data can reflect surface geometry, but in areas with depressions, occlusions, or drastic curvature changes, point cloud data often suffers from sparseness or even missing data. Effectively fusing data from different modalities to complement their respective shortcomings has always been a technical challenge in this field. Simple data overlay cannot solve the problem of spatial inconsistencies in heterogeneous data, nor can it reasonably allocate the contribution weights of different modalities when tissue texture is blurred or geometric information is insufficient.

[0007] In summary, existing livestock and poultry fresh meat cutting technologies have significant shortcomings in terms of standardization, skeletal avoidance capabilities, utilization of muscle fiber orientation, and multimodal data fusion, which limit the quality stability and value enhancement of cut meat products. Summary of the Invention

[0008] One object of the present invention is to solve at least the above-mentioned problems and / or defects, and to provide at least the advantages described below.

[0009] Another objective of this invention is to provide a method and system for livestock and poultry meat tissue identification and intelligent segmentation based on multimodal vision.

[0010] It solves the problems of low standardization in manual segmentation and difficulty in simultaneously taking into account bone avoidance and muscle fiber orientation.

[0011] This invention addresses the problem of unreasonable weight allocation and poor fusion quality caused by blurred textures or sparse point clouds during the fusion of multispectral images and 3D point cloud data.

[0012] This solves the problem that some spatial units in the fused multimodal feature map have inconsistent confidence levels due to differences in data quality, which in turn affects the accuracy of subsequent recognition.

[0013] Therefore, the technical solution provided by this invention is as follows: A method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision includes the following steps: 1) Acquire the original multispectral image data and original three-dimensional point cloud data of the same fresh meat carcass of livestock and poultry or its target cutting part, wherein the three-dimensional point cloud data is generated based on structured light scanning; 2) Spatiotemporal registration is performed on the original multispectral image data and the original three-dimensional point cloud data, and feature fusion is performed on the registered multispectral image data and the three-dimensional point cloud data to generate a spatially consistent multimodal feature map containing tissue feature information. 3) Input the multimodal feature map into the trained deep learning model, and use the deep learning model to identify and segment the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation of the carcass or its target cutting part; 4) Based on the vector field map of the skeletal structure mask, the fat distribution area mask, and the direction of the muscle fibers, a multi-objective optimization function is constructed, and the multi-objective optimization function is solved in the three-dimensional space represented by the multi-modal feature map to obtain a three-dimensional segmentation path that runs from a preset starting point to a preset ending point, avoids the skeletal area, and ensures that the meat yield of the target part segmented along the path meets the predetermined optimization criteria and extends along the direction of the muscle fibers. The three-dimensional segmentation path is used to output to an automatic cutting device for execution.

[0014] Preferably, in the method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision, in step 2), an adaptive weighted fusion algorithm is used for fusion. The adaptive weighted fusion algorithm is as follows: for each spatial unit in the multimodal feature map, the tissue texture clarity index and the point cloud density index of the multispectral image in the unit are extracted, and the tissue texture clarity index and the point cloud density index are nonlinearly mapped according to the preset fusion weight function to dynamically calculate and allocate the fusion weight of the unit.

[0015] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, step 2) further includes a confidence assessment and correction step: After generating the multimodal feature map, a fusion confidence score is calculated for each spatial unit; A dynamic threshold is determined based on the statistical distribution of the fusion confidence scores of all spatial units in the current processing batch; When the fusion confidence score is lower than the dynamic threshold, the unit is marked as a low confidence unit, and in the feature recognition process of step 3), the contribution weight of the low confidence unit to the recognition result is reduced.

[0016] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, step 2) of the adaptive weighted fusion algorithm further includes: For each spatial unit, the matching degree index between the multispectral data of that unit and the preset adipose tissue spectral features and preset muscle tissue spectral features is extracted. The preset fusion weight function also performs a comprehensive mapping of the tissue texture clarity index and the point cloud density index based on the matching degree index, so as to dynamically calculate the fusion weight of the unit.

[0017] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, a dynamic illumination correction step is further included between step 1) and step 2). Scene illumination intensity distribution data are acquired simultaneously when obtaining the original multispectral image data; Based on the light intensity distribution data, light unevenness compensation is performed on each pixel in the original multispectral image data to generate light-corrected multispectral image data. The illumination-corrected multispectral image data is used for subsequent spatiotemporal registration and feature fusion.

[0018] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, a three-dimensional geometric model reconstruction and hole filling step is included between step 2) and step 3). Based on the multimodal feature map after registration and fusion, the geometric contour information of the three-dimensional point cloud data is extracted to reconstruct the three-dimensional geometric model of the carcass or its target cutting part; Identify hole regions in the three-dimensional geometric model, where the hole regions correspond to spatial units where the original three-dimensional point cloud data is missing or the confidence level is below a preset threshold. Obtain the multispectral texture features of the hole region at the corresponding position in the multimodal feature map; Based on the multispectral texture features, semantically guided geometric filling is performed on the hole region to generate a complete and continuous three-dimensional geometric model. The complete three-dimensional geometric model is used for the planning and generation of the three-dimensional segmentation path in step 4).

[0019] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, a post-processing step for recognition results and a path constraint generation step is included between steps 3) and 4). Obtain the vector field map of the skeletal structure mask, the main fat distribution area mask, and the muscle fiber orientation output by the deep learning model, and simultaneously obtain the prediction confidence score output by the deep learning model for each mask unit. Based on the predicted confidence score, the boundary regions of the skeletal structure mask and the main fat distribution area mask are subjected to confidence-weighted morphological smoothing to generate an optimized mask with continuous closed boundaries. Based on the predicted confidence score, a confidence-weighted streamline filter is applied to the vector field map of the muscle fiber orientation to remove abnormal vectors with confidence scores below a preset threshold, thereby generating an optimized vector field map with spatial continuity. Based on the optimized mask and the optimized vector field map, a set of hard and soft constraints on the three-dimensional segmentation path is generated, wherein the hard constraints include insurmountable bone restricted areas and insulated fat preservation areas, and the soft constraints include a preference for cutting directions that preferentially follow the direction of muscle fibers. The hard and soft constraints are used for setting constraints and pruning the search space when constructing the multi-objective optimization function in step 4).

[0020] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, step 4) is followed by a step of cutting feasibility verification and path correction. Obtain the 3D segmentation path generated in step 4) and use it as the initial path to be verified; Based on the bone structure mask identified in step 3), bone interference detection is performed on the initial path to determine whether the initial path spatially intersects with the bone region. If spatial intersection occurs, then based on the muscle fiber orientation vector field map identified in step 3), while maintaining the overall direction along the muscle fiber, the local path segments where the intersection occurs are fine-tuned to generate a first corrected path that avoids the skeletal region. Based on the preset cutting tool geometry model, the tool reachability simulation is performed on the first correction path to determine whether there are any spatial points on the first correction path that the cutting tool cannot reach. If there are unreachable points, the first modified path is fine-tuned a second time according to the kinematic constraints of the cutting tool geometry model to generate a second modified path that simultaneously satisfies the requirements of skeleton avoidance and tool reachability. The second corrected path is used as the final output 3D segmentation path, which is then output to an automated cutting device for execution.

[0021] Preferably, in the multimodal vision-based method for livestock and poultry meat tissue recognition and intelligent segmentation, step 1) further includes a dynamic target adaptive acquisition and registration compensation step: During the acquisition of the original multispectral image data and the original three-dimensional point cloud data, the motion state data of the fresh meat carcass of livestock and poultry or its target cutting part are collected simultaneously. The motion state data includes at least one of displacement velocity, rotation angle and vibration frequency. Based on the motion state data, the acquisition timing of the original multispectral image data and the original three-dimensional point cloud data is dynamically adjusted so that the acquisition time of the multispectral camera and the structured light scanning device is synchronized with the moment of minimum displacement in the motion state data. For the acquired multispectral image data and 3D point cloud data, the motion offset of each pixel or each point cloud point is calculated based on the motion state data, and motion compensation correction is performed on the multispectral image data and 3D point cloud data before registration based on the motion offset to generate motion-compensated multispectral image data and 3D point cloud data. Motion-compensated multispectral image data and 3D point cloud data are used for spatiotemporal registration and feature fusion in step 2).

[0022] A multimodal vision-based system for livestock and poultry meat tissue recognition and intelligent segmentation includes: The data acquisition module is used to acquire the original multispectral image data and original three-dimensional point cloud data of the same fresh meat carcass of livestock and poultry or its target cutting part, wherein the three-dimensional point cloud data is generated based on structured light scanning; The registration and fusion module is used to perform spatiotemporal registration of the original multispectral image data and the original three-dimensional point cloud data, and to perform feature fusion of the registered multispectral image data and the three-dimensional point cloud data to generate a spatially consistent multimodal feature map containing tissue feature information. The feature recognition module is used to input the multimodal feature map into a trained deep learning model, and to identify and segment the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation of the carcass or its target cutting part through the deep learning model. The path planning module is used to construct a multi-objective optimization function based on the vector field map of the skeletal structure mask, the fat distribution area mask, and the direction of muscle fibers, and solve the multi-objective optimization function in the three-dimensional space represented by the multimodal feature map to obtain a three-dimensional segmentation path from a preset starting point to a preset ending point, avoiding the skeletal region, such that the meat yield of the target part segmented along the path meets a predetermined optimization criterion, and extends along the direction of muscle fibers. The three-dimensional segmentation path is used to output to an automatic cutting device for execution.

[0023] Preferably, in the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system, the registration and fusion module includes an adaptive weighted fusion unit. The adaptive weighted fusion unit is used to: extract the tissue texture clarity index and the point cloud density index of the multispectral image within each spatial unit of the multimodal feature map, and perform nonlinear mapping on the tissue texture clarity index and the point cloud density index according to a preset fusion weight function, so as to dynamically calculate and allocate the fusion weight of the unit.

[0024] Preferably, in the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system, the registration and fusion module further includes a confidence assessment and correction unit, used for: calculating a fusion confidence score for each spatial unit after generating the multimodal feature map; determining a dynamic threshold based on the statistical distribution of the fusion confidence scores of all spatial units in the current processing batch; and marking the unit as a low-confidence unit when the fusion confidence score is lower than the dynamic threshold, and reducing the contribution weight of the low-confidence unit to the recognition result when the feature recognition module performs recognition.

[0025] Preferably, in the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system, the adaptive weighted fusion unit is further used to: extract the matching degree index between the multispectral data of each spatial unit and the preset adipose tissue spectral features and the preset muscle tissue spectral features; the preset fusion weight function also performs a comprehensive mapping of the tissue texture clarity index and the point cloud density index according to the matching degree index, so as to dynamically calculate the fusion weight of the unit.

[0026] Preferably, the livestock and poultry meat tissue recognition and intelligent segmentation system based on multimodal vision further includes an illumination correction module, which is located between the data acquisition module and the registration and fusion module. This module is used to: simultaneously acquire scene illumination intensity distribution data while acquiring the original multispectral image data; compensate for uneven illumination in each pixel of the original multispectral image data based on the illumination intensity distribution data to generate illumination-corrected multispectral image data; and provide the illumination-corrected multispectral image data to the registration and fusion module.

[0027] Preferably, the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system further includes a three-dimensional geometric model reconstruction module, located between the registration and fusion module and the feature recognition module. This module is used to: extract geometric contour information from the three-dimensional point cloud data based on the registered and fused multimodal feature map, and reconstruct the three-dimensional geometric model of the carcass or its target cutting portion; identify hole regions in the three-dimensional geometric model, where the hole regions correspond to spatial units that are missing from the original three-dimensional point cloud data or have a confidence level below a preset threshold; obtain the multispectral texture features of the hole regions at the corresponding positions in the multimodal feature map; perform semantically guided geometric filling of the hole regions based on the multispectral texture features to generate a complete and continuous three-dimensional geometric model; and provide the complete three-dimensional geometric model to the path planning module for planning and generating three-dimensional segmentation paths.

[0028] Preferably, the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system further includes a post-processing module for recognition results, located between the feature recognition module and the path planning module, for: Obtain the vector field map of the skeletal structure mask, the main fat distribution area mask, and the muscle fiber orientation output by the deep learning model, and simultaneously obtain the prediction confidence score output by the deep learning model for each mask unit. Based on the predicted confidence score, the boundary regions of the skeletal structure mask and the main fat distribution area mask are subjected to confidence-weighted morphological smoothing to generate an optimized mask with continuous closed boundaries. Based on the predicted confidence score, a confidence-weighted streamline filter is applied to the vector field map of the muscle fiber orientation to remove abnormal vectors with confidence scores below a preset threshold, thereby generating an optimized vector field map with spatial continuity. Based on the optimized mask and the optimized vector field map, a set of hard and soft constraints on the three-dimensional segmentation path is generated, wherein the hard constraints include insurmountable bone restricted areas and insulated fat preservation areas, and the soft constraints include a preference for cutting directions that preferentially follow the direction of muscle fibers. The hard and soft constraints are provided to the path planning module for setting constraints and pruning the search space when constructing the multi-objective optimization function.

[0029] Preferably, the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system further includes a path verification and correction module, which is located after the path planning module and is used for: Obtain the 3D segmentation path generated by the path planning module and use it as the initial path to be verified; Based on the skeletal structure mask identified by the feature recognition module, skeletal interference detection is performed on the initial path to determine whether the initial path spatially intersects with the skeletal region. If spatial intersection occurs, based on the muscle fiber orientation vector field map identified by the feature recognition module, while maintaining the overall direction along the muscle fiber, the local path segments where the intersection occurs are fine-tuned to generate a first corrected path that avoids the skeletal region. Based on the preset cutting tool geometry model, the tool reachability simulation is performed on the first correction path to determine whether there are any spatial points on the first correction path that the cutting tool cannot reach. If there are unreachable points, the first modified path is fine-tuned a second time according to the kinematic constraints of the cutting tool geometry model to generate a second modified path that simultaneously satisfies the requirements of skeleton avoidance and tool reachability. The second corrected path is then used as the final output 3D segmentation path, which is then output to an automatic cutting device for execution.

[0030] Preferably, in the multimodal vision-based livestock and poultry meat tissue recognition and intelligent segmentation system, the data acquisition module further includes a dynamic target adaptive acquisition and registration compensation unit, used for: During the acquisition of the original multispectral image data and the original three-dimensional point cloud data, the motion state data of the fresh meat carcass of livestock and poultry or its target cutting part are collected simultaneously. The motion state data includes at least one of displacement velocity, rotation angle and vibration frequency. Based on the motion state data, the acquisition timing of the original multispectral image data and the original three-dimensional point cloud data is dynamically adjusted so that the acquisition time of the multispectral camera and the structured light scanning device is synchronized with the moment of minimum displacement in the motion state data. For the acquired multispectral image data and 3D point cloud data, the motion offset of each pixel or each point cloud point is calculated based on the motion state data, and motion compensation correction is performed on the multispectral image data and 3D point cloud data before registration based on the motion offset to generate motion-compensated multispectral image data and 3D point cloud data. The motion-compensated multispectral image data and the three-dimensional point cloud data are then provided to the registration and fusion module.

[0031] Compared with the prior art, the present invention has the following significant advantages: This invention acquires multispectral images and three-dimensional point cloud data, performs registration and fusion, then uses a deep learning model to identify the direction of bones, fat and muscle fibers, and finally generates a three-dimensional segmentation path through multi-objective optimization. This provides a complete technical foundation for the automated and precise segmentation of fresh meat from livestock and poultry, and can effectively overcome the shortcomings of manual segmentation, such as low standardization and difficulty in meeting multiple requirements.

[0032] This invention employs an adaptive weighted fusion algorithm based on texture clarity and point cloud density, which enables the fusion weights to be dynamically adjusted according to the local data quality. It reasonably allocates weights in areas with low texture clarity or sparse point clouds, which helps to improve the fusion quality of multimodal feature maps and the accuracy of subsequent recognition.

[0033] This invention introduces a confidence assessment and correction step after fusion to mark spatial units with low fusion quality and reduce their subsequent contribution weight, which can effectively suppress the interference of low-quality fusion regions on the recognition results and enhance the robustness of the recognition process.

[0034] This invention introduces a spectral feature matching index into the calculation of fusion weights, enabling the fusion weights to reflect the spectral similarity between the unit and fat and muscle tissues. Even when texture and point cloud information is insufficient, the weights can still be correctly allocated based on spectral information, significantly improving the semantic fidelity of the fused image.

[0035] This invention introduces a dynamic illumination correction step after data acquisition to compensate for image distortion caused by uneven ambient illumination. This makes the subsequently extracted indicators such as texture clarity and spectral matching degree more realistically reflect the physical characteristics of the tissue itself, providing a high-quality data foundation for accurate identification.

[0036] This invention introduces a semantically guided hole-filling step, using multispectral texture features to guide geometric filling, so that the filled 3D model is not only continuous on the surface, but also conforms to the anatomical features of the tissue in the region, avoiding erroneous geometric fusion across tissue boundaries, and providing a high-fidelity geometric basis for path planning.

[0037] This invention introduces a prediction confidence score into the post-processing of the recognition results, smooths the mask boundary, and performs streamline filtering on the vector field. This effectively eliminates abnormal vectors and generates a continuous and closed optimized mask and vector field. At the same time, the results are transformed into hard / soft constraints, providing high-quality input and a clear constraint framework for subsequent optimization solutions.

[0038] This invention introduces a cutting feasibility verification and correction step after path planning, performs skeletal interference detection and tool reachability simulation on the generated path, and performs local fine-tuning and secondary correction based on the detection results, ensuring that the final output path is physically executable and avoiding the dilemma of being mathematically optimal but physically infeasible.

[0039] This invention introduces a dynamic target adaptive acquisition and registration compensation step in the data acquisition process, synchronously acquiring motion state data and adjusting the acquisition timing and performing motion compensation correction accordingly. This reduces data misalignment and distortion caused by body swaying from the source, laying a solid foundation for subsequent spatiotemporal registration and feature fusion.

[0040] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Detailed Implementation

[0041] The present invention will now be described in further detail so that those skilled in the art can implement it based on the description.

[0042] It should be noted that, unless otherwise specified, the experimental methods described in the following implementation plan are all conventional methods, and the reagents and materials described are all commercially available unless otherwise specified.

[0043] According to one embodiment of the present invention, a method for livestock and poultry meat tissue identification and intelligent segmentation based on multimodal vision includes the following steps: 1) Acquire the original multispectral image data and original three-dimensional point cloud data of the same fresh meat carcass of livestock and poultry or its target cutting part, wherein the three-dimensional point cloud data is generated based on structured light scanning; 2) Spatiotemporal registration is performed on the original multispectral image data and the original three-dimensional point cloud data, and feature fusion is performed on the registered multispectral image data and the three-dimensional point cloud data to generate a spatially consistent multimodal feature map containing tissue feature information. 3) Input the multimodal feature map into the trained deep learning model, and use the deep learning model to identify and segment the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation of the carcass or its target cutting part; 4) Based on the vector field map of the skeletal structure mask, the fat distribution area mask, and the direction of the muscle fibers, a multi-objective optimization function is constructed, and the multi-objective optimization function is solved in the three-dimensional space represented by the multi-modal feature map to obtain a three-dimensional segmentation path that runs from a preset starting point to a preset ending point, avoids the skeletal area, and ensures that the meat yield of the target part segmented along the path meets the predetermined optimization criteria and extends along the direction of the muscle fibers. The three-dimensional segmentation path is used to output to an automatic cutting device for execution.

[0044] This embodiment provides a method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision. The method is implemented according to the following steps.

[0045] The first step involves acquiring raw multispectral image data and raw 3D point cloud data of the same fresh livestock carcass. In this step, the slaughtered and preliminarily processed pig carcass is suspended on a conveyor chain, allowing it to pass stably through the data acquisition station. A multispectral camera and a structured light scanning device are positioned at this station. The multispectral camera acquires images of the carcass surface at multiple specific wavelengths, while the structured light scanning device projects an coded grating onto the carcass and receives reflected light to generate raw 3D point cloud data. To ensure data synchronization, the acquisition control system simultaneously activates both the multispectral camera and the structured light scanning device upon receiving a trigger signal from the conveyor chain encoder, ensuring that the data acquired by both correspond to the carcass at the same spatial location.

[0046] Next, step two involves spatiotemporal registration and feature fusion of the original multispectral image data and the original 3D point cloud data. The acquired original multispectral image and original 3D point cloud have different coordinate systems and resolutions. First, using pre-determined transformation parameters from a calibration board, they are mapped to the same spatial coordinate system, achieving spatial registration. Based on this, an adaptive weighted fusion algorithm is further executed to generate a spatially consistent multimodal feature map containing tissue feature information. During the fusion process, the registered 3D space is divided into multiple voxel units. For each unit, the tissue texture sharpness index of the multispectral image and the point cloud density index of the 3D point cloud are calculated. The tissue texture sharpness index is obtained by calculating the mean of the local gray-level gradient magnitude, while the point cloud density index is determined based on the ratio of the number of points in the point cloud to the unit volume. These two indices are nonlinearly mapped according to a preset fusion weight function, dynamically calculating the fusion weight of the unit. Units with higher tissue texture sharpness or higher point cloud density receive higher fusion weights. Based on the weighted allocation, the spectral reflectance features of the multispectral image and the geometric depth features of the 3D point cloud are weighted and combined to generate a multimodal feature vector for each voxel unit. The feature vectors of all units together constitute a multimodal feature map.

[0047] Step three involves inputting the multimodal feature maps into a trained deep learning model. This model identifies and segments the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation from the carcass. The deep learning model employs an encoder-decoder architecture. The encoder extracts high-level semantic features from the multimodal feature maps through 3D convolution, while the decoder restores spatial resolution and outputs predictions for three branches through deconvolution and skip connections. The skeletal structure mask is represented as a binary voxel map, where regions with a voxel value of 1 represent bone locations. The main fat distribution area mask is also a binary voxel map, marking concentrated areas of fat tissue. The vector field map of muscle fiber orientation outputs a 3D unit vector at each voxel location, indicating the direction of muscle fiber extension. A large amount of labeled multimodal data from livestock and poultry carcasses is used as training samples when training this deep learning model. The annotations include measured data on the skeleton, fat regions, and muscle fiber orientation. The measured data of muscle fiber orientation are obtained through any of the following methods: obtaining the data through three-dimensional reconstruction of microscopic images after continuous tissue sections of the sample meat; or obtaining the data through non-destructive scanning of the sample meat using an ultrasonic anisotropy detection device; or obtaining the data by establishing a muscle fiber orientation template based on a known livestock and poultry anatomy database and mapping it to a specific sample carcass using a non-rigid registration algorithm.

[0048] Finally, in step four, a multi-objective optimization function is constructed based on the skeletal structure mask, the fat distribution area mask, and the vector field map of muscle fiber orientation. This multi-objective optimization function is then solved in the three-dimensional space represented by the multimodal feature map to obtain a three-dimensional segmentation path. The preset starting point of this path is set on the side of the carcass spine line near the neck, and the preset ending point is set at the corresponding position of the carcass buttocks. The design of the multi-objective optimization function comprehensively considers three objectives: the path length should be as short as possible to avoid unnecessary cutting strokes; the path should completely avoid the area marked by the skeletal structure mask in space to prevent tool damage; and the tangential direction at each point on the path should be as consistent as possible with the direction of the muscle fiber orientation vector field at that point to ensure cutting along the muscle fibers. Specifically, the multi-objective optimization function is constructed using a weighted summation: Cost = w1•L + w2•B + w3•F, where L is the path length term, B is the skeletal interference penalty term (maximum if the path point falls into the skeletal region, otherwise 0), and F is the cosine of the angle between the path tangent and the muscle fiber direction or the squared error term. The weight coefficients w1, w2, and w3 are pre-set through calibration experiments based on preset process requirements (e.g., increasing w3 to prioritize meat quality, decreasing w1 to prioritize yield) or configured online by the operator. When multiple objectives conflict, the weight coefficients are adjusted to achieve a trade-off between different objectives. An improved method based on the fast traversal algorithm is used in the solution process, propagating the path cost outward from the starting point in the three-dimensional grid space. The cost function is defined by the multi-objective optimization function, and finally, the path with the lowest cumulative cost is obtained by backtracking at the termination point. This path is a three-dimensional segmentation path that avoids bones, follows the direction of muscle fibers, and meets the predetermined optimization criteria for meat yield. The data format of this path is converted into control instructions for the automatic cutting equipment and output to drive the robotic arm to carry the cutting blade to perform the actual segmentation operation.

[0049] Existing livestock and poultry meat processing technologies employ a scheme combining color imaging and ultrasonic testing. This scheme first uses a color camera to acquire images of the carcass surface, then moves an ultrasonic probe across the carcass surface to probe deep bone boundaries. Finally, the two sets of data are superimposed, and the cutting path is manually planned. While this existing technology can obtain a rough location of the bones, ultrasonic testing is a contact measurement, which is inefficient and cannot cover the entire curved surface of the carcass, resulting in limited resolution of the bone boundaries. Furthermore, this scheme only outputs bone location and cannot provide information on fat distribution and muscle fiber orientation. The cutting path still needs to be determined manually based on experience, failing to achieve automatic path optimization. Compared to this invention, the existing technology lacks deep fusion of multispectral images and 3D point clouds, cannot obtain high-resolution bone masks and muscle fiber vector fields, and path planning relies on manual intervention, failing to guarantee optimal meat yield along muscle fiber orientation. This invention generates high-quality multimodal feature maps by spatiotemporal registration and adaptive fusion of multispectral and point cloud data, automatically identifies various tissue features with the help of deep learning, and automatically generates three-dimensional segmentation paths that meet multiple requirements through multi-objective optimization. It is superior to the existing technology in terms of automation, bone avoidance accuracy and utilization of muscle fiber orientation.

[0050] This embodiment first acquires raw multispectral image data and raw 3D point cloud data of the same fresh livestock and poultry carcass, providing a complete data foundation containing surface spectral information and 3D geometric information for subsequent analysis. By performing spatiotemporal registration and feature fusion on these raw data, a spatially consistent multimodal feature map containing tissue feature information is generated, solving the problem of spatial inconsistency between heterogeneous data and enabling subsequent recognition to utilize both spectral and geometric features simultaneously. Next, the multimodal feature map is input into a deep learning model to identify skeletal structure masks, main fat distribution area masks, and muscle fiber orientation vector field maps, achieving non-contact automatic perception of key internal tissue structures of the carcass and providing quantitative input for path planning. Finally, based on the recognition results, a multi-objective optimization function is constructed and a 3D segmentation path is solved. This path automatically avoids skeletal regions, extends along muscle fiber orientation, and meets the meat yield optimization criterion, directly converting the perception results into executable segmentation instructions. This achieves a complete closed loop from multimodal data acquisition to optimal segmentation path generation, solving the problems of low standardization in manual segmentation and difficulty in simultaneously considering skeletal avoidance and muscle fiber orientation, significantly improving the automation level and cutting quality of livestock and poultry meat segmentation.

[0051] According to one embodiment of the present invention, in the method for identification and intelligent segmentation of livestock and poultry meat tissue based on multimodal vision, step 2) employs an adaptive weighted fusion algorithm. The adaptive weighted fusion algorithm is as follows: for each spatial unit in the multimodal feature map, extract the tissue texture clarity index and the point cloud density index of the multispectral image within the unit, and perform nonlinear mapping on the tissue texture clarity index and the point cloud density index according to a preset fusion weight function to dynamically calculate and allocate the fusion weight of the unit.

[0052] This embodiment further defines the fusion process using an adaptive weighted fusion algorithm based on step two. During the generation of the multimodal feature map, the registered 3D space is first divided into regular voxel units, each corresponding to a spatial location. For each voxel unit, the tissue texture sharpness index of the multispectral image and the point cloud density index of the 3D point cloud are extracted. The tissue texture sharpness index is obtained by calculating the average gradient magnitude of the local region of the multispectral image corresponding to the unit, reflecting the clarity of the tissue texture at that location; the point cloud density index is obtained by statistically analyzing the ratio of the number of point clouds within the unit to the unit volume, reflecting the completeness of the geometric data. Subsequently, a nonlinear mapping is performed on these two indices according to a preset fusion weight function. This function uses a sigmoid form to map the texture sharpness index and the point cloud density index to weight coefficients between 0 and 1, respectively. The final fusion weight of the unit is then obtained through weighted product or weighted summation. In regions with high texture clarity and high point cloud density, the fusion weights are automatically increased, allowing the spectral features of the multispectral image and the geometric features of the point cloud to dominate the fusion process. In regions with blurred textures or sparse point clouds, the fusion weights are automatically decreased to avoid excessive interference from low-quality data in the fusion results. Based on the calculated dynamic weights, the spectral reflectance values ​​of the multispectral image and the depth values ​​of the 3D point cloud are weighted and combined to generate a multimodal feature vector for each voxel unit. The feature vectors of all units together constitute a high-quality multimodal feature map.

[0053] This embodiment further defines the specific implementation of the adaptive weighted fusion algorithm based on the above embodiments. By evaluating the texture sharpness of the multispectral image and the point cloud density of the 3D point cloud in each spatial unit, and dynamically calculating the fusion weight based on these two indicators, the problem of fixed-weight fusion being unable to adapt to local data quality changes is solved. When the texture sharpness of a certain region is low or the point cloud is sparse, the fusion weight of that region is automatically reduced, reducing the contamination of the fusion result by low-quality data; when both indicators are high, the fusion weight is increased, allowing high-quality data to be fully utilized. This adaptive weight allocation mechanism significantly improves the fusion quality of multimodal feature maps, enabling subsequent deep learning models to perform tissue recognition based on more accurate input, thereby enhancing the robustness and accuracy of the entire segmentation method.

[0054] According to one embodiment of the present invention, the method for livestock and poultry meat tissue identification and intelligent segmentation based on multimodal vision, step 2) further includes a confidence assessment and correction step: After generating the multimodal feature map, a fusion confidence score is calculated for each spatial unit; A dynamic threshold is determined based on the statistical distribution of the fusion confidence scores of all spatial units in the current processing batch; When the fusion confidence score is lower than the dynamic threshold, the unit is marked as a low confidence unit, and in the feature recognition process of step 3), the contribution weight of the low confidence unit to the recognition result is reduced.

[0055] In this embodiment, after generating the multimodal feature map, a fusion confidence score is calculated for each spatial unit. This score can be derived based on the quality indicators of the original data used in the fusion process, such as texture clarity, point cloud density, and the stability of the fusion weights. Alternatively, an independent evaluation network can be used to score the quality of the fused feature vector. After calculating the confidence scores for all spatial units in the current batch, the distribution of these scores is statistically analyzed, for example, by calculating their mean and standard deviation. The mean minus one standard deviation is used as a dynamic threshold. Spatial units with a fusion confidence score lower than this dynamic threshold are marked as low-confidence units. When the multimodal feature map is subsequently input into the deep learning model for feature recognition, the feature vectors corresponding to these low-confidence units are multiplied by a decay coefficient less than 1, reducing their contribution weight to the model's loss function and the final recognition result. In this way, regions with poor fusion quality are intentionally suppressed during the recognition process, and the model focuses more on the features of high-confidence regions, thereby reducing the interference of low-quality fusion units on the recognition result.

[0056] This technical solution calculates a confidence score for each fused spatial unit and determines a dynamic threshold based on statistical distribution. Units with confidence scores below the threshold are marked as low-confidence and their weights are reduced in subsequent recognition. This measure solves the problem of inconsistent confidence scores among some units in the fused feature map due to differences in data quality, preventing low-quality fused regions from negatively impacting the recognition results. By suppressing the influence of unreliable features, the deep learning model can more accurately extract information on bone, fat, and muscle fiber orientation, providing more reliable input for subsequent path planning, thereby improving the stability and accuracy of the entire segmentation system.

[0057] According to one embodiment of the present invention, in step 2) of the multimodal vision-based livestock and poultry meat tissue identification and intelligent segmentation method, the adaptive weighted fusion algorithm further includes: For each spatial unit, the matching degree index between the multispectral data of that unit and the preset adipose tissue spectral features and preset muscle tissue spectral features is extracted. The preset fusion weight function also performs a comprehensive mapping of the tissue texture clarity index and the point cloud density index based on the matching degree index, so as to dynamically calculate the fusion weight of the unit.

[0058] This embodiment extracts tissue texture sharpness and point cloud density indices for each spatial unit, and also extracts matching degree indices between the unit's multispectral data and preset adipose tissue spectral features and preset muscle tissue spectral features. Specifically, the reflectance curves of typical adipose and muscle tissues under different multispectral bands are experimentally determined beforehand as reference spectral features. For each spatial unit, the correlation between its multispectral data and the adipose reference spectrum is calculated to obtain the adipose matching degree; the correlation between it and the muscle reference spectrum is calculated to obtain the muscle matching degree. These two matching degree indices reflect the degree of similarity between the unit and adipose or muscle in spectral attributes. The preset fusion weight function, based on the original nonlinear mapping based on texture sharpness and point cloud density, adds a comprehensive consideration of the matching degree indices. For example, the texture sharpness index, point cloud density index, and matching degree index are jointly determined by weighted product or adaptive weighted summation to ensure that in areas with insufficient texture and point cloud information, if the spectral features highly match adipose or muscle, a high fusion weight can still be obtained, thus preserving valuable semantic information. The final fusion weights take into account both the geometric and texture quality of the local data and the fidelity of the spectral semantics, so that the fused multimodal feature map can accurately reflect the organizational attributes in both organizational boundaries and homogeneous regions.

[0059] This technical solution ensures that the calculation of fusion weights depends not only on the texture clarity and point cloud density of the local data, but also on the degree of matching between the spectral attributes of the unit and the target tissue type. Even when the texture is blurry or the point cloud is sparse, but the spectral features are highly consistent with fat or muscle, the fusion weights can still maintain a high level, thus ensuring that key semantic information is not discarded. Through this comprehensive mapping, the fused multimodal feature map accurately reflects tissue attributes in both tissue boundaries and homogeneous regions, providing a purer and more reliable input for deep learning models to identify bone, fat, and muscle fiber orientations, further improving the performance of the entire segmentation method.

[0060] According to one embodiment of the present invention, the method for livestock and poultry meat tissue identification and intelligent segmentation based on multimodal vision further includes an illumination dynamic correction step between step 1) and step 2). Scene illumination intensity distribution data are acquired simultaneously when obtaining the original multispectral image data; Based on the light intensity distribution data, light unevenness compensation is performed on each pixel in the original multispectral image data to generate light-corrected multispectral image data. The illumination-corrected multispectral image data is used for subsequent spatiotemporal registration and feature fusion.

[0061] In this embodiment, scene illumination intensity distribution data is simultaneously acquired when obtaining raw multispectral image data. Specifically, an array of illuminance sensors or the metering module built into the multispectral camera is deployed at the data acquisition station to record the spatial distribution of illumination intensity on the carcass surface in real time. After obtaining the illumination intensity distribution data, for each pixel in the raw multispectral image, illumination unevenness compensation is performed based on the illumination intensity value corresponding to that pixel location. The compensation algorithm employs a correction method based on an illumination reflection model, dividing the pixel value by the relative illumination intensity value at that point, or using a polynomial fitting method to eliminate the influence of illumination gradients, generating illumination-corrected multispectral image data. After correction, overly bright or dark areas in the image that were originally caused by illumination unevenness regain normal spectral reflectance characteristics, making the spectral response of the same tissue tend to be consistent under different illumination conditions. This illumination-corrected multispectral image data is then used for spatiotemporal registration and feature fusion in step two, thereby ensuring that subsequent processing is based on image data that reflects the tissue's own properties.

[0062] This technical solution addresses the problem of image data distortion caused by changes in ambient lighting, which affects the accuracy of subsequent processing, by simultaneously acquiring illumination intensity distribution data during multispectral image acquisition and compensating for uneven illumination in each pixel. After correction, the spectral responses of the same tissue at different locations in the image tend to be consistent, making the extracted tissue texture clarity and spectral matching indices more representative of the tissue's characteristics. This improves the rationality of adaptive weighted fusion and the accuracy of deep learning model recognition. Illumination correction provides robust data preprocessing for the entire segmentation method, enhancing its adaptability to complex real-world lighting environments.

[0063] According to one embodiment of the present invention, the method for livestock and poultry meat tissue identification and intelligent segmentation based on multimodal vision further includes a three-dimensional geometric model reconstruction and hole filling step between step 2) and step 3). Based on the multimodal feature map after registration and fusion, the geometric contour information of the three-dimensional point cloud data is extracted to reconstruct the three-dimensional geometric model of the carcass or its target cutting part; Identify hole regions in the three-dimensional geometric model, where the hole regions correspond to spatial units where the original three-dimensional point cloud data is missing or the confidence level is below a preset threshold. Obtain the multispectral texture features of the hole region at the corresponding position in the multimodal feature map; Based on the multispectral texture features, semantically guided geometric filling is performed on the hole region to generate a complete and continuous three-dimensional geometric model. The complete three-dimensional geometric model is used for the planning and generation of the three-dimensional segmentation path in step 4).

[0064] This embodiment first extracts the geometric contour information of the 3D point cloud data based on the registered and fused multimodal feature map, and generates an initial 3D geometric model of the carcass using a surface reconstruction algorithm. Then, it identifies hole regions in this 3D geometric model. These holes correspond to spatial units that are missing from the original 3D point cloud data or have a confidence level below a preset threshold, such as areas that cannot be scanned due to occlusion or surface reflection. For each hole region, it acquires the multispectral texture features of the corresponding location in the multimodal feature map, including the multispectral reflectance distribution and texture pattern around the region. Based on these multispectral texture features, semantically guided geometric infilling is performed, specifically using deep learning methods or interpolation algorithms. It utilizes known geometric information about the hole boundaries and internal multispectral texture cues to predict the missing geometric shape within the hole. The semantic guidance is implemented as follows: First, the multispectral texture features of the hole region are input into a pre-trained semantic segmentation network to predict the probability map of the soft tissue category (e.g., muscle, fat) to which the hole region belongs. Then, based on the probability map, the corresponding prior geometric template is retrieved from a pre-set tissue anatomy shape library, or a category-constrained surface fitting algorithm is used to ensure that the filled geometry conforms to the physiological characteristics of the region. During the filling process, it is ensured that the generated geometric surface is consistent with the anatomical features of the surrounding tissue to avoid erroneous fusion across tissue boundaries. Finally, a complete and continuous three-dimensional geometric model is generated. This model is used for the planning and generation of the three-dimensional segmentation path in step four, ensuring that the path is optimized based on geometric integrity.

[0065] This technical solution addresses the problem of incomplete geometry in the original 3D point cloud, which is affected by occlusion or scanning blind spots, by identifying hole regions in the 3D geometric model and using the corresponding multispectral texture features for semantically guided geometric filling. The filled 3D geometric model is not only continuous and complete but also conforms to the anatomical characteristics of the tissue in the region, avoiding geometric errors across tissue boundaries. This complete geometric model provides accurate spatial constraints for subsequent multi-objective optimization to solve the 3D segmentation path, ensuring the physical feasibility of the generated path and further improving the practicality and security of the entire segmentation method.

[0066] According to one embodiment of the present invention, the method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision further includes a post-processing step for recognition results and a path constraint generation step between steps 3) and 4). Obtain the vector field map of the skeletal structure mask, the main fat distribution area mask, and the muscle fiber orientation output by the deep learning model, and simultaneously obtain the prediction confidence score output by the deep learning model for each mask unit. Based on the predicted confidence score, the boundary regions of the skeletal structure mask and the main fat distribution area mask are subjected to confidence-weighted morphological smoothing to generate an optimized mask with continuous closed boundaries. Based on the predicted confidence score, a confidence-weighted streamline filter is applied to the vector field map of the muscle fiber orientation to remove abnormal vectors with confidence scores below a preset threshold, thereby generating an optimized vector field map with spatial continuity. Based on the optimized mask and the optimized vector field map, a set of hard and soft constraints on the three-dimensional segmentation path is generated, wherein the hard constraints include insurmountable bone restricted areas and insulated fat preservation areas, and the soft constraints include a preference for cutting directions that preferentially follow the direction of muscle fibers. The hard and soft constraints are used for setting constraints and pruning the search space when constructing the multi-objective optimization function in step 4).

[0067] This embodiment first obtains the vector field maps of the skeletal structure mask, the main fat distribution region mask, and the muscle fiber orientation output by the deep learning model, and simultaneously obtains the prediction confidence score output by the model for each mask unit. For the skeletal structure mask and the main fat distribution region mask, morphological smoothing processing with confidence weighting is performed on the boundary regions based on the prediction confidence score. For example, in the dilatational erosion operation, the weights of the structuring elements are adjusted according to the confidence score, so that the influence of low-confidence boundary points on the shape is reduced, and finally an optimized mask with continuous closed boundaries is generated. For the vector field map of the muscle fiber orientation, confidence weighted streamline filtering is performed. Each voxel is traversed, and if the vector confidence of the voxel is lower than a preset threshold, the abnormal vector is discarded, and it is repaired by using surrounding high-confidence vectors through interpolation or streamline tracing, generating an optimized vector field map with spatial continuity. Based on the optimized mask and vector field map, a set of hard constraints and soft constraints on the 3D segmentation path is generated. Hard constraints include insurmountable skeletal no-go zones and incuttable fat-preserving zones, which are treated as obstacles in the path search. Soft constraints include a preference for cutting directions that follow the muscle fiber direction, reflected in the optimization function as a penalty term for the angle between the path tangent and the muscle fiber vector. These constraints are used for setting constraints and pruning the search space when constructing the multi-objective optimization function in step four, making the optimization process more efficient and the results more in line with practical requirements.

[0068] This technical solution utilizes the prediction confidence scores output by a deep learning model to perform confidence-weighted boundary smoothing on the bone and fat masks, and confidence-weighted streamline filtering on the muscle fiber vector field. This effectively eliminates low-confidence outlier predictions, generating a continuously closed optimized mask and a continuously smooth optimized vector field. Based on this, hard and soft constraints are further extracted, providing a clear constraint framework and pruning basis for subsequent multi-objective optimization functions. This step solves problems such as boundary noise and outlier vectors in the original recognition results, improving the quality of the input information. Simultaneously, the constraint generation simplifies the complexity of the optimization problem, making the final generated 3D segmentation path more accurate and reliable, and able to meet the requirements of practical cutting processes.

[0069] According to one embodiment of the present invention, the method for livestock and poultry meat tissue identification and intelligent segmentation based on multimodal vision further includes a cutting feasibility verification and path correction step after step 4): Obtain the 3D segmentation path generated in step 4) and use it as the initial path to be verified; Based on the bone structure mask identified in step 3), bone interference detection is performed on the initial path to determine whether the initial path spatially intersects with the bone region. If spatial intersection occurs, then based on the muscle fiber orientation vector field map identified in step 3), while maintaining the overall direction along the muscle fiber, the local path segments where the intersection occurs are fine-tuned to generate a first corrected path that avoids the skeletal region. Based on the preset cutting tool geometry model, the tool reachability simulation is performed on the first correction path to determine whether there are any spatial points on the first correction path that the cutting tool cannot reach. If there are unreachable points, the first modified path is fine-tuned a second time according to the kinematic constraints of the cutting tool geometry model to generate a second modified path that simultaneously satisfies the requirements of skeleton avoidance and tool reachability. The second corrected path is used as the final output 3D segmentation path, which is then output to an automated cutting device for execution.

[0070] This embodiment adds a cutting feasibility verification and path correction step after step four. First, the 3D segmentation path obtained through multi-objective optimization in step four is acquired and used as the initial path to be verified. Then, based on the skeletal structure mask identified in step three, skeletal interference detection is performed on the initial path. The interference detection uses a spatial intersection algorithm, traversing each path point on the initial path and determining whether the point falls within the voxel range occupied by the skeletal mask. If a point falls within the skeletal region, a spatial intersection is determined. When an intersection is detected, based on the muscle fiber orientation vector field map identified in step three, the local path segments where the intersection occurs are fine-tuned while maintaining the overall path along the muscle fiber orientation. The fine-tuning method uses the original path as a basis, searching for bypass paths near the intersection region to ensure that the new path segments avoid the skeletal mask region, while ensuring that the angle between the tangential direction of the bypass segment and the muscle fiber vector field is as small as possible, generating a first corrected path. After completing the skeletal avoidance correction, based on a preset cutting tool geometry model, tool reachability simulation is performed on the first corrected path. The tool geometry model includes the blade shape and size, as well as the kinematic parameters of the robotic arm's end effector. The simulation process checks whether each point on the path is within the tool's reachable space, i.e., whether the tool can reach that point without interference and complete the cutting action. If points that the tool cannot reach are found, the first corrected path is fine-tuned based on the kinematic constraints of the tool geometry model. This includes adjusting the orientation or local offset of path points to ensure the new path meets the tool's movement range and obstacle avoidance requirements, generating a second corrected path. Finally, the second corrected path is used as the final output 3D segmentation path, which is converted into control commands and sent to the automated cutting equipment for execution.

[0071] This technical solution identifies and corrects physical conflicts in the initial path by performing skeletal interference detection and tool accessibility simulation. This solves the problem that mathematically optimal paths may fail in actual cutting due to tool interference or inaccessibility. By fine-tuning the generated final path to simultaneously satisfy skeletal avoidance, muscle fiber orientation, and tool accessibility, the solution ensures that the output path is not only theoretically optimal but also physically executable. This improves the operational reliability and cutting quality of automated cutting equipment and reduces tool wear and production interruptions caused by path issues.

[0072] According to one embodiment of the present invention, the method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision further includes a dynamic target adaptive acquisition and registration compensation step in step 1): During the acquisition of the original multispectral image data and the original three-dimensional point cloud data, the motion state data of the fresh meat carcass of livestock and poultry or its target cutting part are collected simultaneously. The motion state data includes at least one of displacement velocity, rotation angle and vibration frequency. Based on the motion state data, the acquisition timing of the original multispectral image data and the original three-dimensional point cloud data is dynamically adjusted so that the acquisition time of the multispectral camera and the structured light scanning device is synchronized with the moment of minimum displacement in the motion state data. For the acquired multispectral image data and 3D point cloud data, the motion offset of each pixel or each point cloud point is calculated based on the motion state data, and motion compensation correction is performed on the multispectral image data and 3D point cloud data before registration based on the motion offset to generate motion-compensated multispectral image data and 3D point cloud data. Motion-compensated multispectral image data and 3D point cloud data are used for spatiotemporal registration and feature fusion in step 2).

[0073] In this embodiment, motion state data of fresh livestock and poultry carcasses are simultaneously acquired during the acquisition of raw multispectral image data and raw 3D point cloud data. This motion state data is acquired in real time using displacement sensors, angle encoders, and vibration accelerometers installed near the carcass suspension chain. It includes the carcass's displacement velocity, rotation angle, and vibration frequency during transport. Based on this motion state data, the acquisition timing of the raw multispectral image data and raw 3D point cloud data is dynamically adjusted. Specifically, when the motion state data shows that the carcass's displacement velocity is high or its vibration amplitude is large, the control system automatically adjusts the trigger times of the multispectral camera and structured light scanning equipment to synchronize them with the moment of minimum displacement and weakest vibration in the motion state data. This allows data acquisition to be completed at the instant the carcass is relatively stationary, reducing motion blur and misalignment. For multispectral images and 3D point cloud data that have been acquired but whose timing has not been adjusted in time, the motion offset of each pixel or each point cloud point is calculated based on the motion state data. The offset is obtained by multiplying the time difference between the acquisition time and the ideal time by the instantaneous velocity vector. To improve compensation accuracy, the motion state data also includes rotation angle and vibration frequency, based on which a rigid body motion model of the torso is established. Based on this rigid body motion model, the spatial transformation matrix (including translation and rotation components) of each pixel or point cloud point relative to the ideal reference time at the acquisition time is calculated, and motion compensation correction is performed on the data according to the transformation matrix. Then, motion compensation correction is performed on the pre-registration multispectral image data and 3D point cloud data based on these offsets, for example, by moving image pixels or point cloud points in the opposite direction of their offset to restore them to their ideal positions when there is no motion. After motion compensation correction, motion-compensated multispectral image data and 3D point cloud data are generated. These data are then used for spatiotemporal registration and feature fusion in step two, ensuring that the data used for fusion has eliminated distortion caused by motion.

[0074] This embodiment solves the problems of multispectral image blurring, point cloud distortion, and intermodal spatial misalignment caused by carcass motion by synchronously acquiring the carcass's motion state data during the data acquisition process, dynamically adjusting the acquisition timing to capture the moment of minimum motion, and performing motion offset compensation on the acquired data. The data after motion compensation and correction accurately reflects the true geometric and spectral information of the carcass, laying a reliable foundation for subsequent spatiotemporal registration and feature fusion. This step significantly improves the quality of data acquired on a dynamic production line, enhances the system's adaptability to fluctuations in conveyor speed and carcass swaying, thereby ensuring the accuracy of subsequent identification and path planning.

[0075] According to one embodiment of the present invention, a multimodal vision-based system for identifying and intelligently segmenting livestock and poultry meat tissue includes: The data acquisition module is used to acquire the original multispectral image data and original three-dimensional point cloud data of the same fresh meat carcass of livestock and poultry or its target cutting part, wherein the three-dimensional point cloud data is generated based on structured light scanning; The registration and fusion module is used to perform spatiotemporal registration of the original multispectral image data and the original three-dimensional point cloud data, and to perform feature fusion of the registered multispectral image data and the three-dimensional point cloud data to generate a spatially consistent multimodal feature map containing tissue feature information. The feature recognition module is used to input the multimodal feature map into a trained deep learning model, and to identify and segment the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation of the carcass or its target cutting part through the deep learning model. The path planning module is used to construct a multi-objective optimization function based on the vector field map of the skeletal structure mask, the fat distribution area mask, and the direction of muscle fibers, and solve the multi-objective optimization function in the three-dimensional space represented by the multimodal feature map to obtain a three-dimensional segmentation path from a preset starting point to a preset ending point, avoiding the skeletal region, such that the meat yield of the target part segmented along the path meets a predetermined optimization criterion, and extends along the direction of muscle fibers. The three-dimensional segmentation path is used to output to an automatic cutting device for execution.

[0076] This embodiment provides a multimodal vision-based system for livestock and poultry meat tissue recognition and intelligent segmentation. The system includes a data acquisition module, a registration and fusion module, a feature recognition module, and a path planning module. The data acquisition module acquires raw multispectral image data and raw 3D point cloud data of the same fresh livestock and poultry carcass, where the 3D point cloud data is generated based on structured light scanning. This module integrates a multispectral camera and a structured light scanning device, and is equipped with a synchronous trigger controller to ensure that the two types of data correspond to the same carcass at the same time. The registration and fusion module is connected to the data acquisition module and performs spatiotemporal registration of the raw multispectral image data and the raw 3D point cloud data. It then performs feature fusion on the registered data to generate a spatially consistent multimodal feature map containing tissue feature information. Spatiotemporal registration achieves coordinate system transformation through calibration parameters, and feature fusion employs an adaptive weighting algorithm, dynamically allocating weights based on local texture clarity and point cloud density. The feature recognition module, connected to the registration and fusion module, inputs multimodal feature maps into a trained deep learning model. This model identifies and segments the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation from the carcass. The deep learning model is pre-trained using labeled samples and can output prediction results for three branches. The path planning module, connected to the feature recognition module, constructs a multi-objective optimization function based on the vector field map of skeletal structure mask, fat distribution area mask, and muscle fiber orientation. This function is then solved in the three-dimensional space represented by the multimodal feature map to obtain a three-dimensional segmentation path from a preset start point to a preset end point, avoiding skeletal regions, ensuring the meat yield of the segmented target parts meets predetermined optimization criteria, and extending along the muscle fiber orientation. This three-dimensional segmentation path is ultimately output to the automated cutting equipment for execution. The entire system integrates data acquisition, fusion processing, tissue recognition, and path planning functions, forming a complete automated segmentation solution.

[0077] This system comprises a data acquisition module, a registration and fusion module, a feature recognition module, and a path planning module. Through the collaborative work of these modules, the system can automatically acquire and fuse multispectral images and 3D point cloud data, identify the orientation of bones, fat, and muscle fibers, and plan the optimal 3D segmentation path. This system solves the problems of existing equipment relying on manual operation, being unable to automatically fuse multimodal data, and struggling to accurately avoid cutting along bones and muscle fibers. By integrating perception, recognition, and planning, this system achieves full automation of the livestock and poultry meat segmentation process, improving meat yield, cutting standardization, and the level of intelligent production, providing the slaughtering and processing industry with efficient and reliable technical equipment.

[0078] According to one embodiment of the present invention, the registration and fusion module includes an adaptive weighted fusion unit, which is used to: extract the tissue texture clarity index and the point cloud density index of the multispectral image of the multimodal feature map for each spatial unit in the multimodal feature map, and perform nonlinear mapping on the tissue texture clarity index and the point cloud density index according to a preset fusion weight function to dynamically calculate and allocate the fusion weight of the unit.

[0079] The registration and fusion module further includes a confidence assessment and correction unit, used for: calculating a fusion confidence score for each spatial unit after generating the multimodal feature map; determining a dynamic threshold based on the statistical distribution of the fusion confidence scores of all spatial units in the current processing batch; and marking the unit as a low-confidence unit when the fusion confidence score is lower than the dynamic threshold, and reducing the contribution weight of the low-confidence unit to the recognition result when the feature recognition module performs recognition.

[0080] The adaptive weighted fusion unit is further configured to: extract the matching degree index between the multispectral data of each spatial unit and the preset adipose tissue spectral features and the preset muscle tissue spectral features; the preset fusion weight function is further configured to comprehensively map the tissue texture clarity index and the point cloud density index according to the matching degree index, so as to dynamically calculate the fusion weight of the unit.

[0081] It also includes an illumination correction module, which is located between the data acquisition module and the registration and fusion module, and is used to: synchronously acquire scene illumination intensity distribution data when acquiring the original multispectral image data; based on the illumination intensity distribution data, perform illumination unevenness compensation on each pixel in the original multispectral image data to generate illumination-corrected multispectral image data, and provide the illumination-corrected multispectral image data to the registration and fusion module.

[0082] It also includes a 3D geometric model reconstruction module, located between the registration and fusion module and the feature recognition module, for: extracting geometric contour information of the 3D point cloud data based on the multimodal feature map after registration and fusion, and reconstructing the 3D geometric model of the carcass or its target cutting part; identifying hole regions in the 3D geometric model, the hole regions corresponding to spatial units that are missing or have a confidence level below a preset threshold in the original 3D point cloud data; obtaining the multispectral texture features of the hole regions at the corresponding positions in the multimodal feature map; performing semantically guided geometric filling of the hole regions based on the multispectral texture features to generate a complete and continuous 3D geometric model; and providing the complete 3D geometric model to the path planning module for the planning and generation of 3D segmentation paths.

[0083] It also includes a post-processing module for recognition results, located between the feature recognition module and the path planning module, for: Obtain the vector field map of the skeletal structure mask, the main fat distribution area mask, and the muscle fiber orientation output by the deep learning model, and simultaneously obtain the prediction confidence score output by the deep learning model for each mask unit. Based on the predicted confidence score, the boundary regions of the skeletal structure mask and the main fat distribution area mask are subjected to confidence-weighted morphological smoothing to generate an optimized mask with continuous closed boundaries. Based on the predicted confidence score, a confidence-weighted streamline filter is applied to the vector field map of the muscle fiber orientation to remove abnormal vectors with confidence scores below a preset threshold, thereby generating an optimized vector field map with spatial continuity. Based on the optimized mask and the optimized vector field map, a set of hard and soft constraints on the three-dimensional segmentation path is generated, wherein the hard constraints include insurmountable bone restricted areas and insulated fat preservation areas, and the soft constraints include a preference for cutting directions that preferentially follow the direction of muscle fibers. The hard and soft constraints are provided to the path planning module for setting constraints and pruning the search space when constructing the multi-objective optimization function.

[0084] It also includes a path verification and correction module, which is located after the path planning module, and is used for: Obtain the 3D segmentation path generated by the path planning module and use it as the initial path to be verified; Based on the skeletal structure mask identified by the feature recognition module, skeletal interference detection is performed on the initial path to determine whether the initial path spatially intersects with the skeletal region. If spatial intersection occurs, based on the muscle fiber orientation vector field map identified by the feature recognition module, while maintaining the overall direction along the muscle fiber, the local path segments where the intersection occurs are fine-tuned to generate a first corrected path that avoids the skeletal region. Based on the preset cutting tool geometry model, the tool reachability simulation is performed on the first correction path to determine whether there are any spatial points on the first correction path that the cutting tool cannot reach. If there are unreachable points, the first modified path is fine-tuned a second time according to the kinematic constraints of the cutting tool geometry model to generate a second modified path that simultaneously satisfies the requirements of skeleton avoidance and tool reachability. The second corrected path is then used as the final output 3D segmentation path, which is then output to an automatic cutting device for execution.

[0085] The data acquisition module further includes a dynamic target adaptive acquisition and registration compensation unit, used for: During the acquisition of the original multispectral image data and the original three-dimensional point cloud data, the motion state data of the fresh meat carcass of livestock and poultry or its target cutting part are collected simultaneously. The motion state data includes at least one of displacement velocity, rotation angle and vibration frequency. Based on the motion state data, the acquisition timing of the original multispectral image data and the original three-dimensional point cloud data is dynamically adjusted so that the acquisition time of the multispectral camera and the structured light scanning device is synchronized with the moment of minimum displacement in the motion state data. For the acquired multispectral image data and 3D point cloud data, the motion offset of each pixel or each point cloud point is calculated based on the motion state data, and motion compensation correction is performed on the multispectral image data and 3D point cloud data before registration based on the motion offset to generate motion-compensated multispectral image data and 3D point cloud data. The motion-compensated multispectral image data and the three-dimensional point cloud data are then provided to the registration and fusion module.

[0086] This embodiment provides a multimodal vision-based system for livestock and poultry meat tissue recognition and intelligent segmentation. The system integrates multiple functional modules, enabling full automation from data acquisition to path execution. The system includes a data acquisition module, an illumination correction module, a registration and fusion module, a 3D geometric model reconstruction module, a feature recognition module, a recognition result post-processing module, a path planning module, and a path verification and correction module.

[0087] The data acquisition module includes a dynamic target adaptive acquisition and registration compensation unit. This unit simultaneously acquires motion state data of the carcass, including displacement velocity, rotation angle, and vibration frequency, while acquiring raw multispectral image data and raw 3D point cloud data of the same fresh livestock carcass. Based on this motion state data, the unit dynamically adjusts the acquisition timing to synchronize the acquisition time of the multispectral camera and structured light scanning device with the moment of minimum displacement in the motion state data, thus completing data acquisition at the instant when the carcass is relatively stationary. For data that has been acquired but is affected by motion, the unit calculates the motion offset of each pixel or point cloud point based on the motion state data and performs motion compensation correction, generating motion-compensated multispectral image data and 3D point cloud data to ensure that subsequent processing is based on high-quality data without motion distortion.

[0088] The illumination correction module is positioned between the data acquisition module and the registration and fusion module. This module receives the motion-compensated raw multispectral image data and simultaneously acquires the illumination intensity distribution data of the scene being acquired. Based on the illumination intensity distribution data, it compensates for uneven illumination in each pixel of the multispectral image, eliminating image distortion caused by changes in ambient lighting, and generating illumination-corrected multispectral image data. This provides an image that accurately reflects the spectral characteristics of the tissue for subsequent feature extraction.

[0089] The registration and fusion module receives illumination-corrected multispectral image data and motion-compensated 3D point cloud data. It first performs spatiotemporal registration, mapping both data to the same spatial coordinate system. Internally, the module includes an adaptive weighted fusion unit and a confidence assessment and correction unit. The adaptive weighted fusion unit extracts the tissue texture sharpness index of the multispectral image and the point cloud density index of the 3D point cloud for each voxel in the registered 3D space. It also extracts the matching degree index between the multispectral data of that unit and preset adipose tissue spectral features and preset muscle tissue spectral features. Based on a preset fusion weight function, these indices are nonlinearly mapped, dynamically calculating the fusion weight of each unit. The multispectral features are then weighted and fused with the point cloud geometric features to generate a multimodal feature map. After generating the multimodal feature map, the confidence assessment and correction unit calculates a fusion confidence score for each spatial unit. Based on the statistical distribution of scores across all units in the current batch, a dynamic threshold is determined. Units with scores below the threshold are marked as low-confidence units, and this information is passed to the subsequent feature recognition module to reduce their contribution weight to the recognition results.

[0090] The 3D geometric model reconstruction module is positioned between the registration and fusion module and the feature recognition module. Based on the registered and fused multimodal feature map, this module extracts the geometric contour information of the 3D point cloud to reconstruct the initial 3D geometric model of the carcass. It identifies hole regions in the model, which correspond to missing or low-confidence units in the original point cloud. For each hole region, it obtains the multispectral texture features of its corresponding position in the multimodal feature map. Using a semantically guided geometric infilling algorithm, it fills the holes based on the multispectral texture features, generating a complete and continuous 3D geometric model, providing an accurate geometric basis for subsequent path planning.

[0091] The feature recognition module receives multimodal feature maps and a 3D geometric model, and inputs the multimodal feature maps into a trained deep learning model. The deep learning model identifies and segments the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation of the carcass, and simultaneously outputs the predicted confidence score for each mask unit. These recognition results, along with the confidence information, are output together.

[0092] The post-processing module for the recognition results is located between the feature recognition module and the path planning module. This module receives the skeletal mask, fat mask, muscle fiber vector field, and corresponding confidence scores output by the deep learning model. Based on the predicted confidence scores, the boundary regions of the skeletal and fat masks undergo confidence-weighted morphological smoothing to generate an optimized mask with continuous closed boundaries. The muscle fiber orientation vector field map undergoes confidence-weighted streamline filtering to remove abnormal vectors with confidence scores below a preset threshold, generating an optimized vector field map with spatial continuity. Based on the optimized mask and optimized vector field map, a set of hard and soft constraints on the 3D segmentation path is generated. Hard constraints include insurmountable skeletal no-go zones and insulated fat preservation zones, while soft constraints include a preference for cutting directions along the muscle fiber orientation. These constraints are then passed to the path planning module.

[0093] The path planning module constructs a multi-objective optimization function based on optimized mask, optimized vector field, and hard and soft constraints. It solves the function in the space represented by the three-dimensional geometric model to obtain a three-dimensional segmentation path that runs from a preset starting point to a preset ending point, avoids the skeletal region, and ensures that the meat yield of the target part segmented along the path meets the predetermined optimization criteria and extends along the direction of the muscle fibers.

[0094] The path verification and correction module is located after the path planning module. This module acquires the 3D segmentation path generated by the path planning module as the initial path to be verified. Based on the skeletal structure mask identified by the feature recognition module, skeletal interference detection is performed on the initial path. If interference occurs, local path segments are fine-tuned according to the muscle fiber orientation vector field map, while maintaining the overall alignment along the muscle fiber orientation, to generate a first corrected path. Based on a preset cutting tool geometry model, tool reachability simulation is performed on the first corrected path. If there are points that the tool cannot reach, a second fine-tuning is performed according to tool kinematic constraints to generate a second corrected path. Finally, the second corrected path is used as the final output 3D segmentation path and output to the automatic cutting equipment for execution.

[0095] By setting an adaptive weighted fusion unit in the registration and fusion module, the fusion weights are dynamically allocated based on local texture clarity and point cloud density, solving the problem that fixed-weight fusion cannot adapt to changes in local data quality and improving the fusion quality of multimodal feature maps. By adding a confidence evaluation and correction unit, low-confidence fusion units are marked and their recognition contribution is reduced, suppressing the interference of low-quality regions on the recognition results and enhancing the robustness of recognition. Introducing a spectral feature matching index into the adaptive weighted fusion allows the fusion weights to reflect the spectral semantic information of the tissue, accurately preserving the tissue type even when texture and point cloud information is insufficient, thus improving the semantic fidelity of the fused image. By setting an illumination correction module, uneven illumination during acquisition is compensated for, eliminating the influence of ambient light on image data and providing a data foundation that truly reflects tissue characteristics for subsequent processing. By adding a 3D geometric model reconstruction module and using multispectral texture features for semantically guided hole filling, the problem of geometric incompleteness caused by missing point clouds is solved, providing a continuous and accurate geometric model for path planning. By setting up a post-processing module for the recognition results, the mask boundary and vector field are optimized using prediction confidence, and structured constraints are generated. This solves the problem of noise and abnormal vectors affecting path planning in the recognition results, while simplifying the optimization solution process. By adding a path verification and correction module, the planned path undergoes skeletal interference detection and tool reachability simulation and correction, solving the problem of the mathematically optimal path being physically infeasible and ensuring the path's practical executability. By setting up a dynamic target adaptive acquisition and registration compensation unit in the data acquisition module, adaptive acquisition and motion compensation of moving targets are performed, reducing motion artifacts and modal misalignment at the source and providing high-quality data input for the entire system. These claims work together to enable the system to adapt to complex production environments, output high-quality, executable 3D segmentation paths, and significantly improve the automation level and cutting quality of livestock and poultry meat segmentation.

[0096] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.

Claims

1. A method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision, characterized in that, Includes the following steps: 1) Acquire the original multispectral image data and original three-dimensional point cloud data of the same fresh meat carcass of livestock and poultry or its target cutting part, wherein the three-dimensional point cloud data is generated based on structured light scanning; 2) Spatiotemporal registration is performed on the original multispectral image data and the original three-dimensional point cloud data, and feature fusion is performed on the registered multispectral image data and the three-dimensional point cloud data to generate a spatially consistent multimodal feature map containing tissue feature information. 3) Input the multimodal feature map into the trained deep learning model, and use the deep learning model to identify and segment the skeletal structure mask, fat distribution area mask and muscle fiber orientation vector field map of the carcass or its target cutting part; 4) Based on the vector field map of the skeletal structure mask, the fat distribution area mask, and the direction of the muscle fibers, a multi-objective optimization function is constructed, and the multi-objective optimization function is solved in the three-dimensional space represented by the multi-modal feature map to obtain a three-dimensional segmentation path that runs from a preset starting point to a preset ending point, avoids the skeletal area, and ensures that the meat yield of the target part segmented along the path meets the predetermined optimization criteria and extends along the direction of the muscle fibers. The three-dimensional segmentation path is used to output to an automatic cutting device for execution.

2. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 1, characterized in that, In step 2), the fusion adopts an adaptive weighted fusion algorithm. The adaptive weighted fusion algorithm is as follows: for each spatial unit in the multimodal feature map, extract the tissue texture clarity index and the point cloud density index of the multispectral image in the unit, and perform nonlinear mapping on the tissue texture clarity index and the point cloud density index according to the preset fusion weight function to dynamically calculate and allocate the fusion weight of the unit.

3. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 1 or 2, characterized in that, Step 2) also includes confidence assessment and correction steps: After generating the multimodal feature map, a fusion confidence score is calculated for each spatial unit; A dynamic threshold is determined based on the statistical distribution of the fusion confidence scores of all spatial units in the current processing batch; When the fusion confidence score is lower than the dynamic threshold, the unit is marked as a low confidence unit, and in the feature recognition process of step 3), the contribution weight of the low confidence unit to the recognition result is reduced.

4. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 2, characterized in that, In step 2), the adaptive weighted fusion algorithm also includes: For each spatial unit, the matching degree index between the multispectral data of that unit and the preset adipose tissue spectral features and preset muscle tissue spectral features is extracted. The preset fusion weight function also performs a comprehensive mapping of the tissue texture clarity index and the point cloud density index based on the matching degree index, so as to dynamically calculate the fusion weight of the unit.

5. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 1, characterized in that, Between step 1) and step 2) there is also a dynamic illumination correction step: Scene illumination intensity distribution data are acquired simultaneously when obtaining the original multispectral image data; Based on the light intensity distribution data, light unevenness compensation is performed on each pixel in the original multispectral image data to generate light-corrected multispectral image data. The illumination-corrected multispectral image data is used for subsequent spatiotemporal registration and feature fusion.

6. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 5, characterized in that, Between steps 2) and 3) there is also a step of three-dimensional geometric model reconstruction and hole filling: Based on the multimodal feature map after registration and fusion, the geometric contour information of the three-dimensional point cloud data is extracted to reconstruct the three-dimensional geometric model of the carcass or its target cutting part; Identify hole regions in the three-dimensional geometric model, where the hole regions correspond to spatial units where the original three-dimensional point cloud data is missing or the confidence level is below a preset threshold. Obtain the multispectral texture features of the hole region at the corresponding position in the multimodal feature map; Based on the multispectral texture features, semantically guided geometric filling is performed on the hole region to generate a complete and continuous three-dimensional geometric model. The complete three-dimensional geometric model is used for the planning and generation of the three-dimensional segmentation path in step 4).

7. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in any one of claims 3 or 6, characterized in that, Between steps 3) and 4) there is also a post-processing step for the recognition results and a path constraint generation step: Obtain the vector field map of the skeletal structure mask, the main fat distribution area mask, and the muscle fiber orientation output by the deep learning model, and simultaneously obtain the prediction confidence score output by the deep learning model for each mask unit. Based on the predicted confidence score, the boundary regions of the skeletal structure mask and the main fat distribution area mask are subjected to confidence-weighted morphological smoothing to generate an optimized mask with continuous closed boundaries. Based on the predicted confidence score, a confidence-weighted streamline filter is applied to the vector field map of the muscle fiber orientation to remove abnormal vectors with confidence scores below a preset threshold, thereby generating an optimized vector field map with spatial continuity. Based on the optimized mask and the optimized vector field map, a set of hard and soft constraints on the three-dimensional segmentation path is generated, wherein the hard constraints include insurmountable bone restricted areas and insulated fat preservation areas, and the soft constraints include a preference for cutting directions that preferentially follow the direction of muscle fibers. The hard and soft constraints are used for setting constraints and pruning the search space when constructing the multi-objective optimization function in step 4).

8. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 1, characterized in that, Step 4) is followed by steps for verifying the feasibility of cutting and revising the path: Obtain the 3D segmentation path generated in step 4) and use it as the initial path to be verified; Based on the bone structure mask identified in step 3), bone interference detection is performed on the initial path to determine whether the initial path spatially intersects with the bone region. If spatial intersection occurs, then based on the muscle fiber orientation vector field map identified in step 3), while maintaining the overall direction along the muscle fiber, the local path segments where the intersection occurs are fine-tuned to generate a first corrected path that avoids the skeletal region. Based on the preset cutting tool geometry model, the tool reachability simulation is performed on the first correction path to determine whether there are any spatial points on the first correction path that the cutting tool cannot reach. If there are unreachable points, the first modified path is fine-tuned a second time according to the kinematic constraints of the cutting tool geometry model to generate a second modified path that simultaneously satisfies the requirements of skeleton avoidance and tool reachability. The second corrected path is used as the final output 3D segmentation path, which is then output to an automated cutting device for execution.

9. The method for livestock and poultry meat tissue recognition and intelligent segmentation based on multimodal vision as described in claim 1, characterized in that, Step 1) also includes a dynamic target adaptive acquisition and registration compensation step: During the acquisition of the original multispectral image data and the original three-dimensional point cloud data, the motion state data of the fresh meat carcass of livestock and poultry or its target cutting part are collected simultaneously. The motion state data includes at least one of displacement velocity, rotation angle and vibration frequency. Based on the motion state data, the acquisition timing of the original multispectral image data and the original three-dimensional point cloud data is dynamically adjusted so that the acquisition time of the multispectral camera and the structured light scanning device is synchronized with the moment of minimum displacement in the motion state data. For the acquired multispectral image data and 3D point cloud data, the motion offset of each pixel or each point cloud point is calculated based on the motion state data, and motion compensation correction is performed on the multispectral image data and 3D point cloud data before registration based on the motion offset to generate motion-compensated multispectral image data and 3D point cloud data. Motion-compensated multispectral image data and 3D point cloud data are used for spatiotemporal registration and feature fusion in step 2).

10. A multimodal vision-based system for identifying and intelligently segmenting livestock and poultry meat tissue, characterized in that, include: The data acquisition module is used to acquire the original multispectral image data and original three-dimensional point cloud data of the same fresh meat carcass of livestock and poultry or its target cutting part, wherein the three-dimensional point cloud data is generated based on structured light scanning; The registration and fusion module is used to perform spatiotemporal registration of the original multispectral image data and the original three-dimensional point cloud data, and to perform feature fusion of the registered multispectral image data and the three-dimensional point cloud data to generate a spatially consistent multimodal feature map containing tissue feature information. The feature recognition module is used to input the multimodal feature map into a trained deep learning model, and to identify and segment the skeletal structure mask, the main fat distribution area mask, and the vector field map of muscle fiber orientation of the carcass or its target cutting part through the deep learning model. The path planning module is used to construct a multi-objective optimization function based on the vector field map of the skeletal structure mask, the fat distribution area mask, and the direction of muscle fibers, and solve the multi-objective optimization function in the three-dimensional space represented by the multimodal feature map to obtain a three-dimensional segmentation path from a preset starting point to a preset ending point, avoiding the skeletal region, such that the meat yield of the target part segmented along the path meets a predetermined optimization criterion, and extends along the direction of muscle fibers. The three-dimensional segmentation path is used to output to an automatic cutting device for execution.