Deboning method, system, medium, and device based on deboning key points

By employing a deboning method based on multi-stage convolutional neural networks and attention mechanisms, and utilizing images acquired by a color industrial camera to determine key deboning points, the problem of high labor consumption and high error rates associated with manual deboning is solved, thus achieving automated and efficient deboning of meat and bone in livestock and poultry products.

CN115272281BActive Publication Date: 2025-12-05SHANDONG UNIV +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210980543.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2025-12-05
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

Existing technologies for separating meat from bones in livestock and poultry products rely on manual operation, which is labor-intensive, prone to errors, and costly. Furthermore, it is difficult to accurately identify key deboning points, thus hindering the large-scale development of the industry.

Method used

A deboning method based on multi-stage convolutional neural networks and attention mechanisms is adopted. Three color industrial cameras are used to acquire orthogonal color images. The key points of deboning are determined by convolutional pooling and attention parameters to achieve automated deboning.

Benefits of technology

It improves the efficiency and accuracy of deboning, reduces energy consumption and costs, and meets the automation and real-time requirements of the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272281B_ABST
    Figure CN115272281B_ABST
Patent Text Reader

Abstract

The disclosure provides a deboning method, system, medium and equipment based on key points of deboning, including collecting a large number of orthogonal color images, establishing a data set; dividing a region of interest, determining 11 key points of deboning; the image of the region of interest is taken as input to perform a first convolution pooling operation, and the image is converted in the spatial domain; the teaching point corresponding to the output data and the attention parameter of the teaching point are obtained, the second convolution pooling is performed by using the attention parameter, and the network running rate is accelerated through the attention mechanism, the three orthogonal color images can well reflect the spatial information of the deboning part, the key points of deboning can be obtained more accurately, the recognition efficiency of the key points of deboning is improved, and the cost of the production line is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of color image processing technology, and specifically to a bone removal system and method based on a multi-stage convolutional neural network and attention mechanism. Background Technology

[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.

[0003] For a long time, the key technologies and equipment for separating meat from bones in general livestock and poultry products have been outdated. The entire production process is completed manually, which consumes a lot of labor and the technical bottlenecks seriously restrict the development of the industry on a large scale.

[0004] Existing technologies generally use manual methods for separation, resulting in a large number of workers in typical processing enterprises. This leads to relatively high investment in worker protection equipment and other inputs, resulting in insignificant economic benefits.

[0005] Furthermore, the working environment is quite smelly, and the production environment involves meat and bone cutting, which may have a certain impact on people's physical and mental health in the long run. In the use of manual meat and bone cutting, it is necessary for people to judge and find the key points of deboning based on certain experience, and then determine the key points of deboning before completing the deboning action. The manual identification method has a certain degree of error and the recognition rate is relatively low. The visual acquisition solution is more expensive, time-consuming, and energy-intensive, especially since the grasp of the key points of deboning is not particularly accurate. In addition, the inventors found that the existing deboning technology for animals cannot well reflect the spatial information of the deboning part, and there is a certain degree of error in the acquisition of the key points of deboning. Summary of the Invention

[0006] To address the aforementioned problems, this disclosure proposes a deboning method, system, medium, and device based on key deboning points. By utilizing orthogonal color images and combining them with multi-stage neural network training using an attention mechanism, the correspondence between orthogonal color images and key deboning points is obtained, thus determining the key deboning points, completing the deboning action, and improving the deboning efficiency for livestock.

[0007] According to some embodiments, the present disclosure adopts the following technical solutions:

[0008] Deboning methods based on key deboning points include:

[0009] Collect orthogonal color images of the production site and build a dataset;

[0010] Acquire the image to be processed, preprocess it, and delineate the region of interest;

[0011] The image of the region of interest is used as input for the first convolutional pooling operation, which transforms the image in the spatial domain;

[0012] Obtain the matrix and attention parameters corresponding to the output data of the first-stage convolutional pooling module, and use the attention parameters to perform bone removal operation on the spatial point coordinates of the output of the second convolutional pooling module.

[0013] The bone removal operation is performed using the output spatial point coordinates.

[0014] According to other embodiments, the present disclosure adopts the following technical solutions:

[0015] Osteotomy systems based on key osteotomy points include:

[0016] The data acquisition and receiving module is configured to acquire a set of orthogonal color images, which are captured by three fixed-position color industrial cameras on site.

[0017] The data processing module is configured to define the region of interest;

[0018] The first-stage convolutional pooling module is configured to take the image of the region of interest as input and perform a first convolutional pooling operation to transform the image in the spatial domain;

[0019] The second-stage attention module is configured to obtain the matrix and attention parameters corresponding to the output data of the first-stage convolutional pooling module, and use the attention parameters to perform the spatial point coordinates of the second convolutional pooling output.

[0020] Furthermore, it also includes a data output module, configured to perform bone removal operations using the output spatial point coordinates.

[0021] According to other embodiments, this disclosure also adopts the following technical solutions:

[0022] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the bone removal method based on bone removal key points.

[0023] According to other embodiments, this disclosure also adopts the following technical solutions:

[0024] A terminal device is characterized in that it includes a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, which are adapted to be loaded and executed by the processor for the bone removal method based on bone removal key points.

[0025] Compared with the prior art, the beneficial effects of this disclosure are as follows:

[0026] This invention utilizes three color industrial cameras to acquire image data, which has a much lower cost and energy consumption than other sensors such as linear lasers and X-rays. The three orthogonal color images acquired can well reflect the spatial information of the area to be deboned, and can be used to accurately obtain the key points for deboning.

[0027] The use of attention mechanisms in neural networks offers greater scalability compared to traditional methods. It can improve recognition accuracy by expanding the dataset and significantly increase the network's operating speed during production. This method can effectively help meat processing production lines achieve automation while meeting the requirements of accuracy and real-time performance. Attached Figure Description

[0028] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.

[0029] Figure 1 This is a flowchart of the pig hind leg deboning scheme in Embodiment 1 of this disclosure;

[0030] Figure 2 This is a diagram of the training process of a multi-stage convolutional neural network that incorporates an attention mechanism in Embodiment 1 of this disclosure.

[0031] Figure 3 This is a schematic diagram of the one-stage convolutional pooling module described in the embodiments of this disclosure;

[0032] Figure 4 This is a schematic diagram of the two-stage convolutional pooling module described in the embodiments of this disclosure. Detailed Implementation

[0033] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0034] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0035] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0036] For a long time, my country's key technologies and equipment for separating meat from bones in livestock and poultry products have lagged behind, with the entire production process being done manually, consuming a large amount of labor. This technological bottleneck severely restricts the industry's ability to scale up. On typical pig hind leg deboning production lines, manual deboning is used, increasing labor consumption in this stage. Taking pig hind leg deboning as an example, the inventors discovered that three orthogonal color images can effectively represent the spatial information of the pig hind leg, enabling more accurate identification of key deboning points.

[0037] Attention mechanisms are special structures embedded in machine learning models to automatically learn and calculate the contribution of input data to output data.

[0038] Example 1

[0039] One embodiment of this disclosure provides a deboning method based on key deboning points, taking the deboning of a pig's hind leg as an example, including:

[0040] S101: Collect orthogonal color images of the production site and establish a dataset;

[0041] S102: Acquire the image to be processed, preprocess it, and delineate the region of interest;

[0042] S103: The image of the region of interest is used as input to perform the first convolutional pooling operation, and the image is transformed in the spatial domain;

[0043] S104: Obtain the teaching point corresponding to the output data and the attention parameter of the teaching point, use the attention parameter to perform a second convolution pooling, and output the spatial coordinates of the bone removal part;

[0044] S105: Perform bone removal operation using the output spatial point coordinates.

[0045] Furthermore, a dataset was created by acquiring a set of color images of pig hind legs from the production site using three fixed-position color industrial cameras.

[0046] Define the region of interest and fix the captured image to 1024*1024*3.

[0047] Orthogonal color images of pig hind legs are acquired, and regions of interest (ROIs) are defined. These ROIs are then used as input to a neural network, which undergoes multiple convolutional pooling operations. First, the image of the ROI is used as input for the first convolutional pooling operation, which involves spatial transformation, such as... Figure 2As shown, a set of color images undergoes spatial domain transformation. The image undergoing the first convolutional pooling operation is a 1024*1024*3 image, which is then transformed into a 64*64*64 image in the spatial domain and output. Then, the teaching points on the output image and the attention parameters corresponding to the teaching points are obtained. The obtained attention parameters are used to perform a second convolutional pooling operation to output the spatial coordinates of the bone removal area. The bone removal action is then performed using the spatial coordinates.

[0048] Specifically, 11 teaching points were obtained, and the process of obtaining these teaching points is as follows:

[0049] The process of obtaining teaching points involves mimicking the actions of human experts, breaking down the bone-removing action, and obtaining key points of the trajectory. These key points can be used to interpolate and obtain the bone-removing trajectory.

[0050] The teaching point is obtained by installing a tool on the collaborative robot, setting the TCP to the tool tip, and manually operating the robot to make the tool tip touch the teaching point. The coordinates of the teaching point in the robot coordinate system can then be obtained.

[0051] The teaching trajectory is obtained by interpolation algorithm. Circular interpolation is used for circumferential trajectories, and linear interpolation is used for other trajectories.

[0052] Next, the deboning process was completed according to the taught trajectory. First, a simulated experimental environment, mimicking the production line environment, was established, and the pig's hind leg was fixed in place. Second, under the guidance of skilled workers, a collaborative robot was operated to complete the deboning work. Finally, the taught motion was optimized. There were two optimization methods: one was to adjust the position of the taught points based on the experimental results to achieve better deboning results; the other was to reduce some taught points based on geometric relationships (e.g., if point A can be derived from points B and C, then point A can be deleted), ultimately resulting in 11 independent taught points. The number of taught points was reduced.

[0053] The final teaching points are as follows:

[0054] The pig's hind leg, to be deboned, is suspended on a track. The first cut severs the hind elbow muscle to expose the bone. The second cut rotates around the lower side of the tibia and fibula, dissecting the muscle below the hind elbow. The third cut cuts the muscle above the hind elbow along the outer side of the joint, then turns around the knee joint to cut the thigh muscle. The fourth cut cuts the muscle above the hind elbow along the inner side of the joint, then turns to cut the thigh muscle. The output spatial coordinates are the coordinates of ordered points, i.e., the coordinates of 11 ordered points (P1-P11) are learned and output through a neural network.

[0055] The 11 key points for deboning are as follows: The first cut is a vertical incision, requiring only two points: the initial cut and the stopping point. The remaining nine points are grouped into threes, distributed at the entry point, turning point, and stopping point. At each position, the three points are located on either side and below the bone, thus defining four surfaces that form the trajectory of the final two cuts. The stopping point of the first cut is also the starting point of the second circular cut. The radius of the circular arc is obtained from the distance between this point and a point below the bone at the entry point. The coordinates of these key deboning points are obtained through robot teaching.

[0056] After spatially transforming the acquired orthogonal color images, the coordinates of the obtained bone-removing key points are used as input. Attention parameters are obtained through an attention mechanism, specifically a spatial attention mechanism, which is a 64*64 attention parameter matrix corresponding to the data output from the first convolutional pooling. Different channels share the same attention parameter matrix. The value of the attention parameter represents the degree of attention given to that point, gradually decreasing outwards from the teaching point. Backpropagation is used to iterate the attention parameters. This method sets the attention parameters of points farther from the teaching point to zero, reducing computational load.

[0057] The attention parameters correspond to 11 teaching points, and the initial configuration of the attention parameters is 1.

[0058] The results of the first convolutional pooling are filtered using the attention parameters, and a second convolutional pooling is performed to output the spatial coordinates of the bone removal site. After the second convolutional pooling outputs the coordinates, the robot is then operated to perform the action.

[0059] Example 2

[0060] One embodiment of this disclosure provides a deboning system based on deboning key points, comprising:

[0061] The data acquisition and receiving module is configured to acquire a set of orthogonal color images, which are captured by three fixed-position color industrial cameras on site.

[0062] The data processing module is configured to define the region of interest;

[0063] The first-stage convolutional pooling module is configured to take the image of the region of interest as input and perform a first convolutional pooling operation to transform the image in the spatial domain;

[0064] The second-stage attention module is configured to obtain the 64*64*64 matrix output by the first-stage convolutional pooling module and the attention parameters corresponding to the output data. After filtering the 64*64*64 matrix using the attention parameters, a second convolutional pooling is performed to output the spatial coordinates of the bone removal site.

[0065] Specifically, such as Figure 3As shown, the first-stage convolutional pooling module is configured to take the image of the region of interest as input and perform multi-layer convolutional pooling operations to transform the image in the spatial domain, compressing the input 1024*1024 image to 64*64, and expanding the original RGB3 channels to 64 channels.

[0066] The second-stage attention module is configured to obtain the 64*64*64 matrix output by the first-stage convolutional pooling module and the attention parameters corresponding to the output data. After filtering the 64*64*64 matrix using the attention parameters, a second convolutional pooling is performed. The convolutional pooling process is as follows: Figure 4 As shown, the space is compressed from 64*64 to 1*1, and the channel is further expanded to 256. Then, the spatial coordinates of the bone removal site are obtained through high-level information.

[0067] It also includes a data output module, configured to perform bone removal operations using the output spatial point coordinates.

[0068] Dataset collection: A set of orthogonal color images were captured using three industrial cameras, and the spatial coordinates of 11 key points were obtained through robot teaching. A large amount of data was collected to build a dataset.

[0069] Network training: Input the dataset into the network for training, with orthogonal color images as input and keypoint spatial coordinates as output;

[0070] Production line application: A set of orthogonal color images is collected online, input into a trained network to obtain the coordinates of the key points for deboning, and then transmitted to the robot to complete the deboning action.

[0071] The following methods are specifically executed using the above system:

[0072] Collect orthogonal color images of the production site and build a dataset;

[0073] Acquire the image to be processed, preprocess it, and delineate the region of interest;

[0074] The image of the region of interest is used as input for the first convolutional pooling operation, which transforms the image in the spatial domain;

[0075] Obtain the 64*64*64 matrix output by the first-stage convolutional pooling module and the attention parameters corresponding to the output data, and use the attention parameters to perform a second convolutional pooling;

[0076] The bone removal operation is performed using the output spatial point coordinates.

[0077] Specifically, 11 teaching points were obtained, and the process of obtaining these teaching points is as follows:

[0078] First, a simulated experimental environment was established to mimic the production line environment, and the pig's hind legs were fixed in place. Second, under the guidance of skilled workers, a collaborative robot was operated to complete the deboning work. Finally, the teaching actions were optimized, and the number of teaching points was reduced.

[0079] The final teaching points are as follows:

[0080] The pig's hind leg, to be deboned, is suspended on a track. The first cut severs the hind elbow muscle to expose the bone. The second cut rotates around the lower side of the tibia and fibula, dissecting the muscle below the hind elbow. The third cut cuts the muscle above the hind elbow along the outer side of the joint, then turns around the knee joint to cut the thigh muscle. The fourth cut cuts the muscle above the hind elbow along the inner side of the joint, then turns to cut the thigh muscle. The output spatial coordinates are the coordinates of ordered points, i.e., the coordinates of 11 ordered points (P1-P11) are learned and output through a neural network.

[0081] The 11 key points for deboning are as follows: The first cut is a vertical incision, requiring only two points: the initial cut and the stopping point. The remaining nine points are grouped into threes, distributed at the entry point, turning point, and stopping point. At each position, the three points are located on either side and below the bone, thus defining four surfaces that form the trajectory of the final two cuts. The stopping point of the first cut is also the starting point of the second circular cut. The radius of the circular arc is obtained from the distance between this point and a point below the bone at the entry point. The coordinates of these key deboning points are obtained through robot teaching.

[0082] After spatially transforming the acquired orthogonal color images, the coordinates of the obtained bone-removing key points are used as input, and attention parameters are obtained through an attention mechanism.

[0083] The attention parameters correspond to 11 teaching points, and the initial configuration of the attention parameters is 1.

[0084] The attention parameters are used to perform a second convolutional pooling, which outputs the spatial coordinates of the bone removal site. After outputting the coordinates through the second convolutional pooling, the robot is then operated to perform the action.

[0085] Example 3

[0086] One embodiment of this disclosure provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the bone removal method based on bone removal key points.

[0087] Example 4

[0088] One embodiment of this disclosure provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, which are adapted to be loaded and executed by the processor for the bone removal method based on bone removal key points.

[0089] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0090] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0091] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0093] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.

Claims

1. A deboning method based on deboning key points, characterized in that, The method comprises the following steps: Collecting orthogonal color images of a production site to establish a data set; wherein a set of orthogonal color images of the production site are obtained by three fixed-position color industrial cameras; Obtaining an image to be processed, preprocessing and demarcating a region of interest; Conducting a first convolutional pooling operation on the image of the region of interest as input, and converting the image in the spatial domain; Obtaining a matrix output by the one-stage convolutional pooling module and attention parameters corresponding to the output data, and conducting a second convolutional pooling operation on the spatial point coordinates output by the second convolutional pooling operation to perform a bone removal operation; The first knife is a vertical cut, including a lower knife and a stop point, the stop point of the first knife is also the starting point of the second knife, and the radius of the circular cut is obtained from the distance between the starting point and the point below the cut position; in addition to the lower knife point and the stop point, the remaining 9 points are divided into 3 groups, which are distributed at the cut position, the turning position and the stop position, and each position has three points on both sides and below the bone, and four surfaces are determined, which are the trajectories of the last two knives.

2. The debone keypoint-based deboning method of claim 1, wherein, The image subjected to the first convolutional pooling operation is a 1024*1024*3 image, which is converted into a 64*64*64 image in the spatial domain. 3.The debone keypoint-based deboning method of claim 1, wherein, The obtained teaching points are 11, and the attention parameters are initially configured as 1.

4. The debone keypoint-based deboning method of claim 1, wherein, The spatial point coordinates output are the coordinates of the points in sequence.

5. A deboning system based on deboning key points, specifically implementing the deboning method based on deboning key points as claimed in any one of claims 1-4, characterized in that, The method comprises the following steps: The data acquisition receiving module is configured to obtain a set of orthogonal color images by three fixed-position color industrial cameras in the field; The data processing module is configured to demarcate a region of interest; The one-stage convolutional pooling module is configured to conduct a first convolutional pooling operation on the image of the region of interest as input, and convert the image in the spatial domain; The two-stage attention module is configured to obtain a matrix output by the one-stage convolutional pooling module and attention parameters corresponding to the output data, filter the matrix by using the attention parameters, and conduct a second convolutional pooling operation to output the spatial point coordinates of the bone removal part.

6. The debone keypoint-based deboning system of claim 5, wherein, Further comprising a data output module configured to perform a bone removal operation by using the output spatial point coordinates.

7. A computer readable storage medium characterized in that, A plurality of instructions are stored in the computer readable storage medium, and the instructions are suitable for being loaded and executed by the processor of the terminal device to implement the bone removal method based on the bone removal key point.

8. A terminal device, comprising: The computer readable storage medium is used to store a plurality of instructions, and the instructions are suitable for being loaded and executed by the processor to implement the bone removal method based on the bone removal key point.

Citation Information

Patent Citations

  • Unhealthy image differentiating method based on robust visual attention feature and sparse representation

    CN102034107A

  • End-to-end behavior recognition method and system based on self-adaptive space-time attention mechanism

    CN111401177A