A method and device for estimating the weight of a target object based on a structured light image

By using a structured light image-based method, multiple image features and geometric dimensions of the target object are acquired and fused, solving the problem of inaccurate weight estimation of 2D images in existing technologies and achieving high-precision target object weight estimation.

CN120853156BActive Publication Date: 2026-08-04HEFEI LASSETER ROBOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI LASSETER ROBOT TECH CO LTD
Filing Date
2025-07-11
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing methods for estimating the weight of target objects based on 2D images cannot accurately reflect the true weight of the target object, as they ignore three-dimensional features such as depth and thickness, resulting in inaccurate weight estimation results.

Method used

A structured light image-based approach is adopted, which acquires multiple images of a structured light scene, performs image segmentation and feature extraction, combines grayscale images and grayscale mask images, and uses a weight prediction model to determine the weight of the target object, fusing image texture features and spatial geometric features.

Benefits of technology

It achieves high-precision target object weight estimation, improves the accuracy and stability of weight estimation, and is applicable to target objects with different postures and shapes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853156B_ABST
    Figure CN120853156B_ABST
Patent Text Reader

Abstract

The application provides a target object weight estimation method and device based on a structured light image, relates to the field of image processing, and solves the problem that the weight estimation result is not comprehensive and accurate due to the same pixel area of different target objects on a 2D image and the difference in body shape. The method comprises the following steps: acquiring multiple images collected by a camera under different structured light scenes, performing image segmentation on a target object in a first image to obtain a segmented image of the segmented target object, and determining a segmented gray image and a gray mask image based on the segmented image; determining the width of the target object based on a second image, determining the length of the target object based on a third image, and determining the height of the target object based on the second image and the third image; and inputting the gray image, the gray mask image, the length, the width and the height of the target object into a weight prediction model to determine the weight of the target object. The application is used in the target object weight estimation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and in particular to a method and apparatus for weight estimation of target objects based on structured light images. Background Technology

[0002] Existing weight estimation methods are mainly based on 2D images, while the actual target object is three-dimensional. Estimating weight solely based on the pixel area of ​​a 2D image ignores features such as depth and thickness. Even if different target objects have the same pixel area in a 2D image, their actual weight may differ due to variations in body shape (e.g., size, length). This makes the weight estimation results incomplete and inaccurate, failing to accurately reflect the true weight of the target object. Therefore, a more precise method for estimating the weight of target objects is urgently needed. Summary of the Invention

[0003] This application provides a target object weight estimation method and apparatus based on structured light images, which solves the technical problem of inaccurate weight estimation results when using 2D images to estimate the weight of target objects in the prior art.

[0004] To achieve the above objectives, this application adopts the following technical solution:

[0005] In a first aspect, a target object weight estimation method based on structured light images is provided, comprising: acquiring multiple images captured by a camera under different structured light scenes, the multiple images including: a first image of the target object captured in a scene where the structured light source is not turned on; a second image of the target object captured when two structured lights are respectively projected onto a first preset position and a second preset position of the target object; the first preset position and the second preset position are parallel to the width direction of the target object; a third image of the target object captured when the structured light is projected along the central axis of the target object; performing image segmentation on the target object in the first image to obtain a segmented image of the target object, and determining a segmented grayscale image and a grayscale mask image based on the segmented image; determining the width of the target object based on the second image, determining the length of the target object based on the third image, and determining the height of the target object based on the second image and the third image; and inputting the grayscale image, the grayscale mask image, the length, width, and height of the target object into a weight prediction model to determine the weight of the target object.

[0006] In conjunction with the first aspect above, in one possible implementation, determining the segmented grayscale image and grayscale mask image based on the segmented image includes: setting pixels other than those corresponding to the target object in the segmented image to a first pixel value to obtain a grayscale image; and setting the pixel values ​​of pixels in the grayscale image that are not the first pixel value to a second pixel value to obtain a grayscale mask image.

[0007] In conjunction with the first aspect described above, in one possible implementation, the first preset position is the position where the width of the front side of the target object satisfies the first condition, and the second preset position is the position where the width of the rear side of the target object satisfies the second condition. Determining the width of the target object based on the second image includes: acquiring the first endpoint and the second endpoint of the first preset position, and acquiring the third endpoint and the fourth endpoint of the second preset position; calculating the first actual distance between the first endpoint and the second endpoint based on the actual vertical distance from the first endpoint and the second endpoint to the camera, the pixel distance between the first endpoint and the second endpoint, and the camera calibration parameters; calculating the second actual distance between the third endpoint and the fourth endpoint based on the actual vertical distance from the third endpoint and the fourth endpoint to the camera, the pixel distance between the third endpoint and the fourth endpoint, and the camera calibration parameters; and determining the width of the target object based on the first actual distance and the second actual distance.

[0008] In conjunction with the first aspect mentioned above, in one possible implementation, the actual vertical distances from the first endpoint, the second endpoint, and the camera satisfy the following formula:

[0009]

[0010] Among them, l i H represents the distance of the corresponding pixel from the camera. camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. ' The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y' represents the vertical pixel distance of a structured light pixel from the image center. This indicates the camera's vertical field of view.

[0011] In conjunction with the first aspect above, in one possible implementation, the first actual distance between the first endpoint and the second endpoint, and the second actual distance between the third endpoint and the fourth endpoint, satisfy the following formula:

[0012]

[0013] Where, d j ε represents the distance between two corresponding target points in actual space, and ε represents the camera's pixel size. j This represents the horizontal pixel distance between two pixels, Q' represents the width of the camera target area, and l i This indicates the distance of the corresponding pixel from the camera. This indicates the camera's horizontal field of view.

[0014] In conjunction with the first aspect above, in one possible implementation, determining the length of the target object based on the third image includes: acquiring the fifth endpoint and the sixth endpoint, which are the two endpoints on the target object projected onto the central axis; calculating the third actual distance between the fifth endpoint and the sixth endpoint based on the actual vertical distance from the fifth endpoint and the sixth endpoint to the camera, the pixel distance between the fifth endpoint and the sixth endpoint, and the camera calibration parameters, and determining the length of the target object.

[0015] In conjunction with the first aspect mentioned above, in one possible implementation, determining the height of the target object based on the second and third images includes:

[0016] The average vertical distance from the first endpoint, second endpoint, third endpoint, fourth endpoint, fifth endpoint, and sixth endpoint to the camera is determined as the vertical actual distance from the target object to the camera; the difference between the camera installation height and the vertical actual distance from the target object to the camera is determined as the height of the target object.

[0017] In conjunction with the first aspect above, in one possible implementation, the grayscale image, grayscale mask image, length, width, and height of the target object are input into the weight prediction model to determine the weight of the target object, including:

[0018] Feature extraction and fusion are performed on the grayscale image, grayscale mask image, and the width and length of the target object to obtain the first fused feature; the height of the target object is linearly transformed, and the extracted feature is fused with the first fused feature to obtain the second fused feature; linear regression prediction is performed on the second fused feature to determine the weight of the target object.

[0019] In conjunction with the first aspect mentioned above, in one possible implementation, features are extracted and fused from the grayscale image, the grayscale mask image, and the width and length of the target object to obtain a first fused feature, including:

[0020] Based on the length of the target object, the grayscale image and the grayscale mask image are unified to the same scale as the length of the target object to obtain the first grayscale image and the first grayscale mask image; the first grayscale image and the first grayscale mask image are input into the feature extraction model to extract features to obtain grayscale image fusion features; the length and width of the target object are respectively linearly transformed and fused with the grayscale image fusion features to obtain the first fusion feature.

[0021] In a second aspect, an electronic device is provided, comprising: a communication unit and a processing unit;

[0022] A communication unit is used to acquire multiple images captured by the camera under different structured light scenarios. These multiple images include: a first image of the target object captured in a scenario where the structured light source is not activated; a second image of the target object captured when two structured lights are projected onto a first preset position and a second preset position of the target object, respectively; the first and second preset positions are parallel to the width direction of the target object; and a third image of the target object captured when the structured light is projected along the central axis of the target object. A processing unit performs image segmentation on the target object in the first image to obtain a segmented image of the target object, and determines the segmented grayscale image and grayscale mask image based on the segmented image; determines the width of the target object based on the second image, the length of the target object based on the third image, and the height of the target object based on the second and third images; and inputs the grayscale image, grayscale mask image, length, width, and height of the target object into a weight prediction model to determine the weight of the target object.

[0023] Thirdly, this application provides an electronic device, including: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to implement the methods described in the first aspect and any possible implementation thereof. This electronic device may be an electronic device or a chip within an electronic device.

[0024] Fourthly, this application provides a target object weight estimation system, including: a camera and an electronic device; wherein the camera is used to acquire multiple images under different structured light scenarios, the multiple images including: a first image of the target object acquired in a scenario where the structured light source is not turned on; a second image of the target object acquired when two structured lights are respectively projected onto a first preset position and a second preset position of the target object; the first preset position and the second preset position are parallel to the width direction of the target object; a third image of the target object acquired when the structured light is projected along the central axis of the target object; the electronic device performs image segmentation on the target object in the first image to obtain a segmented image of the segmented target object, and determines a segmented grayscale image and a grayscale mask image based on the segmented image; determines the width of the target object based on the second image, determines the length of the target object based on the third image, and determines the height of the target object based on the second image and the third image; and inputs the grayscale image, the grayscale mask image, the length, width, and height of the target object into a weight prediction model to determine the weight of the target object.

[0025] Fifthly, this application provides a computer-readable storage medium storing instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first aspect and any possible implementation thereof.

[0026] In a sixth aspect, this application provides a computer program product containing instructions that, when run on an electronic device, cause the electronic device to perform the methods described in the first aspect and any possible implementation thereof.

[0027] This application provides a target object weight estimation method based on structured light images. It utilizes multi-line structured light to acquire multiple images of the target object at different projection angles, extracting its grayscale image, segmentation mask image, and target object dimensions (length, width, and height). This information is then input into a weight prediction model to achieve high-precision target weight estimation. This method fuses image texture features with spatial geometric features. Structured light is used to accurately acquire the three-dimensional shape of the target object, image segmentation improves the model's robustness, and size parameters enhance the discriminative ability of weight prediction. This achieves more accurate, stable, and universal intelligent weight estimation, solving the technical problem of inaccurate weight estimation results when using 2D images for target object weight estimation in existing technologies.

[0028] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0029] Figure 1 A system architecture diagram of a target object weighting system based on structured light images provided in this application embodiment;

[0030] Figure 2 A flowchart illustrating a target object weight estimation method based on structured light images provided in this application embodiment;

[0031] Figure 3 A schematic diagram of the camera mounting structure for another target object weight estimation method based on structured light images provided in this application embodiment;

[0032] Figure 4 A schematic diagram of the target object for another target object weighting method based on structured light images provided in this application embodiment;

[0033] Figure 5 A flowchart illustrating another target object weight estimation method based on structured light images provided in this application embodiment;

[0034] Figure 6 A flowchart illustrating another target object weight estimation method based on structured light images provided in this application embodiment;

[0035] Figure 7 A flowchart illustrating another target object weight estimation method based on structured light images provided in this application embodiment;

[0036] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0037] Figure 9 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application; Detailed Implementation

[0038] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0039] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0040] The target object weight estimation method provided in this application embodiment can be used for, for example... Figure 1 In the target object weighting system shown, such as Figure 1 As shown, the system includes a camera 101 and an electronic device 102.

[0041] The camera 101 is used to acquire multiple images under different structured light scenarios. These images include: a first image of the target object acquired in a scenario where the structured light source is not activated; a second image of the target object acquired when two structured lights are projected onto a first preset position and a second preset position, respectively; the first and second preset positions are parallel to the width direction of the target object; and a third image of the target object acquired when the structured light is projected along the central axis of the target object. The electronic device 102 is used to perform image segmentation on the target object in the first image, obtaining a segmented image of the target object, and determining a segmented grayscale image and a grayscale mask image based on the segmented image; determining the width of the target object based on the second image, the length of the target object based on the third image, and the height of the target object based on the second and third images; and inputting the grayscale image, the grayscale mask image, and the length, width, and height of the target object into a weight prediction model to determine the weight of the target object.

[0042] Figure 2 This is a flowchart illustrating the target object weight estimation method provided in the embodiments of this application, as shown below. Figure 2 As shown, the method includes:

[0043] Step 201: The camera acquires multiple images under different structured light scenarios.

[0044] The multiple images include: a first image of the target object captured in a scene where the structured light source is not turned on; a second image of the target object captured when two structured lights are projected onto a first preset position and a second preset position of the target object, respectively; the first preset position and the second preset position are parallel to the width direction of the target object; and a third image of the target object captured when the structured light is projected along the central axis of the target object.

[0045] As an example, this embodiment selects two RGB wide-angle cameras with relatively low distortion, mounted side-by-side. Camera 1 has no filter, while camera 2 has a filter. Two symmetrical structured light sources are mounted vertically on camera 2, with an angle θ between the structured light sources and the camera's optical axis; this embodiment does not limit the specific angle. The three-view diagram of the camera assembly structure in this embodiment is shown below. Figure 3 As shown.

[0046] In one possible implementation, the camera is adjusted so that the direction of the line structured light is perpendicular to the central axis of the head and tail of the standing target object. Without turning on the light source, camera 1 is activated to acquire a first image, which is the original image of the target object. Then, the light source is turned on, and camera 2 is activated to acquire a second image, which is the image of the target object with a calculated width. Finally, the camera is rotated 90° along the optical axis so that one of the structured lights is projected along the central axis of the target object, and a third image is acquired, which is the image of the target object with a calculated length.

[0047] As an example, in an embodiment of this application, such as Figure 3 Taking pigs as an example, the first, second, and third images of the pigs were collected in the manner described above. Structured light rays were used to image the pigs to obtain their length and width.

[0048] Step 202: The electronic device performs image segmentation on the target object in the first image to obtain a segmented image of the segmented target object, and determines the segmented grayscale image and grayscale mask image based on the segmented image.

[0049] In one possible implementation, the electronic device performs image segmentation on the target object in the first image to obtain a segmented image representing the target object region. Based on the segmented image, the electronic device generates a segmented grayscale image and a grayscale mask image. The grayscale image is used to represent the image region of the target object, preserving the detailed information of the target object. The grayscale mask image is used to represent the boundary region of the target object, providing clear boundary information.

[0050] As an example, in an embodiment of this application, the electronic device determines the segmented grayscale image and grayscale mask image based on the segmented image in the following specific implementation: the electronic device sets the pixels other than the pixels corresponding to the target object in the segmented image to the first pixel value to obtain the grayscale image; the electronic device sets the pixel values ​​of the pixels in the grayscale image that are not the first pixel value to the second pixel value to obtain the grayscale mask image.

[0051] Step 203: The electronic device determines the width of the target object based on the second image, the length of the target object based on the third image, and the height of the target object based on the second and third images.

[0052] In one possible implementation, structured light illuminates the back of the target object and the ground, forming bright lines in the image. The electronic device selects two lateral endpoints of the structured light at the first and second positions on the target object, respectively, and calculates their pixel spacing. Based on the camera calibration parameters, structured light device parameters, and imaging geometry, combined with the actual vertical distance from the target point to the camera, trigonometric geometric formulas are applied to determine the average lateral distance of the target object, i.e., its width. Based on the third image, the electronic device identifies the two endpoints of the structured light bright line on the central axis of the target object and obtains its longitudinal pixel distance. Combining the camera calibration parameters and structured light projection position parameters, the actual length of this segment is calculated as the length of the pig. The electronic device identifies four endpoints of the target object width marker in the second image and two endpoints of the target object length marker in the third image. The difference between the actual vertical distances from the six endpoints to the camera and the actual vertical distance between the camera and the ground is calculated as the height of the target object.

[0053] It should be noted that, in the embodiments of this application, pixel distance refers to the distance between two pixels in the image coordinate system, while actual distance refers to the actual physical distance from the target corresponding to the pixel in the image to another target in the real world.

[0054] Step 204: The electronic device inputs the grayscale image, grayscale mask image, length, width and height of the target object into the weight prediction model to determine the weight of the target object.

[0055] In one possible implementation, the electronic device stitches together a grayscale image and a grayscale mask image, which are processed to the same scale as the actual length, width, and height of the target object, as two channels to form a dual-channel image. This image is then input into a fusion network for feature extraction. The fused features are then fused with the length and width of the target object to obtain a first fused feature. The first fused feature is then fused with the height of the target object to obtain a second fused feature. The weight of the target object is then predicted based on the second fused feature.

[0056] As an example, some application embodiments use splicing or neural networks to fuse image features. In this application embodiment, the shape, texture and edge features of the target area under different poses of the target object are effectively extracted. The MobileNetV4 model is used to fuse grayscale images and grayscale mask images. This application embodiment does not limit this.

[0057] This application provides a target object weight estimation method. By fusing 2D feature information such as the target object's shape, texture, and edges with body shape feature information such as length, width, and height, the method effectively extracts relevant feature information for target object weight estimation, enabling accurate weight estimation of the target object under different postures and body shape conditions. This solves the technical problem of inaccurate weight estimation results when using 2D images to estimate the weight of target objects in the prior art.

[0058] In one possible implementation, combining Figure 2 ,like Figure 5 The specific implementation steps for determining the width of the target object based on the second image in step 203 above can be achieved through steps 501-504, which are explained in detail below:

[0059] Step 501: The electronic device acquires the first endpoint and the second endpoint of the first preset position, and acquires the third endpoint and the fourth endpoint of the second preset position.

[0060] In one possible implementation, the electronic device acquires a first endpoint and a second endpoint at a first preset position, and acquires a third endpoint and a fourth endpoint at a second preset position. The first preset position is located in the front region of the target object's back, and the second preset position is located in the rear region of the target object's back. Based on the structured light image or image segmentation results, the electronic device identifies the edge points of the structured light projection in the corresponding regions, and determines the coordinate positions of the first, second, third, and fourth endpoints in the image.

[0061] As an example, in this application embodiment, pigs are used as an example, such as... Figure 4 As shown in the second image of the pig, structured light is projected onto the pig's back to form an image. Points A and B in the image are selected as the two endpoints of the pig's first preset position, and points C and D are the two endpoints of the pig's second preset position.

[0062] Step 502: The electronic device calculates the first actual distance between the first endpoint and the second endpoint based on the actual vertical distances from the first endpoint and the second endpoint to the camera, the pixel distance between the first endpoint and the second endpoint, and the camera calibration parameters.

[0063] As an example, the actual vertical distance l from the first endpoint to the camera A Satisfy the following formula:

[0064]

[0065] Among them, H camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. ' The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y... ' A This represents the vertical pixel distance from the first endpoint of the structured light to the image center. This indicates the camera's vertical field of view.

[0066] As an example, the actual vertical distance l from the second endpoint to the camera B Satisfy the following formula:

[0067]

[0068] Among them, H camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. ' The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y... ' B This represents the vertical pixel distance from the second endpoint of the structured light to the image center. This indicates the camera's vertical field of view.

[0069] As an example, the first actual distance d from the first endpoint to the second endpoint AB Satisfy the following formula:

[0070]

[0071] Where ε represents the camera's pixel size, pix w_AB The horizontal pixel distance between the corresponding pixels at the first and second endpoints is represented by Q', where Q' represents the width of the camera target surface. AB This represents the average actual vertical distance between the first and second endpoints and the camera. This indicates the camera's horizontal field of view.

[0072] Step 503: The electronic device calculates the second actual distance between the third endpoint and the fourth endpoint based on the actual vertical distances from the third endpoint and the fourth endpoint to the camera, the pixel distance between the third endpoint and the fourth endpoint, and the camera's calibration parameters.

[0073] As an example, the actual vertical distance l from the third endpoint to the camera C Satisfy the following formula:

[0074]

[0075] Among them, H camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. ' The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y... ' C This represents the vertical pixel distance from the third endpoint of the structured light to the image center. This indicates the camera's vertical field of view.

[0076] As an example, the actual vertical distance l from the fourth endpoint to the camera D Satisfy the following formula:

[0077]

[0078] Among them, H camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. ' The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y... ' D This represents the vertical pixel distance from the fourth endpoint of the structured light to the image center. This indicates the camera's vertical field of view.

[0079] As an example, the second actual distance d from the third endpoint to the fourth endpoint CD Satisfy the following formula:

[0080]

[0081] Where ε represents the camera's pixel size, pix w_CD The horizontal pixel distance between the corresponding pixels at the third and fourth endpoints is represented by Q', where Q' represents the width of the camera target surface. CD This represents the average actual vertical distance between the third and fourth endpoints and the camera. This indicates the camera's horizontal field of view.

[0082] Step 504: The electronic device determines the width of the target object based on the first actual distance and the second actual distance.

[0083] As an example, in this embodiment, the average of the first actual distance and the second actual distance is used as the width of the target object, and the width w of the target object satisfies the following formula:

[0084]

[0085] Where, d AB d represents the first actual distance. CD This indicates the second actual distance.

[0086] It should be noted that, in this embodiment, the calculation method for the length of the target object is similar to the calculation method for the width of the target object, specifically as follows:

[0087] The electronic device acquires the fifth and sixth endpoints, which are the two endpoints of the central axis projected onto the target object. Based on the actual vertical distances from the fifth and sixth endpoints to the camera, the pixel distance between the fifth and sixth endpoints, and the camera's calibration parameters, the electronic device calculates a third actual distance between the fifth and sixth endpoints to determine the length of the target object.

[0088] As an example, the actual vertical distance l from the fifth endpoint to the camera E Satisfy the following formula:

[0089]

[0090] Among them, H camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. 'The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y... ' E This represents the vertical pixel distance from the fifth endpoint of the structured light to the image center. This indicates the camera's vertical field of view.

[0091] As an example, the actual vertical distance l from the sixth endpoint to the camera D Satisfy the following formula:

[0092]

[0093] Among them, H camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. ' The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y... ' F This represents the vertical pixel distance from the sixth endpoint of the structured light to the image center. This indicates the camera's vertical field of view.

[0094] As an example, the length L of the target object pig Satisfy the following formula:

[0095]

[0096] Where ε represents the camera's pixel size, pix w_EF This represents the lateral pixel distance between the corresponding pixels at the fifth and sixth endpoints, where Q' represents the width of the camera target surface. EF This represents the average actual vertical distance between the fifth and sixth endpoints and the camera. Indicates the horizontal field of view of the camera

[0097] In this embodiment, the method of calculating the actual length and width of a target object based on structured light images enables non-contact, high-precision dimensional measurement of the target object. Structured light forms stable light stripes on the surface of the target object. The electronic device, combining the pixel distance between key points in the image, the structured light projection angle, and camera calibration parameters, accurately reconstructs the true physical dimensions of the corresponding area based on trigonometric geometric relationships. This method effectively avoids measurement errors caused by image distortion, viewing angle deviation, or environmental complexity, improving the stability and robustness of dimensional measurement. Simultaneously, the acquired actual length and width information serves as an important input to the weight estimation model, enhancing the model's accurate representation of the target object's body shape and improving the accuracy and generalization ability of weight prediction.

[0098] In one possible implementation, combining Figure 5 ,like Figure 6 Determining the height of the target object can be achieved through the following steps 601-602, which are explained in detail below:

[0099] Step 601: The electronic device determines the average value of the actual vertical distances from the first endpoint, the second endpoint, the third endpoint, the fourth endpoint, the fifth endpoint, and the sixth endpoint to the camera, which is the actual vertical distance from the target object to the camera.

[0100] In one possible implementation, in this embodiment, the electronic device calculates the actual vertical distances from the first endpoint, second endpoint, third endpoint, fourth endpoint, fifth endpoint, and sixth endpoint to the camera, and takes the average of the six actual vertical distances as the actual vertical distance from the target object as a whole to the camera. The positions of the above six endpoints are the endpoint positions used to calculate the length and width of the target object. By using multi-point averaging, errors caused by changes in the target object's posture, local undulations, or abnormal structured light can be effectively reduced, improving the stability and accuracy of the actual vertical distance calculation.

[0101] As an example, the actual vertical distance from the target object to the camera satisfies the following formula:

[0102]

[0103] Step 602: The electronic device determines the difference between the camera installation height and the actual vertical distance from the target object to the camera, i.e., the height of the target object.

[0104] In one possible implementation, the electronic device in this application determines the camera's installation height and, based on the actual vertical distance from the target object to the camera calculated in the aforementioned steps, calculates the difference between the two as the height of the target object. The camera installation height is a fixed physical height where the camera lens optical axis is perpendicular to the ground, and is typically calibrated during system deployment. By subtracting the camera installation height from the actual vertical distance from the target object to the camera, the height of the target object relative to the ground can be accurately obtained.

[0105] In this embodiment, the height value, as one of the important spatial parameters for measuring the body shape of the target object, can provide key input features for the subsequent weight estimation model. Combined with the length and width of the target object, it provides effective body shape information of the target object and improves the accuracy and stability of weight estimation.

[0106] In one possible implementation, combining Figure 2 ,like Figure 7In step 204 above, the grayscale image, grayscale mask image, length, width, and height of the target object are input into the weight prediction model to determine the weight of the target object. The specific implementation steps for this are steps 701-703, which are described in detail below:

[0107] Step 701: The electronic device extracts and fuses features from the grayscale image, the grayscale mask image, and the width and length of the target object to obtain the first fused feature.

[0108] In one possible implementation, the electronic device unifies the grayscale image and grayscale mask image to the same scale as the target object length, obtaining a first grayscale image and a first grayscale mask image. The first grayscale image and the first grayscale mask image are then input into a feature extraction model to extract features, resulting in grayscale image fusion features. The target object length and target object width are then linearly transformed and fused with the grayscale image fusion features to obtain a first fusion feature.

[0109] As an example, in this embodiment of the application, the first grayscale image and the first grayscale mask image are concatenated, and features are extracted using the MobilenetV4 model. The specific implementation includes the following:

[0110] As an example, the input image is first downsampled and its features extracted to quickly reduce the spatial resolution and extract basic visual features such as edges and textures at a low level.

[0111] As an example, the F1 feature extracted by the initial downsampling satisfies the following formula:

[0112]

[0113] in, This indicates downsampling using a 3×3 standard convolution with a stride of 2, BN() represents batch normalization, σ represents the ReLU() activation function, and X... input This represents the input image.

[0114] As an example, in this embodiment of the application, multiple general inverted bottleneck blocks are then used to progressively extract mid-to-high-level semantic features.

[0115] As an example, mid-to-high-level semantic features F N Satisfy the following formula:

[0116] F i+1 =UIB i (F i ), i = 1, 2, ..., N

[0117] One of the basic formulas in the UIB module is:

[0118]

[0119] Z2=σ(BN(DW k×k (Z1)))

[0120]

[0121] F i+1 =F i +Z3

[0122] Among them, PW 1×1 It is pointwise convolution, DW k×k It is a depthwise convolution, σ() is the activation function GELU(), and BN() represents batch normalization.

[0123] As an example, in this embodiment, an attention mechanism is finally introduced to enhance the response of key target regions, suppress invalid background, improve the model's ability to model spatial context, and use an optimized attention mechanism to maintain computational efficiency and extract fused features.

[0124] As an example, the fused feature F a Satisfy the following formula:

[0125] F a =MobileMQA(F N )

[0126] MobileMQA refers to the space reduction achieved by adding DW convolutions with a stride of 2 to K and V in the attention mechanism. MobileMQA only uses a single Key and Value header, which is shared by all queries.

[0127] As an example, in this embodiment of the application, the width and length of the target object are extracted using a fully connected layer, and then fused with the resulting feature F. a splicing.

[0128] As an example, the target object width feature W p Satisfy the following formula:

[0129] W p =W w W pig +b w

[0130] Among them, W w The weight matrix W represents the width. pig Indicates the width of the target object, b w This represents the bias vector.

[0131] As an example, the target object length feature Lp Satisfy the following formula:

[0132] L p =W L L pig +b L

[0133] Among them, W L The weight matrix L represents the length. pig b represents the length of the target object. L This represents the bias vector.

[0134] As an example, the first fusion feature F out Satisfy the following formula:

[0135] F out =concat(F a W p ,L p )

[0136] Among them, F a W represents the features obtained by fusing the first grayscale image and the first grayscale mask image. p L represents the width feature of the target object. p This indicates the length characteristic of the target object.

[0137] Step 702: The electronic device performs a linear transformation on the height of the target object and fuses the extracted features with the first fusion feature to obtain the second fusion feature.

[0138] In one possible implementation, this embodiment of the application performs a linear transformation on the height value of the target object obtained from the aforementioned calculation, mapping the original numerical height feature to an embedding vector consistent with the image feature dimension. Specifically, the height value is transformed through a fully connected neural network layer to adapt it to the representation structure of the subsequent depth feature space. The transformed height vector is then concatenated and fused with the first fusion feature output by the image feature extraction module as structural information to obtain a second fusion feature containing target geometric size information and image texture semantic information.

[0139] As an example, the height vector H p Satisfy the following formula:

[0140] H p =W H H pig +b H

[0141] Among them, W L The weight matrix H represents the height. pig b represents the height of the target object. H This represents the bias vector.

[0142] As an example, the second fusion feature F fin Satisfy the following formula:

[0143] F fin =concat(F out H p )

[0144] Among them, F out H represents the first fusion feature. p This represents the height vector.

[0145] Step 703: Perform linear regression prediction on the second fusion feature to determine the weight of the target object.

[0146] In one possible implementation, the second fusion feature in this embodiment includes a high-dimensional feature representation obtained by fusing image features and structural size features. This fusion feature is input to a linear regression module, and the output is a scalar value representing the predicted weight of the target object.

[0147] As an example, the predicted weight y' satisfies the following formula:

[0148] y′=wf vec +b

[0149] f vec =GAP(F fin )

[0150] Where w represents the weight vector, b represents the bias term, and f vec This represents the vector obtained by averaging the second fused feature. GAP() is the function that averages the input features along the spatial dimension.

[0151] In this embodiment, by jointly modeling deep semantic features and geometric dimensional information, the two-dimensional image information of the target object (such as 2D features like shape, edges, and texture) is effectively fused with the target object's dimensional information (such as length, width, and height) calculated by structured light. This fusion method not only preserves the image features' ability to depict the details of the target's appearance but also introduces spatial structural constraints at the physical scale, thereby enhancing the model's overall ability to express the target's body structure. Furthermore, the system uses linear regression to model and predict the fused features, achieving accurate estimation of the target object's weight. This method improves weight estimation accuracy while exhibiting good robustness and adaptability, making it applicable to target objects with different postures and body shapes. It also possesses practical value and deployment efficiency in large-scale scenarios.

[0152] The foregoing mainly describes the solutions of the embodiments of this application from the perspective of device implementation. It is understood that each device, such as an electronic device, includes at least one of the hardware structures and software modules corresponding to the execution of each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software-driven hardware manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0153] This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0154] When using integrated units, Figure 8 A possible structural schematic diagram of the electronic device (referred to as electronic device 80) involved in the above embodiments is shown. The electronic device 80 includes a processing unit 801 and a communication unit 802, and may also include a storage unit 803. Figure 8 The structural diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.

[0155] when Figure 8 The schematic diagram shown is used to illustrate the structure of the electronic device involved in the above embodiments. The processing unit 801 is used to control and manage the operation of the electronic device, the communication unit 802 is used for the electronic device to communicate with other devices, and the storage unit 803 is used to store the program code and data of the electronic device.

[0156] For example, communication unit 802 is used to acquire multiple images captured by the camera under different structured light scenarios. These multiple images include: a first image of the target object captured in a scenario where the structured light source is not activated; a second image of the target object captured when two structured lights are projected onto a first preset position and a second preset position of the target object, respectively; the first and second preset positions are parallel to the width direction of the target object; and a third image of the target object captured when the structured light is projected along the central axis of the target object. Processing unit 801 is used to perform image segmentation on the target object in the first image to obtain a segmented image of the target object, and to determine the segmented grayscale image and grayscale mask image based on the segmented image; to determine the width of the target object based on the second image, the length of the target object based on the third image, and the height of the target object based on the second and third images; and to input the grayscale image, grayscale mask image, length, width, and height of the target object into a weight prediction model to determine the weight of the target object.

[0157] In one possible implementation, the processing unit 801 is further configured to set the pixels other than the pixels corresponding to the target object in the segmented image to the first pixel value to obtain a grayscale image; and set the pixel values ​​of the pixels in the grayscale image that are not the first pixel value to the second pixel value to obtain a grayscale mask image.

[0158] In one possible implementation, the first preset position is the position where the width of the front side of the target object satisfies a first condition, and the second preset position is the position where the width of the rear side of the target object satisfies a second condition. Determining the width of the target object based on the second image includes: acquiring the first endpoint and the second endpoint of the first preset position, and acquiring the third endpoint and the fourth endpoint of the second preset position; calculating the first actual distance between the first endpoint and the second endpoint based on the actual vertical distance from the first endpoint and the second endpoint to the camera, the pixel distance between the first endpoint and the second endpoint, and the camera calibration parameters; calculating the second actual distance between the third endpoint and the fourth endpoint based on the actual vertical distance from the third endpoint and the fourth endpoint to the camera, the pixel distance between the third endpoint and the fourth endpoint, and the camera calibration parameters; and determining the width of the target object based on the first actual distance and the second actual distance.

[0159] In one possible implementation, the actual vertical distances from the first endpoint, the second endpoint, and the camera satisfy the following formula:

[0160]

[0161] Among them, l i H represents the distance of the corresponding pixel from the camera. camera_light H represents the distance H between the light-emitting aperture of the line structured light source and the optical axis of the camera. 'The height represents the camera target surface size, θ represents the angle between the optical axis of the line structured light source and the camera imaging optical axis, ε represents the camera pixel size, and Y' represents the vertical pixel distance of a structured light pixel from the image center. This indicates the camera's vertical field of view.

[0162] In one possible implementation, the first actual distance between the first endpoint and the second endpoint, and the second actual distance between the third endpoint and the fourth endpoint, satisfy the following formula:

[0163]

[0164] Where, d j ε represents the distance between two corresponding target points in actual space, and ε represents the camera's pixel size. j This represents the horizontal pixel distance between two pixels, Q' represents the width of the camera target area, and l i This indicates the distance of the corresponding pixel from the camera. This indicates the camera's horizontal field of view.

[0165] In one possible implementation, the processing unit 801 is further configured to determine the length of the target object based on the third image, including:

[0166] Obtain the fifth and sixth endpoints, which are the two endpoints of the central axis projected onto the target object; based on the actual vertical distances from the fifth and sixth endpoints to the camera, the pixel distance between the fifth and sixth endpoints, and the camera calibration parameters, calculate the third actual distance between the fifth and sixth endpoints to determine the length of the target object.

[0167] In one possible implementation, the processing unit 801 is further configured to determine the height of the target object based on the second image and the third image, including:

[0168] The average vertical distance from the first endpoint, second endpoint, third endpoint, fourth endpoint, fifth endpoint, and sixth endpoint to the camera is determined as the vertical actual distance from the target object to the camera; the difference between the camera installation height and the vertical actual distance from the target object to the camera is determined as the height of the target object.

[0169] In one possible implementation, the processing unit 801 is further configured to input the grayscale image, the grayscale mask image, and the length, width, and height of the target object into the weight prediction model to determine the weight of the target object, including:

[0170] Feature extraction and fusion are performed on the grayscale image, grayscale mask image, and the width and length of the target object to obtain the first fused feature; the height of the target object is linearly transformed, and the extracted feature is fused with the first fused feature to obtain the second fused feature; linear regression prediction is performed on the second fused feature to determine the weight of the target object.

[0171] In one possible implementation, the processing unit 801 is further configured to extract and fuse features from the grayscale image, the grayscale mask image, and the width and length of the target object to obtain a first fused feature, including:

[0172] Based on the length of the target object, the grayscale image and the grayscale mask image are unified to the same scale as the length of the target object to obtain the first grayscale image and the first grayscale mask image; the first grayscale image and the first grayscale mask image are input into the feature extraction model to extract features to obtain grayscale image fusion features; the length and width of the target object are respectively linearly transformed and fused with the grayscale image fusion features to obtain the first fusion feature.

[0173] The processing unit 801 can be a processor or a controller, and the communication unit 802 can be a communication interface, transceiver, transceiver circuit, transceiver device, etc. The term "communication interface" is a general term and may include one or more interfaces. The storage unit 803 can be a memory. When the electronic device 80 is a chip, the processing unit 801 can be a processor or a controller, and the communication unit 802 can be an input interface and / or an output interface, pins, or circuits, etc. The storage unit 803 can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip (e.g., read-only memory (ROM), random access memory (RAM, etc.).

[0174] The communication unit can also be called a transceiver unit. The antenna and control circuit with transceiver functions in the electronic device 80 can be considered as the communication unit 802 of the electronic device 80, and the processor with processing functions can be considered as the processing unit 801 of the electronic device 80. Optionally, the device in the communication unit 802 used to implement the receiving function can be considered as the communication unit. The communication unit is used to execute the receiving steps in the embodiments of this application, and the communication unit can be a receiver, a receiver circuit, etc. The device in the communication unit 802 used to implement the transmitting function can be considered as the transmitting unit. The transmitting unit is used to execute the transmitting steps in the embodiments of this application, and the transmitting unit can be a transmitter, a transmitter, a transmitting circuit, etc.

[0175] Figure 8If the integrated units in the process are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of software products. These computer software products are stored in a storage medium and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. Storage media for storing computer software products include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0176] Figure 8 The units in the process can also be called modules; for example, a processing unit can be called a processing module.

[0177] This application also provides a hardware structure diagram of an electronic device (referred to as electronic device 90), see [link to diagram]. Figure 9 The electronic device 00 includes a processor 901, and optionally, a memory 902 connected to the processor 901.

[0178] In the first possible implementation, see Figure 9 The electronic device 90 also includes a transceiver 903. The processor 901, memory 902, and transceiver 903 are connected via a bus. The transceiver 903 is used to communicate with other devices or communication networks. Optionally, the transceiver 903 may include a transmitter and a receiver. The device in the transceiver 903 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the transceiver 903 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.

[0179] Based on the first possible implementation method Figure 9 The structural diagram shown can be used to illustrate the structure of the electronic device involved in the above embodiments.

[0180] in, Figure 9 This can also be illustrated by a system chip in an electronic device. In this case, the actions performed by the aforementioned electronic device can be implemented by this system chip; the specific actions performed can be found above and will not be repeated here.

[0181] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0182] The processor in this application may include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., and other computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor may be a standalone semiconductor chip or integrated with other circuits into a single semiconductor chip. For example, it may form a System-on-a-Chip (SoC) with other circuits (such as encoding / decoding circuits, hardware acceleration circuits, or various bus and interface circuits), or it may be integrated as a built-in processor within an ASIC. The ASIC with the integrated processor may be packaged separately or together with other circuits. In addition to the cores for executing software instructions to perform calculations or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), or logic circuits that implement dedicated logic operations.

[0183] The memory in the embodiments of this application may include at least one of the following types: read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; or electrically erasable programmable-only memory (EEPROM). In some scenarios, the memory may also be a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0184] This application also provides a computer-readable storage medium including instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0185] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0186] This application also provides a chip including a processor and an interface circuit. The interface circuit is coupled to the processor. The processor is used to run computer programs or instructions to implement the above-described method. The interface circuit is used to communicate with other modules outside the chip.

[0187] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0188] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, the disclosure, and the appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0189] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of this application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

Claims

1. A method for estimating weight of a target object based on a structured light image, characterized in that, include: The camera acquires multiple images under different structured light scenarios, including: a first image of the target object acquired in a scenario where the structured light source is not turned on; a second image of the target object acquired when two structured lights are projected onto a first preset position and a second preset position of the target object, respectively; the first preset position and the second preset position are parallel to the width direction of the target object; and a third image of the target object acquired when the structured light is projected along the central axis of the target object. The target object in the first image is segmented to obtain a segmented image of the target object, and a segmented grayscale image and a grayscale mask image are determined based on the segmented image; The width of the target object is determined based on the second image, the length of the target object is determined based on the third image, and the height of the target object is determined based on the second image and the third image; The grayscale image, the grayscale mask image, and the length, width, and height of the target object are input into the weight prediction model to determine the weight of the target object; The first preset position is the position where the width of the front side of the target object meets the first condition, and the second preset position is the position where the width of the rear side of the target object meets the second condition; Determining the width of the target object based on the second image includes: Obtain the first and second endpoints of the first preset position, and obtain the third and fourth endpoints of the second preset position; Based on the actual vertical distances from the first endpoint and the second endpoint to the camera, the pixel distance between the first endpoint and the second endpoint, and the calibration parameters of the camera, calculate the first actual distance between the first endpoint and the second endpoint; Based on the actual vertical distances from the third endpoint and the fourth endpoint to the camera, the pixel distance between the third endpoint and the fourth endpoint, and the calibration parameters of the camera, calculate the second actual distance between the third endpoint and the fourth endpoint; The width of the target object is determined based on the first actual distance and the second actual distance; The actual vertical distances from the first endpoint, the second endpoint, and the camera satisfy the following formula: in, This indicates the distance of the corresponding pixel from the camera. This indicates the distance between the light-emitting aperture of the line structured light source and the optical axis of the camera. Indicates the height of the camera target surface. This represents the angle between the optical axis of the line structured light source and the imaging optical axis of the camera. Indicates the camera's pixel size. This represents the vertical pixel distance of a specific pixel in the structured light source from the center of the image. Indicates the camera's vertical field of view; The first actual distance between the first endpoint and the second endpoint, and the second actual distance between the third endpoint and the fourth endpoint, satisfy the following formula: in, This represents the distance between two pixels corresponding to the target in actual space. Indicates the camera's pixel size. This represents the horizontal pixel distance between two pixels. The width of the camera target surface indicates its size. This indicates the distance of the corresponding pixel from the camera. This indicates the camera's horizontal field of view.

2. The method according to claim 1, characterized in that, The step of determining the segmented grayscale image and grayscale mask image based on the segmented image includes: The pixels in the segmented image other than the pixels corresponding to the target object are set to the first pixel value to obtain the grayscale image; The pixel values ​​of pixels in the grayscale image that are not the first pixel value are set to the second pixel value to obtain the grayscale mask image.

3. The method according to claim 1, characterized in that, Determining the length of the target object based on the third image includes: Obtain the fifth and sixth endpoints, which are the two endpoints on which the central axis is projected onto the target object; Based on the actual vertical distances from the fifth endpoint and the sixth endpoint to the camera, the pixel distance between the fifth endpoint and the sixth endpoint, and the calibration parameters of the camera, the third actual distance between the fifth endpoint and the sixth endpoint is calculated to determine the length of the target object.

4. The method according to claim 3, characterized in that, Determining the height of the target object based on the second image and the third image includes: The average vertical actual distance from the first endpoint, the second endpoint, the third endpoint, the fourth endpoint, the fifth endpoint, and the sixth endpoint to the camera is determined as the vertical actual distance from the target object to the camera. The difference between the camera's installation height and the actual vertical distance from the target object to the camera is determined as the height of the target object.

5. The method according to claim 1, characterized in that, The step of inputting the grayscale image, the grayscale mask image, and the length, width, and height of the target object into the weight prediction model to determine the weight of the target object includes: The grayscale image, the grayscale mask image, and the width and length of the target object are subjected to feature extraction and fusion to obtain the first fused feature; The height of the target object is linearly transformed, and the extracted features are fused with the first fusion feature to obtain the second fusion feature; The weight of the target object is determined by performing linear regression prediction on the second fusion feature.

6. The method according to claim 5, characterized in that, The step of extracting and fusing features from the grayscale image, the grayscale mask image, and the width and length of the target object to obtain a first fused feature includes: Based on the length of the target object, the grayscale image and the grayscale mask image are unified to the same scale as the length of the target object to obtain the first grayscale image and the first grayscale mask image; The first grayscale image and the first grayscale mask image are input into the feature extraction model to extract features and obtain grayscale image fusion features; The length and width of the target object are linearly transformed and then fused with the grayscale image fusion feature to obtain the first fusion feature.

7. A target object weight estimation device based on structured light images, characterized in that, The device includes: a communication unit and a processing unit; The communication unit is used to acquire multiple images captured by the camera under different structured light scenarios. The multiple images include: a first image of the target object captured in a scenario where the structured light source is not turned on; a second image of the target object captured when two structured lights are respectively projected onto a first preset position and a second preset position of the target object; the first preset position and the second preset position are parallel to the width direction of the target object; and a third image of the target object captured when the structured light is projected along the central axis of the target object. The processing unit performs image segmentation on the target object in the first image to obtain a segmented image of the target object, and determines a segmented grayscale image and a grayscale mask image based on the segmented image; determines the width of the target object based on the second image, determines the length of the target object based on the third image, and determines the height of the target object based on the second image and the third image; and inputs the grayscale image, the grayscale mask image, the length, width, and height of the target object into a weight prediction model to determine the weight of the target object. The first preset position is the position where the width of the front side of the target object meets the first condition, and the second preset position is the position where the width of the rear side of the target object meets the second condition; Determining the width of the target object based on the second image includes: Obtain the first and second endpoints of the first preset position, and obtain the third and fourth endpoints of the second preset position; Based on the actual vertical distances from the first endpoint and the second endpoint to the camera, the pixel distance between the first endpoint and the second endpoint, and the calibration parameters of the camera, calculate the first actual distance between the first endpoint and the second endpoint; Based on the actual vertical distances from the third endpoint and the fourth endpoint to the camera, the pixel distance between the third endpoint and the fourth endpoint, and the calibration parameters of the camera, calculate the second actual distance between the third endpoint and the fourth endpoint; The width of the target object is determined based on the first actual distance and the second actual distance; The actual vertical distances from the first endpoint, the second endpoint, and the camera satisfy the following formula: in, This indicates the distance of the corresponding pixel from the camera. This indicates the distance between the light-emitting aperture of the line structured light source and the optical axis of the camera. Indicates the height of the camera target surface. This represents the angle between the optical axis of the line structured light source and the imaging optical axis of the camera. Indicates the camera's pixel size. This represents the vertical pixel distance of a specific pixel in the structured light source from the center of the image. Indicates the camera's vertical field of view; The first actual distance between the first endpoint and the second endpoint, and the second actual distance between the third endpoint and the fourth endpoint, satisfy the following formula: in, This represents the distance between two pixels corresponding to the target in actual space. Indicates the camera's pixel size. This represents the horizontal pixel distance between two pixels. The width of the camera target surface indicates its size. This indicates the distance of the corresponding pixel from the camera. This indicates the camera's horizontal field of view.