Point cloud estimation model training method, object volume measurement method and device

By training the point cloud estimation model and using images and point cloud features for point cloud estimation, the problems of time-consuming point cloud acquisition and data loss in the existing technology are solved, and the efficiency and accuracy of point cloud data acquisition are improved.

CN119992253APending Publication Date: 2025-05-13SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072494.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the three-dimensional measurement of objects, the point cloud acquisition method is time-consuming and data is missing, which affects the measurement accuracy; there are also data missing problems during the depth map generation process, which affects the measurement accuracy.

Method used

By obtaining the image to be trained and the point cloud information, the two-dimensional coordinate information of each point cloud is determined, and the point cloud estimation model is trained based on this information. The model includes a feature extraction module and a backbone network, which uses images and point cloud features to estimate point clouds and generate point cloud information for specified areas.

Benefits of technology

It improves the efficiency of point cloud data acquisition, reduces the long point cloud collection process, enhances the accuracy of point cloud information, and thus improves the accuracy of object volume measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992253A_ABST
    Figure CN119992253A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud estimation model training method and an object volume measurement method and device, and the method comprises the steps: obtaining a to-be-trained image and point cloud information, and determining the two-dimensional coordinate information of each point cloud projection in the to-be-trained image; and training a point cloud estimation model based on the to-be-trained image, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of a to-be-estimated position, the to-be-estimated position being a position of expected point cloud information estimation selected from the to-be-trained image. According to the method, the point cloud estimation model is trained based on the to-be-trained image, the three-dimensional coordinate information and the two-dimensional coordinate information of the point cloud and the two-dimensional coordinate information of the to-be-estimated position, the model can capture the global features of the image and the local features of the point cloud, and the accuracy of subsequent point cloud information prediction of the model is improved. And the point cloud information of the specified area position is generated by using the point cloud estimation model, so that a long point cloud collection process is not needed, and the point cloud data collection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a point cloud estimation model training method, an object volume measurement method and a device. Background Art

[0002] In the three-dimensional measurement of objects, point cloud information of the object is usually collected for calculation. Point cloud information is obtained by scanning or sensing the surface of the object to obtain a set of points located in three-dimensional space. By processing these point cloud information, operations such as object size measurement, volume calculation, and shape reconstruction can be performed. However, common point cloud acquisition methods usually rely on technologies such as laser scanning, structured light, and stereo vision, which often take a long time to accumulate data and are time-consuming. In addition, due to the limitations of scanning technology, point cloud information often cannot cover the entire surface of the object. Some areas may not be able to effectively collect enough point cloud information due to limitations of viewing angle, occlusion, or equipment accuracy, resulting in missing data in some areas. The missing areas of the point cloud will directly affect the accuracy of subsequent calculations.

[0003] To solve this problem, depth maps can be used to assist in the acquisition and completion of point clouds. A depth map is an image acquired by a sensor (such as a depth camera), and each pixel value represents the depth information from the camera to different points on the surface of the object. The depth map provides the distance value from the surface of the object to the camera, which can reflect the shape and position of the object. Compared with point cloud information, the data density of the depth map is usually much higher. However, in the process of depth map generation, some areas of the object may lack depth information due to perspective problems, occlusion, sensor accuracy, or surface characteristics of the object (such as poorly reflective surfaces). These missing areas will affect the integrity of the depth map, and thus affect the subsequent measurement accuracy. Summary of the invention

[0004] Based on this, it is necessary to provide a point cloud estimation model training method, object volume measurement method and device for the above-mentioned technical problems, so as to solve at least one of the above-mentioned technical problems.

[0005] An embodiment of the present invention provides a method for training a point cloud estimation model, comprising:

[0006] Acquire the image to be trained and point cloud information, and determine the two-dimensional coordinate information of each point cloud projected on the image to be trained, wherein the point cloud information includes the three-dimensional coordinate information of all point clouds;

[0007] Based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated, a point cloud estimation model is trained, wherein the position to be estimated is the position of the expected estimated point cloud information selected in the image to be trained.

[0008] Optionally, according to a training method for a point cloud estimation model provided by an embodiment of the present invention, the point cloud estimation model includes a first feature extraction module, a second feature extraction module and a backbone network;

[0009] The first feature extraction module includes a first feedforward neural network and a first convolutional neural network; used to extract features from the image to be trained and the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to each of the point clouds;

[0010] The second feature extraction module includes a second feedforward neural network; used to extract features from the two-dimensional coordinate information of any position to be estimated;

[0011] The backbone network is used to perform point cloud estimation on the image to be trained, the first feature information output by the second feature extraction module, and the second feature information output by the first feature extraction module.

[0012] Optionally, according to a training method for a point cloud estimation model provided by an embodiment of the present invention, the backbone network includes a preset number of cascaded transformer layers, wherein the output of each transformer layer serves as the input of the next transformer layer;

[0013] The second convolutional neural network, where the output of the second convolutional neural network is converted into a one-dimensional feature vector:

[0014] Multi-layer cascaded third feed-forward neural network and attention layer;

[0015] The output of the second convolutional neural network, the output of the transformer layer, and the output of the third feedforward neural network are used as inputs of the attention layer;

[0016] The output of the attention layer is used as the input of the next layer of the third feedforward neural network until the last layer of the third feedforward neural network is reached, wherein the result output by the last layer of the third feedforward neural network is used as the final result output by the point cloud estimation model.

[0017] Optionally, according to a method for training a point cloud estimation model provided by an embodiment of the present invention, the training process of the point cloud estimation model is as follows:

[0018] Inputting the image to be trained into a second convolutional neural network for encoding to obtain image encoding feature information;

[0019] Converting the image coding feature information into a one-dimensional feature vector, and mapping the one-dimensional feature vector to obtain a first key-value pair feature, wherein the first key-value pair feature includes a first K value feature and a first V value feature;

[0020] Inputting the first feature information corresponding to any of the point clouds into the cascaded transformer layer to obtain a second key-value pair feature output by each transformer layer, wherein the second key-value pair feature includes a second K value feature and a second V value feature;

[0021] Using the second feature information corresponding to any of the positions to be estimated as the input of the third feedforward neural network of the first layer, so as to encode and map the input feature information using the third feedforward neural network to obtain query feature information;

[0022] Inputting the first key-value pair feature, the query feature information and the second key-value pair feature into the attention layer for feature fusion to obtain target feature information;

[0023] The target feature information is used as the input of the third feedforward neural network of the next layer, so as to return to the step of encoding and mapping the input feature information by using the third feedforward neural network to obtain the query feature information, until the output result of the last layer of the third feedforward neural network is obtained, and the output result is used as the point cloud coordinate information corresponding to the position to be estimated;

[0024] The point cloud estimation model is trained based on the point cloud coordinate information corresponding to the position to be estimated and the preset point cloud labels.

[0025] Optionally, according to a method for training a point cloud estimation model provided by an embodiment of the present invention, the first feature information is obtained based on the following steps:

[0026] For any point cloud:

[0027] Inputting the three-dimensional coordinate information corresponding to the point cloud and the two-dimensional coordinate information corresponding to the point cloud into the first feedforward neural network to obtain a point cloud feature tensor;

[0028] Based on the two-dimensional coordinate information corresponding to the point cloud, an image region having a preset pixel side length is selected in the image to be trained;

[0029] Encoding the image region using the first convolutional neural network to obtain an image feature tensor;

[0030] The point cloud feature tensor and the image feature tensor are concatenated to obtain first feature information.

[0031] Optionally, according to a training method for a point cloud estimation model provided by an embodiment of the present invention, the second feature information is obtained based on the following steps:

[0032] The two-dimensional coordinate information of any of the positions to be estimated is input into the second feedforward neural network to obtain the second feature information.

[0033] The present invention also provides a method for measuring the volume of an object, comprising:

[0034] Obtain target point cloud information and object image of the object to be identified;

[0035] Determine image coordinate information corresponding to a target area, wherein the target area is a region position of desired estimated point cloud information selected from the object image;

[0036] Based on the target point cloud information, the object image and the image coordinate information, point cloud estimation is performed using a point cloud estimation model to obtain predicted point cloud information of the target area, wherein the point cloud estimation model is trained according to the training method of the point cloud estimation model;

[0037] Based on the predicted point cloud information and the target point cloud information, the object volume of the object to be identified is determined.

[0038] Optionally, according to an object volume measurement method provided by an embodiment of the present invention, determining the object volume of the object to be identified based on the predicted point cloud information and the target point cloud information includes:

[0039] Based on the predicted point cloud information and the target point cloud information, determining the minimum circumscribed hexahedron of all point clouds;

[0040] The minimum circumscribed hexahedron is used as the object volume of the object to be identified.

[0041] The present invention also provides a training device for a point cloud estimation model, comprising:

[0042] A first acquisition module is used to acquire the image to be trained and point cloud information, and determine the two-dimensional coordinate information of each point cloud projected on the image to be trained, wherein the point cloud information includes the three-dimensional coordinate information of all point clouds;

[0043] A training module is used to train a point cloud estimation model based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated, wherein the position to be estimated is the position of the expected estimated point cloud information selected in the image to be trained.

[0044] The present invention also provides an object volume measuring device, comprising:

[0045] The second acquisition module is used to acquire the target point cloud information and the object image of the object to be identified;

[0046] A first determination module is used to determine image coordinate information corresponding to a target area, wherein the target area is a region position of desired estimated point cloud information selected from the object image;

[0047] A point cloud estimation module, configured to perform point cloud estimation processing using a point cloud estimation model based on the target point cloud information, the object image and the image coordinate information, to obtain predicted point cloud information of the target area;

[0048] The second determination module is used to determine the object volume of the object to be identified based on the predicted point cloud information and the target point cloud information.

[0049] The present invention also provides a computer device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the above-mentioned point cloud estimation model training method or object volume measurement method when executing the computer-readable instructions.

[0050] The present invention also provides one or more readable storage media storing computer-readable instructions, which, when executed by a processor, implement the above-mentioned point cloud estimation model training method or object volume measurement method.

[0051] The training method of the point cloud estimation model, the object volume measurement method and device include: obtaining the image to be trained and the point cloud information, and determining the two-dimensional coordinate information of each point cloud projected in the image to be trained, wherein the point cloud information includes the three-dimensional coordinate information of all point clouds; based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated, the point cloud estimation model is trained, wherein the position to be estimated is the position of the expected estimated point cloud information selected in the image to be trained. The present invention trains the point cloud estimation model by using the three-dimensional coordinate information and the two-dimensional coordinate information of the image to be trained, the point cloud, and the two-dimensional coordinate information of the position to be estimated. The model can capture the global features of the image and the local features of the point cloud, thereby learning more diverse features and improving the accuracy of the model's subsequent prediction of the point cloud information. In addition, the point cloud estimation model is used to generate point cloud information of a specified area position, without the need for a long point cloud collection process, thereby improving the efficiency of point cloud data acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0053] Figure 1 is a flow chart of a method for training a point cloud estimation model in one embodiment of the present invention;

[0054] Figure 2 is a model structure diagram of a first feature extraction module provided by an embodiment of the present invention;

[0055] Figure 3 is a model structure diagram of a second feature extraction module provided by an embodiment of the present invention;

[0056] Figure 4 is a model structure diagram of a backbone network provided by an embodiment of the present invention;

[0057] Figure 5 is a schematic diagram of a flow chart of a method for measuring the volume of an object in one embodiment of the present invention;

[0058] Figure 6 is a structural schematic diagram of a training device for a point cloud estimation model in one embodiment of the present invention;

[0059] Figure 7 is a structural schematic diagram of an object volume measuring device in one embodiment of the present invention;

[0060] Figure 8 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] The terms used in one or more embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the present invention. The singular forms of "a", "said" and "the" used in one or more embodiments of the present invention are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of the present invention refers to and includes any or all possible combinations of one or more associated listed items.

[0063] In one embodiment, specifically, Figure 1 As shown, Figure 1 is a flow chart of a method for training a point cloud estimation model in an embodiment of the present invention. The present invention provides a method for training a point cloud estimation model, comprising the following steps:

[0064] Step S11, obtaining the image to be trained and the point cloud information, and determining the two-dimensional coordinate information of each point cloud projected in the image to be trained;

[0065] Specifically, the image to be trained and point cloud information are obtained, the point cloud information includes the three-dimensional coordinate information of all point clouds, and then according to the three-dimensional coordinate information of each point cloud, each point cloud is projected into the two-dimensional image to be trained to obtain the projection coordinates of the point cloud in the image to be trained (that is, the two-dimensional coordinate information in this embodiment).

[0066] Step S12, training a point cloud estimation model based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated.

[0067] It should be noted that the position to be estimated is the position of the desired estimated point cloud information selected in the image to be trained, that is, the pixel position where point cloud prediction is required is selected in the two-dimensional image to be trained.

[0068] It should be noted that the point cloud estimation model includes a first feature extraction module, a second feature extraction module and a backbone network; Figure 2 , Figure 2 It is a model structure diagram of the first feature extraction module provided by an embodiment of the present invention. The first feature extraction module includes a first convolutional neural network and a multi-layer cascaded first feedforward neural network. The number of the first feedforward neural network can be set according to actual conditions, for example, 2 layers or 3 layers are set, wherein the first feature extraction module is used to extract features of the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to the training image and each point cloud to obtain the first feature information. The specific feature extraction process is specifically described in the following embodiments and will not be repeated here.

[0069] In addition, refer to Figure 3 , Figure 3 It is a model structure diagram of the second feature extraction module provided by one embodiment of the present invention. The second feature extraction module includes a multi-layer cascaded second feedforward neural network; the number can be set according to actual conditions, for example, 2 or 3 layers are set, and the second feature extraction module is used to extract features of the two-dimensional coordinate information of any position to be estimated to obtain second feature information.

[0070] In addition, refer to Figure 4 , Figure 4It is a model structure diagram of a backbone network provided by an embodiment of the present invention, wherein the backbone network includes a preset number of cascaded transformer layers, wherein the output of each transformer layer is used as the input of the next transformer layer; the backbone network also includes a second convolutional neural network, wherein the output of the second convolutional neural network is converted into a one-dimensional feature vector: the backbone network also includes a multi-layer cascaded third feedforward neural network and an attention layer; the output of the second convolutional neural network, the output of the transformer layer, and the output of the third feedforward neural network are used as the input of the attention layer; it can be understood that the output of the second convolutional neural network, the output of the i-th transformer layer, and the output of the i-th third feedforward neural network are used as the input of the i-th attention layer. The output of the attention layer is used as the input of the next third feedforward neural network until the last third feedforward neural network is reached, wherein the output of the last third feedforward neural network is used as the final result output by the point cloud estimation model. Then, based on the point cloud coordinate information corresponding to the position to be estimated in the final result and the preset point cloud label, the point cloud estimation model is trained, and the specific training process is specifically described in the following embodiments, which will not be repeated here.

[0071] The present invention trains a point cloud estimation model by using the three-dimensional coordinate information and two-dimensional coordinate information of the image to be trained, the point cloud, and the two-dimensional coordinate information of the position to be estimated. The model can capture the global features of the image and the local features of the point cloud, thereby learning more diverse features and improving the accuracy of the model's subsequent prediction of point cloud information. In addition, the point cloud estimation model is used to generate point cloud information at a specified area location, which does not require a long point cloud collection process, thereby improving the efficiency of point cloud data collection.

[0072] In one embodiment of the present invention, the training process of the point cloud estimation model is as follows:

[0073] Inputting the image to be trained into the second convolutional neural network for encoding to obtain image encoding feature information;

[0074] Converting the image coding feature information into a one-dimensional feature vector, and mapping the one-dimensional feature vector to obtain a first key-value pair feature;

[0075] Specifically, the image to be trained is input into the second convolutional neural network for encoding to obtain image encoding feature information, which includes the global features of the image, and then the high-dimensional image encoding feature information is converted into a one-dimensional feature vector using the flatten layer. Further, the one-dimensional feature vector is mapped into a first key-value pair feature, wherein the first key-value pair feature includes a first K value feature and a first V value feature; wherein K represents "key", which is a label of a feature, or a mapping feature obtained after certain operations. V represents "value", which usually corresponds to a feature value.

[0076] Input the first feature information corresponding to any point cloud into the cascaded transformer layer to obtain the second key-value pair feature output by each transformer layer;

[0077] Specifically, for the first feature information corresponding to any point cloud, the following steps are performed: the first feature information is input into the cascade-connected transformer layer, wherein the output of each transformer layer is used as the input of the next transformer layer, thereby obtaining the second key-value pair feature output by each transformer layer, and optionally, the model adopts a 20-layer transformer structure, and the second key-value pair feature includes a second K-value feature and a second V-value feature, wherein K represents "key", which is the label of a feature, or a mapping feature obtained after certain operations. V represents "value", which usually corresponds to the feature value.

[0078] The second feature information corresponding to any position to be estimated is used as the input of the third feedforward neural network of the first layer, so as to encode and map the input feature information by using the third feedforward neural network to obtain query feature information;

[0079] Input the first key-value pair feature, the query feature information and the second key-value pair feature into the attention layer for feature fusion to obtain the target feature information;

[0080] Specifically, for any second feature information corresponding to a position to be estimated, the following steps are performed: the second feature information corresponding to any position to be estimated is used as the input of the third feedforward neural network of the first layer, so as to use the third feedforward neural network to encode and map the input feature information to obtain query feature information. Furthermore, the first key-value pair feature, the query feature information, and the second key-value pair feature belonging to the same layer as the query feature information are input into the attention layer for feature fusion to obtain the target feature information. It can be understood that the number of query feature information and the second key-value pair feature is the same, combined with Figure 4 , each time the attention layer is input with the query feature information at the same level and its associated second key-value pair feature, for example, the query feature information query queries the K and V of the second key-value pair feature respectively. Optionally, the attention layer is a cross attention layer, and the attention layer operation of each layer is:

[0081]

[0082] Where: Q l Represents the query feature information output by the third feedforward neural network at layer l; V represents the K value output by the lth transformer layer;l V represents the output value of the lth transformer layer; V img Represents the V value in the first key-value pair feature corresponding to the image to be trained; represents the K value in the first key-value pair feature corresponding to the image to be trained, and T represents the preset matrix transpose.

[0083] The target feature information is used as the input of the third feedforward neural network of the next layer, so as to return to the step of encoding and mapping the input feature information by using the third feedforward neural network to obtain the query feature information, until the output result of the last layer of the third feedforward neural network is obtained, and the output result is used as the point cloud coordinate information corresponding to the position to be estimated;

[0084] The point cloud estimation model is trained based on the point cloud coordinate information corresponding to the position to be estimated and the preset point cloud labels.

[0085] Specifically, combined Figure 4 , the target feature information is used as the input of the third feedforward neural network of the next layer, so as to return to the step of encoding and mapping the input feature information using the third feedforward neural network to obtain the query feature information, until the output result of the last layer of the third feedforward neural network is obtained, and the output result is used as the point cloud coordinate information corresponding to the position to be estimated. Further, based on the point cloud coordinate information corresponding to the position to be estimated and the preset point cloud label, the model loss value is calculated, and then the point cloud estimation model is trained according to the model loss value.

[0086] The embodiment of the present invention encodes the training image and uses the attention layer to fuse the multi-modal feature information, so that the model can capture the global features of the image and the local features of the point cloud, thereby learning more diverse features and improving the accuracy of the model's subsequent prediction of point cloud information. The model can then be used to quickly predict the point cloud information of the specified area, improving the efficiency of point cloud data collection.

[0087] In one embodiment of the present invention, the first feature information is obtained based on the following steps:

[0088] For any point cloud: input the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to the point cloud into the first feedforward neural network to obtain the point cloud feature tensor; based on the two-dimensional coordinate information corresponding to the point cloud, select an image area with a preset pixel side length in the image to be trained; use the first convolutional neural network to encode the image area to obtain the image feature tensor; splice the point cloud feature tensor and the image feature tensor to obtain the first feature information.

[0089] Specifically, the following steps are performed for any point cloud:

[0090] Taking the two-dimensional coordinate information of the point cloud in the image to be trained as the center, an image area with a preset pixel side length is selected in the image to be trained. For example, an image area with a pixel side length of 20*20 is intercepted, and then the image area is input into the first convolutional neural network for encoding to obtain a one-dimensional feature vector, and the one-dimensional feature vector is used as the image feature tensor.

[0091] In addition, the three-dimensional coordinate information corresponding to the point cloud and the two-dimensional coordinate information corresponding to the point cloud are input into the first feedforward neural network of the first layer for encoding and mapping. For example, the three-dimensional coordinate information corresponding to the point cloud is (x, y, z), and the two-dimensional coordinate information of the point cloud projected on the image is (px, py). Based on the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to the point cloud, a one-dimensional tensor (x, y, z, px, py) is constructed, and the one-dimensional tensor (x, y, z, px, py) is input into the first feedforward neural network of the first layer, and the output result of the first feedforward neural network of the first layer is used as the input of the first feedforward neural network of the next layer, so that the output of the first feedforward neural network of the last layer is used as the point cloud feature tensor. In a specific example, the first feedforward neural network is a fully connected layer of 512 units, with a total of 3 layers.

[0092] Furthermore, the point cloud feature tensor and the image feature tensor are spliced ​​to obtain the first feature information. For example, the first feature extraction module also includes a connection layer, and the point cloud feature tensor and the image feature tensor are spliced ​​using the connection layer.

[0093] The embodiment of the present invention constructs the first feature information by combining the three-dimensional coordinates and two-dimensional projection coordinates of the point cloud and the global image. The model can capture the global features of the image and the local features of the point cloud, thereby learning more diverse features and improving the accuracy of the subsequent point cloud information of the model.

[0094] In one embodiment of the present invention, the second characteristic information is obtained based on the following steps:

[0095] The two-dimensional coordinate information of any position to be estimated is input into the second feedforward neural network to obtain second feature information.

[0096] Specifically, for any two-dimensional coordinate information of the position to be estimated, the following steps are performed: the two-dimensional coordinate information of the position to be estimated is input into the second feedforward neural network of the first layer for encoding mapping, and the output result of each layer of the second feedforward neural network is used as the input of the second feedforward neural network of the next layer, so that the output of the last layer of the second feedforward neural network is used as the second feature information.

[0097] The embodiment of the present invention inputs the image pixel coordinates corresponding to the position to be estimated into a second feedforward neural network to extract second feature information, and then inputs the second feature information into the model to predict the point cloud information corresponding to the position to be estimated, so as to train the point cloud estimation model in combination with the point cloud information corresponding to the position to be estimated, so that the model can be used to quickly predict the point cloud information of the specified area position, thereby improving the efficiency of point cloud data collection.

[0098] In one embodiment, specifically, Figure 5 As shown, Figure 5 1 is a flow chart of a method for measuring the volume of an object in an embodiment of the present invention. The present invention provides a method for measuring the volume of an object, comprising the following steps:

[0099] Step S21, obtaining target point cloud information and object image of the object to be identified;

[0100] Step S22, determining image coordinate information corresponding to a target area, wherein the target area is a region position of desired estimated point cloud information selected from the object image;

[0101] Step S23, based on the target point cloud information, the object image and the image coordinate information, point cloud estimation is performed using a point cloud estimation model to obtain predicted point cloud information of the target area;

[0102] Step S24: determining the object volume of the object to be identified based on the predicted point cloud information and the target point cloud information.

[0103] Specifically, the target point cloud information and object image of the object to be identified are obtained, and then the area position for point cloud information prediction is selected in the object image to obtain the target area. Then the image coordinate information corresponding to the target area is extracted, and further, the target point cloud information, the object image and the image coordinate information are input into the point cloud estimation model to perform point cloud estimation using the point cloud estimation model to obtain the predicted point cloud information of the target area, that is, the projection coordinates of each point cloud in the target point cloud information, the object image and the target point cloud information in the object image are input into the first feature extraction module in the point cloud estimation model, and the image coordinate information corresponding to the target area is input into the second feature extraction module in the point cloud estimation model, and then according to the object image, the output result of the first feature extraction module and the output result of the second feature extraction module, the backbone network in the point cloud estimation model is used to perform point cloud estimation to obtain the predicted point cloud information of the target area. Among them, the training process of the cloud estimation model is specifically described in the above embodiment and will not be repeated here. Furthermore, based on the predicted point cloud information of the target area and the original sampled target point cloud information, the minimum circumscribed hexahedron of all point clouds is calculated, and then the minimum circumscribed hexahedron is used as the object volume of the object to be identified.

[0104] Through the above steps, the embodiment of the present invention realizes the point cloud estimation model based on the object image, the pixel coordinates of the designated area position and the sparse point cloud data collected by the device, and generates the point cloud information of the designated area position, without the need for a long point cloud collection process, thereby improving the efficiency of point cloud data collection. Furthermore, the point cloud information collected by the device and the point cloud information generated by the model are used to measure the volume of the object, effectively improving the accuracy of the object volume estimation.

[0105] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0106] In one embodiment, a training device for a point cloud estimation model is provided, and the training device for the point cloud estimation model corresponds one-to-one to the training method for the point cloud estimation model in the above embodiment. Figure 6 As shown, Figure 6 : is a structural schematic diagram of a training device for a point cloud estimation model in one embodiment of the present invention, the training device for a point cloud estimation model comprises:

[0107] A first acquisition module 31 is used to acquire the image to be trained and the point cloud information, and determine the two-dimensional coordinate information of each point cloud projected in the image to be trained, wherein the point cloud information includes the three-dimensional coordinate information of all point clouds;

[0108] The training module 32 is used to train a point cloud estimation model based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated, wherein the position to be estimated is the position of the expected estimated point cloud information selected in the image to be trained.

[0109] The training device for the point cloud estimation model also includes:

[0110] The point cloud estimation model includes a first feature extraction module, a second feature extraction module and a backbone network;

[0111] The first feature extraction module includes a first feedforward neural network and a first convolutional neural network; used to extract features from the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to the training image and each point cloud;

[0112] The second feature extraction module includes a second feedforward neural network; used for extracting features from the two-dimensional coordinate information of any position to be estimated;

[0113] The backbone network is used to perform point cloud estimation on the training image, the first feature information output by the second feature extraction module, and the second feature information output by the first feature extraction module.

[0114] The training device for the point cloud estimation model also includes:

[0115] The backbone network consists of a preset number of cascaded transformer layers, where the output of each transformer layer serves as the input of the next transformer layer;

[0116] The second convolutional neural network, where the output of the second convolutional neural network is converted into a one-dimensional feature vector:

[0117] Multi-layer cascaded third feed-forward neural network and attention layer;

[0118] The output of the second convolutional neural network, the output of the transformer layer, and the output of the third feedforward neural network are used as the input of the attention layer;

[0119] The output of the attention layer is used as the input of the next layer of the third feedforward neural network until it reaches the last layer of the third feedforward neural network, where the output result of the last layer of the third feedforward neural network is used as the final result of the point cloud estimation model output.

[0120] The training module 32 is also used to:

[0121] Inputting the image to be trained into the second convolutional neural network for encoding to obtain image encoding feature information;

[0122] Converting the image coding feature information into a one-dimensional feature vector, and mapping the one-dimensional feature vector to obtain a first key-value pair feature, wherein the first key-value pair feature includes a first K value feature and a first V value feature;

[0123] Input the first feature information corresponding to any point cloud into the cascaded transformer layer to obtain the second key-value pair feature output by each transformer layer, wherein the second key-value pair feature includes a second K value feature and a second V value feature;

[0124] The second feature information corresponding to any position to be estimated is used as the input of the third feedforward neural network of the first layer, so as to encode and map the input feature information by using the third feedforward neural network to obtain query feature information;

[0125] Input the first key-value pair feature, the query feature information and the second key-value pair feature into the attention layer for feature fusion to obtain the target feature information;

[0126] The target feature information is used as the input of the third feedforward neural network of the next layer, so as to return to the step of encoding and mapping the input feature information by using the third feedforward neural network to obtain the query feature information, until the output result of the last layer of the third feedforward neural network is obtained, and the output result is used as the point cloud coordinate information corresponding to the position to be estimated;

[0127] The point cloud estimation model is trained based on the point cloud coordinate information corresponding to the position to be estimated and the preset point cloud labels.

[0128] The training module 32 is also used to:

[0129] For any point cloud:

[0130] Inputting the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to the point cloud into the first feedforward neural network to obtain the point cloud feature tensor;

[0131] Based on the two-dimensional coordinate information corresponding to the point cloud, an image region with a preset pixel side length is selected in the image to be trained;

[0132] Encode the image region using the first convolutional neural network to obtain an image feature tensor;

[0133] The point cloud feature tensor and the image feature tensor are concatenated to obtain the first feature information.

[0134] The training module 32 is also used to:

[0135] The two-dimensional coordinate information of any position to be estimated is input into the second feedforward neural network to obtain second feature information.

[0136] In one embodiment, a device for measuring the volume of an object is provided, and the device for measuring the volume of an object corresponds one-to-one to the method for measuring the volume of an object in the above embodiment. Figure 7 As shown, Figure 7 1 is a schematic diagram of a structure of an object volume measuring device according to an embodiment of the present invention, wherein the object volume measuring device comprises:

[0137] The second acquisition module 41 is used to acquire target point cloud information and object image of the object to be identified;

[0138] A first determination module 42 is used to determine image coordinate information corresponding to a target area, wherein the target area is a region position of desired estimated point cloud information selected from an object image;

[0139] The point cloud estimation module 43 is used to perform point cloud estimation processing using a point cloud estimation model based on target point cloud information, object image and image coordinate information to obtain predicted point cloud information of the target area;

[0140] The second determination module 44 is used to determine the object volume of the object to be identified based on the predicted point cloud information and the target point cloud information.

[0141] The second determining module 44 is further used for:

[0142] Based on the predicted point cloud information and the target point cloud information, the minimum circumscribed hexahedron of all point clouds is determined;

[0143] The smallest circumscribed hexahedron is taken as the object volume of the object to be identified.

[0144] For the specific definition of the training device of the point cloud estimation model, please refer to the definition of the training method of the point cloud estimation model above, which will not be repeated here. Each module in the above-mentioned training device of the point cloud estimation model can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0145] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, Figure 8 : is a schematic diagram of a computer device in one embodiment of the present invention. The computer device includes a processor, a memory, a network interface and a database connected by a device bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium and an internal memory. The readable storage medium stores an operating device, a computer-readable instruction and a database. The internal memory provides an environment for the operation of the operating device and the computer-readable instructions in the readable storage medium. The database of the computer device is used to store data involved in the training method of the point cloud estimation model. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a training method for a point cloud estimation model is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0146] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, and a network interface connected through a device bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a training method for a point cloud estimation model or an object volume measurement method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.

[0147] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the training method of the point cloud estimation model or the object volume measurement method are implemented.

[0148] In one embodiment, a readable storage medium is provided, the readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the training method steps of the point cloud estimation model described above are implemented. A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0149] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0150] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.

Claims

1. A training method for a point cloud estimation model, characterized in that: include: Acquire the image to be trained and point cloud information, and determine the two-dimensional coordinate information of each point cloud projected on the image to be trained, wherein the point cloud information includes the three-dimensional coordinate information of all point clouds; Based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated, a point cloud estimation model is trained, wherein the position to be estimated is the position of the expected estimated point cloud information selected in the image to be trained.

2. The method for training a point cloud estimation model according to claim 1, characterized in that: The point cloud estimation model includes a first feature extraction module, a second feature extraction module and a backbone network; The first feature extraction module includes a first feedforward neural network and a first convolutional neural network; used to extract features from the image to be trained and the three-dimensional coordinate information and the two-dimensional coordinate information corresponding to each of the point clouds; The second feature extraction module includes a second feedforward neural network; Used to extract features from the two-dimensional coordinate information of any position to be estimated; The backbone network is used to perform point cloud estimation on the image to be trained, the first feature information output by the second feature extraction module, and the second feature information output by the first feature extraction module.

3. The training method of the point cloud estimation model according to claim 2, characterized in that: The backbone network includes a preset number of cascaded transformer layers, wherein the output of each transformer layer serves as the input of the next transformer layer; The second convolutional neural network, where the output of the second convolutional neural network is converted into a one-dimensional feature vector: Multi-layer cascaded third feed-forward neural network and attention layer; The output of the second convolutional neural network, the output of the transformer layer, and the output of the third feedforward neural network are used as inputs of the attention layer; The output of the attention layer is used as the input of the next layer of the third feedforward neural network until the last layer of the third feedforward neural network is reached, wherein the result output by the last layer of the third feedforward neural network is used as the final result output by the point cloud estimation model.

4. The method for training a point cloud estimation model according to claim 3, characterized in that: The training process of the point cloud estimation model is as follows: Inputting the image to be trained into a second convolutional neural network for encoding to obtain image encoding feature information; Converting the image coding feature information into a one-dimensional feature vector, and mapping the one-dimensional feature vector to obtain a first key-value pair feature, wherein the first key-value pair feature includes a first K value feature and a first V value feature; Inputting the first feature information corresponding to any of the point clouds into the cascaded transformer layer to obtain a second key-value pair feature output by each transformer layer, wherein the second key-value pair feature includes a second K value feature and a second V value feature; Using the second feature information corresponding to any of the positions to be estimated as the input of the third feedforward neural network of the first layer, so as to encode and map the input feature information using the third feedforward neural network to obtain query feature information; Inputting the first key-value pair feature, the query feature information and the second key-value pair feature into the attention layer for feature fusion to obtain target feature information; The target feature information is used as the input of the third feedforward neural network of the next layer, so as to return to the step of encoding and mapping the input feature information by using the third feedforward neural network to obtain the query feature information, until the output result of the last layer of the third feedforward neural network is obtained, and the output result is used as the point cloud coordinate information corresponding to the position to be estimated; The point cloud estimation model is trained based on the point cloud coordinate information corresponding to the position to be estimated and the preset point cloud labels.

5. The method for training a point cloud estimation model according to claim 2, characterized in that: The first characteristic information is obtained based on the following steps: For any point cloud: Inputting the three-dimensional coordinate information corresponding to the point cloud and the two-dimensional coordinate information corresponding to the point cloud into the first feedforward neural network to obtain a point cloud feature tensor; Based on the two-dimensional coordinate information corresponding to the point cloud, an image region having a preset pixel side length is selected in the image to be trained; Encoding the image region using the first convolutional neural network to obtain an image feature tensor; The point cloud feature tensor and the image feature tensor are concatenated to obtain first feature information.

6. The method for training a point cloud estimation model according to claim 2, characterized in that: The second characteristic information is obtained based on the following steps: The two-dimensional coordinate information of any of the positions to be estimated is input into the second feedforward neural network to obtain the second feature information.

7. A method for measuring the volume of an object, characterized in that: include: Obtain target point cloud information and object image of the object to be identified; Determine image coordinate information corresponding to a target area, wherein the target area is a region position of desired estimated point cloud information selected from the object image; Based on the target point cloud information, the object image and the image coordinate information, point cloud estimation is performed using a point cloud estimation model to obtain predicted point cloud information of the target area, wherein the point cloud estimation model is trained according to the training method of the point cloud estimation model according to any one of claims 1 to 6; Based on the predicted point cloud information and the target point cloud information, the object volume of the object to be identified is determined.

8. The object volume measurement method according to claim 7, characterized in that: The determining the object volume of the object to be identified based on the predicted point cloud information and the target point cloud information includes: Based on the predicted point cloud information and the target point cloud information, determining the minimum circumscribed hexahedron of all point clouds; The minimum circumscribed hexahedron is used as the object volume of the object to be identified.

9. A training device for a point cloud estimation model, characterized in that: include: A first acquisition module is used to acquire the image to be trained and point cloud information, and determine the two-dimensional coordinate information of each point cloud projected on the image to be trained, wherein the point cloud information includes the three-dimensional coordinate information of all point clouds; A training module is used to train a point cloud estimation model based on the image to be trained, the three-dimensional coordinate information and the two-dimensional coordinate information of any point cloud, and the two-dimensional coordinate information of the position to be estimated, wherein the position to be estimated is the position of the expected estimated point cloud information selected in the image to be trained.

10. An object volume measuring device, characterized in that: include: The second acquisition module is used to acquire the target point cloud information and the object image of the object to be identified; A first determination module is used to determine image coordinate information corresponding to a target area, wherein the target area is a region position of desired estimated point cloud information selected from the object image; A point cloud estimation module, configured to perform point cloud estimation processing using a point cloud estimation model based on the target point cloud information, the object image and the image coordinate information, to obtain predicted point cloud information of the target area; The second determination module is used to determine the object volume of the object to be identified based on the predicted point cloud information and the target point cloud information.