Training method for pose model, point cloud registration method, device and storage medium
By training the target pose model and combining point cloud and image information, the problem of obtaining pose matrix in point cloud registration is solved, and more efficient and accurate point cloud registration is achieved.
Patent Information
- Application Number
- CN202311140320.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-05
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-09-05
AI Technical Summary
The prior art is difficult to effectively solve the problem of pose matrix acquisition in point cloud registration, resulting in poor point cloud registration effect.
By acquiring the first sample point cloud and the second sample point cloud, the first pose matrix is called based on the initial pose model, combined with the depth image and the intensity image, the loss value is calculated and the initial pose model is trained, and the target pose model is obtained to determine the pose matrix between the point clouds.
It improves the accuracy and robustness of point cloud registration, making the acquisition of pose matrix more accurate and suitable for dense and sparse point clouds.
Smart Images

Figure CN117197201B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of point cloud technology, and particularly to a method for training an attitude model, a point cloud registration method, a device, and a storage medium. Background Art
[0002] Point cloud registration is a basic task in computer vision, which focuses on finding the rigid alignment between two point clouds. With the continuous popularization of 3D (Three Dimensions) acquisition devices, point cloud registration has been applied in many fields such as autonomous driving, augmented reality, and 3D reconstruction.
[0003] In order to register two point clouds, it is necessary to obtain the attitude matrix between the images corresponding to the two point clouds. Therefore, there is an urgent need for a method for training an attitude model to obtain the attitude model, and then obtain the attitude matrix between the images corresponding to the two point clouds according to the attitude model. Summary of the Invention
[0004] Embodiments of the present application provide a method for training an attitude model, a point cloud registration method, a device, and a storage medium. The technical solutions are as follows:
[0005] In a first aspect, embodiments of the present application provide a method for training an attitude model, the method comprising:
[0006] Obtain a first sample point cloud and a second sample point cloud, both the first sample point cloud and the second sample point cloud are collected at a first position at a first time and a second time respectively, and the first sample point cloud includes the relevant information of a plurality of first sample points, the second sample point cloud includes the relevant information of a plurality of second sample points, and the relevant information of any sample point includes the position information and depth information of the any sample point;
[0007] According to the relevant information of the plurality of first sample points, obtain a first sample image of the first position at the first time; according to the relevant information of the plurality of second sample points, obtain a second sample image of the first position at the second time;
[0008] Call an initial attitude model to determine a first attitude matrix based on the first sample image and the second sample image, the first attitude matrix is used to register the first sample point cloud and the second sample point cloud;
[0009] According to the first attitude matrix, a first depth image, and a second depth image, obtain a third depth image, the first depth image is determined based on the first sample image, the second sample image, and a camera image at the first position at the first time, and the second depth image is obtained based on the depth information of the plurality of second sample points;
[0010] Determine a first loss value between the second depth image and the third depth image;
[0011] Train the initial pose model according to the first loss value to obtain a target pose model, where the target pose model is used to determine a pose matrix between images corresponding to two point clouds, and the pose matrix is used to register the two point clouds.
[0012] In a possible implementation, the relevant information of any sample point further includes the intensity information of the any sample point; the method further includes:
[0013] Obtain a third intensity image according to the first pose matrix, the first intensity image and the second intensity image, where the first intensity image is determined based on the first sample image, the second sample image and the camera image at the first position at the first time, and the second intensity image is obtained based on the intensity information of the plurality of second sample points;
[0014] Determine a second loss value between the second intensity image and the third intensity image;
[0015] The training of the initial pose model according to the first loss value to obtain a target pose model includes:
[0016] Train the initial pose model according to the first loss value and the second loss value to obtain the target pose model.
[0017] In a possible implementation, the method further includes:
[0018] Determine a fourth depth image according to the depth information of the plurality of first sample points;
[0019] Determine a third loss value between the fourth depth image and the first depth image;
[0020] Determine a fourth intensity image according to the intensity information of the plurality of first sample points;
[0021] Determine a fourth loss value between the fourth intensity image and the first intensity image;
[0022] The training of the initial pose model according to the first loss value and the second loss value to obtain the target pose model includes:
[0023] Train the initial pose model according to the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target pose model.
[0024] In a possible implementation, training the initial pose model according to the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target pose model includes:
[0025] Determine the total loss value according to the first loss value, the second loss value, the third loss value, and the fourth loss value;
[0026] Based on the fact that the total loss value is not in a converged state, adjust the parameters of the initial pose model to obtain an adjusted pose model;
[0027] Determine an updated total loss value according to the adjusted pose model, the first sample image, and the second sample image;
[0028] Based on the fact that the updated total loss value is in a converged state, use the adjusted pose model as the target pose model.
[0029] In a possible implementation, calling the initial pose model to determine a first pose matrix based on the first sample image and the second sample image includes:
[0030] Obtain a first combined image according to the first sample image and the second sample image, where the first combined image includes the first sample image and the second sample image;
[0031] Input the first combined image into the initial pose model to obtain a feature vector of the first combined image, where the feature vector of the first combined image is used to represent the first combined image;
[0032] Call a pose decoder to decode the feature vector of the first combined image to obtain the first pose matrix.
[0033] In a possible implementation, obtaining a third depth image according to the first pose matrix, the first depth image, and the second depth image includes:
[0034] Obtain a plurality of target position information according to the position information, depth information, and the first pose matrix of each pixel point in the second depth image, where each pixel point in the second depth image corresponds to one target position information;
[0035] Use the depth information of the pixel points in the first depth image that have the same positions as the respective target position information as the depth information corresponding to the respective target position information;
[0036] Generate the third depth image according to the position information of each pixel point in the second depth image and the depth information corresponding to the target position information corresponding to the position information of each pixel point.
[0037] In a second aspect, an embodiment of the present application provides a point cloud registration method, and the method includes:
[0038] Obtain a first point cloud and a second point cloud. The first point cloud and the second point cloud are both collected at a second position at a third time and a fourth time respectively. The first point cloud includes relevant information of a plurality of first points, and the second point cloud includes relevant information of a plurality of second points. The relevant information of any point includes the position information and depth information of the any point;
[0039] According to the relevant information of the plurality of first points, obtain a first image of the second position at the third time; according to the relevant information of the plurality of second points, obtain a second image of the second position at the fourth time;
[0040] Call a target pose model to determine a second pose matrix based on the first image and the second image. The second pose matrix is used to register the first point cloud and the second point cloud, and the target pose model is trained based on the method described in the first aspect above.
[0041] In a possible implementation manner, the calling the target pose model to determine the second pose matrix based on the first image and the second image includes:
[0042] According to the first image and the second image, obtain a second combined image, and the second combined image includes the first image and the second image;
[0043] Input the second combined image into the target pose model to obtain a feature vector of the second combined image, and the feature vector of the second combined image is used to characterize the second combined image;
[0044] Call a pose decoder to decode the feature vector of the second combined image to obtain the second pose matrix.
[0045] In a third aspect, an embodiment of the present application provides a training device for a pose model, and the device includes:
[0046] An acquisition module, configured to acquire a first sample point cloud and a second sample point cloud. The first sample point cloud and the second sample point cloud are both collected at a first position at a first time and a second time respectively. The first sample point cloud includes relevant information of a plurality of first sample points, and the second sample point cloud includes relevant information of a plurality of second sample points. The relevant information of any sample point includes the position information and depth information of the any sample point;
[0047] The obtaining module is further configured to obtain a first sample image of the first position at the first time according to the relevant information of the multiple first sample points; and obtain a second sample image of the first position at the second time according to the relevant information of the multiple second sample points;
[0048] The determining module is configured to call an initial pose model to determine a first pose matrix based on the first sample image and the second sample image, and the first pose matrix is used for registering the first point cloud and the second point cloud;
[0049] The obtaining module is further configured to obtain a third depth image according to the first pose matrix, the first depth image and the second depth image, where the first depth image is determined based on the first sample image, the second sample image and a camera image at the first position at the first time, and the second depth image is obtained based on the depth information of the multiple second sample points;
[0050] The determining module is further configured to determine a first loss value between the second depth image and the third depth image;
[0051] The training module is configured to train the initial pose model according to the first loss value to obtain a target pose model, and the target pose model is used to determine a pose matrix between images corresponding to two point clouds, and the pose matrix is used for registering the two point clouds.
[0052] In a possible implementation manner, the relevant information of any sample point further includes the intensity information of the any sample point;
[0053] The obtaining module is further configured to obtain a third intensity image according to the first pose matrix, the first intensity image and the second intensity image, where the first intensity image is determined based on the first sample image, the second sample image and a camera image at the first position at the first time, and the second intensity image is obtained based on the intensity information of the multiple second sample points;
[0054] The determining module is further configured to determine a second loss value between the second intensity image and the third intensity image;
[0055] The training module is configured to train the initial pose model according to the first loss value and the second loss value to obtain the target pose model.
[0056] In a possible implementation, the determining module is further configured to determine a fourth depth image according to the depth information of the multiple first sample points; determine a third loss value between the fourth depth image and the first depth image; determine a fourth intensity image according to the intensity information of the multiple first sample points; determine a fourth loss value between the fourth intensity image and the first intensity image;
[0057] The training module is configured to train the initial pose model according to the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target pose model.
[0058] In a possible implementation, the training module is configured to determine a total loss value according to the first loss value, the second loss value, the third loss value, and the fourth loss value; adjust the parameters of the initial pose model based on the fact that the total loss value is not in a converged state to obtain an adjusted pose model; determine an updated total loss value according to the adjusted pose model, the first sample image, and the second sample image; and use the adjusted pose model as the target pose model based on the fact that the updated total loss value is in a converged state.
[0059] In a possible implementation, the determining module is configured to obtain a first combined image according to the first sample image and the second sample image, where the first combined image includes the first sample image and the second sample image; input the first combined image into the initial pose model to obtain a feature vector of the first combined image, where the feature vector of the first combined image is used to characterize the first combined image; and call a pose decoder to decode the feature vector of the first combined image to obtain the first pose matrix.
[0060] In a possible implementation, the obtaining module is configured to obtain multiple target position information according to the position information, depth information of each pixel point in the second depth image, and the first pose matrix, where each pixel point in the second depth image has a corresponding target position information; use the depth information of the pixel points in the first depth image with the same positions corresponding to the respective target position information as the depth information corresponding to the respective target position information; and generate the third depth image according to the position information of each pixel point in the second depth image and the depth information corresponding to the target position information corresponding to each pixel point's position information.
[0061] Fourthly, an embodiment of the present application provides a point cloud registration device, and the device includes:
[0062] An acquisition module, configured to acquire a first point cloud and a second point cloud, both the first point cloud and the second point cloud are acquired at a second location at a third time and a fourth time respectively, and the first point cloud includes relevant information of a plurality of first points, the second point cloud includes relevant information of a plurality of second points, and the relevant information of any point includes the position information and depth information of the any point;
[0063] The acquisition module is further configured to obtain a first image of the second location at the third time according to the relevant information of the plurality of first points; and obtain a second image of the second location at the fourth time according to the relevant information of the plurality of second points;
[0064] A determination module, configured to call a target pose model to determine a second pose matrix based on the first image and the second image, the second pose matrix is used to register the first point cloud and the second point cloud, and the target pose model is trained based on the device described in the above third aspect.
[0065] In a possible implementation manner, the determination module is configured to obtain a second combined image according to the first image and the second image, the second combined image includes the first image and the second image; input the second combined image into the target pose model to obtain a feature vector of the second combined image, the feature vector of the second combined image is used to characterize the second combined image; call a pose decoder to decode the feature vector of the second combined image to obtain the second pose matrix.
[0066] In a fifth aspect, an embodiment of the present application provides a computer device, the computer device includes a processor and a memory, and at least one program code is stored in the memory, and the at least one program code is loaded and executed by the processor to enable the computer device to implement the training method of the pose model described in the above first aspect, or to enable the computer device to implement the point cloud registration method described in the above second aspect.
[0067] In a sixth aspect, a computer-readable storage medium is further provided, and at least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to enable a computer to implement the training method of the pose model described in the above first aspect, or to enable a computer device to implement the point cloud registration method described in the above second aspect.
[0068] In a seventh aspect, a computer program or a computer program product is further provided, and at least one computer instruction is stored in the computer program or the computer program product, and the at least one computer instruction is loaded and executed by a processor to enable a computer to implement the training method of the pose model described in the above first aspect, or to enable a computer to implement the point cloud registration method described in the above second aspect.
[0069] The technical solution provided by the embodiment of the present application at least brings the following beneficial effects:
[0070] The technical solution provided by the embodiment of the present application does not require screening in the point cloud. Therefore, it is applicable to both dense point clouds and sparse point clouds, has stronger robustness and a wider range of usage scenarios. Moreover, this method not only considers the position information of the sample points, but also considers the depth information of the sample points, and the considered information is relatively comprehensive. As a result, the accuracy of the trained target pose model is higher. Using the target pose model with higher accuracy to determine the pose matrix between the images corresponding to the two point clouds makes the determined pose matrix more accurate. Furthermore, using the pose matrix with higher accuracy to register the two point clouds makes the registration effect of the two point clouds better and the registration accuracy higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0072] Figure 1 It is a schematic diagram of the implementation environment of a method for training a pose model and a method for point cloud registration provided by the embodiment of the present application;
[0073] Figure 2 It is a flowchart of a method for training a pose model provided by the embodiment of the present application;
[0074] Figure 3 It is a schematic diagram of the determination process of a first pose matrix provided by the embodiment of the present application;
[0075] Figure 4 It is a flowchart of the training process of a pose model provided by the embodiment of the present application;
[0076] Figure 5 It is a flowchart of a method for point cloud registration provided by the embodiment of the present application;
[0077] Figure 6 It is a picture of point cloud registration provided by the embodiment of the present application;
[0078] Figure 7 It is a flowchart of a method for point cloud registration provided by the embodiment of the present application;
[0079] Figure 8 It is a schematic diagram of the structure of a device for training a pose model provided by the embodiment of the present application;
[0080] Figure 9 It is a schematic structural diagram of a point cloud registration device provided by an embodiment of the present application;
[0081] Figure 10 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application;
[0082] Figure 11 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific embodiments
[0083] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0084] Figure 1 It is a schematic diagram of the implementation environment of a method for training an attitude model and a method for registering point clouds provided by an embodiment of the present application. As Figure 1 shown, the implementation environment includes: a computer device 101. The computer device 101 can be a terminal device or a server, and the embodiments of the present application do not limit this. The computer device 101 is used to execute the method for training an attitude model and the method for registering point clouds provided by the embodiments of the present application.
[0085] Optionally, the computer device 101 is a terminal device. The terminal device can be any electronic device product that can perform human-computer interaction with a user in one or more ways such as a keyboard, a touchpad, a remote control, voice interaction, or a handwriting device. For example, a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart car machine, a smart TV, a smart speaker, a smart watch, etc.
[0086] The terminal device can generally refer to one of multiple terminal devices. This embodiment only uses the terminal device as an example for illustration. Those skilled in the art can know that the number of the above terminal devices can be more or less. For example, the above terminal device can be only one, or the above terminal devices can be dozens or hundreds, or more. The embodiments of the present application do not limit the number and type of the terminal devices.
[0087] When the computer device 101 is a server, the server can be a single server, a server cluster composed of multiple servers, or any one of a cloud computing platform and a virtualization center. The embodiments of the present application do not limit this. The server and the terminal device are communicatively connected through a wired network or a wireless network. The server has a data receiving function, a data processing function, and a data sending function. Of course, the server can also have other functions, and the embodiments of the present application do not limit this.
[0088] Those skilled in the art should understand that the above terminal devices and servers are only for illustrative purposes. Other existing or future terminal devices or servers that can be applied to the present application should also be included in the protection scope of the present application and are hereby incorporated by reference.
[0089] The embodiments of the present application provide a method for training a pose model. This method can be applied to the above Figure 1 shown implementation environment. Taking Figure 2 the flowchart of a method for training a pose model provided by the embodiments of the present application shown as an example, this method can be executed by the Figure 1 computer device 101 in. As Figure 2 shown, this method includes the following steps 201 to step 206.
[0090] In step 201, a first sample point cloud and a second sample point cloud are obtained.
[0091] Among them, the first sample point cloud and the second sample point cloud are both collected at the first position at the first time and the second time respectively. The first sample point cloud includes the relevant information of multiple first sample points, and the second sample point cloud includes the relevant information of multiple second sample points. The relevant information of any sample point includes the position information and depth information of any sample point.
[0092] The first time and the second time are different. The first time can be earlier than the second time or later than the second time. The embodiments of the present application do not limit this. Exemplarily, the first sample point cloud is collected at the first position at the first time, and the second sample point cloud is collected at the first position at the second time.
[0093] In a possible implementation, a lidar is installed in a vehicle, and the lidar is used to collect point clouds. Exemplarily, a first sample point cloud is collected at a first position at a first time, and a second sample point cloud is collected at the first position at a second time. After the lidar collects the first sample point cloud at the first position at the first time, the lidar sends the first sample point cloud to a computer device so that the computer device can obtain the first sample point cloud. After the lidar collects the second sample point cloud at the first position at the second time, the lidar sends the second sample point cloud to the computer device so that the computer device can obtain the second sample point cloud.
[0094] Optionally, the positions corresponding to the first sample point cloud and the second sample point cloud may not be the same position. For example, the first sample point cloud is collected at a first position at a first time, and the second sample point cloud is collected at a reference position at a second time. The distance between the first position and the reference position is less than a distance threshold. The distance threshold is set based on experience or adjusted according to the implementation environment, and the embodiments of the present application do not limit this. Exemplarily, the distance threshold is 1 meter.
[0095] In step 202, according to the relevant information of multiple first sample points, a first sample image of the first position at the first time is obtained; according to the relevant information of multiple second sample points, a second sample image of the first position at the second time is obtained.
[0096] In a possible implementation, the process of obtaining the first sample image of the first position at the first time according to the relevant information of multiple first sample points is similar to the process of obtaining the second sample image of the first position at the second time according to the relevant information of multiple second sample points. The embodiments of the present application only take the process of obtaining the first sample image of the first position at the first time according to the relevant information of multiple first sample points as an example for illustration.
[0097] Optionally, the process of obtaining the first sample image of the first position at the first time according to the relevant information of multiple first sample points includes: adjusting the relevant information of multiple first sample points according to a transformation matrix to obtain the adjusted relevant information of multiple first sample points, and obtaining the first sample image of the first position at the first time according to the adjusted relevant information of multiple first sample points. The transformation matrix is the external parameter between the lidar and the camera. The first sample point cloud and the second sample point cloud are collected by the lidar, and the camera image at the first position at the first time is collected by the camera.
[0098] A lidar and a camera are installed in the vehicle. Among them, the lidar is used to collect point clouds, and the camera is used to collect camera images. When the lidar collects the point cloud at the first position at the first time, the camera also collects the camera image at the first position at the first time. When the lidar collects the point cloud at the first position at the second time, the camera also collects the camera image at the first position at the second time. The extrinsic parameters between the lidar and the camera are known, and the extrinsic parameters are used to convert the point cloud collected by the lidar into an image.
[0099] Since the relevant information of any sample point includes the position information and depth information of any sample point, therefore, the first sample image includes a depth image obtained according to the depth information of each first sample point. Optionally, the relevant information of any sample point further includes the intensity information of any sample point, then the first sample image further includes an intensity image obtained according to the intensity information of each first sample point. That is, the first sample image includes a depth image and an intensity image.
[0100] In step 203, call the initial pose model to determine the first pose matrix based on the first sample image and the second sample image.
[0101] Among them, the first pose matrix is used to register the first sample point cloud and the second sample point cloud.
[0102] Optionally, the process of calling the initial pose model to determine the first pose matrix based on the first sample image and the second sample image includes: according to the first sample image and the second sample image, obtain a first combined image, the first combined image includes the first sample image and the second sample image; input the first combined image into the initial pose model to obtain a feature vector of the first combined image, and the feature vector of the first combined image is used to represent the first combined image; call the pose decoder to decode the feature vector of the first combined image to obtain the first pose matrix.
[0103] Among them, the process of obtaining the first combined image according to the first sample image and the second sample image includes: superimpose the first sample image and the second sample image together to obtain the first combined image.
[0104] In a possible implementation, the initial pose model (PoseNet) is based on the feature map output by R-18, and uses a multi-layer transformer structure with num_queries = 1 to output a pose matrix, and the pose matrix includes 6 degrees of freedom (three Euler angles and translations in three directions).
[0105] As Figure 3 is a schematic diagram of a process for determining a first pose matrix provided by an embodiment of the present application. In Figure 3Among them, the pose decoder includes a Norm layer, an MLP layer, a Muti-Head Attention layer, a Cross-Attention layer, a linear layer, a sigmoid activation function layer, and a softmax non-linear transformation function layer. The feature vector of the first combined image is input into the pose decoder, and each layer included in the pose decoder decodes the feature vector of the first combined image to obtain the first pose matrix.
[0106] In step 204, according to the first pose matrix, the first depth image, and the second depth image, a third depth image is obtained.
[0107] Among them, the first depth image is determined based on the first sample image, the second sample image, and the camera image at the first position at the first time, and the second depth image is obtained based on the depth information of multiple second sample points.
[0108] Optionally, the process of obtaining the first depth image according to the first sample image, the second sample image, and the camera image at the first position at the first time includes: superimposing the first sample image, the second sample image, and the camera image at the first position at the first time to obtain a superimposed image; inputting the superimposed image into an initial depth model to obtain the first depth image. Among them, the initial depth model is an Encoder-Decoder structure based on R-18 similar to the Hourglass Network (pose estimation network), and is used to predict the dense depth attribute from the image.
[0109] Optionally, since the second sample image includes a depth image and an intensity image, the depth image included in the second sample image can also be directly used as the second depth image.
[0110] In a possible implementation manner, the process of obtaining the third depth image according to the first pose matrix, the first depth image, and the second depth image includes: obtaining a plurality of target position information according to the position information, depth information, and the first pose matrix of each pixel point in the second depth image, and each pixel point in the second depth image corresponds to a target position information; using the depth information of the pixel points in the first depth image with the same position corresponding to each target position information as the depth information corresponding to each target position information; generating a third depth image according to the position information of each pixel point in the second depth image and the depth information corresponding to the target position information corresponding to each pixel point's position information. The target position information corresponding to the position information of any pixel point in the second depth image is the position information of any pixel point in the second depth image mapped to the coordinate system of the first depth image.
[0111] Optionally, according to the position information, depth information, and the first pose matrix of each pixel point in the second depth image, the target position information corresponding to the position information of each pixel point is obtained according to the following formula (1).
[0112] P t = KTD r K -1 P r Formula (1)
[0113] In the above formula (1), P t is the target position information corresponding to the position information of any pixel point in the second depth image, K is the camera internal parameter, T is the first pose matrix, and D r is the depth information of any pixel point in the second depth image, K -1 is the transpose of the camera internal parameter, and P r is the position information of any pixel point in the second depth image.
[0114] Exemplarily, the second depth image includes 3 pixel points. Among them, the position information of the first pixel point in the second depth image is (x1, y1), the position information of the second pixel point in the second depth image is (x2, y2), and the position information of the third pixel point in the second depth image is (x3, y3). The target position information corresponding to the position information of the first pixel point in the second depth image obtained through the above formula (1) is (x4, y4), the target position information corresponding to the position information of the second pixel point in the second depth image is (x5, y5), and the target position information corresponding to the position information of the third pixel point in the second depth image is (x6, y6). Determine that the depth information of the pixel point with the position information (x4, y4) in the first depth image is D1, the depth information of the pixel point with the position information (x5, y5) in the first depth image is D2, and the depth information of the pixel point with the position information (x6, y6) in the first depth image is D3. According to (x1, y1), (x2, y2), (x3, y3), D1, D2, D3, generate a third depth image, where the depth information of the pixel point (x1, y1) in the third depth image is D1, the depth information of the pixel point (x2, y2) is D2, and the depth information of the pixel point (x3, y3) is D3.
[0115] In step 205, determine the first loss value between the second depth image and the third depth image.
[0116] Optionally, after obtaining the third depth image in the above steps, the process of determining the first loss value between the second depth image and the third depth image includes: determining the first loss value between the second depth image and the third depth image according to the depth information of each pixel point in the second depth image and the depth information of each pixel point in the third depth image.
[0117] Exemplarily, the first loss value between the second depth image and the third depth image is determined according to the following formula (2).
[0118]
[0119] In the above formula (2), loss1 is the first loss value between the second depth image and the third depth image, n is the number of pixel points in the second depth image, is the depth information of the i-th pixel point in the second depth image, is the depth information of the i-th pixel point in the third depth image. The number of pixel points in the second depth image and the third depth image is the same.
[0120] In step 206, the initial pose model is trained according to the first loss value to obtain the target pose model.
[0121] Optionally, based on the fact that the first loss value is in a non-converged state, the parameters of the initial pose model are adjusted to obtain an adjusted pose model; according to the adjusted pose model, the first sample image, and the second sample image, an updated loss value is determined. Based on the fact that the updated loss value is in a converged state, the adjusted pose model is used as the target pose model. If the updated loss value is still in a non-converged state, the parameters of the updated pose model are continuously adjusted until a converged loss value is obtained, and the pose model at the time of the converged loss value is used as the target pose model. The target pose model is used to determine the pose matrix between the images corresponding to the two point clouds, and the pose matrix is used for registering the two point clouds.
[0122] Optionally, based on the fact that the first loss value is not less than the loss threshold, the parameters of the initial pose model are adjusted to obtain an adjusted pose model; according to the adjusted pose model, the first sample image, and the second sample image, an updated loss value is determined; based on the fact that the updated loss value is less than the loss threshold, the adjusted pose model is used as the target pose model. If the updated loss value is still not less than the loss threshold, the parameters of the updated pose model are continuously adjusted until a loss value less than the loss threshold is obtained, and the pose model at the time of the loss value less than the loss threshold is used as the target pose model. The loss threshold is set based on experience or adjusted according to the implementation environment, and the embodiments of the present application do not limit this.
[0123] In a possible implementation, the relevant information of any sample point further includes the intensity information of any sample point. The training method of the pose model further includes: obtaining a third intensity image according to a first pose matrix, a first intensity image, and a second intensity image, where the first intensity image is determined based on a first sample image, a second sample image, and a camera image at a first position at a first time, and the second intensity image is obtained based on the intensity information of a plurality of second sample points. Determining a second loss value between the second intensity image and the third intensity image. The process of training the initial pose model according to the first loss value to obtain the target pose model includes: training the initial pose model according to the first loss value and the second loss value to obtain the target pose model.
[0124] Wherein, the process of obtaining the first intensity image according to the first sample image, the second sample image, and the camera image at the first position at the first time includes: superimposing the first sample image, the second sample image, and the camera image at the first position at the first time to obtain a superimposed image; inputting the superimposed image into an initial intensity model to obtain the first intensity image. The initial intensity model is based on an Encoder-Decoder structure similar to Hourglass Network of R-18 and is used to predict the dense intensity attribute from the image.
[0125] Optionally, since the second sample image includes a depth image and an intensity image, the intensity image included in the second sample image can be directly used as the second intensity image.
[0126] In a possible implementation, the process of obtaining the third intensity image according to the first pose matrix, the first intensity image, and the second intensity image includes: obtaining a plurality of reference position information according to the position information, depth information, and the first pose matrix of each pixel point in the second intensity image, and each pixel point in the second intensity image corresponds to a reference position information; using the intensity information of the pixel points with the same position as the corresponding reference position information in the first intensity image as the intensity information corresponding to each reference position information; generating the third intensity image according to the position information of each pixel point in the second intensity image and the intensity information corresponding to the reference position information corresponding to each pixel point's position information.
[0127] Wherein, the depth information of the pixel points in the second intensity image is used as the depth information of the pixel points in the second intensity image, which is the same as the position information of the pixel points in the second intensity image. The process of obtaining a plurality of reference position information according to the position information, depth information, and the first pose matrix of each pixel point in the second intensity image is similar to the process of obtaining a plurality of target position information according to the position information, depth information, and the first pose matrix of each pixel point in the second depth image in the above process, and will not be elaborated here.
[0128] Exemplarily, the second intensity image includes 3 pixel points. The position information of the first pixel point in the second intensity image is (x7, y7), the position information of the second pixel point in the second intensity image is (x8, y8), and the position information of the third pixel point in the second intensity image is (x9, y9). The reference position information corresponding to the position information of the first pixel point is obtained as (x10, y10) through the above formula (3), the reference position information corresponding to the position information of the second pixel point is (x11, y11), and the reference position information corresponding to the position information of the third pixel point is (x12, y12). Determine that the intensity information of the pixel point with the position information (x10, y10) in the first intensity image is I1, the intensity information of the pixel point with the position information (x11, y11) in the first intensity image is I2, and the intensity information of the pixel point with the position information (x12, y12) in the first intensity image is I3. Generate a third intensity image according to (x7, y7), (x8, y8), (x9, y9), I1, I2, and I3. In the third intensity image, the intensity information of the pixel point (x7, y7) is I1, the intensity information of the pixel point (x8, y8) is I2, and the intensity information of the pixel point (x9, y9) is I3.
[0129] Optionally, the process of determining the second loss value between the second intensity image and the third intensity image includes: determining the second loss value between the second intensity image and the third intensity image according to the intensity information of each pixel point in the second intensity image and the intensity information of each pixel point in the third intensity image.
[0130] Exemplarily, the second loss value between the second intensity image and the third intensity image is determined according to the following formula (3).
[0131]
[0132] In the above formula (3), loss2 is the second loss value between the second intensity image and the third intensity image, n is the number of pixel points in the second intensity image, is the intensity information of the i-th pixel point in the second intensity image, is the intensity information of the i-th pixel point in the third intensity image. The number of pixel points in the second intensity image and the third intensity image is the same.
[0133] Optionally, the cross-entropy loss function can also be used to determine the second loss value between the second intensity image and the third intensity image, and the cross-entropy loss function can also be used to determine the first loss value between the second depth image and the third depth image. The embodiments of the present application do not limit the determination methods of the second loss value and the first loss value.
[0134] In a possible implementation, the method for training the pose model further includes: determining a fourth depth image according to the depth information of a plurality of first sample points; determining a third loss value between the fourth depth image and the first depth image; determining a fourth intensity image according to the intensity information of the plurality of first sample points; and determining a fourth loss value between the fourth intensity image and the first intensity image. The process of training the initial pose model according to the first loss value and the second loss value to obtain the target pose model includes: training the initial pose model according to the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target pose model.
[0135] Among them, the process of determining the fourth depth image according to the depth information of the plurality of first sample points includes: generating a fourth depth image according to the position information of the plurality of first sample points and the depth information of each first sample point, where the depth at the position information of each first sample point in the fourth depth image is the depth information of each first sample point. The process of determining the fourth intensity image according to the intensity information of the plurality of first sample points includes: generating a fourth intensity image according to the position information of the plurality of first sample points and the intensity information of each first sample point, where the intensity at the position information of each first sample point in the fourth intensity image is the intensity information of each first sample point.
[0136] The process of determining the third loss value between the fourth depth image and the first depth image is similar to the process of determining the first loss value between the second depth image and the third depth image described above, and will not be elaborated here. The process of determining the fourth loss value between the fourth intensity image and the first intensity image is similar to the process of determining the second loss value between the second intensity image and the third intensity image described above, and will not be elaborated here either.
[0137] In a possible implementation, the process of training the initial pose model according to the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target pose model includes: determining a total loss value according to the first loss value, the second loss value, the third loss value, and the fourth loss value; based on the fact that the total loss value is not in a converged state, adjusting the parameters of the initial pose model to obtain an adjusted pose model, and determining an updated total loss value according to the adjusted pose model, the first sample image, and the second sample image; and based on the fact that the updated total loss value is in a converged state, using the adjusted pose model as the target pose model.
[0138] The process of determining the total loss value according to the first loss value, the second loss value, the third loss value, and the fourth loss value includes: using the sum of the first loss value, the second loss value, the third loss value, and the fourth loss value as the total loss value.
[0139] When the total loss value is not in a convergent state, the parameters of the initial pose model are adjusted to obtain an adjusted pose model. The initial depth model and the initial intensity model can also be adjusted to obtain an adjusted depth model and an adjusted intensity model. According to the adjusted pose model, the adjusted depth model, the adjusted intensity model, the first sample image, and the second sample image, the updated total loss value is determined.
[0140] The average error and standard deviation of the pose matrix determined by the target pose model obtained by the pose model training method provided by the embodiments of the present application in each dimension are shown in Table 1 below, where the angle unit is degrees and the translation unit is meters.
[0141] Table 1
[0142] Quantitative index Quantitative result Angle average error [0.1042,0.1455,0.1566] Translation average error [0.0975,0.0756,0.1936] Angle standard deviation [0.2691,0.1969,0.1868] Translation standard deviation [0.23,0.2668,0.2999]
[0143] As can be seen from Table 1 above, the average angle error of the pose matrix in the x direction is 0.1042, the average angle error in the y direction is 0.1455, and the average angle error in the z direction is 0.1566. The average translation error, angle standard deviation, and translation standard deviation of the pose matrix in the x direction, y direction, and z direction are shown in Table 1 above and will not be elaborated here. It can be seen from Table 1 that the average error and standard deviation of the target pose model trained by the embodiments of the present application in each dimension are smaller than those of other deep learning solutions. Therefore, the accuracy of the target pose model trained by the embodiments of the present application is relatively high and it is friendly for deployment.
[0144] The above method does not require screening in the point cloud. Therefore, it is applicable to both dense point clouds and sparse point clouds, and has stronger robustness and a wider range of usage scenarios. Moreover, this method not only considers the position information of the sample points, but also considers the depth information of the sample points, and the considered information is relatively comprehensive. As a result, the accuracy of the trained target pose model is higher. Using the target pose model with higher accuracy to determine the pose matrix between the images corresponding to the two point clouds makes the determined pose matrix more accurate. Furthermore, using the more accurate pose matrix to register the two point clouds makes the registration effect of the two point clouds better and the registration accuracy higher.
[0145] As Figure 4 is a flowchart of the training process of a pose model provided by the embodiments of the present application. In Figure 4 , the first sample point cloud and the second sample point cloud are obtained, the first sample image is generated according to the first sample point cloud, and the second sample image is generated according to the second sample point cloud; according to the first sample image and the second sample image, a first combined image is obtained; the first combined image is input into the initial pose model to obtain the feature vector of the first combined image; the feature vector of the first combined image is input into the pose decoder to obtain the first pose matrix.
[0146] Obtain a camera image at a first position at a first time; obtain a superimposed image according to a first sample image, a second sample image, and the camera image at the first position at the first time; input the superimposed image into an initial depth model to obtain a first depth image; input the superimposed image into an initial intensity model to obtain a first intensity image.
[0147] Obtain a first intensity image and a first depth image according to a first sample point cloud; obtain a second intensity image and a second depth image according to a second sample point cloud. Obtain a third depth image according to the first depth image, the second depth image, and a first pose matrix. Obtain a third intensity image according to the first intensity image, the second intensity image, and the first pose matrix.
[0148] Determine a first loss value between the second depth image and the third depth image, determine a second loss value between the second intensity image and the third intensity image, determine a third loss value between the first depth image and a fourth depth image, determine a fourth loss value between the first intensity image and a fourth intensity image; determine a loss sum value among the first loss value, the second loss value, the third loss value, and the fourth loss value, and train an initial pose model, an initial depth model, and an initial intensity model according to the loss sum value to obtain a target pose model, a target depth model, and a target intensity model.
[0149] Figure 5 is a flowchart of a point cloud registration method provided by an embodiment of the present application, and this method can be executed by Figure 1 the computer device 101 in. As Figure 5 shown, this method includes the following steps 501 to step 503.
[0150] In step 501, obtain a first point cloud and a second point cloud.
[0151] Among them, the first point cloud and the second point cloud are both collected at a second position at a third time and a fourth time respectively, and the first point cloud includes relevant information of a plurality of first points, the second point cloud includes relevant information of a plurality of second points, and the relevant information of any point includes the position information and depth information of any point. The third time and the fourth time are different. The third time can be earlier than the fourth time or later than the fourth time. The embodiment of the present application does not limit this.
[0152] Optionally, a lidar is installed in the vehicle. When the lidar collects the first point cloud at the second position at the third time, the lidar sends the first point cloud to the computer device so that the computer device can obtain the first point cloud. When the lidar collects the second point cloud at the second position at the fourth time, the lidar sends the second point cloud to the computer device so that the computer device can obtain the second point cloud.
[0153] In step 502, according to the relevant information of multiple first points, obtain a first image of the second position at the third time, and according to the relevant information of multiple second points, obtain a second image of the second position at the fourth time.
[0154] In a possible implementation, the process of obtaining the first image of the second position at the third time according to the relevant information of multiple first points is similar to the process of obtaining the second image of the second position at the fourth time according to the relevant information of multiple second points, and the process of obtaining the first image of the second position at the third time according to the relevant information of multiple first points is similar to the process of obtaining the first sample image of the first position at the first time according to the relevant information of multiple first sample points described above, which will not be elaborated here.
[0155] In step 503, call the target pose model to determine the second pose matrix based on the first image and the second image.
[0156] Among them, the second pose matrix is used for registering the first point cloud and the second point cloud, and the target pose model is trained based on the method embodiments shown above Figure 2 and obtained.
[0157] In a possible implementation, the process of calling the target pose model to determine the second pose matrix based on the first image and the second image includes: obtaining a second combined image according to the first image and the second image, where the second combined image includes the first image and the second image; inputting the second combined image into the target pose model to obtain a feature vector of the second combined image, and the feature vector of the second combined image is used to represent the second combined image; calling a pose decoder to decode the feature vector of the second combined image to obtain the second pose matrix.
[0158] Among them, the first image and the second image are superimposed to obtain the second combined image.
[0159] Optionally, after determining the second pose matrix, when the external parameters between the lidar and the camera are known, the first point cloud and the second point cloud can also be registered according to the second pose matrix. The registration process includes: adjusting the second image according to the second pose matrix to obtain a third image; registering the third image and the first image to obtain the different regions between the third image and the first image, and the different regions are the changed regions in the second position. In addition, the pose relationship between the first point cloud and the second point cloud can be calculated according to the second pose matrix and the external parameters between the lidar and the camera, so as to realize the registration of the two point clouds and the detection of the changed region of the point cloud.
[0160] As Figure 6 is a picture of point cloud registration provided by an embodiment of the present application. Figure 6 In (1) of it, it is the registration map between the second images corresponding to the first point cloud and the second point cloud,Figure 6 Among them, (2) is the registration map between the first point cloud and the third image, and the third image is the image obtained by processing the second image corresponding to the second point cloud through the above-mentioned second pose matrix. From Figure 6 It can be seen that Figure 6 the registration effect of (2) in
[0161] The above method does not require screening in the point cloud. Therefore, it is applicable to both dense point clouds and sparse point clouds, with stronger robustness and a wider range of usage scenarios. Moreover, the accuracy of the target pose model used in this method is higher. Using a target pose model with higher accuracy to determine the pose matrix between the images corresponding to the first point cloud and the second point cloud makes the determined pose matrix more accurate. Furthermore, using a more accurate pose matrix to register the first point cloud and the second point cloud results in better registration effect and higher registration accuracy for the two point clouds. In addition, this method only involves one forward propagation of the neural network and does not require a process of multiple iterations for refinement, making the efficiency of point cloud registration higher.
[0162] Figure 7 is a flowchart of a point cloud registration method provided by an embodiment of the present application. The method includes: obtaining a first point cloud and a second point cloud. Generating a first image according to the first point cloud; generating a second image according to the second point cloud. Generating a second combined image according to the first image and the second image. Inputting the second combined image into a target pose model to obtain a feature vector of the second combined image. Inputting the feature vector of the second combined image into a pose decoder to obtain a second pose matrix.
[0163] Figure 8 shows a schematic structural diagram of a training device for a pose model provided by an embodiment of the present application. As Figure 8 shown, the device includes:
[0164] An acquisition module 801, configured to acquire a first sample point cloud and a second sample point cloud. The first sample point cloud and the second sample point cloud are both acquired at a first position at a first time and a second time respectively, and the first sample point cloud includes relevant information of a plurality of first sample points, and the second sample point cloud includes relevant information of a plurality of second sample points. The relevant information of any sample point includes the position information and depth information of any sample point;
[0165] The acquisition module 801 is further configured to acquire a first sample image of the first position at the first time according to the relevant information of the plurality of first sample points; acquire a second sample image of the first position at the second time according to the relevant information of the plurality of second sample points;
[0166] A determination module 802, configured to call an initial pose model to determine a first pose matrix based on the first sample image and the second sample image. The first pose matrix is used to register the first sample point cloud and the second sample point cloud;
[0167] The obtaining module 801 is further configured to obtain a third depth image according to the first pose matrix, the first depth image, and the second depth image, where the first depth image is determined based on the first sample image, the second sample image, and the camera image at the first position at the first time, and the second depth image is obtained based on the depth information of multiple second sample points;
[0168] The determining module 802 is further configured to determine a first loss value between the second depth image and the third depth image;
[0169] The training module 803 is configured to train the initial pose model according to the first loss value to obtain a target pose model, where the target pose model is used to determine the pose matrix between the images corresponding to two point clouds, and the pose matrix is used to register the two point clouds.
[0170] In a possible implementation manner, the relevant information of any sample point further includes the intensity information of any sample point;
[0171] The obtaining module 801 is further configured to obtain a third intensity image according to the first pose matrix, the first intensity image, and the second intensity image, where the first intensity image is determined based on the first sample image, the second sample image, and the camera image at the first position at the first time, and the second intensity image is obtained based on the intensity information of multiple second sample points;
[0172] The determining module 802 is further configured to determine a second loss value between the second intensity image and the third intensity image;
[0173] The training module 803 is configured to train the initial pose model according to the first loss value and the second loss value to obtain a target pose model.
[0174] In a possible implementation manner, the determining module 802 is further configured to determine a fourth depth image according to the depth information of multiple first sample points; determine a third loss value between the fourth depth image and the first depth image; determine a fourth intensity image according to the intensity information of multiple first sample points; determine a fourth loss value between the fourth intensity image and the first intensity image;
[0175] The training module 803 is configured to train the initial pose model according to the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain a target pose model.
[0176] In a possible implementation, a training module 803 is configured to determine a total loss value according to a first loss value, a second loss value, a third loss value, and a fourth loss value; adjust parameters of an initial pose model based on that the total loss value is not in a converged state to obtain an adjusted pose model; determine an updated total loss value according to the adjusted pose model, a first sample image, and a second sample image; and use the adjusted pose model as a target pose model based on that the updated total loss value is in a converged state.
[0177] In a possible implementation, a determination module 802 is configured to obtain a first combined image according to a first sample image and a second sample image, where the first combined image includes the first sample image and the second sample image; input the first combined image into the initial pose model to obtain a feature vector of the first combined image, where the feature vector of the first combined image is used to represent the first combined image; and call a pose decoder to decode the feature vector of the first combined image to obtain a first pose matrix.
[0178] In a possible implementation, an acquisition module 801 is configured to obtain a plurality of target position information according to position information, depth information, and the first pose matrix of each pixel point in a second depth image, where each pixel point in the second depth image corresponds to one piece of target position information; use the depth information of the pixel points in the first depth image with the same positions corresponding to the respective target position information as the depth information corresponding to the respective target position information; and generate a third depth image according to the position information of each pixel point in the second depth image and the depth information corresponding to the target position information corresponding to each pixel point's position information.
[0179] The above device does not need to perform screening in the point cloud. Therefore, it is applicable to both dense point clouds and sparse point clouds, and has stronger robustness and a wider range of usage scenarios. Moreover, not only the position information of the sample points is considered, but also the depth information of the sample points is considered, and the considered information is relatively comprehensive. As a result, the accuracy of the trained target pose model is higher. Using the target pose model with higher accuracy to determine the pose matrix between the images corresponding to the two point clouds makes the determined pose matrix more accurate. Furthermore, using the pose matrix with higher accuracy to register the two point clouds makes the registration effect of the two point clouds better and the registration accuracy higher.
[0180] Figure 9 is a schematic structural diagram of a point cloud registration device provided by an embodiment of the present application, as Figure 9 shown, the device includes:
[0181] An acquisition module 901, configured to acquire a first point cloud and a second point cloud. The first point cloud and the second point cloud are respectively acquired at a third time and a fourth time at a second position. The first point cloud includes information related to a plurality of first points, and the second point cloud includes information related to a plurality of second points. The information related to any point includes the position information and depth information of any point.
[0182] The acquisition module 901 is further configured to obtain a first image of the second position at the third time according to the information related to the plurality of first points; and obtain a second image of the second position at the fourth time according to the information related to the plurality of second points.
[0183] A determination module 902, configured to call a target pose model to determine a second pose matrix based on the first image and the second image. The second pose matrix is used for registering the first point cloud and the second point cloud. The target pose model is trained based on the device embodiments shown above. Figure 8 shown in the device embodiment.
[0184] In a possible implementation manner, the determination module 902 is configured to obtain a second combined image according to the first image and the second image. The second combined image includes the first image and the second image; input the second combined image into the target pose model to obtain a feature vector of the second combined image. The feature vector of the second combined image is used to characterize the second combined image; call a pose decoder to decode the feature vector of the second combined image to obtain a second pose matrix.
[0185] The above device does not need to perform screening on the point cloud. Therefore, it is applicable to both dense point clouds and sparse point clouds, and has stronger robustness and a wider range of usage scenarios. Moreover, the accuracy of the target pose model used is higher. Using a target pose model with higher accuracy to determine the pose matrix between the images corresponding to the first point cloud and the second point cloud makes the determined pose matrix more accurate. Furthermore, using a more accurate pose matrix to register the first point cloud and the second point cloud makes the registration effect of the two point clouds better and the registration accuracy higher. In addition, this method only involves one forward propagation of the neural network and does not require a process of multiple iterations for refinement, so the efficiency of point cloud registration is higher.
[0186] It should be understood that when the above provided device implements its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0187] Figure 10The block diagram of the terminal device 1000 provided by an exemplary embodiment of the present application is shown. The terminal device 1000 may be a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart vehicle head unit, a smart TV, a smart speaker, a smart watch, etc.
[0188] Generally, the terminal device 1000 includes: a processor 1001 and a memory 1002.
[0189] The processor 1001 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1001 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1001 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1001 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1001 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0190] The memory 1002 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1002 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1001 to implement the Figure 2 training method of the pose model provided by the method embodiment shown in the present application, or to implement the Figure 5 point cloud registration method provided by the method embodiment shown in the present application.
[0191] In some embodiments, the terminal device 1000 may further optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, the memory 1002, and the peripheral device interface 1003 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1003 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1009.
[0192] The peripheral device interface 1003 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1001 and the memory 1002. In some embodiments, the processor 1001, the memory 1002, and the peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1001, the memory 1002, and the peripheral device interface 1003 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0193] The radio frequency circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1004 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1004 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1004 may communicate with other terminal devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1004 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0194] The display screen 1005 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1005 is a touch display screen, the display screen 1005 also has the ability to collect touch signals on or above the surface of the display screen 1005. The touch signals can be input as control signals to the processor 1001 for processing. At this time, the display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, which is provided on the front panel of the terminal device 1000; in other embodiments, there may be at least two display screens 1005, which are respectively provided on different surfaces of the terminal device 1000 or are in a foldable design; in other embodiments, the display screen 1005 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal device 1000. Even further, the display screen 1005 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1005 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0195] The camera module 1006 is used to capture images or videos. Optionally, the camera module 1006 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal device 1000, and the rear camera is provided on the back of the terminal device 1000. In some embodiments, there are at least two rear cameras, which can be any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, to achieve functions such as the fusion of the main camera and the depth-of-field camera to achieve background blurring, the fusion of the main camera and the wide-angle camera to achieve panoramic shooting, and VR (Virtual Reality) shooting functions or other fusion shooting functions. In some embodiments, the camera module 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0196] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1001 for processing, or input to the radio frequency circuit 1004 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal device 1000. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1007 may further include a headphone jack.
[0197] The power supply 1009 is used to supply power to each component in the terminal device 1000. The power supply 1009 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1009 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0198] In some embodiments, the terminal device 1000 further includes one or more sensors 1010. The one or more sensors 1010 include but are not limited to: an acceleration sensor 1011, a gyroscope sensor 1012, a pressure sensor 1013, an optical sensor 1015, and a proximity sensor 1016.
[0199] The acceleration sensor 1011 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal device 1000. For example, the acceleration sensor 1011 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1001 can control the display screen 1005 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1011. The acceleration sensor 1011 can also be used for game or user motion data collection.
[0200] The gyroscope sensor 1012 can detect the body direction and rotation angle of the terminal device 1000. The gyroscope sensor 1012 can cooperate with the acceleration sensor 1011 to collect the 3D actions of the user on the terminal device 1000. Based on the data collected by the gyroscope sensor 1012, the processor 1001 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0201] The pressure sensor 1013 can be disposed on the side frame of the terminal device 1000 and / or the lower layer of the display screen 1005. When the pressure sensor 1013 is disposed on the side frame of the terminal device 1000, it can detect the holding signal of the user on the terminal device 1000, and the processor 1001 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1013. When the pressure sensor 1013 is disposed on the lower layer of the display screen 1005, the processor 1001 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1005. The operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0202] The optical sensor 1015 is used to collect the ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 according to the ambient light intensity collected by the optical sensor 1015. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera module 1006 according to the ambient light intensity collected by the optical sensor 1015.
[0203] The proximity sensor 1016, also known as the distance sensor, is usually disposed on the front panel of the terminal device 1000. The proximity sensor 1016 is used to collect the distance between the user and the front of the terminal device 1000. In one embodiment, when the proximity sensor 1016 detects that the distance between the user and the front of the terminal device 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from the lit state to the off state; when the proximity sensor 1016 detects that the distance between the user and the front of the terminal device 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from the off state to the lit state.
[0204] Those skilled in the art can understand that Figure 10 the structure shown in does not limit the terminal device 1000, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0205] Figure 11The following is a schematic structural diagram of the server provided by the embodiment of the present application. The server 1100 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Unit, CPU) 1101 and one or more memories 1102. Among them, at least one program code is stored in the one or more memories 1102, and the at least one program code is loaded and executed by the one or more processors 1101 to implement the above-mentioned Figure 2 training method of the pose model provided by the method embodiment shown, or to implement the above-mentioned Figure 5 point cloud registration method provided by the method embodiment shown. Of course, the server 1100 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server 1100 may also include other components for implementing the functions of the device, which will not be elaborated here.
[0206] In an exemplary embodiment, a computer-readable storage medium is also provided. At least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to enable a computer to implement the above-mentioned Figure 2 training method of the pose model provided by the method embodiment shown, or to implement the above-mentioned Figure 5 point cloud registration method provided by the method embodiment shown.
[0207] Optionally, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0208] In an exemplary embodiment, a computer program or a computer program product is also provided. At least one computer instruction is stored in the computer program or the computer program product, and the at least one computer instruction is loaded and executed by a processor to enable a computer to implement the above-mentioned Figure 2 training method of the pose model provided by the method embodiment shown, or to implement the above-mentioned Figure 5 point cloud registration method provided by the method embodiment shown.
[0209] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the point clouds and camera images involved in this application are obtained under full authorization.
[0210] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0211] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0212] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for a pose model, characterized in that, The method includes: Obtaining a first sample point cloud and a second sample point cloud, where both the first sample point cloud and the second sample point cloud are collected at a first position at a first time and a second time respectively, and the first sample point cloud includes information related to a plurality of first sample points, the second sample point cloud includes information related to a plurality of second sample points, and the information related to any sample point includes the position information and depth information of the any sample point; Obtaining a first sample image of the first position at the first time according to the information related to the plurality of first sample points; obtaining a second sample image of the first position at the second time according to the information related to the plurality of second sample points; Invoking an initial pose model to determine a first pose matrix based on the first sample image and the second sample image, where the first pose matrix is used to register the first sample point cloud and the second sample point cloud; Obtaining a third depth image according to the first pose matrix, a first depth image and a second depth image, where the first depth image is determined based on the first sample image, the second sample image and a camera image at the first position at the first time, and the second depth image is obtained based on the depth information of the plurality of second sample points; Determining a first loss value between the second depth image and the third depth image; Training the initial pose model according to the first loss value to obtain a target pose model, where the target pose model is used to determine a pose matrix between images corresponding to two point clouds, and the pose matrix is used to register the two point clouds.
2. The method according to claim 1, characterized in that, The information related to any sample point further includes the intensity information of the any sample point; the method further includes: Obtaining a third intensity image according to the first pose matrix, a first intensity image and a second intensity image, where the first intensity image is determined based on the first sample image, the second sample image and a camera image at the first position at the first time, and the second intensity image is obtained based on the intensity information of the plurality of second sample points; Determining a second loss value between the second intensity image and the third intensity image; The training the initial pose model according to the first loss value to obtain the target pose model includes: Training the initial pose model according to the first loss value and the second loss value to obtain the target pose model.
3. The method according to claim 2, wherein The method further includes: Determining a fourth depth image according to the depth information of the plurality of first sample points; Determining a third loss value between the fourth depth image and the first depth image; Determining a fourth intensity image according to the intensity information of the plurality of first sample points; Determining a fourth loss value between the fourth intensity image and the first intensity image; The training the initial pose model according to the first loss value and the second loss value to obtain the target pose model includes: Training the initial pose model according to the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target pose model.
4. The method according to claim 3, wherein Training the initial pose model according to the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target pose model includes: Determining the total loss value according to the first loss value, the second loss value, the third loss value, and the fourth loss value; Based on the fact that the total loss value is not in a converged state, adjusting the parameters of the initial pose model to obtain an adjusted pose model; Determining an updated total loss value according to the adjusted pose model, the first sample image, and the second sample image; Based on the fact that the updated total loss value is in a converged state, using the adjusted pose model as the target pose model.
5. The method according to any one of claims 1 to 4, characterized in that The calling the initial pose model to determine a first pose matrix based on the first sample image and the second sample image includes: Obtaining a first combined image according to the first sample image and the second sample image, where the first combined image includes the first sample image and the second sample image; Inputting the first combined image into the initial pose model to obtain a feature vector of the first combined image, where the feature vector of the first combined image is used to represent the first combined image; Calling a pose decoder to decode the feature vector of the first combined image to obtain the first pose matrix.
6. The method according to any one of claims 1 to 4, characterized in that The obtaining a third depth image according to the first pose matrix, the first depth image, and the second depth image includes: Obtaining a plurality of target position information according to the position information, depth information, and the first pose matrix of each pixel point in the second depth image, where each pixel point in the second depth image corresponds to one target position information; Using the depth information of the pixel points in the first depth image that have the same position as the target position information as the depth information corresponding to the respective target position information; Generating the third depth image according to the position information of each pixel point in the second depth image and the depth information corresponding to the target position information corresponding to each pixel point's position information.
7. A point cloud registration method, characterized in that, The method includes: Obtaining a first point cloud and a second point cloud, where both the first point cloud and the second point cloud are collected at a second position at a third time and a fourth time respectively, and the first point cloud includes relevant information of a plurality of first points, the second point cloud includes relevant information of a plurality of second points, and the relevant information of any point includes the position information and depth information of the any point; Obtaining a first image of the second position at the third time according to the relevant information of the plurality of first points; obtaining a second image of the second position at the fourth time according to the relevant information of the plurality of second points; Calling a target pose model to determine a second pose matrix based on the first image and the second image, where the second pose matrix is used for registering the first point cloud and the second point cloud, and the target pose model is trained according to the method described in any one of claims 1 to 6 above.
8. The method according to claim 7, wherein The calling the target pose model to determine a second pose matrix based on the first image and the second image includes: Obtain a second combined image according to the first image and the second image, where the second combined image includes the first image and the second image; Input the second combined image into the target pose model to obtain a feature vector of the second combined image, where the feature vector of the second combined image is used to characterize the second combined image; Call a pose decoder to decode the feature vector of the second combined image to obtain the second pose matrix.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program code is stored in the memory. The at least one program code is loaded and executed by the processor so that the computer device implements the pose model training method according to any one of claims 1 to 6, or so that the computer device implements the point cloud registration method according to claim 7 or 8.
10. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium. The at least one program code is loaded and executed by a processor so that a computer device implements the pose model training method according to any one of claims 1 to 6, or so that the computer device implements the point cloud registration method according to claim 7 or 8.
Citation Information
Patent Citations
Head posture detection method and system based on RGB-D image
CN111414798A
Visual repositioning method and system
CN112132900A