Communication method and apparatus

By ensuring the neighborhood matching degree of parallel imaging planes in multimodal sensing, the problem of low neighborhood matching degree in multimodal fusion sensing is solved, thereby improving sensing performance and reducing computational and communication overhead.

WO2026067107A1PCT designated stage Publication Date: 2026-04-02HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

In existing multimodal sensing schemes, there are still challenges in improving the performance of multimodal fusion sensing, especially when the target object imaging results are greatly deformed, resulting in low neighborhood matching degree and easy misidentification as noise.

Method used

By acquiring the first and second neighborhoods and ensuring their imaging planes are parallel, multimodal fusion sensing technology is used, combined with the determination of neighborhood matching degree and noise identification, to reduce computational overhead and improve the accuracy of multimodal fusion sensing.

Benefits of technology

It effectively reduces the deformation of target imaging results, improves the performance of multimodal fusion sensing, reduces computational and communication overhead, and enhances the accuracy of neighborhood matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025121368_02042026_PF_FP_ABST
    Figure CN2025121368_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of communications, and provides a communication method and an apparatus, which are used for improving multi-modal sensing performance. In the method, a first apparatus acquires a first neighborhood i and a second neighborhood i on the basis of N pieces of point cloud data, and performs multi-modal fusion sensing on the basis of the first neighborhood i and the second neighborhood i, wherein a first image used for determining the first neighborhood i is parallel to an imaging plane of a second image used for determining the second neighborhood i, such that an imaging result of a target object in the first neighborhood i and the second neighborhood i involves small deformation. In this way, when the neighborhood matching degree of i-th point cloud data is determined on the basis of the first neighborhood i and the second neighborhood i, the misidentification of the i-th point cloud data as noise on the basis of the neighborhood matching degree because the neighborhood matching degree of the i-th point cloud data is low due to large deformation of the imaging result of the target object in the first neighborhood i and the second neighborhood i can be avoided, thereby improving the performance of multi-modal fusion sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Communication method and apparatus

[0001] The present application claims priority from the Chinese patent application No. 202411369046.0 filed on September 27, 2024, and entitled "Communication method and apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present application relates to the field of communication, in particular to a communication method and apparatus. BACKGROUND

[0003] Perception can be divided into two categories: two-dimensional (2D) or three-dimensional (3D) perception. Among them, 2D perception can capture a scene onto a two-dimensional plane to generate a plane image. Common 2D perception schemes include optical perception, radar perception, etc. 3D perception can capture the shape and structure of a scene in three-dimensional space, i.e., generate a stereoscopic image. Common 3D perception schemes include radio frequency perception, computer tomography, etc.

[0004] Currently, 2D and 3D perception schemes can be integrated to improve perception performance through multi-dimensional data, i.e., multi-modal fusion perception. For example, radio frequency perception and optical perception multi-modal information can be integrated to improve perception performance. On this basis, how to improve multi-modal perception performance is a hot issue currently discussed. SUMMARY

[0005] Embodiments of the present application provide a communication method and apparatus to improve multi-modal perception performance.

[0006] To achieve the above object, the present application adopts the following technical solutions:

[0007] In a first aspect, a communication method is provided. The method can be performed by a first device, a component of the first device, such as a processor, a chip, a chip system, or a circuit of the first device, or a logic module or software that can implement all or part of the functions of the first device. The method includes: obtaining a first neighborhood i and a second neighborhood i according to N point cloud data, and performing multi-modal fusion perception according to the first neighborhood i and the second neighborhood i; wherein the N point cloud data are three-dimensional point cloud data obtained by sensing a target object, N is an integer greater than 1, i is an integer traversing 1 to N, the first neighborhood i is a region where an i-th point cloud data in the N point cloud data is located after being projected onto a first image, the second neighborhood i is a region where the i-th point cloud data is located after being projected onto a second image, the first image and the second image both include sensing results of the target object, an imaging plane of the first image and an imaging plane of the second image are parallel, and the first image and the second image are different two-dimensional images.

[0008] According to the method of the first aspect, the imaging plane of the first image used to determine the first neighborhood i is parallel to the imaging plane of the second image used to determine the second neighborhood i, which can make the imaging result of the target object on the first image and the imaging result of the target object on the second image have less deformation, so that the imaging result of the target object in the first neighborhood i and the second neighborhood i has less deformation. In this way, when determining the neighborhood matching degree of the i-th point cloud data based on the first neighborhood i and the second neighborhood i, it can be avoided that the neighborhood matching degree of the i-th point cloud data is too low due to the large deformation of the imaging result of the target object in the first neighborhood i and the second neighborhood i, so that the i-th point cloud data is mistakenly recognized as noise based on the neighborhood matching degree. Alternatively, a more accurate boundary and / or corner point of the target object can be determined based on the first neighborhood i and the second neighborhood i, and the multi-modal fusion perception is performed using the boundary and / or corner point of the target object. In this way, the performance of the multi-modal fusion perception can be improved.

[0009] In a possible design, the first neighborhood i is obtained according to the N point cloud data, including: sending a first message to a second device, the first message being used to request a region where each point cloud data in the N point cloud data is located after being projected onto a target image, the target image being an image obtained by two-dimensional sensing of the target object by the second device and converted according to a first parameter, and the first parameter being an internal and external parameter associated with the target image; and receiving the first neighborhood i from the second device. That is, the first device can request the second device to provide the first neighborhood i to the first device by sending the first message to the second device. In this way, the first device does not need to calculate the first neighborhood i, so that the computational overhead of the first device can be reduced.

[0010] Optionally, the internal and external parameters comprise at least one of a focal length, a resolution, a field of view, a position, or a pointing direction. The focal length, the resolution, the field of view, the position, and the pointing direction can be understood as the focal length, the resolution, the field of view, the position, and the pointing direction of the device associated with the target image. For example, the internal and external parameters can be the focal length, the resolution, the field of view, the position, and the pointing direction of the device associated with the image.

[0011] Optionally, the first message comprises at least one of a spatial range of each of the N point cloud data, a normal of each of the N point cloud data, or the first parameter, the spatial range of each of the N point cloud data is used to indicate the size and / or shape of the space where each of the N point cloud data is located, and the normal of each of the N point cloud data is used to determine the rotation matrix corresponding to each of the N point cloud data. It can be understood that the spatial range of each of the N point cloud data can be pre-set or protocol predefined, in which case the first message can not carry the spatial range of each of the N point cloud data, and the specific setting can be flexible according to the actual situation, which is not limited.

[0012] Optionally, the first neighborhood i is an area where the i-th point cloud data is located after being projected onto the first image, comprising: the first neighborhood i is an area where the i-th point cloud data is located in the first space i projected onto the first image; the second neighborhood i is an area where the i-th point cloud data is located after being projected onto the second image, comprising: the second neighborhood i is an area where the first space i is projected onto the second image; the first message is further used to request a scaling parameter, the scaling parameter being a scaling ratio between the imaging results of the objects in the two-dimensional perception; the first neighborhood i from the second device comprises: the first neighborhood i and the first scaling parameter from the second device, the first scaling parameter being used to indicate the scaling ratio between the imaging results of the target objects in the first image; the method of the first aspect further comprises: determining a third neighborhood i according to the first neighborhood i, the first scaling parameter, and a second scaling parameter, the second scaling parameter being used to indicate the scaling ratio between the imaging results of the target objects in the second image, the third neighborhood i comprising the same number of pixel points as the second neighborhood i; performing multi-modal fusion perception according to the first neighborhood i and the second neighborhood i, comprising: performing multi-modal fusion perception according to the third neighborhood i and the second neighborhood i.

[0013] It can be understood that when the first neighborhood i is a region of the first space i where the i-th point cloud data is projected onto the first image, and the second neighborhood i is a region of the first space i where the i-th point cloud data is projected onto the second image, the number of pixel points in the first neighborhood i and the number of pixel points in the second neighborhood i can be different. In this case, the number of pixel points in the first neighborhood i can be changed by the first scaling parameter and the second scaling parameter, so that the number of pixel points in the changed first neighborhood i (i.e., the third neighborhood i) is the same as the number of pixel points in the second neighborhood i, thereby facilitating determining the neighborhood matching degree of the i-th point cloud data by using the element-by-element comparison similarity method.

[0014] Further, according to the third neighborhood i and the second neighborhood i, multi-modal fusion perception is performed, including: determining the neighborhood matching degree of the i-th point cloud data according to the similarity of the third neighborhood i and the second neighborhood i; determining the noise in the N point cloud data according to the neighborhood matching degree of the i-th point cloud data; and performing multi-modal fusion perception based on the point cloud data in the N point cloud data except the noise.

[0015] Further, the first space i is spherical or cubic. Of course, the first space i can also be other shapes, which can be flexibly set according to actual conditions, and is not limited.

[0016] Further, before sending the first message to the second device, the method of the first aspect further includes: receiving a first capability parameter from the second device, the first capability parameter being used to indicate that the second device supports providing the scaling parameter; and sending the first message to the second device, including: sending the first message to the second device according to the first capability parameter. It can be understood that the first capability parameter can also be used to indicate that the second device supports providing the scaling parameter (to the first device). That is, the second device has the capability of providing the scaling parameter (to the first device). In this way, it can avoid the failure of the first device to request the scaling parameter from the second device, thereby causing additional communication overhead.

[0017] Optionally, the first message is also used to indicate an angle range, the pointing direction of the second device belongs to the angle range, and the two-dimensional image obtained by the device belonging to the angle range for perceiving the target object can be used to determine the two-dimensional image for projecting the N point cloud data. It can be understood that the pointing direction of the second device belonging to the angle range can ensure that there is an effective spatial range between the second device and the device associated with the first image. In this way, it can ensure that the first device obtains the related parameters and images of the two-dimensional perception of the object by the second device from the second device, thereby avoiding the failure of the first device to obtain the related parameters and images of the two-dimensional perception of the object by the second device, and generating additional communication overhead.

[0018] In a possible design, the first neighborhood i is a region where the i th point cloud data is located after being projected onto the first image, including: the first neighborhood i is a region where the i th point cloud data is located in the first space i projected onto the first image; the second neighborhood i is a region where the i th point cloud data is located after being projected onto the second image, including: the second neighborhood i is a region where the first space i is projected onto the second image; the method of the first aspect further includes: determining a third neighborhood i according to the first neighborhood i, a first scaling parameter and a second scaling parameter, the first scaling parameter being used to indicate a scaling ratio between an object and an imaging result of the object in the first image, the second scaling parameter being used to indicate a scaling ratio between the object and the imaging result of the object in the second image, the third neighborhood i including the same number of pixel points as the second neighborhood i; performing multi-modal fusion perception according to the first neighborhood i and the second neighborhood i, including: performing multi-modal fusion perception according to the third neighborhood i and the second neighborhood i.

[0019] It can be understood that when the first neighborhood i is a region where the i th point cloud data is located in the first space i projected onto the first image, and the second neighborhood i is a region where the i th point cloud data is located in the first space i projected onto the second image, the number of pixel points in the first neighborhood i and the number of pixel points in the second neighborhood i can be different. In this case, the number of pixel points in the first neighborhood i can be changed through the first scaling parameter and the second scaling parameter, so that the number of pixel points in the changed first neighborhood i (i.e., the third neighborhood i) is the same as the number of pixel points in the second neighborhood i, thereby facilitating determining the neighborhood matching degree of the i th point cloud data in a manner of comparing similarity element by element.

[0020] Optionally, the multi-modal fusion perception based on the first neighborhood i and the second neighborhood i includes: determining the neighborhood matching degree of the i th point cloud data according to the similarity degree of the third neighborhood i and the second neighborhood i; determining noise in the N point cloud data according to the neighborhood matching degree of the i th point cloud data; and performing multi-modal fusion perception based on the point cloud data in the N point cloud data except the noise.

[0021] Optionally, the method of the first aspect further includes: sending a second message to a second device, the second message being used to request a scaling parameter, the scaling parameter being a scaling ratio between an object and an imaging result of the object for two-dimensional perception; and receiving the first scaling parameter from the second device. In this way, the first device can obtain the scaling parameter from the second device in time when the scaling parameter of the second device is needed, thereby facilitating the first device to determine the third neighborhood i.

[0022] Further, the second message comprises at least one of: a spatial range of each of the N point cloud data, or a normal of each of the N point cloud data, the spatial range of each of the N point cloud data being used to indicate a size and / or a shape of a space where each of the N point cloud data is located, the normal of each of the N point cloud data being used to determine a rotation matrix corresponding to each of the N point cloud data.

[0023] Further, before sending the second message to the second device, the method of the first aspect further comprises: receiving a second capability parameter from the second device, the second capability parameter being used to indicate that the second device supports providing the scaling parameter; and sending the second message to the second device, comprising: sending the second message to the second device according to the second capability parameter. In this way, the first device can avoid failure in requesting the scaling parameter from the second device, thereby causing additional communication overhead.

[0024] Optionally, the method of the first aspect further comprises: sending a fifth message to a third device, the fifth message being used to request a scaling parameter, the scaling parameter being a scaling ratio between an object and an imaging result of the object for two-dimensional perception; and receiving the second scaling parameter from the third device. In this way, the first device can timely obtain the scaling parameter from the third device when the scaling parameter of the third device is needed, thereby facilitating the first device to determine the third neighborhood i.

[0025] Further, the fifth message comprises at least one of: a spatial range of each of the N point cloud data, or a normal of each of the N point cloud data, the spatial range of each of the N point cloud data being used to indicate a size and / or a shape of a space where each of the N point cloud data is located, the normal of each of the N point cloud data being used to determine a rotation matrix corresponding to each of the N point cloud data.

[0026] Further, before sending the fifth message to the third device, the method of the first aspect further comprises: receiving a fifth capability parameter from the third device, the fifth capability parameter being used to indicate that the third device supports providing the scaling parameter; and sending the fifth message to the third device, comprising: sending the fifth message to the second device according to the fifth capability parameter. In this way, the first device can avoid failure in requesting the scaling parameter from the third device, thereby causing additional communication overhead.

[0027] In a possible design, the first neighborhood i is obtained according to the N point cloud data, comprising: obtaining a second parameter and a third image, the second parameter being used to convert the third image into the first image; and determining the first neighborhood i according to the N point cloud data, the second parameter and the third image. It can be understood that the third image is an image obtained by two-dimensional perception of the target object, i.e., the third image comprises an imaging result of the target object. The third image can be used to determine the first image. In this way, the first device can convert the third image into the first image using the second parameter, and determine the first neighborhood i based on the first image.

[0028] Optionally, the obtaining the second parameter and the third image comprises: receiving the second parameter and the third image from the second device, the second parameter being a parameter related to the two-dimensional perception of the target object by the second device. That is, the first device can obtain the second parameter and the third image from the second device.

[0029] Further, the second parameter comprises at least one of: a focal length, a resolution, a field of view, a position, or a pointing direction. For example, the second parameter comprises the focal length, the resolution, the field of view, the position, and the pointing direction of the second device.

[0030] Further, the method of the first aspect further comprises: sending a third message to the second device, the third message being used to request the parameter related to the two-dimensional perception of the target object by the second device and the image. In this way, the first device can obtain the second parameter and the third image in time. Of course, the second device can also actively report the parameter related to the two-dimensional perception of the target object by the second device to the first device, such as periodically reporting the parameter to the first device. The specific reporting manner can be set according to actual conditions, and is not limited.

[0031] Further, the method of the first aspect further comprises: receiving a third capability parameter from the second device, the third capability parameter being used to indicate that the second device supports providing the parameter related to the two-dimensional perception of the target object and the image; and sending the third message to the second device, comprising: sending the third message to the second device according to the third capability parameter. In this way, the first device can avoid failure in requesting the parameter related to the two-dimensional perception of the target object by the second device and the image from the second device, thereby causing additional communication overhead.

[0032] In a possible design, the imaging plane of the first image and the imaging plane of the second image are parallel, and the method comprises: the positions of the second virtual device and the third virtual device are different, the third parameter of the second virtual device and the third virtual device is the same, the second virtual device is associated with the second image, the third virtual device is associated with the third image, and the third parameter is used to indicate the parameter related to the two-dimensional perception. In other words, when the positions of the second virtual device and the third virtual device are different and the third parameter is the same, the imaging plane #1 and the imaging plane #2 are parallel.

[0033] Optionally, the third parameter comprises at least one of: an angle or an internal parameter, and the internal parameter comprises at least one of: a focal length, a field of view, or a resolution.

[0034] In a possible design, the first neighborhood is a region where the i-th point cloud data is located after being projected onto the first image, and the second neighborhood is a region where the i-th point cloud data is located after being projected onto the second image. The i-th point cloud data is located in a first space i, which belongs to a first effective perception space. The first effective perception space is an intersection of a viewing pyramid of a second virtual device and a viewing pyramid of a second device. The second virtual device is associated with the first image, and the first image is determined according to an image obtained by two-dimensional perception of the second device. The second virtual device is located at the same position as the second device. The i-th point cloud data is located in the first space i, which belongs to a second effective perception space. The second effective perception space is an intersection of a viewing pyramid of a third virtual device and a viewing pyramid of a third device. The third virtual device is associated with the second image, and the second image is determined according to an image obtained by two-dimensional perception of the third device. The third virtual device is located at the same position as the third device.

[0035] In a possible design, the second neighborhood i is obtained according to the N point cloud data, including: sending, to the third device, a fourth message, where the fourth message is used to request a region where each point cloud data in the N point cloud data is located after being projected onto a target image, and the target image is an image obtained by two-dimensional perception of the target object by the third device, and the target image is converted according to a first parameter, and the first parameter is an internal-external parameter associated with the target image; and receiving the second neighborhood i from the third device. That is, the first device can request the third device to provide the second neighborhood i to the first device by sending the fourth message to the third device. In this way, the first device does not need to calculate the second neighborhood i, and therefore, the calculation overhead of the first device can be reduced.

[0036] Optionally, the internal-external parameter includes at least one of the following: a focal length, a resolution, a field of view, a position, or a pointing direction.

[0037] Optionally, the fourth message includes at least one of the following: a spatial range of each of the N point cloud data, a normal of each of the N point cloud data, or the first parameter. The spatial range of each of the N point cloud data is used to indicate a size and / or shape of a space where each of the N point cloud data is located. The normal of each of the N point cloud data is used to determine a rotation matrix corresponding to each of the N point cloud data.

[0038] Optionally, the first neighborhood i is a region where the i-th point cloud data is located after being projected onto the first image, including: the first neighborhood i is a region where the i-th point cloud data is located in the first space i projected onto the first image; the second neighborhood i is a region where the i-th point cloud data is located after being projected onto the second image, including: the second neighborhood i is a region where the first space i is projected onto the second image; the fourth message is further used to request a scaling parameter, the scaling parameter being a scaling ratio between the objects for two-dimensional perception and imaging results of the objects; the second neighborhood i and the second scaling parameter are received from the third device, the second scaling parameter being used to indicate a scaling ratio between the target objects and imaging results of the target objects in the second image; the method of the first aspect further includes: determining a third neighborhood i according to the first neighborhood i, the first scaling parameter and the second scaling parameter, the first scaling parameter being used to indicate a scaling ratio between the target objects and imaging results of the target objects in the first image, the third neighborhood i including the same number of pixel points as the second neighborhood i; performing multi-modal fusion perception according to the first neighborhood i and the second neighborhood i, including: performing multi-modal fusion perception according to the third neighborhood i and the second neighborhood i.

[0039] Further, the multi-modal fusion perception is performed according to the third neighborhood i and the second neighborhood i, including: determining a neighborhood matching degree of the i-th point cloud data according to a similarity degree of the third neighborhood i and the second neighborhood i; determining noise in the N point cloud data according to the neighborhood matching degree of the i-th point cloud data; and performing multi-modal fusion perception based on the point cloud data in the N point cloud data except the noise.

[0040] Further, the first space i is spherical or cubic.

[0041] Further, before sending the fourth message to the third device, the method of the first aspect further includes: receiving a fourth capability parameter from the third device, the fourth capability parameter being used to indicate that the third device supports providing the scaling parameter; and sending the fourth message to the third device, including: sending the fourth message to the third device according to the fourth capability parameter. In this way, the first device can avoid failure in requesting the scaling parameter from the third device, thereby causing additional communication overhead.

[0042] Optionally, the fourth message is further used to indicate an angle range, a direction of the third device belonging to the angle range, and a two-dimensional image obtained by the device belonging to the angle range for perceiving the target object can be used to determine the two-dimensional image for projecting the N point cloud data. It can be understood that the direction of the third device belonging to the angle range can ensure that the third device and the device associated with the second image exist in an effective spatial range. In this way, the first device can obtain the relevant parameters and the image of the object for two-dimensional perception by the third device from the third device, thereby avoiding failure of the first device in obtaining the relevant parameters and the image of the object for two-dimensional perception by the second device, and causing additional communication overhead.

[0043] In a possible design, the second neighborhood i is acquired according to the N pieces of point cloud data, including: acquiring a fourth parameter and a fourth image, the fourth parameter being used for converting the fourth image into a second image; and determining the second neighborhood i according to the N pieces of point cloud data, the fourth parameter, and the fourth image. In this way, the first device can convert the fourth image into the second image using the fourth parameter, and determine the second neighborhood i based on the second image.

[0044] Optionally, the acquiring of the fourth parameter and the fourth image includes: receiving the fourth parameter and the fourth image from a third device, the fourth parameter being a parameter related to two-dimensional sensing of the target object by the third device. That is, the first device can acquire the fourth parameter and the fourth image from the third device.

[0045] Optionally, the fourth parameter includes at least one of the following: focal length, resolution, field of view, position, or pointing direction. For example, the second parameter includes the focal length, the resolution, the field of view, the position, and the pointing direction of the third device.

[0046] Further, the method of the first aspect further includes: sending a sixth message to the third device, the sixth message being used for requesting a parameter related to two-dimensional sensing of the target object and an image by the third device. In this way, the first device can acquire the fourth parameter and the fourth image in time. Of course, the third device can also actively report the parameter related to two-dimensional sensing of the target object and the image to the first device, for example, periodically reporting the parameter and the image to the first device, which can be set according to actual conditions and is not limited.

[0047] Further, the method of the first aspect further includes: receiving a sixth capability parameter from the third device, the sixth capability parameter being used for indicating that the third device supports providing a parameter related to two-dimensional sensing of the target object and an image; and sending the sixth message to the third device, including: sending the sixth message to the second device according to the sixth capability parameter. In this way, the first device can avoid failure in requesting the parameter related to two-dimensional sensing of the target object and the image from the third device, thereby causing additional communication overhead.

[0048] In a second aspect, a communication method is provided. The method can be performed by a second device, or by a component of the second device, such as a processor, a chip, a chip system, or a circuit of the second device, or by a logic module or software that can implement all or part of the function of the second device. The method is described below by way of example with the method being performed by the second device. The method comprises: receiving a first message from a first device, and sending, to the first device, a first neighborhood i according to the first message; wherein the first message is used to request a region in which each point cloud data of N point cloud data is located after being projected to a target image, the N point cloud data is three-dimensional point cloud data obtained by sensing a target object, the target image is an image obtained by two-dimensional sensing of the target object after being converted according to a first parameter, the first parameter is an internal and external parameter of the target image, N is an integer greater than 1, the first neighborhood i is a region in which an i-th point cloud data of the N point cloud data is located after being projected to a first image, the first image comprises a sensing result of the target object, and i is an integer traversing from 1 to N.

[0049] In a possible design, the first message comprises at least one of: a spatial range of each of the N point cloud data, a normal of each of the N point cloud data, or the first parameter, the spatial range of each of the N point cloud data is used to indicate a size and / or shape of a space in which each of the N point cloud data is located, and the normal of each of the N point cloud data is used to determine a rotation matrix corresponding to each of the N point cloud data.

[0050] In a possible design, the first neighborhood i is a region in which the i-th point cloud data is located after being projected to the first image, and the first message is further used to request a scaling parameter, the scaling parameter being a scaling ratio between an object subjected to two-dimensional sensing and an imaging result of the object; and the sending, to the first device, of the first neighborhood i comprises: sending, to the first device, the first neighborhood i and the first scaling parameter, the first scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the first image.

[0051] Optionally, the first space i is spherical or cubic.

[0052] In a possible design, before receiving the first message from the first device, the method of the second aspect further comprises: sending, to the first device, a first capability parameter, the first capability parameter being used to indicate that a scaling parameter is supported.

[0053] In a possible design, the first message is further used to indicate an angle range, and a two-dimensional image obtained by sensing of the target object by a device belonging to the angle range can be used to determine a two-dimensional image used for projection of the N point cloud data.

[0054] In addition, the technical effects of the method of the second aspect can also refer to the technical effects of the method of the first aspect, which will not be repeated here.

[0055] In a third aspect, a communication method is provided. The method can be performed by a second device, or by a component of the second device, such as a processor, a chip, a chip system, or a circuit of the second device, or by a logic module or software that can implement all or part of the functions of the second device. Hereinafter, the method is described by taking the second device as an example. The method comprises: receiving a second message from a first device, and sending a first scaling parameter to the first device according to the second message; wherein the second message is used to request a scaling parameter, the scaling parameter is a scaling ratio between an object and an imaging result of the object in two-dimensional perception, and the first scaling parameter is used to indicate a scaling ratio between a target object and an imaging result of the target object in a first image, and the first image comprises a perception result of the target object.

[0056] In a possible design, the second message comprises at least one of: a spatial range of each of the N point cloud data, or a normal line of each of the N point cloud data, the spatial range of each of the N point cloud data is used to indicate a size and / or shape of a space where each of the N point cloud data is located, and the normal line of each of the N point cloud data is used to determine a rotation matrix corresponding to each of the N point cloud data.

[0057] In a possible design, before receiving the second message from the first device, the method of the third aspect further comprises: sending a second capability parameter to the first device, the second capability parameter is used to indicate that the scaling parameter is supported.

[0058] In addition, the technical effects of the method of the third aspect can also refer to the technical effects of the method of the first aspect, which will not be repeated here.

[0059] In a fourth aspect, a communication method is provided. The method can be performed by a second device, or by a component of the second device, such as a processor, a chip, a chip system, or a circuit of the second device, or by a logic module or software that can implement all or part of the functions of the second device. Hereinafter, the method is described by taking the second device as an example. The method comprises: obtaining a second parameter and a third image, and sending the second parameter and the third image to a first device; wherein the second parameter is used to indicate a related parameter of two-dimensional perception of a target object, and the third image is an image obtained by two-dimensional perception of the target object.

[0060] In a possible design, the second parameter comprises at least one of: a focal length, a resolution, a field of view, a position, or a pointing direction.

[0061] In a possible design, the method of the fourth aspect further includes: receiving a third message from the first device, the third message being used to request the related parameters and the image for the two-dimensional perception of the target object; and sending the second parameters and the third image to the first device, including: sending the second parameters and the third image to the first device according to the third message.

[0062] In a possible design, before the first parameters and the third image are acquired, the method of the fourth aspect further includes: sending a third capability parameter to the first device, the third capability parameter being used to indicate that the provision of the related parameters and the image for the two-dimensional perception of the object is supported.

[0063] In addition, the technical effects of the method of the fourth aspect can also refer to the technical effects of the method of the first aspect, which will not be repeated here.

[0064] In a fifth aspect, a communication method is provided. The method can be executed by a third device, or by a component of the third device, such as a processor, a chip, a chip system, or a circuit of the third device, or by a logic module or software that can implement all or part of the functions of the third device. Hereinafter, the method is taken as an example for description. The method includes: receiving a fourth message from a first device, and sending a second neighborhood i to the first device according to the fourth message; wherein the fourth message is used to request a region where each point cloud data in N point cloud data is located after being projected to a target image, the N point cloud data is three-dimensional point cloud data obtained by perceiving a target object, the target image is an image obtained by two-dimensional perception of the target object after being converted according to a first parameter, the first parameter is an internal and external parameter of the target image, N is an integer greater than 1, the second neighborhood i is a region where an i-th point cloud data in the N point cloud data is located after being projected to a second image, the second image includes a perception result of the target object, and i is an integer traversing from 1 to N.

[0065] In a possible design, the fourth message includes at least one of the following: a spatial range of each of the N point cloud data, a normal line of each of the N point cloud data, or the first parameter, the spatial range of each of the N point cloud data being used to indicate a size and / or shape of a space where each of the N point cloud data is located, and the normal line of each of the N point cloud data being used to determine a rotation matrix corresponding to each of the N point cloud data.

[0066] In a possible design, the second neighborhood i is a region where the i th point cloud data is located after being projected onto the second image, and the first neighborhood i is a region where the i th point cloud data is located in the first space i projected onto the second image, and the first message is further used to request a scaling parameter, and the scaling parameter is a scaling ratio between an object performing two-dimensional sensing and imaging results of objects, and the method further includes: sending, to the first device, the second neighborhood i, including: sending, to the first device, the second neighborhood i and a second scaling parameter, and the second scaling parameter is used to indicate a scaling ratio between the target object and imaging results of the target object in the second image.

[0067] Optionally, the first space i is a sphere or a cube.

[0068] In a possible design, before receiving the fourth message from the first device, the method of the fifth aspect further includes: sending, to the first device, a fourth capability parameter, and the fourth capability parameter is used to indicate that the scaling parameter is supported.

[0069] In a possible design, the fourth message is further used to indicate an angle range, and a two-dimensional image obtained by sensing the target object by a device belonging to the angle range can be used to determine the two-dimensional image for projection of the N point cloud data.

[0070] In addition, the technical effects of the method of the fifth aspect can also refer to the technical effects of the method of the first aspect, which will not be described herein again.

[0071] In a sixth aspect, a communication method is provided, which can be executed by a third device, or by components of the third device, such as a processor, a chip, a chip system, or a circuit of the third device, or by a logic module or software capable of realizing all or part of the functions of the third device. The following takes the method executed by the third device as an example. The method includes: receiving a fifth message from a first device, and sending, to the first device, a second scaling parameter according to the fifth message, wherein the fifth message is used to request a scaling parameter, the scaling parameter is a scaling ratio between an object performing two-dimensional sensing and imaging results of objects, and the second scaling parameter is used to indicate a scaling ratio between the target object and imaging results of the target object in a second image, and the second image includes a sensing result of the target object.

[0072] In a possible design, the fifth message includes at least one of the following: a spatial range of each of the N point cloud data, or a normal line of each of the N point cloud data, the spatial range of each of the N point cloud data is used to indicate a size and / or shape of a space where each of the N point cloud data is located, and the normal line of each of the N point cloud data is used to determine a rotation matrix corresponding to each of the N point cloud data.

[0073] In a possible design, before receiving the fifth message from the first device, the method of the fifth aspect further includes: sending, to the first device, a fifth capability parameter, where the fifth capability parameter is used to indicate that the support for providing the scaling parameter is supported.

[0074] In addition, the technical effect of the method of the sixth aspect can also refer to the technical effect of the method of the first aspect, which will not be repeated here.

[0075] In a seventh aspect, a communication method is provided. The method can be executed by a third device, or by a component of the third device, such as a processor, a chip, a chip system, or a circuit of the third device, or by a logic module or software that can implement all or part of the function of the third device. Hereinafter, the method is taken as an example for description. The method includes: obtaining a fourth parameter and a fourth image, and sending the fourth parameter and the fourth image to a first device, where the fourth parameter is used to indicate a related parameter for two-dimensional perception of a target object, and the fourth image is an image obtained by two-dimensional perception of the target object.

[0076] In a possible design, the fourth parameter includes at least one of the following: focal length, resolution, field of view, position, or direction.

[0077] In a possible design, the method of the seventh aspect further includes: receiving a sixth message from the first device, where the sixth message is used to request the related parameter for two-dimensional perception of the target object and the image; and sending, to the first device, the fourth parameter and the fourth image includes: sending, to the first device, the fourth parameter and the fourth image according to the sixth message.

[0078] In a possible design, before obtaining the fourth parameter and the fourth image, the method of the seventh aspect further includes: sending, to the first device, a sixth capability parameter, where the sixth capability parameter is used to indicate that the support for providing the related parameter for two-dimensional perception of the target object and the image is supported.

[0079] In addition, the technical effect of the method of the seventh aspect can also refer to the technical effect of the method of the first aspect, which will not be repeated here.

[0080] In an eighth aspect, a communication method is provided. The method includes: a first device executing the method of the first aspect, and a second device executing the method of the second aspect; or the first device executing the method of the first aspect, and the second device executing the method of the third aspect; or the first device executing the method of the first aspect, and the second device executing the method of the fourth aspect.

[0081] Optionally, the method of the eighth aspect further includes: the third device executing the method of the fifth aspect; or the third device executing the method of the sixth aspect; or the third device executing the method of the seventh aspect.

[0082] In a ninth aspect, a communication apparatus is provided. The communication apparatus includes a module or unit (e.g., a chip, or a chip system, or a circuit) for performing the method / operation / step / action of any one of the first aspect to the sixth aspect, e.g., a transceiver module and a processing module. For example, the transceiver module is configured to perform the transceiving function of the communication apparatus, and the processing module is configured to perform the function of the communication apparatus other than the transceiving function.

[0083] Optionally, the transceiver module can include a sending module and a receiving module. The sending module is configured to perform the sending function of the communication apparatus of the ninth aspect, and the receiving module is configured to perform the receiving function of the communication apparatus of the ninth aspect.

[0084] Optionally, the communication apparatus of the ninth aspect can further include a storage module. The storage module stores a program or instruction. When the processing module executes the program or instruction, the communication apparatus can perform the method of any one of the first aspect to the sixth aspect.

[0085] It can be understood that the communication apparatus of the ninth aspect can be a terminal device or a network device, or a chip (system) or other components or assemblies that can be arranged in the terminal device or the network device, or an apparatus including the terminal device or the network device, and the present application does not limit the same.

[0086] In addition, the technical effects of the communication apparatus of the ninth aspect can refer to the technical effects of the method of any one of the first aspect to the sixth aspect, which will not be repeated here.

[0087] In a tenth aspect, a communication apparatus is provided. The communication apparatus includes a processor. When the processor executes a computer instruction, the communication apparatus performs the method of any one of the first aspect to the sixth aspect.

[0088] In a possible design, the communication apparatus of the tenth aspect can further include a transceiver. The transceiver can be a transceiver circuit or an interface circuit. The transceiver can be configured to enable the communication apparatus of the tenth aspect to communicate with other communication apparatuses.

[0089] In a possible design, the communication apparatus of the tenth aspect can further include a memory. The memory can be integrated with the processor, or can be separately arranged. The memory can be configured to store a computer program and / or data related to the method of any one of the first aspect to the sixth aspect.

[0090] In embodiments of the present application, the communication apparatus of the tenth aspect can be the terminal device or the network device of any one of the first aspect to the sixth aspect, or a chip (system) or other components or assemblies that can be arranged in the terminal device or the network device, or an apparatus including the terminal device or the network device.

[0091] In addition, the technical effects of the communication apparatus of the tenth aspect can refer to the technical effects of the method of any one of the first aspect to the sixth aspect, which will not be described here.

[0092] The eleventh aspect provides a communication apparatus. The communication apparatus includes a processor coupled with a memory, and the processor is configured to execute computer programs or instructions stored in the memory, so that the communication apparatus performs the method of any one of the possible implementation manners of the first aspect to the sixth aspect.

[0093] In a possible design, the communication apparatus of the eleventh aspect can further include a transceiver. The transceiver can be a transceiver circuit or an interface circuit. The transceiver can be used for the communication apparatus of the eleventh aspect to communicate with other communication apparatuses.

[0094] In embodiments of the present application, the communication apparatus of the eleventh aspect can be the terminal device or the network device of any one of the first aspect to the sixth aspect, or a chip (system) or other components or assemblies that can be arranged in the terminal device or the network device, or an apparatus including the terminal device or the network device.

[0095] In addition, the technical effects of the communication apparatus of the eleventh aspect can refer to the technical effects of the method of any one of the first aspect to the sixth aspect, which will not be described here.

[0096] The twelfth aspect provides a communication apparatus, including a processor and a memory. The memory is configured to store computer programs or instructions, and when the processor executes the computer programs or instructions, the communication apparatus performs the method of any one of the implementation manners of the first aspect to the sixth aspect.

[0097] In a possible design, the communication apparatus of the twelfth aspect can further include a transceiver. The transceiver can be a transceiver circuit or an interface circuit. The transceiver can be used for the communication apparatus of the twelfth aspect to communicate with other communication apparatuses.

[0098] In embodiments of the present application, the communication apparatus of the twelfth aspect can be the terminal device or the network device of any one of the first aspect to the sixth aspect, or a chip (system) or other components or assemblies that can be arranged in the terminal device or the network device, or an apparatus including the terminal device or the network device.

[0099] In addition, the technical effects of the communication device of the twelfth aspect can refer to the technical effects of the method of any one of the first aspect to the sixth aspect, which will not be described here again.

[0100] The thirteenth aspect provides a communication device for implementing the method of any one of the possible implementation manners of the first aspect to the sixth aspect.

[0101] The fourteenth aspect provides a communication chip, comprising: a logic circuit for executing computer instructions, and a communication interface for communication between the communication chip and other devices or chips, so that the method of any one of the first aspect to the sixth aspect is implemented when the logic circuit executes the computer instructions.

[0102] The fifteenth aspect provides a communication system, comprising at least one of: a first device for executing the method of the first aspect, or a second device for executing the method of the second aspect; or, a first device for executing the method of the first aspect, or a second device for executing the method of the third aspect; or, the communication system comprises at least one of: a first device for executing the method of the first aspect, or a second device for executing the method of the fourth aspect.

[0103] Optionally, the communication system of the fifteenth aspect further comprises: a third device for executing the method of the fifth aspect; or, a third device for executing the method of the sixth aspect; or, a third device for executing the method of the seventh aspect.

[0104] The sixteenth aspect provides a computer readable storage medium, comprising: a computer program or instructions; when the computer program or instructions run on a computer, the computer executes the method of any one of the possible implementation manners of the first aspect to the sixth aspect.

[0105] The seventeenth aspect provides a computer program product, comprising a computer program or instructions, when the computer program or instructions run on a computer, the computer executes the method of any one of the possible implementation manners of the first aspect to the sixth aspect. BRIEF DESCRIPTION OF DRAWINGS

[0106] FIG. 1 is a schematic diagram of a radio frequency sensing result provided by an embodiment of the present application;

[0107] FIG. 2 is a schematic diagram of an optical sensing coordinate system provided by an embodiment of the present application;

[0108] FIG. 3 is a schematic diagram of a camera coordinate system to an image coordinate system provided by an embodiment of the present application;

[0109] FIG. 4 is a schematic diagram of the relationship between the image coordinate system and the pixel coordinate system according to an embodiment of the present application;

[0110] FIG. 5 is a schematic diagram of the relationship between the field of view and the focal length according to an embodiment of the present application;

[0111] FIG. 6 is a schematic diagram of the imaging planes of the actual camera and the virtual camera according to an embodiment of the present application;

[0112] FIG. 7 is a schematic diagram of the view frustum according to an embodiment of the present application;

[0113] FIG. 8 is a schematic diagram of the framework of the multi-modal fusion perception scheme according to an embodiment of the present application;

[0114] FIG. 9 is a schematic diagram of the neighborhood matching according to an embodiment of the present application;

[0115] FIG. 10 is a schematic diagram of the image morphing according to an embodiment of the present application;

[0116] FIG. 11 is a schematic diagram of the flow of the communication method according to an embodiment of the present application;

[0117] FIG. 12 is a schematic diagram of the first neighborhood i and the second neighborhood i according to an embodiment of the present application;

[0118] FIG. 13 is a schematic diagram of the simulation configuration and the results according to an embodiment of the present application;

[0119] FIG. 14 is a schematic diagram of the flow of the communication method according to an embodiment of the present application;

[0120] FIG. 15 is a schematic diagram of the flow of the communication method according to an embodiment of the present application;

[0121] FIG. 16 is a schematic diagram of the flow of the communication method according to an embodiment of the present application;

[0122] FIG. 17 is a schematic diagram of the structure of the communication apparatus according to an embodiment of the present application;

[0123] FIG. 18 is a schematic diagram of the structure of the communication apparatus according to an embodiment of the present application. DETAILED DESCRIPTION

[0124] For the convenience of understanding, the technical terms involved in the embodiments of the present application are introduced as follows.

[0125] 1. Perception

[0126] Perception includes a variety of different methods, which can be divided into two categories of 2D or 3D perception. Among them, 2D perception can capture the scene to a two-dimensional plane to generate a plane image. And the plane image generated by 2D perception includes width and height information, no depth information. Common 2D perception schemes include optical perception, radar perception, etc. 3D perception can capture the shape and structure of the scene in three-dimensional space, that is, generate a stereo image. And the stereo image generated by 3D perception includes width, height and depth information. Common 3D perception schemes include radio frequency perception and computer tomography, etc. The following will introduce the above radio frequency perception and the above optical perception respectively.

[0127] 1.1 Radio frequency perception

[0128] Radio frequency perception can achieve perception of the target by receiving the echo of the perceived target. Among them, the perceived target (or called perception target) can reflect signals, diffract signals or scatter signals, etc. As shown in FIG. 1, radio frequency perception can obtain a plurality of 3D scattering points. It can be understood that in the embodiments of the present application, the 3D scattering point can also be referred to as point cloud data, 3D point cloud data, or other possible names, which are not limited in the embodiments of the present application.

[0129] 1.2 Optical perception

[0130] Optical perception uses image sensors to detect light waves and generate images. The parameters of the optical perception node are divided into intrinsic parameters and extrinsic parameters. Among them, the intrinsic parameters are related to the characteristics of the camera itself, such as: focal length, field of view, resolution, etc.; the extrinsic parameters of the camera refer to the position and rotation direction of the camera. The imaging process of optical perception involves four coordinate systems (world, camera, image and pixel coordinate systems) and the conversion of the four coordinate systems. The following will explain the four coordinate systems in combination with FIG. 2.

[0131] As shown in FIG. 2, the optical perception coordinate system involves the world coordinate system, the camera coordinate system, the image coordinate system and the pixel coordinate system.

[0132] The world coordinate system is the three-dimensional coordinate system of the objective world. The optical perception node is placed in the three-dimensional space, and the world coordinate system is used as the reference to describe the position of the optical perception node as P (X W ,Y W ,Z W ).

[0133] The camera coordinate system is a coordinate system established with the optical perception node (such as the optical center of the camera) as the coordinate origin and the optical axis of the camera as the Z axis. In the camera coordinate system, the position of the optical perception node can be represented as (X C ,Y C ,Z C). It can be understood that the conversion from the world coordinate system to the camera coordinate system can be achieved by rotation and translation operations, and vice versa.

[0134] The image coordinate system takes the center of the camera image sensor as the coordinate origin, and the X and Y axes are parallel to the two vertical edges of the image sensor, with (x, y) representing the coordinate values thereof. The image coordinate system generally uses physical units (such as millimeters (mm)) to represent the positions of pixels in the image.

[0135] The pixel coordinate system takes the top-left corner of the image sensor as the origin, and the X and Y axes are parallel to the X and Y axes of the image coordinate system, with (u, v) representing the coordinate values thereof. Each image contains M rows and N columns of elements, each element is called a pixel, and the pixel coordinate system is in units of pixels. Among them, the row corresponds to the horizontal direction of the image sensor, and the column corresponds to the vertical direction of the image sensor.

[0136] The conversion process of the above four coordinate systems is as follows:

[0137] (1) World coordinate system→camera coordinate system: 3D to 3D projection, which can transform the object coordinates from the world coordinate system to the camera coordinate system through rotation and translation. The camera position refers to the position of the camera in the world coordinate system O c , and the camera pointing refers to the direction of the three axes (generally x, y, z axes) of the camera coordinate system in the world coordinate system [X axis , Y axis , Z axis ]. When any two coordinate axes in the camera pointing direction are determined, the direction of the other axis can be obtained by calculation, such as Y axis = Z axis × X axis . The rotation matrix from the world coordinate system to the camera coordinate system is a 3×3 orthogonal matrix, represented as ||·|| is the norm of the vector, and the translation vector is a 3×1 vector, represented as t = -RO c . It can be seen that the rotation and translation of the camera are determined by the camera extrinsic parameters (such as camera position and pointing), and the extrinsic parameter matrix of the camera is defined as

[0138] (2) Camera coordinate system→image coordinate system: 3D to 2D projection, which can be calculated according to the principle of similar triangles. As shown in FIG. 3, where f represents the focal length of the camera, which is represented in matrix form as follows: It can be seen that the scaling factor of the length of the line segment on the object and the length of the line segment in the imaging result is , that is, the ratio of the value of Z c to the focal length.

[0139] (3) Image Coordinate System → Pixel Coordinate System: Converting physical units to pixel units. Both the pixel coordinate system and the image coordinate system lie on the imaging plane of the optical sensing node, but their origins and units of measurement differ. As shown in Figure 4, the image coordinate system (x, y) generally uses the center of the image sensor as its origin, and the unit is a physical unit, such as mm. The pixel coordinate system (u, v) generally uses the top-left corner of the image sensor as its origin, and the unit is pixels. The number of pixels determines the camera's resolution. For example, in Figure 4, the camera's pixel count is M*N, meaning that the number of pixels per row is M, and the number of pixels per column is N, corresponding to a resolution of M*N. The conversion relationship between the pixel coordinate system coordinates [u, v] and the image coordinate system coordinates [x, y] is as follows: Where dx and dy represent the number of millimeters represented by each column and each row of pixels, respectively, and u m and v m The coordinates of the pixel point corresponding to the center of the image sensor are equal to half of the horizontal and vertical resolutions, respectively. Represented in matrix form as follows:

[0140] The above content introduces the transformation process of four coordinate systems involved in the optical sensing imaging process. It can be understood that, combining the transformation of the second coordinate system (i.e., (2) camera coordinate system → image coordinate system) and the transformation of the third coordinate system (i.e., (3) image coordinate system → pixel coordinate system), the camera's intrinsic parameter matrix can be defined as: in, This indicates the number of pixels corresponding to the focal length in the horizontal direction. This represents the number of pixels corresponding to the focal length in the vertical direction. The horizontal field of view is known to be θ. u The relationship between focal length and field of view is shown in Figure 5, that is... Similarly, θ v This represents the field of view in the vertical direction. Therefore, when the focal length and resolution are known, or when the field of view and resolution are known, the camera's intrinsic parameter matrix can be determined.

[0141] It is understood that in the embodiments of this application, "camera" and "optical sensing node" both refer to objects that are capable of (or support) optical sensing, that is, in some cases, "camera" and "optical sensing node" can be used interchangeably.

[0142] The above content introduced the imaging process of optical sensing, which involves four coordinate systems and the transformation between these four coordinate systems. The following describes the transformation process of the optical sensing results.

[0143] When the positions of the multiple cameras are different, the pointing directions and the intrinsic parameters (focal length, field of view angle and resolution) are the same, the imaging planes of the cameras are parallel. That is, at this time, the imaging results of the multiple cameras for objects of the same depth only have translation and scaling, that is, the size and position of the object imaging change. The following takes two cameras (camera 1 and camera 2) as an example for illustration.

[0144] For a point in the world coordinate system The pixel points imaged on the camera 1 and the camera 2 are p1 = K1(R1P w +t1) and p2 = K2(R2P w +t2) respectively, where K1 is the intrinsic parameter matrix of the camera 1, R1 is the rotation matrix of the camera 1, t1 is the translation vector of the camera 1, K2 is the intrinsic parameter matrix of the camera 2, R2 is the rotation matrix of the camera 2, and t2 is the translation vector of the camera 2. The pointing directions of the camera 1 and the camera 2 are the same, so R1 = R2; the intrinsic parameters of the camera 1 and the camera 2 are the same, so K1 = K2. In this case, the relationship of the imaging results is: p1-p2 = K1(R1P w +t1)-K2(R2P w +t2) = K1(t1-t2), that is, p1 and p2 are translated by K1(t1-t2). And the scaling relationship is shown in the above FIG. 4, that is, the scaling relationship satisfies: It can be seen that the scaling factor will affect the scaling of the imaged object, the larger the scaling factor is, the smaller the object imaging area is.

[0145] In practice, it is difficult to have the pointing directions and the intrinsic parameters of two cameras completely the same. In this case, the two original cameras can be converted (such as rotation and re-projection) to generate two virtual cameras, so that the two virtual cameras have the same pointing direction and intrinsic parameter, that is, the imaging planes of the two virtual cameras are parallel. It can be understood that the positions of the original camera and the virtual camera corresponding to the original camera are the same, and the pointing directions and the intrinsic parameters can be the same or different. For example, as shown in FIG. 6, the imaging plane 1A and the imaging plane 2A are the imaging planes of two different original cameras, the imaging plane 1B and the imaging plane 2B are the imaging planes of the virtual cameras corresponding to the two different original cameras, the imaging plane 1A and the imaging plane 2A are not parallel, and the imaging plane 1B and the imaging plane 2B are parallel.

[0146] The conversion relationship from the real camera image pixel to the virtual camera image pixel is: Where p real is the pixel point of a point in a world coordinate system on the image of the original camera, p vir represents the pixel point of the same point in the world coordinate system on the image of the corresponding virtual camera, E real and Evirtual Let K represent the extrinsic parameter matrices of the original camera and the virtual camera, respectively. real and K virtual These represent the intrinsic parameter matrices of the original camera and the virtual camera, respectively.

[0147] After the original camera is transformed into a virtual camera, the effective perceptual space is the intersection of the effective perceptual spaces of the original camera and the virtual camera. This can be understood as objects located in the effective perceptual space being projected onto the imaging planes of both the real and virtual cameras; or, in other words, objects located in the effective perceptual space, after being projected onto the imaging plane of the original camera, can have their imaging results transformed onto the imaging plane of the virtual camera. The effective perceptual space of a camera can be represented by its field of view (fov). As shown in Figure 7, the field of view angle range is defined by the field of view angle fov, and the depth range is determined by the near and far planes. The near plane refers to the visible plane closest to the camera's optical Oc, and the far plane refers to the visible plane farthest from the camera's optical Oc.

[0148] 2. Multimodal fusion sensing

[0149] Multimodal fusion sensing can fuse radio frequency (RF) and optical multimodal information. For example, multimodal fusion sensing can determine the surface where a target is located based on the 3D scattering points obtained from RF sensing; and then project the optical imaging results back onto the surface determined by RF sensing, using the intersection points to determine the target shape. As shown in Figure 8, the implementation idea of ​​multimodal fusion sensing is described in detail below:

[0150] (1) Regarding optical sensing:

[0151] First, obtain the extrinsic and intrinsic parameter matrices of the camera; then, based on the intrinsic parameter matrix, project the 2D image back onto the camera coordinate system to obtain the set R of ray equations for key points (such as boundary pixels) in the image. E ={γ0*(X0,Y0,1),γ1*(X1,Y1,1),…,γ K *(X K Z K ,1)}. The specific back projection process is represented as follows:

[0152] It is understandable that, due to the unknown depth, the ray equation of the key points in the image (such as boundary pixels) is γ0*(X0,Y0,1), where (u0,v0) represents the coordinates in the pixel coordinate system, K is the camera intrinsic parameter matrix, and γ0 represents the Z-axis coordinate in the camera coordinate system, i.e., the depth information, which is the parameter to be estimated.

[0153] (2) Regarding radio frequency sensing:

[0154] Based on 3D scattering point P S{(X0, Y0, Z0), (X1, Y1, Z1), …, (Xn, Yn, Zn)}, determine the surface F(x, y, z) = 0 of the object. N ,Y N ,Z N )}, determine the surface F(x, y, z) = 0 of the object.

[0155] (3) Multi-modal fusion:

[0156] The ray equation R E where the key point is brought into the surface equation F(x, y, z) = 0, and the depth γ is solved to determine the key point coordinates and the object shape.

[0157] It can be understood that the above describes the implementation idea of multi-modal fusion perception. In the process of multi-modal fusion perception, the multiple 3D scattering points obtained by radio frequency perception of the target object can be denoised, and the denoised 3D scattering points can be used for multi-modal fusion perception, thereby improving the multi-modal perception performance.

[0158] For example, each 3D scattering point in the above multiple 3D scattering points can be projected onto N 2D images, a neighborhood is taken near the projection point, the matching degree of the neighborhood is calculated according to the pixel value of the neighborhood, and it is judged whether the 3D scattering point is a noise point according to the matching degree, so as to retain or remove the 3D scattering point. It can be understood that the N 2D images are images obtained by two-dimensional perception (such as optical perception) of the target object, and N is an integer greater than or equal to 2. The following describes this process with an example of a 3D scattering point.

[0159] (1) As shown in FIG. 9, the 3D scattering point (P point in FIG. 9) obtained by radio frequency perception is projected onto two different 2D images obtained by optical perception to obtain projection point 1 (P1 point in FIG. 9) and projection point 2 (P2 point in FIG. 9). The two different 2D images are obtained by cameras C1 and C2, respectively. That is, the above projection from the world coordinate system to the pixel coordinate system is performed on the 3D scattering point. The projection process is shown in the following formula:

[0160] Where [X P ,Y P ,Z P ] is the coordinate of point cloud data #1, [u p ,v p ] is the coordinate of point cloud data #1 projected into the pixel coordinate system, is the extrinsic parameter matrix of the camera, is the intrinsic parameter matrix of the camera coordinate system converted to the pixel coordinate system.

[0161] (2) performing a neighborhood extraction operation on the projection point 1 and the projection point 2 to obtain a neighborhood #1 corresponding to the projection point 1 and a neighborhood #2 corresponding to the projection point 2. For example, the neighborhood #1 is obtained by taking the neighborhood around the projection point 1, and the neighborhood #2 is obtained by taking the neighborhood around the projection point 2. It can be understood that the neighborhood #1 and the neighborhood #2 each include a plurality of pixel points (u, v).

[0162] (3) determining a neighborhood matching degree of the 3D scattering point according to the neighborhood #1 and the neighborhood #2. The neighborhood matching degree can represent the image similarity degree of the neighborhood #1 and the neighborhood #2. In other words, the neighborhood matching degree can be used to measure the closeness of the neighborhood pixel values of the neighborhood #1 and the neighborhood #2, which can be understood as the pixel value corresponding to each pixel point in the neighborhood.

[0163] The neighborhood matching degree can be calculated in various ways. For example, the neighborhood matching degree is determined by normalized sum of squared difference (NSSD), and the expression for calculating the neighborhood matching degree by NSSD is as follows: wherein A u,v is the pixel value corresponding to the pixel point (u, v) in the neighborhood #1, and B u,v is the pixel value corresponding to the pixel point (u, v) in the neighborhood #2. Alternatively, the distribution difference of the pixel values corresponding to the pixel points in the neighborhood #1 and the neighborhood #2 is compared in the form of a cumulative distribution function (CDF), and the expression for calculating the neighborhood matching degree by CDF is as follows: wherein F is a density function, which is expressed as h(x) represents the frequency distribution of the pixel value, and h(x) = ∑ u,v δ(A(u, v) - x), δ is a dirac δ function, and x is a pixel value. The neighborhood matching degree is calculated in the above two ways. The closer the result of the neighborhood matching degree is to 1, the higher the neighborhood matching degree is.

[0164] (4) determining whether the 3D scattering point is a noise point according to the neighborhood matching degree. For example, the neighborhood matching degree is calculated by NSSD or CDF, and the neighborhood matching degree is compared with a preset threshold value; if the neighborhood matching degree of the 3D scattering point is greater than or equal to the preset threshold value, it indicates that the 3D scattering point is not a noise point, i.e., the 3D scattering point is a real 3D scattering point; if the neighborhood matching degree of the 3D scattering point is less than the preset threshold value, it indicates that the 3D scattering point is a noise point. The preset threshold value can be set according to actual conditions, without limitation.

[0165] It can be understood that, as shown in FIG. 9, after the real 3D scattering point is projected onto multiple two-dimensional images (such as the imaging planes in FIG. 9), the areas where the respective projection points of the 3D scattering point are located correspond to the imaging results of similar (or the same) areas of the object, and therefore the similarity of the areas where the respective projection points of the 3D scattering point are located is relatively high; and after the noise 3D scattering point is projected onto multiple two-dimensional images, the areas where the respective projection points of the noise 3D scattering point are located correspond to the imaging results of different areas, and therefore the similarity of the areas where the respective projection points of the noise 3D scattering point are located is relatively low. It can be understood that the similar areas described above can be understood as two areas on the target object that include real point cloud data, and the two areas have a large overlap, or in other words, the two areas have a small difference. The different areas described above can be understood as completely different areas, or can be understood as two areas with a large difference.

[0166] It can also be understood that the “real 3D scattering point” and the “noise 3D scattering point” in the embodiments of the present application are only an exemplary way of expression, and can be replaced by “real point cloud data” or “real scattering point”, and so on, and the “noise 3D scattering point” can be replaced by “noise”, “noise point”, or “noise point cloud data”, and so on, without limitation.

[0167] It is found through research that when the configurations (such as the positions, orientations, internal parameters, and so on) of the optical perception nodes are different, the imaging results have a large deformation. In this case, the elements corresponding to the neighborhood change, which can cause the matching degree of the neighborhood corresponding to the real 3D scattering point to decrease, so that the real 3D scattering point is identified as a noise point. In this way, the number of real 3D scattering points used for multi-modal perception is reduced, and therefore the multi-modal perception performance is reduced.

[0168] For example, as shown in FIG. 10, the upper graph in FIG. 10 describes a simulation scenario, that is, a building is considered, and the 3D coordinate range of a certain wall surface of the building is x = -57, the value range of y is [-103, 103], the value range of z is [0, 20], and the material is brick. The two lower graphs in FIG. 10 are images obtained by optical perception of the building by two optical perception nodes. The neighborhood of point P (-57, -90, 10) on the building in the two 2D images is marked with a black frame. It can be seen that because the images obtained by optical perception of the building by the two optical perception nodes have a large deformation, the matching degree of the two neighborhoods is low, and therefore point P can be misidentified as noise.

[0169] In view of the above technical problems, the embodiments of the present application propose the following technical solutions to improve the multi-modal perception performance.

[0170] The technical solutions in the present application will be described below with reference to the accompanying drawings.

[0171] The technical solutions of the embodiments of the present application can be applied to various communication systems, for example, a 4th generation (4G) mobile communication system such as a long term evolution (LTE) system, a 5th generation (5G) mobile communication system such as a new radio (NR) system, and a communication system evolved after 5G, such as a future communication network system, and can also be applied to a wireless fidelity (WiFi) system, a vehicle to everything (V2X) communication system, a device-to-device (D2D) communication system, a vehicle networking communication system, and the like.

[0172] The present application will present various aspects, embodiments or features around a system that can include a plurality of devices, components, modules, etc. It should be understood and appreciated that each system can include additional devices, components, modules, etc., and / or can not include all of the devices, components, modules, etc. discussed in connection with the figures. Furthermore, combinations of these aspects can also be used.

[0173] In addition, in the embodiments of the present application, the words "example", "for example", and the like are used to mean serving as an example, instance, or illustration. Any embodiment or design presented as an "example" in the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the word "example" is used to present concepts in a concrete manner.

[0174] In the embodiments of the present application, "information", "signal", "message", "channel", and "signaling" can be used interchangeably at times, and it should be pointed out that when the distinction is not emphasized, the meanings expressed are matched. "Of", "corresponding", and "corresponding" can be used interchangeably at times, and it should be pointed out that when the distinction is not emphasized, the meanings expressed are matched. In addition, the " / " mentioned in the present application can be used to represent the "or" relationship.

[0175] The network architecture and service scenarios described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of network architecture and the appearance of new service scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0176] For the convenience of understanding the embodiments of the present application, first, a communication system applicable to the embodiments of the present application is described.

[0177] The communication system comprises a first device and a second device.

[0178] The first device can be used for communication and data processing, for example, the first device can receive three-dimensional point cloud data from other devices and process the three-dimensional point cloud data. The first device can also be used for radio frequency sensing of objects. The first device can be a terminal device, or a communication module, a circuit with communication function, a chip, a chip system or other components or assemblies in the terminal device. The first device can also be a network device, or a communication module, a circuit with communication function, a chip, a chip system or other components or assemblies in the network device.

[0179] The second device can be used for communication, for example, the second device can receive two-dimensional images from other devices. The second device can also be used for data processing, for example, the second device can process two-dimensional images. The second device can also be used for optical sensing of objects. The second device can be a terminal device, or a communication module, a circuit with communication function, a chip, a chip system or other components or assemblies in the terminal device. The second device can also be a network device, or a communication module, a circuit with communication function, a chip, a chip system or other components or assemblies in the network device. It can be understood that the second device can be one or more, which can be flexibly set according to actual conditions, and is not limited. It can be understood that when the first device is a communication module, a circuit with communication function, a chip, a chip system or other components or assemblies in the network device, and the second device is a communication module, a circuit with communication function, a chip, a chip system or other components or assemblies in the network device, the first device and the second device can be deployed in the same network device.

[0180] The terminal device can also be referred to as user equipment (UE), mobile station (MS), mobile terminal (MT), user device, access terminal, subscriber unit, subscriber station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device, etc., or a device used to provide voice or data connectivity to a user, and can also be an Internet of Things device. For example, a terminal device includes a handheld device having wireless connection capability, a car-mounted device, etc. Currently, the terminal can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a notebook computer, a palm computer, a mobile Internet device (MID), a wearable device (e.g., a smart watch, a smart bracelet, a pedometer, etc.), a vehicle-mounted device (e.g., a car, a bicycle, an electric vehicle, an airplane, a ship, a train, a high-speed rail, etc.), a satellite terminal, a virtual reality (VR) device, an augmented reality (AR) device, a smart point of sale (POS) machine, a customer-premises equipment (CPE), a wireless terminal in industrial control, a smart home device (e.g., a refrigerator, a television, an air conditioner, an electricity meter, etc.), a smart robot, a mechanical arm, a workshop device, a wireless terminal in unmanned driving, a wireless terminal in telemedicine, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, or a wireless terminal in a smart home, a flight device (e.g., a smart robot, a hot air balloon, a drone, an airplane), etc. The terminal device can also be other devices with terminal functions, for example, the terminal device can also be a device that plays a terminal function in device to device (D2D) communication.

[0181] The network device can be a base station, an evolved NodeB (eNodeB), a transmitting and receiving point (TRP), a transmitting point (TP), a next generation NodeB (gNB), a next generation NodeB in a future communication system, a base station in a future mobile communication system, a satellite, or an access point (AP) in a WiFi system, such as a home gateway, a router, a server, a switch, a bridge, and the like, an integrated access and backhaul (IAB) node, a network device in a non-terrestrial network (NTN) communication system, i.e., can be deployed on a high-altitude platform or a satellite, and the like. The network device can be a macro base station, a micro base station, or an indoor station, a relay node or a donor node, or a wireless controller in a cloud-radio access network (C-RAN) scenario. The network device can also be a device that plays a base station function in D2D communication, vehicle-to-everything (V2X) communication, unmanned aerial vehicle communication, and machine communication. Alternatively, the network device can also be a server, a wearable device, a vehicle or a vehicle-mounted device, and the like. For example, an access network device in V2X technology can be a road side unit (RSU). In addition, the network device can also be a core network element, such as a policy control function (PCF), an access and mobility management function (AMF), a location management function (LMF), a sensing management function (SMF), a sensing function (SF), and the like.

[0182] It can be understood that the "first device" and the "second device" in the embodiments of the present application are only an exemplary expression, and both can be replaced by any possible expression, for example, the "first device" can be replaced by an "optical sensing node" or a "first node", and the like, and the "second device" can be replaced by a "central processing node", a "radio frequency sensing node", or a "second node", and the like, and the embodiments of the present application do not limit this.

[0183] In the communication system, the first device obtains the first neighborhood i and the second neighborhood i of the i th point cloud data in the N point cloud data in the first image and the second image respectively. Since the first image and the second image are two different two-dimensional images including the sensing result of the target object, and the imaging plane of the first image and the imaging plane of the second image are parallel, the imaging result of the target object on the first image and the imaging result of the target object on the second image have small deformation, so that the imaging result of the target object in the first neighborhood i and the second neighborhood i has small deformation. In this way, when the first device performs multi-modal fusion sensing based on the first neighborhood i and the second neighborhood i, the performance of multi-modal fusion sensing can be improved. For example, the first device determines the neighborhood matching degree of the i th point cloud data based on the first neighborhood i and the second neighborhood i, and determines whether the i th point cloud data is real point cloud data based on the neighborhood matching degree, and then performs multi-modal fusion sensing on the real point cloud data in the N point cloud data. Alternatively, the boundary and / or corner point of the target object is determined based on the first neighborhood i and the second neighborhood i, and then the determined boundary and / or corner point of the target object is used to perform multi-modal fusion sensing.

[0184] In addition, the above-mentioned communication system can further include other network devices and / or other terminal devices, which can be flexibly set according to actual conditions without limitation.

[0185] For the convenience of understanding, the communication method provided by the embodiment of the application will be described in detail below in combination with FIG. 11.

[0186] For example, FIG. 11 is a flowchart of the communication method provided by the embodiment of the application. The method can be applied to the interaction between the first device and the second device in the above-mentioned communication system. Hereinafter, the second device includes two devices, i.e., the second device and the third device, as an example.

[0187] As shown in FIG. 11, the flow of the communication method is as follows:

[0188] S1101, the first device obtains the first neighborhood i and the second neighborhood i according to the N point cloud data.

[0189] S1102, the first device performs multi-modal fusion sensing according to the first neighborhood i and the second neighborhood i.

[0190] The above steps will be described respectively.

[0191] For S1101:

[0192] The N pieces of point cloud data are three-dimensional point cloud data obtained by perceiving the target object, or the N pieces of point cloud data are point cloud data obtained by three-dimensionally perceiving the target object, and N is an integer greater than 1. For example, the N pieces of point cloud data are three-dimensional point cloud data obtained by radio frequency sensing the target object. The N pieces of point cloud data can be used to generate a stereoscopic image of the target object. In addition, the N pieces of point cloud data can be data obtained by the first device three-dimensionally perceiving the target object, that is, the first device can obtain the data from an upper layer, or the first device can obtain the data from other devices, such as a radio frequency sensing node, without limitation.

[0193] It can be understood that the "point cloud data" in the embodiments of the present application is only an exemplary expression, and the "point cloud data" can be replaced by any possible expression, such as "three-dimensional scattering points" or "three-dimensional point cloud data", and the embodiments of the present application do not limit this.

[0194] The first neighborhood i is a region where the i-th point cloud data in the N pieces of point cloud data is located after being projected (or projected, or mapped) to a first image (described below), or the first neighborhood i is a region where a point (denoted as a first projection point i) on which the i-th point cloud data is located after being projected to the first image, that is, the first neighborhood includes the first projection point i, and i is an integer traversing 1 to N. The first neighborhood i can be a rectangular region with a certain pixel point (such as the first projection point i) as the center, and the size of the rectangular region can be determined by the width and height of the first neighborhood i. The first neighborhood i can also be a circular region with a certain pixel point (such as the first projection point i) as the center, and the pixel points in the circular region are a set of all pixel points whose Euclidean distance from the center pixel point is less than or equal to the radius of the circle.

[0195] For example, the expression of the first neighborhood i can be: Nr(u, v) = {(u+x, v+y) | -r≤x≤r, -k≤y≤k}, (u, v) is the coordinate of the point on which the i-th point cloud data is mapped to the first image, that is, the coordinate of the first projection point i, (2r+1) is greater than or equal to 1, and (2k+1) is greater than or equal to 1; at this time, the size of the first neighborhood i is (2r+1) x (2k+1). It can be understood that | is a separator, and (u+x, v+y) before | is used to indicate the pixel coordinates of the first neighborhood i, and the value range of x and y after |. It can be understood that when r is greater than the horizontal resolution of the first image, the first neighborhood i can be filled by zero (or other values); when k is greater than the vertical resolution of the first image, the first neighborhood i can be filled by zero (or other values).

[0196] For another example, the expression of the first neighborhood i can be: Nr(u, v) = {(u+x, v+y) | -r≤x≤r, -k≤y≤k}, (u, v) is the coordinate of the point on which the i-th point cloud data is mapped to the first image, that is, the coordinate of the first projection point i, (2r+1) is greater than or equal to 1, and (2k+1) is greater than or equal to 1; at this time, the size of the first neighborhood i is (2r+1) x (2k+1). It can be understood that | is a separator, and (u+x, v+y) before | is used to indicate the pixel coordinates of the first neighborhood i, and the value range of x and y after |. It can be understood that when r is greater than the horizontal resolution of the first image, the first neighborhood i can be filled by zero (or other values); when k is greater than the vertical resolution of the first image, the first neighborhood i can be filled by zero (or other values). 2 +y2 ≤r 2 Let (u, v) be the coordinates of the i-th point cloud data mapped onto the first image, i.e., the coordinates of the first projected point i, where r is greater than or equal to 1. It can be understood that | is a separator; (u+x, v+y) before | indicates the pixel coordinates of the first neighboring i, and the values ​​after | are the ranges for x and y. It can be understood that when r is greater than the horizontal or vertical resolution of the first image, the first neighboring i can be padded with zeros (or other values).

[0197] In addition, the first neighborhood i can also be a region of other shapes and / or sizes centered on a certain pixel (such as the first projection point i). The specific settings can be flexibly configured according to the actual situation without any restrictions.

[0198] Optionally, the first neighborhood i, which is the region where the i-th point cloud data is projected onto the first image from the N point cloud data, can specifically include: the first neighborhood i is the region projected onto the first image from the first space i where the i-th point cloud data is located. That is, the first device can first determine the first space i where the i-th point cloud is located, such as constructing a three-dimensional region centered on the i-th point cloud data, which is the first space i; then project the first space i onto the first image to obtain the first neighborhood i, that is, the image corresponding to the first space i on the first image is the first neighborhood i. The first space i can be spherical or cubic, such as a cube, cuboid, etc. Of course, the first space i can also be other types of three-dimensional regions, which can be flexibly set according to the actual situation without limitation.

[0199] The second neighborhood i is the region where the i-th point cloud data is projected (or mapped) onto the second image (described below). In other words, the second neighborhood i is the region where the point (denoted as the second projection point i) of the i-th point cloud data is projected onto the second image. The method for determining the second neighborhood i is similar to that for determining the first neighborhood i, and can be understood by referring to the relevant introduction of the first neighborhood i above. It will not be repeated here.

[0200] Optionally, the second neighborhood i, the region where the i-th point cloud data is projected onto the second image, can specifically include: the region where the first space i containing the i-th point cloud data is projected onto the second image. The specific implementation principle of projecting the first space i onto the second image to obtain the second neighborhood i is similar to the specific implementation principle of projecting the first space i onto the first image to obtain the first neighborhood i described above, and can be understood by referring to the above content, and will not be repeated here.

[0201] The first image is a two-dimensional image including the sensing result of the target object. Illustratively, the first image can be an image obtained by two-dimensional sensing of the target object, or can be an image obtained by conversion of a two-dimensional image (denoted as image #1) obtained by two-dimensional sensing of the target object. The second image is a two-dimensional image including the sensing result of the target object. Illustratively, the second image can be an image obtained by two-dimensional sensing of the target object, or can be an image obtained by conversion of a two-dimensional image (denoted as image #2) obtained by two-dimensional sensing of the target object. When the first image is an image obtained by conversion of image #1, and the second image is an image obtained by conversion of image #2, image #1 and image #2 are different. It can be understood that the first image and the second image both include the sensing result of the target object, and the first image and the second image are different two-dimensional images. In addition, an imaging plane (denoted as imaging plane #1) of the first image and an imaging plane (denoted as imaging plane #2) of the second image are parallel.

[0202] The parallelism of imaging plane #1 and imaging plane #2 can be understood as that the included angle between imaging plane #1 and imaging plane #2 belongs to an included angle range. The included angle range can be pre-set or pre-defined by a protocol. For example, the pre-set included angle range is 0-10 degrees (°), or the pre-set included angle range is 0-15°. The specific setting is not limited and can be set according to actual conditions. For example, the included angle between imaging plane #1 and imaging plane #2 is 0°; or the included angle between imaging plane #1 and imaging plane #2 is 5°; or the included angle between imaging plane #1 and imaging plane #2 is 10°.

[0203] The positional relationship between imaging plane #1 and imaging plane #2 is related to the device associated with the first image and the device associated with the second image. Illustratively, the parallelism of imaging plane #1 and imaging plane #2 can specifically include that the second virtual device is different from the third virtual device in position, the second virtual device and the third virtual device have the same third parameter, the second virtual device is associated with the second image, the third virtual device is associated with the third image, and the third parameter is used to indicate a related parameter of two-dimensional sensing. In other words, when the device associated with the first image and the device associated with the second image are different in position, and the third parameter is the same, imaging plane #1 and imaging plane #2 are parallel.

[0204] The third parameter includes at least one of an angle or an internal parameter, and the internal parameter includes at least one of a focal length, a field of view angle or a resolution. For example, the third parameter is an angle (i.e., an angle of perceiving or photographing the target object) and an internal parameter, and the internal parameter is a focal length, a field of view angle and a resolution. It can be understood that the third parameter of the device associated with the first image is the same as the third parameter of the device associated with the second image, and a difference between the third parameter of the device associated with the first image and the third parameter of the device associated with the second image is within a preset range, which can be set according to actual conditions and is not limited. For example, the preset range includes an angle difference of 0-1°; the device associated with the image #a1 and the device associated with the image #a2 are different in position, the internal parameters are the same, and the angles are different by 1°, and the third parameter of the device associated with the image #a1 and the device associated with the image #a2 is the same, that is, the imaging plane of the image #a1 is parallel to the imaging plane of the image #a2.

[0205] It can be understood that when the first image is an image obtained by perceiving the target object in two dimensions, the device associated with the first image is a device obtaining the first image by perceiving the target object in two dimensions; when the first image is an image converted from the image #1, the device associated with the first image can be understood as a virtual device (i.e., a non-real device, i.e., the second virtual device), and the second virtual device is different from the device obtaining the image #1 by perceiving the target object in two dimensions in position, and the third parameter is different or at least partially the same.

[0206] When the first neighborhood i is an area of the first space i in which the i-th point cloud data is located, projected onto the first image, and the first image is an image converted from the image #1, the first space i in which the i-th point cloud data is located belongs to a first effective perception space, the first effective perception space is an intersection of a viewing pyramid of a second virtual device and a viewing pyramid of a second device, the second virtual device is associated with the first image, the first image is determined according to an image (i.e., the image #1) obtained by perceiving the target object in two dimensions by the second device, and the second virtual device is the same as the second device in position.

[0207] In addition, when the second image is an image obtained by perceiving the target object in two dimensions, the device associated with the second image is a device obtaining the second image by perceiving the target object in two dimensions; when the second image is an image converted from the image #2, the device associated with the second image can be understood as a virtual device (referred to as a third virtual device), and the virtual device is different from the device obtaining the image #2 by perceiving the target object in two dimensions in position, and the third parameter is different or at least partially the same.

[0208] When the second neighborhood is a region on the second image to which the first space i in which the i-th point cloud data is located is projected, and the second image is an image converted from image #2, the first space i in which the i-th point cloud data is located belongs to a second effective perception space, the second effective perception space is an intersection of a viewing frustum of a third virtual device and a viewing frustum of the third device, the third virtual device is associated with the second image, the second image is determined according to an image (i.e., image #2) obtained by two-dimensional perception of the third device, and the third virtual device has the same position as the third device.

[0209] It can be understood that the above-mentioned effective perception space can be understood with reference to the related description in the aforementioned “1.2 optical perception”, which will not be repeated here.

[0210] The above-mentioned content will be described below with a specific example.

[0211] For example, as shown in FIG. 12, the i-th point cloud data (S point in FIG. 12) is projected onto the first image to obtain S1 point, and the region (the box in which S1 point is located in FIG. 12) in which the S1 point is located is the first neighborhood i; the i-th point cloud data is projected onto the second image to obtain S2 point, and the region (the box in which S2 point is located in FIG. 12) in which the S2 point is located is the second neighborhood i. The first image and the second image are parallel.

[0212] The above-mentioned content introduces the first neighborhood i and the second neighborhood i. It can be understood that the following takes the first neighborhood i as an example, in which the first neighborhood i is the region in which the i-th point cloud data is projected onto the first image, and the first image is a two-dimensional image converted from image #1, to introduce different ways in which the first device obtains the first neighborhood i.

[0213] Case 11.1a: The first device obtains the first neighborhood i from the second device.

[0214] In this case, the above-mentioned obtaining, by the first device, of the first neighborhood i according to the N point cloud data can specifically include: the first device sends a first message to the second device, and correspondingly, the second device receives the first message from the first device, the first message being used to request a region in which each point cloud data in the N point cloud data is projected onto a target image, the target image being an image converted from an image obtained by two-dimensional perception of a target object by the second device according to a first parameter, the first parameter being an internal and external parameter associated with the target image; the second device sends the first neighborhood i to the first device according to the first message, and correspondingly, the first device receives the first neighborhood i from the second device.

[0215] The first parameter is used for converting the image #1 into a target image, i.e., the second device can convert the image #1 into the first image mentioned above by using the first parameter. The first parameter is an internal and external parameter associated with the target image, which can be understood as an internal and external parameter of a device associated with the target image. The device associated with the target image can be a virtual device. The internal and external parameter includes at least one of the following: focal length, resolution, field of view, position, or pointing direction. The focal length, resolution, field of view, position, and pointing direction can be understood as the focal length, resolution, field of view, position, and pointing direction of the device associated with the target image. For example, the internal and external parameter is the focal length, resolution, field of view, position, and pointing direction of the device associated with the image. It can be understood that the first parameter can be pre-set by the first device, such as setting a commonly used parameter as the first parameter, or can be pre-defined by a protocol, which can be flexibly set according to actual conditions, and is not limited.

[0216] In the embodiment of the present application, the first device can send a first message to the second device to request the second device to provide the first device with the first neighborhood i. In this way, the first device does not need to calculate the first neighborhood i, thereby reducing the calculation overhead of the first device.

[0217] In addition, the second device can project the first i point cloud data to the target image to obtain a projection point i, and then perform a neighborhood extraction operation on the projection point i. Details can be referred to the related description in the foregoing “2. Multi-modal fusion perception”. The second device can also determine the first space i in which the i-th point cloud data is located, and then project the first space i to the target image to obtain the first neighborhood i, i.e., the region on the target image where the first space i is projected is the first neighborhood i.

[0218] Optionally, the first message can include at least one of the following: the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, or the first parameter mentioned above. Each of the parameters is introduced as follows.

[0219] The spatial range of each of the N point cloud data is used to indicate the size and / or shape of the space in which each of the N point cloud data is located, for example, the spatial range of each of the N point cloud data is spherical, and the radius of the sphere is 5 centimeters (cm), for another example, the spatial range of each of the N point cloud data is a cube, and the side length of the cube is 3 cm. The spatial range of each of the N point cloud data can be used to determine the first space i in which the i-th point cloud data is located, for example, when the spatial range of each of the N point cloud data is spherical and the radius of the sphere is 5 cm, the second device determines the center point as the i-th point cloud data and the spherical region with a radius of 5 cm, and the spherical region is the first space i. It can be understood that when the first neighborhood i is the region on the first image projected by the first space i in which the i-th point cloud data is located, the first message can include the spatial range of each of the N point cloud data, so that the second device determines the first space i based on the spatial range of each of the N point cloud data in the first message. Of course, the spatial range of each of the N point cloud data can be pre-set or protocol pre-defined, that is, in this case, the first message can not carry the spatial range of each of the N point cloud data, and the specific setting can be flexibly set according to the actual situation, which is not limited.

[0220] The normal of each of the N point cloud data is used to determine the rotation matrix corresponding to each of the N point cloud data. It can be understood that the first device can preliminarily estimate the plane information in which each of the N point cloud data is located according to the N point cloud data, and determine the corresponding plane normal based on the plane information in which the point cloud data is located, thereby obtaining the normal of each of the N point cloud data. In addition, the normal of each of the N point cloud data can be used as the Z-axis direction, and the second device can determine the rotation matrix corresponding to each of the N point cloud data based on the Z-axis direction. The determination method of the rotation matrix can refer to the related description in the foregoing “1.2 optical perception”, which will not be described here again. In addition, after determining the rotation matrix corresponding to each of the N point cloud data based on the normal of each of the N point cloud data, the second device can project the N point cloud data to the first image based on the rotation matrix.

[0221] For example, the second device can determine the normal of the i-th point cloud data as the Z-axis direction of the camera coordinate system, thereby determining the value of the Z-axis of the i-th point cloud data in the camera coordinate system, and determining the rotation matrix of the i-th point cloud data based on the value; after determining the rotation matrix of the i-th point cloud data, the second device can determine the extrinsic parameter matrix based on the rotation matrix, and thereby project the i-th point cloud data to the first image based on the extrinsic parameter matrix. It can be understood that the camera coordinate system refers to the camera coordinate system corresponding to the device associated with the target image; the extrinsic parameter matrix corresponds to E in the foregoing “1.2 optical perception”. virtualThe implementation principle of projecting the i-th point cloud data to the first image based on the extrinsic parameter matrix can be understood with reference to the foregoing description of the conversion of real camera image pixels to virtual camera image pixels in the “1.2 Optical perception” section, which will not be repeated here.

[0222] In the embodiments of the present application, the first message can include the normal of each of the N point cloud data and the first parameter, or the first message can include the spatial range of each of the N point cloud data, the normal of each of the N point cloud data and the first parameter. The specific configuration can be flexibly set according to actual conditions, and is not limited.

[0223] Optionally, the first message is further used to indicate an angle range, and the pointing direction of the second device belongs to the angle range. The two-dimensional image obtained by the device belonging to the angle range from perceiving the target object can be used to determine the two-dimensional image for projecting the N point cloud data.

[0224] The pointing direction of the second device belonging to the angle range can ensure that the second device and the second virtual device have an effective spatial range, i.e., the image obtained by the second device from two-dimensional perception of the target object can be used to determine the first neighborhood i. The first message can carry the angle range. After receiving the angle range, the second device can first determine whether the pointing direction of the second device belongs to the angle range, and obtain the first neighborhood i when the pointing direction of the second device belongs to the angle range. In this way, the first device can obtain the first neighborhood i from the second device, thereby avoiding the failure of the first device to obtain the first neighborhood i and generating additional communication overhead.

[0225] It can be understood that the angle range can be set according to actual conditions. For example, when the quality requirement of the first image is high, the angle range can be set to be small, so that the pointing angle of the second device and the second virtual device is small, thereby improving the quality of the converted first image, and further improving the perception quality and noise filtering quality. Conversely, when the quality requirement of the first image is low, the angle range can be set to be large, so that the pointing angle of the second device and the second virtual device is large.

[0226] Case 11.2a: The first device converts the third image into the first image using the second parameter, and determines the first neighborhood i based on the first image.

[0227] In this case, the first device obtaining the first neighborhood i according to the N point cloud data can specifically include: the first device obtaining a second parameter and a third image, the second parameter being used for converting the third image into the first image; and the first device determining the first neighborhood i according to the N point cloud data, the second parameter and the third image. The third image is an image obtained by two-dimensionally perceiving the target object, that is, the third image includes an imaging result of the target object. The third image can be used to determine the first image, that is, the first image can be determined based on the third image. It can be understood that the third image is the image #1. In this way, the first device can convert the third image into the first image using the second parameter, and determine the first neighborhood i based on the first image.

[0228] Optionally, the communication method can further include: the second device obtaining a second parameter and a third image, the second parameter being a parameter indicating that the target object is two-dimensionally perceived, and the third image being an image obtained by two-dimensionally perceiving the target object; and the second device sending the second parameter and the third image to the first device. Correspondingly, the first device obtaining the second parameter and the third image can specifically include: the first device receiving the second parameter and the third image from the second device.

[0229] The second parameter can be understood as an internal and external parameter of the second device. The second parameter can include at least one of the following: a focal length, a resolution, a field of view, a position, or a pointing direction. For example, the second parameter includes the focal length, the resolution, the field of view, the position, and the pointing direction of the second device.

[0230] In the embodiments of the present application, the first device can obtain the second parameter and the third image from the second device, to determine the first neighborhood i based on the second parameter and the third image. It can be understood that the first device can also obtain the second parameter and the third image in other ways, which are not limited.

[0231] Further, before the second device sends the second parameter and the third image to the first device, the above communication method further comprises: the first device sends a third message to the second device, and correspondingly, the second device receives the third message from the first device, the third message being used to request the second device to provide the related parameter and the image for the two-dimensional perception of the target object; and the second device sending the second parameter and the third image to the first device can specifically include: the second device sending the second parameter and the third image to the first device according to the third message. That is, the first device can request the second device to provide the related parameter and the image for the two-dimensional perception of the target object, so that the second device provides the second parameter and the third image to the first device based on the request. In this way, the first device can obtain the second parameter and the third image in time. Of course, the second device can also actively report the related parameter and the image for the two-dimensional perception of the target object to the first device, such as periodically reporting the related parameter and the image to the first device, which can be set according to actual conditions and is not limited.

[0232] Further, before the first device sends the third message to the second device, the above communication method can further comprise: the second device sends a third capability parameter to the first device, and correspondingly, the first device receives the third capability parameter from the second device, the third capability parameter being used to indicate that the second device supports providing the related parameter and the image for the two-dimensional perception of the object; and the first device sending the third message to the second device can specifically include: the first device sending the third message to the second device according to the third capability parameter.

[0233] The third capability parameter can also be used to indicate that the second device supports providing the related parameter and the image for the two-dimensional perception of the object (to the first device). Or, the third capability parameter can also be used to indicate that the second device can provide the related parameter and the image for the two-dimensional perception of the object (to the first device). That is, the second device has the capability of providing the related parameter and the image for the two-dimensional perception of the object (to the first device).

[0234] After receiving the third capability parameter, the first device can determine, based on the third capability parameter, that the second device has the capability of providing the related parameter and the image for the two-dimensional perception of the object to the first device. At this time, the first device can (determine) to send the third message to the second device, that is, to request the second device to provide the related parameter and the image for the two-dimensional perception of the target object. In this way, it can be avoided that the first device fails to request the second device to provide the related parameter and the image for the two-dimensional perception of the target object, thereby causing additional communication overhead.

[0235] Further, the third message is also used to indicate an angle range, the pointing direction of the second device belongs to the angle range, and the two-dimensional image obtained by the device belonging to the angle range for perceiving the target object can be used to determine the two-dimensional image for N point cloud data projection.

[0236] The pointing direction of the second device belongs to an angle range which can ensure that the second device and the second virtual device exist in an effective spatial range. Details can be referred to the foregoing description of case 11.1a, which will not be repeated here. The third message can carry the angle range. After receiving the angle range, the second device can first determine whether the pointing direction of the second device belongs to the angle range, and send the relevant parameters and the image of the two-dimensional perception of the object by the second device to the first device when the pointing direction of the second device belongs to the angle range. In this way, the first device can obtain the relevant parameters and the image of the two-dimensional perception of the object by the second device, thereby avoiding the failure of the first device to obtain the relevant parameters and the image of the two-dimensional perception of the object by the second device, and generating additional communication overhead.

[0237] It can be understood that the angle range can be set according to actual conditions. Details can be referred to the foregoing description of case 11.1a, which will not be repeated here.

[0238] It can also be understood that the above describes different ways for the first device to obtain the first neighborhood i through case 11.1a and case 11.2a. It can be understood that when the first neighborhood i is the area where the i th point cloud data is projected to the first image, and the first image is the two-dimensional image converted from the image #1, the second neighborhood i can be the area where the i th point cloud data is projected to the second image, and the second image can be the image obtained by the third device for two-dimensional perception of the target object, or the two-dimensional image converted from the image #2. Taking the second neighborhood i as the area where the i th point cloud data is projected to the second image, and the second image as the two-dimensional image converted from the image #2 as an example, different ways for the first device to obtain the second neighborhood i are introduced.

[0239] Case 11.1b: The first device obtains the second neighborhood i from the third device.

[0240] In this case, the above-mentioned obtaining of the second neighborhood i by the first device according to the N point cloud data can specifically include:

[0241] The first device sends a fourth message to the third device, and correspondingly, the third device receives the fourth message from the first device. The fourth message is used to request the area where each point cloud data in the N point cloud data is projected to the target image, and the target image is the image converted from the image obtained by the third device for two-dimensional perception of the target object according to the first parameter, and the first parameter is the internal and external parameters associated with the target image. The third device sends the second neighborhood i to the first device according to the fourth message, and correspondingly, the first device receives the second neighborhood i from the third device.

[0242] The fourth message can be the same as or different from the first message, such as the fourth message being the same as the first message in content and different in information element (IE). The internal and external parameters include at least one of the following: focal length, resolution, field of view, position, or pointing direction. In addition, the fourth message can refer to the related description of the first message, and the first parameter and the internal and external parameters can refer to the related description above, which will not be repeated here. In the embodiment of the application, the first device can request the third device to provide the second neighborhood i to the first device by sending the fourth message to the third device. In this way, the first device does not need to calculate the second neighborhood i, thereby reducing the calculation overhead of the first device.

[0243] Optionally, the fourth message includes at least one of the following: the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, or the first parameter, which can refer to the related description in the aforementioned "Case 11.1a", and will not be repeated here. In the embodiment of the application, the fourth message can include the normal of each of the N point cloud data and the first parameter, and the fourth message can also include the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, and the first parameter. The specific implementation can be flexibly set according to actual conditions, and is not limited.

[0244] It can be understood that the specific implementation principle of the third device obtaining the second neighborhood i is similar to that of the second device obtaining the first neighborhood i, and can be understood by mutual reference, which will not be repeated here.

[0245] Optionally, the fourth message is also used to indicate an angle range, the pointing direction of the third device belongs to the angle range, and the two-dimensional image obtained by the device belonging to the angle range sensing the target object can be used to determine the two-dimensional image used for projection of the N point cloud data.

[0246] It can be understood that the specific implementation principle of the pointing direction of the third device belonging to the angle range is similar to that of the second device belonging to the angle range in the aforementioned "Case 11.1a", and can be understood by mutual reference, which will not be repeated here. In addition, the angle range can be set according to actual conditions, and can refer to the related description in the aforementioned "Case 11.1a", such as replacing the first image with the second image, replacing the second device with the third device, and replacing the second virtual device with the third virtual device, which will not be repeated here.

[0247] Case 11.2b: The first device converts the fourth image into a second image using a fourth parameter, and determines the second neighborhood i based on the second image.

[0248] In this case, the first device obtaining the second neighborhood i according to the N point cloud data can specifically include: the first device obtaining a fourth parameter and a fourth image, the fourth parameter being used for converting the fourth image into the second image; and the first device determining the second neighborhood i according to the N point cloud data, the fourth parameter and the fourth image. The fourth image is an image obtained by two-dimensionally perceiving the target object, that is, the fourth image includes an imaging result of the target object. The fourth image can be used to determine the second image, that is, the second image can be determined based on the fourth image. It can be understood that the fourth image is the image #2 described above. In this way, the first device can convert the fourth image into the second image using the fourth parameter, and determine the second neighborhood i based on the second image.

[0249] Optionally, the communication method can further include: the third device obtaining a fourth parameter and a fourth image, the fourth parameter being used for indicating a related parameter of two-dimensionally perceiving the target object by the third device, and the fourth image being an image obtained by two-dimensionally perceiving the target object by the third device; and the third device sending the fourth parameter and the fourth image to the first device. Correspondingly, the first device obtaining the fourth parameter and the fourth image can specifically include: the first device receiving the fourth parameter and the fourth image from the third device.

[0250] The fourth parameter can be understood as an internal and external parameter of the third device. The fourth parameter can include at least one of the following: a focal length, a resolution, a field of view, a position, or a pointing direction. For example, the fourth parameter includes the focal length, the resolution, the field of view, the position, and the pointing direction of the third device.

[0251] In the embodiments of the present application, the first device can obtain the fourth parameter and the fourth image from the third device, to determine the second neighborhood i based on the fourth parameter and the fourth image. It can be understood that the first device can also obtain the fourth parameter and the fourth image in other manners, which is not limited.

[0252] Further, before the third device sends the fourth parameter and the fourth image to the first device, the communication method can further include that the first device sends a sixth message to the third device, and correspondingly, the third device receives the sixth message from the first device, the sixth message being used to request the third device to provide the parameter and the image related to the two-dimensional perception of the target object; and the third device sending the fourth parameter and the fourth image to the first device can specifically include that the third device sends the fourth parameter and the fourth image to the first device according to the sixth message. That is, the first device can request the third device to provide the parameter and the image related to the two-dimensional perception of the target object, so that the third device provides the fourth parameter and the fourth image to the first device based on the request. In this way, the first device can obtain the fourth parameter and the fourth image in time. Of course, the third device can also actively report the parameter and the image related to the two-dimensional perception of the target object to the first device, for example, periodically report the parameter and the image to the first device, which can be set according to actual conditions and is not limited.

[0253] Further, before the third device obtains the fourth parameter and the fourth image, the communication method can further include that the third device sends a sixth capability parameter to the first device, and correspondingly, the first device receives the sixth capability parameter from the third device, the sixth capability parameter being used to indicate that the third device supports providing the parameter and the image related to the two-dimensional perception of the object; and the first device sending the sixth message to the third device can specifically include that the first device sends the sixth message to the second device according to the sixth capability parameter.

[0254] The sixth capability parameter can also be used to indicate that the third device supports providing the parameter and the image related to the two-dimensional perception of the object to the first device. The sixth capability parameter is similar to the third capability parameter in the aforementioned "case 11.2a", and can be understood by referring to the related description of the third capability parameter, for example, by replacing the second device with the third device, which will not be described herein again.

[0255] After the first device receives the sixth capability parameter, the first device can determine, based on the sixth capability parameter, that the third device has the capability of providing the parameter and the image related to the two-dimensional perception of the object to the first device. At this time, the first device can (determine to) send the sixth message to the third device, that is, request the third device to provide the parameter and the image related to the two-dimensional perception of the target object. In this way, the first device can avoid failure in requesting the third device to provide the parameter and the image related to the two-dimensional perception of the target object, thereby avoiding additional communication overhead.

[0256] Further, the sixth message is also used to indicate an angle range, and the third device points to the angle range. The two-dimensional image obtained by the device belonging to the angle range perceiving the target object can be used to determine the two-dimensional image for the projection of the N point cloud data.

[0257] The pointing direction of the third device belongs to an angle range that can ensure that the third device and the third virtual device have an effective spatial range, which can be understood with reference to the foregoing “case 11.1a”, and will not be described here again. In addition, the angle range can be set according to actual conditions, which can be understood with reference to the foregoing “case 11.1a”, for example, by replacing the first image with the second image, replacing the second device with the third device, and replacing the second virtual device with the third virtual device, and will not be described here again.

[0258] It can be understood that the specific implementation principle of the first device for determining the second neighborhood i is similar to the specific implementation principle of the first device for determining the first neighborhood i, and can be understood by mutual reference, and will not be described here again.

[0259] In addition, the above describes the acquisition manner of the first neighborhood i and the second neighborhood i. It can be understood that the first device can also acquire other neighborhoods of the i-th point cloud data, such as a third neighborhood 3i, a fourth neighborhood i, and the like. In other words, the first device can acquire two or more neighborhoods of the i-th point cloud data. The specific manner in which the first device acquires a neighborhood of the i-th point cloud data other than the first neighborhood i and the second neighborhood i can be understood with reference to the specific manner in which the first device acquires the first neighborhood i, such as the first device acquiring the third neighborhood 3i of the i-th point cloud data from the fourth device, and the first device acquiring the relevant information of the parameters and the image from the fourth device and acquiring the third neighborhood 3i based on the relevant information, and will not be described here again.

[0260] For S1102:

[0261] After the first device acquires the first neighborhood i and the second neighborhood i, the first device can determine the neighborhood matching degree of the i-th point cloud data based on the first neighborhood i and the second neighborhood i, and determine whether the i-th point cloud data is real point cloud data or noise (point cloud data) based on the neighborhood matching degree.

[0262] For example, after determining the neighborhood matching degree of each of the N point cloud data through the above NSSD or CDF, the first device can compare the neighborhood matching degree of the i-th point cloud data with a matching degree threshold after determining the neighborhood matching degree of the i-th point cloud data. If the neighborhood matching degree of the i-th point cloud data is less than the matching degree threshold, it is determined that the i-th point cloud data is noise point cloud data, and if the neighborhood matching degree of the i-th point cloud data is greater than or equal to the matching degree threshold, it is determined that the i-th point cloud data is real point cloud data.

[0263] It can be understood that when the first device acquires more than two neighborhoods of the i-th point cloud data, the first device can determine the neighborhood matching degree of the i-th point cloud data according to the acquired multiple neighborhoods, such as calculating the neighborhood matching degrees corresponding to each pair of neighborhoods in the multiple neighborhoods, and then calculating the average value based on the multiple neighborhood matching degrees to obtain the final neighborhood matching degree, or such as calculating the neighborhood matching degrees corresponding to each pair of neighborhoods in the multiple neighborhoods, and taking the neighborhood matching degree with the maximum or minimum value in the multiple neighborhood matching degrees as the final neighborhood matching degree of the i-th point cloud data.

[0264] It can be understood that the N point cloud data in the embodiments of the present application includes at least one real point cloud data and / or at least one noise point cloud data. The real point cloud data can be understood as the point cloud data determined by the echo generated by the target object in the radio frequency perception process; or in other words, the real point cloud data can be understood as the point cloud data other than the noise point cloud data. The noise point cloud data can also be referred to as a noise point, a three-dimensional noise scattering point, or other possible names, without limitation.

[0265] In addition, after the real point cloud data is projected onto multiple two-dimensional images, the regions where the projection points corresponding to the point cloud data are located correspond to the imaging results of similar (or the same) regions of the object, so the similarity of the regions where the projection points corresponding to the point cloud data are located is relatively high. After the noise point cloud data is projected onto multiple two-dimensional images, the regions where the projection points corresponding to the noise point cloud data are located correspond to the imaging results of different regions, so the similarity of the regions where the projection points corresponding to the noise point cloud data are located is relatively low.

[0266] In the embodiments of the present application, the imaging plane of the first image and the imaging plane of the second image are parallel, so that the imaging results of the target object on the first image and the imaging results of the target object on the second image have relatively small deformation, thereby making the imaging results of the target object in the first neighborhood i and the second neighborhood i have relatively small deformation. Therefore, the first device can avoid determining (or identifying) the real point cloud data as noise, i.e., can avoid determining the real point cloud data as noise due to the relatively large deformation of the imaging results of the target object in the first neighborhood i and the second neighborhood i. In this way, after the first device determines the real point cloud data in the N point cloud data, using the real point cloud data for multi-modal fusion perception can improve the multi-modal perception performance.

[0267] The above describes one way of multi-modal fusion perception of the first device according to the first neighborhood i and the second neighborhood i. It can be understood that the first device can also determine the boundary and / or corner point of the target object according to the first neighborhood i and the second neighborhood i, and perform multi-modal fusion perception based on the determined boundary and / or corner point of the target object. That is, at this time, the first device can determine more accurate boundary and / or corner point of the target object according to the first neighborhood i and the second neighborhood i, thereby improving the performance of multi-modal fusion perception. In addition, the specific implementation principle of determining the boundary and / or corner point of the target object according to the first neighborhood i and the second neighborhood i, and the specific implementation principle of performing multi-modal fusion perception based on the boundary and / or corner point of the target object can refer to the prior art, which will not be described here.

[0268] In summary, in the embodiments of the present application, the first device obtains the first neighborhood i after the i-th point cloud data in the N point cloud data is projected to the first image, and the second neighborhood i after the i-th point cloud data is projected to the second image. Because the imaging planes of the first image and the second image are parallel, the imaging results of the target object on the first image and the imaging results of the target object on the second image have small deformation, so that the imaging results of the target object in the first neighborhood i and the second neighborhood i have small deformation. In this way, when the first device performs multi-modal fusion perception based on the first neighborhood i and the second neighborhood i, the performance of multi-modal fusion perception can be improved.

[0269] Optionally, in combination with the above embodiments, for the above case 11.1a, the first neighborhood i where the i-th point cloud data is located after being projected onto the first image can specifically include: the first neighborhood i is the region on the first image where the first space i where the i-th point cloud data is located is projected onto; the second neighborhood i where the i-th point cloud data is located after being projected onto the second image can specifically include: the second neighborhood i is the region on the second image where the first space i where the i-th point cloud data is located is projected onto; the first message is further used to request a scaling parameter, the scaling parameter being a scaling ratio between the imaging results of the objects in the two-dimensional perception; the second device sends the first neighborhood i to the first device, and the first scaling parameter used to indicate the scaling ratio between the imaging results of the target objects in the first image can be specifically included; correspondingly, the first device receives the first neighborhood i from the second device can specifically include: the first device receives the first neighborhood i and the first scaling parameter from the second device; the communication method can further include: the first device determines a third neighborhood i according to the first neighborhood i, the first scaling parameter, and a second scaling parameter used to indicate the scaling ratio between the imaging results of the target objects in the second image, the number of pixel points included in the third neighborhood i being the same as the number of pixel points included in the second neighborhood i; the first device performs multi-modal fusion perception according to the first neighborhood i and the second neighborhood i can specifically include: the first device performs multi-modal fusion perception according to the third neighborhood i and the second neighborhood i.

[0270] The first scaling parameter can be determined by the second device, for example, based on the spatial range of the i-th point cloud data and the normal of the i-th point cloud data. For example, the second device determines the value of the Z-axis in the camera coordinate system corresponding to the second device of the i-th point cloud data based on the spatial range of the i-th point cloud data and the normal of the i-th point cloud data, and determines the first scaling parameter based on the value and the focal length of the second device. For another example, the second device determines the value of the Z-axis in the camera coordinate system corresponding to the device associated with the first image (i.e., the second virtual device) of the i-th point cloud data based on the spatial range of the i-th point cloud data and the normal of the i-th point cloud data, and determines the first scaling parameter based on the value and the focal length of the device (i.e., the focal length in the first parameter). The specific implementation principle of calculating the scaling parameter can be referred to the related description of the scaling factor in the aforementioned "1.2 optical perception", which will not be repeated here. The second scaling parameter can be determined by the third device, for example, based on the spatial range of the i-th point cloud data and the normal of the i-th point cloud data. The specific implementation principle of determining the second scaling parameter by the third device is similar to that of determining the first scaling parameter by the second device, and can be understood by mutual reference, which will not be repeated here.

[0271] In the embodiments of the present application, when the first neighborhood i is a region on the first image projected by the first space i where the i-th point cloud data is located, and the second neighborhood i is a region on the second image projected by the first space i where the i-th point cloud data is located, the number of pixel points in the first neighborhood i and the number of pixel points in the second neighborhood i can be different. In this case, the first device can change the number of pixel points in the first neighborhood i through the first scaling parameter and the second scaling parameter, so that the number of pixel points in the changed first neighborhood i (i.e., the third neighborhood i) is the same as the number of pixel points in the second neighborhood i, thereby facilitating the first device to determine the neighborhood matching degree of the i-th point cloud data in the manner of element-by-element comparison of similarity, such as the NSSD or CDF in the aforementioned "2. Multimodal fusion perception", and thereby performing multimodal fusion perception based on the neighborhood matching degree.

[0272] The specific manner in which the first device determines the third neighborhood i according to the first neighborhood i, the first scaling parameter and the second scaling parameter is related to the number of pixel points in the first neighborhood i and the number of pixel points in the second neighborhood i. The following will be specifically explained.

[0273] When the number of pixel points in the first neighborhood i is greater than the number of pixel points in the second neighborhood i, the first device can reduce the pixel points in the first neighborhood i. For example, the first device samples the first neighborhood i at intervals of a first value, and the first value is the ratio of the second scaling parameter to the first scaling parameter.

[0274] When the number of pixel points in the first neighborhood i is less than the number of pixel points in the second neighborhood i, the first device can add pixel points to the first neighborhood i. For example, the first device interpolates the first neighborhood i at a first ratio, and the first ratio is the ratio of the first scaling parameter to the second scaling parameter. It can be understood that the interpolation of the first device to the first neighborhood i can be 0, 1 or other possible values, which can be flexibly set according to actual conditions without limitation.

[0275] Further, the first device performing the multi-modal fusion perception according to the third neighborhood i and the second neighborhood i can specifically include: the first device determining a neighborhood matching degree of the i-th point cloud data according to a similarity degree of the third neighborhood i and the second neighborhood i; the first device determining noise in the N point cloud data according to the neighborhood matching degree of the i-th point cloud data; and the first device performing the multi-modal fusion perception based on the point cloud data in the N point cloud data except the noise. It can be understood that the first device can determine (or calculate) the similarity degree of the third neighborhood i and the second neighborhood i by using the above-mentioned NSSD, CDF or other manners. The specific implementation principle of the first device performing the multi-modal fusion perception according to the third neighborhood i and the second neighborhood i is similar to the specific implementation principle of the first device performing the multi-modal fusion perception according to the first neighborhood i and the second neighborhood i in the above-mentioned “S1102”, and can be understood by mutual reference, which will not be described here.

[0276] Further, before the first device sends the first message to the second device, the above-mentioned communication method can further include: the second device sending a first capability parameter to the first device, and correspondingly, the first device receiving the first capability parameter from the second device, the first capability parameter being used to indicate that the support for providing the scaling parameter; and the first device sending the first message to the second device can specifically include: the first device sending the first message to the second device according to the first capability parameter.

[0277] The first capability parameter can also be used to indicate that the second device supports (to the first device) providing the scaling parameter. Or, the first capability parameter can also be used to indicate that the second device can (to the first device) provide the scaling parameter. That is, the second device has the capability of providing the scaling parameter (to the first device).

[0278] After receiving the first capability parameter, the first device can determine that the second device has the capability of providing the scaling parameter to the first device based on the first capability parameter. At this time, the first device can (determine) to send the first message to the second device, that is, to request the second device to provide the scaling parameter. In this way, it can avoid the failure of the first device requesting the scaling parameter from the second device, thereby causing additional communication overhead.

[0279] Optionally, in combination with the above embodiments, for the above case 11.1b, the first neighborhood i of the i-th point cloud data after being projected onto the first image can specifically include: the first neighborhood i is the region of the i-th point cloud data in the first space i projected onto the first image; the second neighborhood i of the i-th point cloud data after being projected onto the second image can specifically include: the second neighborhood i is the region of the first space i projected onto the second image; the fourth message is further used to request a scaling parameter, the scaling parameter being a scaling ratio between the object in two-dimensional perception and the imaging result of the object; the first device receiving the second neighborhood i from the third device can specifically include: the third device sending the second neighborhood i and the second scaling parameter to the first device, and correspondingly, the first device receiving the second neighborhood i and the second scaling parameter from the third device, the second scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the second image; the communication method can further include: the first device determining a third neighborhood i according to the first neighborhood i, the first scaling parameter and the second scaling parameter, the first scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the first image, the third neighborhood i including the same number of pixel points as the second neighborhood i; the first device performing multi-modal fusion perception according to the first neighborhood i and the second neighborhood i can specifically include: the first device performing multi-modal fusion perception according to the third neighborhood i and the second neighborhood i.

[0280] The first scaling parameter and the second scaling parameter can refer to the above related description, which will not be repeated here. In addition, the specific implementation principles of the first device determining the third neighborhood i according to the first neighborhood i, the first scaling parameter and the second scaling parameter, and the first device performing multi-modal fusion perception according to the third neighborhood i and the second neighborhood i can refer to the foregoing related description, which will not be repeated here.

[0281] Further, the first device performing multi-modal fusion perception according to the third neighborhood i and the second neighborhood i can specifically include: the first device determining the neighborhood matching degree of the i-th point cloud data according to the similarity degree of the third neighborhood i and the second neighborhood i; the first device determining the noise in the N point cloud data according to the neighborhood matching degree of the i-th point cloud data; the first device performing multi-modal fusion perception based on the point cloud data in the N point cloud data except the noise, which can refer to the foregoing related description, which will not be repeated here.

[0282] Further, before the first device sends the fourth message to the third device, the above communication method can further include: the first device sends a fourth capability parameter to the third device, and correspondingly, the first device receives the fourth capability parameter from the third device, the fourth capability parameter being used to indicate that the third device supports providing the scaling parameter; and the above sending of the fourth message by the first device to the third device can specifically include: the first device sends the fourth message to the third device according to the fourth capability parameter. It can be understood that the fourth capability parameter can also be used to indicate that the third device supports providing (to the first device) the scaling parameter, and the fourth capability parameter is similar to the above first capability parameter, and can be understood with reference to the above description of the first capability parameter, for example, by replacing the second device with the third device, and details are not described herein again. In this way, the first device can avoid failure in requesting the scaling parameter from the third device, thereby causing additional communication overhead.

[0283] Optionally, in combination with the above embodiments, the first neighborhood i of the i th point cloud data after being projected onto the first image can specifically include: the first neighborhood i of the i th point cloud data is a region on the first image after the first space i where the i th point cloud data is located is projected onto the first image; and the second neighborhood i of the i th point cloud data after being projected onto the second image can specifically include: the second neighborhood i of the i th point cloud data is a region on the second image after the first space i where the i th point cloud data is located is projected onto the second image; and the above communication method can further include: the first device determines a third neighborhood i according to the first neighborhood i, a first scaling parameter and a second scaling parameter, the first scaling parameter being used to indicate a scaling ratio between the target object and an imaging result of the target object in the first image, the second scaling parameter being used to indicate a scaling ratio between the target object and an imaging result of the target object in the second image, and the third neighborhood i including the same number of pixel points as the second neighborhood i; and the above multi-modal fusion perception by the first device according to the first neighborhood i and the second neighborhood i can specifically include: the first device performs multi-modal fusion perception according to the third neighborhood i and the second neighborhood i, and details can be referred to the above description, which are not described herein again.

[0284] Further, the above multi-modal fusion perception by the first device based on the first neighborhood i and the second neighborhood i can specifically include: the first device determines a neighborhood matching degree of the i th point cloud data according to a similarity degree between the third neighborhood i and the second neighborhood i; the first device determines noise in the N point cloud data according to the neighborhood matching degree of the i th point cloud data; and the first device performs multi-modal fusion perception based on the point cloud data in the N point cloud data except the noise, and details can be referred to the above description, which are not described herein again.

[0285] It can be understood that the first device can acquire the first scaling parameter from the second device and acquire the second scaling parameter from the third device. Details are described below.

[0286] Case 11.1c: the first device acquires the first scaling parameter from the second device.

[0287] In this case, before the first device determines the third neighborhood i according to the first neighborhood i, the first scaling parameter and the second scaling parameter, the communication method can further include that the first device sends a second message to the second device, and correspondingly, the second device receives the second message from the first device, the second message being used to request a scaling parameter for scaling between an object performing two-dimensional perception and an imaging result of the object; the second device sends the first scaling parameter to the first device according to the second message, and correspondingly, the second device receives the first scaling parameter from the second device. That is, the second device can request the first device for the scaling parameter.

[0288] It can be understood that in the case 11.1a, the first device can request the first neighborhood i and the first scaling parameter from the second device through different messages (the first message and the second message); in the case 11.2a, the first device can also not request the scaling parameter from the second device, that is, the first device can determine the first scaling parameter according to the received second parameter. In this way, the first device can timely acquire the scaling parameter from the second device when the scaling parameter of the second device is needed, so as to facilitate the first device to determine the third neighborhood i.

[0289] The second message can include at least one of the following: a spatial range of each of the N point cloud data, or a normal line of each of the N point cloud data, the spatial range of each of the N point cloud data being used to indicate a size and / or shape of a space where each of the N point cloud data is located, and the normal line of each of the N point cloud data being used to determine a rotation matrix corresponding to each of the N point cloud data. The spatial range of each of the N point cloud data and the normal line of each of the N point cloud data can refer to the related description in the case 11.1a, which will not be repeated here.

[0290] In the embodiments of the present application, the second message can include the spatial range of each of the N point cloud data and the normal line of each of the N point cloud data, so that the second device determines the first scaling parameter according to the spatial range of each of the N point cloud data and the normal line of each of the N point cloud data.

[0291] Further, before the first device sends the second message to the second device, the communication method can further include that the second device sends a second capability parameter to the first device, and correspondingly, the first device receives the second capability parameter from the second device, the second capability parameter being used to indicate that the scaling parameter is supported; the first device sending the second message to the second device can specifically include that the first device sends the second message to the second device according to the second capability parameter.

[0292] The second capability parameter can also be used to indicate that the second device supports providing the scaling parameter to the first device. The second capability parameter is similar to the first capability parameter described above, and can be understood by referring to the description of the first capability parameter, which will not be repeated here.

[0293] After receiving the second capability parameter, the first device can determine, based on the second capability parameter, that the second device has the capability of providing the scaling parameter to the first device. At this time, the first device can send (determine to send) a second message to the second device, i.e., request the second device to provide the scaling parameter. In this way, the first device can avoid failure to request the scaling parameter from the second device, thereby causing additional communication overhead.

[0294] Case 11.2c: The first device obtains the second scaling parameter from the third device.

[0295] In this case, before the first device determines the third neighborhood i according to the first neighborhood i, the first scaling parameter, and the second scaling parameter, the communication method can further include: the first device sends a fifth message to the third device, and correspondingly, the third device receives the fifth message from the first device, the fifth message being used to request a scaling parameter, the scaling parameter being a scaling ratio between the imaging result of the object and the object for two-dimensional perception; the third device sends the second scaling parameter to the first device according to the fifth message, and correspondingly, the first device receives the second scaling parameter from the third device.

[0296] It can be understood that in the above case 11.1b, the first device can request the second neighborhood i and the second scaling parameter from the second device through different messages (the fourth message and the fifth message); in the above case 11.2b, the first device can also not request the scaling parameter from the second device, i.e., the first device can determine the second scaling parameter according to the received third parameter. In this way, the first device can obtain the scaling parameter from the third device in time when the scaling parameter of the third device is needed, thereby facilitating the first device to determine the third neighborhood i.

[0297] The fifth message can include at least one of the following: the spatial range of each of the N point cloud data, or the normal of each of the N point cloud data, each parameter can be understood by referring to the foregoing description, which will not be repeated here. For example, the fifth message can include the spatial range of each of the N point cloud data and the normal of each of the N point cloud data, so that the third device determines the second scaling parameter according to the spatial range of each of the N point cloud data and the normal of each of the N point cloud data. It can be understood that the fifth message and the second message described above can be the same or different, for example: the fifth message and the second message can be the same in content and different in information element. The fifth message can be understood by referring to the second message described above, which will not be repeated here.

[0298] Further, before the first device sends the fifth message to the third device, the above communication method can further include that the third device sends a fifth capability parameter to the first device, and correspondingly, the first device receives the fifth capability parameter from the third device, the fifth capability parameter being used to indicate that the third device supports providing the scaling parameter; and the above sending, by the first device, of the fifth message to the third device can specifically include that the first device sends the fifth message to the second device according to the fifth capability parameter.

[0299] The fifth capability parameter can also be used to indicate that the third device supports providing (to the first device) the scaling parameter, and the fifth capability parameter can be understood in a similar manner as the second capability parameter, for example, the second device can be replaced by the third device, and details are not repeated here.

[0300] After receiving the fifth capability parameter, the first device can determine, based on the fifth capability parameter, that the third device has the capability of providing the scaling parameter to the first device. At this time, the first device can (determine to) send the second message to the third device, that is, request the third device to provide the scaling parameter. In this way, the first device can avoid failure in requesting the third device to provide the scaling parameter, thereby causing additional communication overhead.

[0301] The above describes the flow of the communication method provided by the embodiments of the present application in combination with the method embodiments. A specific example is given below to illustrate the above method.

[0302] As shown in (a) of FIG. 13, for a building, the three-dimensional coordinate range of a certain wall surface of the building is x=-57, the value range of y is [-103, 103], and the value range of z is [0, 20]; the material of the building is brick, and there are multiple windows on the wall surface, that is, a window is placed every 10 m horizontally and every 4 m vertically, the length and width of the window are both 2 m, and the material of the window is glass. As shown in (b) of FIG. 13, two second devices obtain two two-dimensional images by optical sensing on the wall surface. As shown in (c) of FIG. 13, two converted two-dimensional images are obtained after conversion on the two two-dimensional images, and the conversion can be understood as converting the two two-dimensional images obtained by optical sensing into two images with parallel imaging planes. As shown in (d) of FIG. 13, when the two two-dimensional images obtained by optical sensing are used for noise identification, the number of real point cloud data identified as noise is 16; when the two converted two-dimensional images are used for noise identification, the number of real point cloud data identified as noise is 0.

[0303] It can be seen that the difference between the wall imaging results in the two two-dimensional images obtained by optical perception of the wall is large, and the difference between the wall imaging results in the two two-dimensional images after conversion is small. Moreover, when noise identification is performed using the two two-dimensional images after conversion, the real point cloud data can be avoided from being identified as noise, so that when multi-modal fusion perception is performed using the determined more real point cloud data, the performance of multi-modal fusion perception can be improved. In other words, converting the two two-dimensional images obtained by optical perception into two images with parallel imaging planes can effectively alleviate the influence of deformation, thereby effectively reducing the probability of real point cloud data being incorrectly identified as noise, and further improving the performance of multi-modal fusion perception.

[0304] The above describes the flow of the communication method provided by the embodiments of the present application in combination with the method embodiment shown in FIG. 11. For the convenience of understanding, the interaction process between the first device and the second device in the communication system will be specifically introduced below in combination with FIGS. 14-16.

[0305] Scenario 1:

[0306] FIG. 14 is a flow diagram of a communication method provided by the embodiments of the present application. In scenario 1, the center processing node obtains the internal and external parameters #1 of the optical perception node #1 and the two-dimensional image #1 obtained by the optical perception node #1 on the target object from the optical perception node #1, and obtains the internal and external parameters #2 of the optical perception node #2 and the two-dimensional image #2 obtained by the optical perception node #2 on the target object from the optical perception node #2; and determines the real point cloud data in the N point cloud data based on the internal and external parameters #1, the two-dimensional image #1, the internal and external parameters #2, and the two-dimensional image #2, and finally performs multi-modal fusion perception based on the real point cloud data. It can be understood that the above center processing node corresponds to the first device in the embodiment shown in FIG. 11, the optical perception node #1 corresponds to the second device in the embodiment shown in FIG. 11, and the optical perception node #2 corresponds to the third device in the embodiment shown in FIG. 11.

[0307] As shown in FIG. 14, the flow of the communication method is as follows:

[0308] S1401, the center processing node sends a multi-modal perception capability request message #1 to the optical perception node #1. Correspondingly, the optical perception node #1 receives the multi-modal perception capability request message #1 from the center processing node.

[0309] The capability request message #1 is used to request the optical perception node #1 to report its capability parameters. Alternatively, the multi-modal perception capability request message #1 is used to request the optical perception node #1 to participate in multi-modal fusion perception. It can be understood that when the center processing node requests the optical perception node #1 to participate in multi-modal fusion perception, the optical perception node #1 needs to report its capability parameters.

[0310] S1402, the central processing node sends a multimodal perception capability request message #2 to the optical perception node #2. Correspondingly, the optical perception node #2 receives the multimodal perception capability request message #2 from the central processing node.

[0311] The multimodal perception capability request message #2 is used to request the optical perception node #2 to report its capability parameters. Alternatively, the multimodal perception capability request message #2 is used to request the optical perception node #2 to participate in multimodal fusion perception. It can be understood that when the central processing node requests the optical perception node #2 to participate in multimodal fusion perception, the optical perception node #2 needs to report its capability parameters.

[0312] It can be understood that the multimodal perception capability request message #1 and the multimodal perception capability request message #2 can be the same or different, such as the contents of the multimodal perception capability request message #1 and the multimodal perception capability request message #2 are the same, and the information elements are different.

[0313] In addition, the order of S1401 and S1402 is not limited in the embodiments of the present application, for example: S1401 can be performed first, and then S1402 is performed; or S1402 is performed first, and then S1401 is performed; or S1401 and S1402 are performed at the same time.

[0314] S1403, the optical perception node #1 sends capability parameters #1 to the central processing node. Correspondingly, the central processing node receives the capability parameters #1 from the optical perception node #1.

[0315] The capability parameters #1 are used to indicate that the optical perception node #1 supports providing the central processing node with related parameters and images for two-dimensional perception of an object. The capability parameters #1 correspond to the third capability parameters.

[0316] It can be understood that the optical perception node #1 can send the capability parameters #1 to the central processing node based on the multimodal perception capability request message #1 after receiving the multimodal perception capability request message #1; the optical perception node #1 can also actively report the capability parameters #1, such as periodically reporting the capability parameters of the optical perception node #1, which can be set according to actual conditions and is not limited.

[0317] S1404, the optical perception node #2 sends capability parameters #2 to the central processing node. Correspondingly, the central processing node receives the capability parameters #2 from the optical perception node #2.

[0318] The capability parameters #2 are used to indicate that the optical perception node #2 supports providing the central processing node with related parameters and images for two-dimensional perception of an object. The capability parameters #2 correspond to the sixth capability parameters.

[0319] It can be understood that the optical perception node #2 can send the capability parameter #2 to the central processing node based on the multimodal perception capability request message #2 after receiving the multimodal perception capability request message #2; the optical perception node #2 can also actively report the capability parameter #2, for example, periodically report the capability parameter of the optical perception node #2, which can be set according to actual conditions and is not limited.

[0320] S1405, the central processing node sends a parameter and image request message #1 to the optical perception node #1. Correspondingly, the optical perception node #1 receives the parameter and image request message #1 from the central processing node.

[0321] The parameter and image request message #1 corresponds to the third message described above.

[0322] S1406, the central processing node sends a parameter and image request message #2 to the optical perception node #2. Correspondingly, the optical perception node #2 receives the parameter and image request message #2 from the central processing node.

[0323] The parameter and image request message #2 corresponds to the sixth message described above.

[0324] S1407, the optical perception node #1 sends internal and external parameters #1 and a two-dimensional image #1 to the central processing node. Correspondingly, the central processing node receives the internal and external parameters #1 and the two-dimensional image #1 from the optical perception node #1.

[0325] The internal and external parameters #1 correspond to the second parameter described above. The two-dimensional image #1 corresponds to the third image described above.

[0326] S1408, the optical perception node #2 sends internal and external parameters #2 and a two-dimensional image #2 to the central processing node. Correspondingly, the central processing node receives the internal and external parameters #2 and the two-dimensional image #2 from the optical perception node #2.

[0327] The internal and external parameters #2 correspond to the fourth parameter described above. The two-dimensional image #2 corresponds to the fourth image described above.

[0328] S1409, the central processing node determines the real point cloud data in the N point cloud data based on the internal and external parameters #1, the two-dimensional image #1, the internal and external parameters #2, and the two-dimensional image #2.

[0329] That is, the central processing node can determine the neighborhood matching degree of each of the N point cloud data based on the internal and external parameters #1, the two-dimensional image #1, the internal and external parameters #2, and the two-dimensional image #2, N being an integer greater than 1; and filter the noise in the N point cloud data based on the neighborhood matching degree of each of the N point cloud data to obtain the real point cloud data, which can be referred to the related description in the foregoing “S1102” and will not be described here.

[0330] S14010, the center processing node performs multi-modal fusion perception based on the real point cloud data.

[0331] The specific implementation principle of the center processing node performing multi-modal fusion perception based on the real point cloud data can refer to the prior art, and will not be described here.

[0332] It can be understood that if the optical perception node (optical perception node #1 or optical perception node #2) does not support the reporting of each type of capability, the corresponding multi-modal perception capability request fails. In this case, the optical perception node can not send the capability parameters, that is, it does not perform the above S1403 and / or S1404; or the optical perception node can send an instruction (such as “fail”) to the optical perception node to indicate that it does not participate in this multi-modal fusion perception, and can not inform the reason. And in this case, the optical perception node and the center processing node can not perform the subsequent interaction process.

[0333] It can also be understood that the above S1401-S14010 can refer to the related description in the foregoing FIG. 11, and will not be described here. It can also be understood that the order of each step of the interaction between the center processing node and the optical perception node #1 and the order of each step of the interaction between the center processing node and the optical perception node #2 are not limited by the embodiments of the present application, for example, the center processing node can first interact with the optical perception node #1 and then interact with the optical perception node #2, that is, first perform S1401, S1403, S1405, S1407, and then perform S1402, S1404, S1406, S1408, or the center processing node can perform the above steps in the order.

[0334] In addition, the above S1401-S1406 are optional steps, for example, the optical perception node #1 and the optical perception node #2 can periodically report the capability information to the first device, and / or the optical perception node #1 and the optical perception node #2 can periodically report the internal and external parameters and the two-dimensional image to the first device.

[0335] Scenario 2:

[0336] FIG. 15 is a flow diagram of a communication method provided by the embodiments of the present application. In scenario 2, the center processing node can obtain the neighborhood #1i and the scaling parameter #1 from the optical perception node #1, and the center processing node can obtain the neighborhood #2i and the scaling parameter #2 from the optical perception node #2, and determine the real point cloud data in the N point cloud data based on the neighborhood #1i, the scaling parameter #1, the neighborhood #2i and the scaling parameter #2, and perform multi-modal fusion perception based on the real point cloud data. It can be understood that the above-mentioned center processing node corresponds to the first device in the embodiment shown in FIG. 11, the optical perception node #1 corresponds to the second device in the embodiment shown in FIG. 11, and the optical perception node #2 corresponds to the third device in the embodiment shown in FIG. 11.

[0337] As shown in FIG. 15, the flow of the communication method is as follows:

[0338] S1501, the center processing node sends a multi-modal perception capability request message #1 to the optical perception node #1. Correspondingly, the optical perception node #1 receives the multi-modal perception capability request message #1 from the center processing node.

[0339] S1502, the center processing node sends a multi-modal perception capability request message #2 to the optical perception node #2. Correspondingly, the optical perception node #2 receives the multi-modal perception capability request message #2 from the center processing node.

[0340] The specific implementation principle of S1501-S1502 can be referred to the related description in the foregoing S1401-S1402, which will not be described here.

[0341] S1503, the optical perception node #1 sends a capability parameter #1 to the center processing node. Correspondingly, the center processing node receives the capability parameter #1 from the optical perception node #1.

[0342] The capability parameter #1 is used to indicate that the optical perception node #1 supports providing the scaling parameter to the center processing node. The capability parameter #1 corresponds to the first capability parameter described above.

[0343] The optical perception node #1 can send the capability parameter #1 to the center processing node based on the multi-modal perception capability request message #1 after receiving the multi-modal perception capability request message #1.

[0344] S1504, the optical perception node #2 sends a capability parameter #2 to the center processing node. Correspondingly, the center processing node receives the capability parameter #2 from the optical perception node #2.

[0345] The capability parameter #2 is used to indicate that the optical perception node #2 supports providing the scaling parameter to the center processing node. The capability parameter #2 corresponds to the fourth capability parameter described above.

[0346] The optical perception node #2 can send the capability parameter #2 to the central processing node based on the multi-modal perception capability request message #2 after receiving the multi-modal perception capability request message #2.

[0347] S1505, the central processing node determines the parameter #1.

[0348] The parameter #1 includes at least one of the following: a virtual camera parameter, a spatial range of each of the N point cloud data, a normal of each of the N point cloud data, or a required original camera pointing range, N being an integer greater than 1. The virtual camera parameter corresponds to the first parameter described above, and the central processing node can convert the first parameter into an intrinsic parameter matrix for transmission. The required original camera pointing range corresponds to the angle range described above.

[0349] It can be understood that the index described above is optional, that is, the central processing node can also not set the index when transmitting the data of the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, and the intrinsic parameter matrix described above.

[0350] S1506, the central processing node sends the message #1 to the optical perception node #1. Correspondingly, the optical perception node #1 receives the message #1 from the central processing node.

[0351] The message #1 is used to request the scaling parameter, and is used to request the region where each of the N point cloud data is located after being projected to a target image, the target image being an image obtained by two-dimensional perception of a target object by the optical perception node #1 and converted according to the virtual camera parameter. The message #1 includes the parameter #1. And the message #1 corresponds to the first message described above.

[0352] It can be understood that when the parameter #1 includes the required original camera pointing range, the optical perception node #1 can determine whether its pointing belongs to the original camera pointing range based on the original camera pointing range. If the optical perception node #1 belongs to the original camera pointing range, the optical perception node #1 continues the subsequent steps; if the optical perception node #1 does not belong to the original camera pointing range, the optical perception node #1 does not respond to the message #1, or the optical perception node #1 sends a signaling (such as a failure (fail)) to the central processing node to indicate that the request fails, at this time the optical perception node #1 can also send the central processing node a failure reason.

[0353] In addition, the data format of the central processing node transmitting the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, and the intrinsic parameter matrix can be as shown in Table 1.

[0354] Table 1

[0355] S1507, the optical perception node #1 determines the scaling parameter #1 and the neighborhood #i1.

[0356] That is, the optical perception node #1 can determine the scaling parameter #1 and the neighborhood #i1 of the i-th point cloud data in the N point cloud data according to the virtual camera parameter, the spatial range of each of the N point cloud data, and the normal of each of the N point cloud data, i is an integer traversing 1 to N, and details can be referred to the foregoing relevant description, which will not be repeated here.

[0357] S1508, the optical perception node #1 sends the scaling parameter #1 and the neighborhood #i1 to the central processing node. Correspondingly, the central processing node receives the scaling parameter #1 and the neighborhood #i1 from the optical perception node #1.

[0358] The data format of the optical perception node #1 sending the scaling parameter #1 and the neighborhood #i1 to the central processing node can be in the form of describing the spatial range of the point cloud data, the scaling parameter corresponding to the point cloud data, and the neighborhood (as shown in Table 2 below), or in the form of index (as shown in Table 3 below). It can be understood that sending the scaling parameter #1 and the neighborhood #i1 in the form of index can reduce the communication overhead of the optical perception node #1.

[0359] Table 2

[0360] Table 3

[0361] S1509, the central processing node sends a message #2 to the optical perception node #2. Correspondingly, the optical perception node #2 receives the message #2 from the central processing node.

[0362] The message #2 is used to request the scaling parameter, and is used to request the region where each point cloud data in the N point cloud data is located after being projected to a target image, the target image being the image converted from the image obtained by two-dimensionally perceiving the target object by the optical perception node #2 according to the virtual camera parameter. The message #2 includes a parameter #2. And the message #2 corresponds to the fourth message.

[0363] It can be understood that the message #2 and the message #1 can be the same or different, such as the content of the message #2 and the message #1 being the same and the information elements being different. The data format of the central processing node sending the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, and the data of the internal parameter matrix can be referred to the relevant description in S1506 above, which will not be repeated here.

[0364] Further, when the parameter #1 includes the original camera pointing range satisfying the requirement, the optical perception node #2 can determine whether its pointing belongs to the original camera pointing range based on the original camera pointing range, which can be understood with reference to the related description in the foregoing “S1506”, and details are not described herein again.

[0365] S15010, the optical perception node #2 calculates the scaling parameter #2 and the neighborhood #i2.

[0366] That is, the optical perception node #2 can determine the scaling parameter #2 and the neighborhood #i2 of the i-th point cloud data in the N point cloud data according to the virtual camera parameter, the spatial range of each of the N point cloud data, and the normal of each of the N point cloud data, where i is an integer traversing 1 to N, and details can be understood with reference to the foregoing related description, and details are not described herein again.

[0367] S15011, the optical perception node #2 sends the scaling parameter #2 and the neighborhood #i2 to the central processing node. Correspondingly, the central processing node receives the scaling parameter #2 and the neighborhood #i2 from the optical perception node #2.

[0368] The data format of the optical perception node #2 sending the scaling parameter #2 and the neighborhood #i2 to the central processing node is similar to the data format of the optical perception node #1 sending the scaling parameter #1 and the neighborhood #i1 to the central processing node, which can be understood with reference to the related description in the foregoing S1508, and details are not described herein again.

[0369] S15012, the central processing node determines the real point cloud data in the N point cloud data based on the neighborhood #1i, the scaling parameter #1, the neighborhood #2i, and the scaling parameter #2.

[0370] That is, the central processing node can determine the neighborhood matching degree of each of the N point cloud data based on the neighborhood #1i, the scaling parameter #1, the neighborhood #2i, and the scaling parameter #2, where N is an integer greater than 1; and filter noise in the N point cloud data based on the neighborhood matching degree of each of the N point cloud data to obtain the real point cloud data, which can be understood with reference to the related description in the foregoing “S1102”, and details are not described herein again.

[0371] S15013, the central processing node performs multi-modal fusion perception based on the real point cloud data.

[0372] The specific implementation principle of S15013 can be understood with reference to the related description of the foregoing S14010, and details are not described herein again.

[0373] It can be understood that the above S1501-S15013 can refer to the related description in the foregoing FIG. 11, and will not be described here again. It can also be understood that the order of each step of the interaction between the center processing node and the optical perception node #1 and the order of each step of the interaction between the center processing node and the optical perception node #2 are not limited by the embodiments of the present application. For example, the center processing node can first interact with the optical perception node #1 and then interact with the optical perception node #2, that is, first perform S1501, S1503, S1506-S1508, and then perform S1502, S1504, S1509-S15011. For another example, the center processing node can perform the above steps in the order.

[0374] In addition, S1501-S1504 are optional steps. For example, the optical perception node #1 and the optical perception node #2 can periodically report the capability information to the first device.

[0375] Scenario 3:

[0376] FIG. 16 is a flow diagram of a communication method provided by the embodiments of the present application. In scenario 3, the optical perception node #1 and the optical perception node #2 support providing the scaling parameter to the center processing node; the center processing node obtains the scaling parameter #1 from the optical perception node #1 and the scaling parameter #2 from the optical perception node #2 in the case that the image used to determine the neighborhood of the N point cloud data has been obtained, and determines the real point cloud data in the N point cloud data based on the scaling parameter #1 and the scaling parameter #2, and finally performs multi-modal fusion perception based on the real point cloud data. It can be understood that the above center processing node corresponds to the first device in the embodiment shown in FIG. 11, the optical perception node #1 corresponds to the second device in the embodiment shown in FIG. 11, and the optical perception node #2 corresponds to the third device in the embodiment shown in FIG. 11.

[0377] As shown in FIG. 16, the flow of the communication method is as follows:

[0378] S1601, the center processing node determines parameter #1.

[0379] The parameter #1 includes at least one of the following: the spatial range of each of the N point cloud data, or the normal line of each of the N point cloud data, N being an integer greater than 1.

[0380] S1602, the center processing node sends a message #1 to the optical perception node #1. Correspondingly, the optical perception node #1 receives the message #1 from the center processing node.

[0381] The message #1 is used to request the scaling parameter from the optical perception node #1. And the message #1 corresponds to the second message.

[0382] S1603, the optical perception node #1 calculates the scaling parameter #1.

[0383] The specific implementation of the optical perception node #1 calculating the scaling parameter according to the spatial range of each of the N point cloud data and the normal line of each of the N point cloud data can refer to the related description in the foregoing embodiment shown in FIG. 11, and will not be described here again.

[0384] S1604, the optical perception node #1 sends the scaling parameter #1 to the central processing node. Correspondingly, the central processing node receives the scaling parameter #1 from the optical perception node #1.

[0385] S1605, the central processing node sends a message #2 to the optical perception node #2. Correspondingly, the optical perception node #2 receives the message #2 from the central processing node.

[0386] The message #2 is used to request the scaling parameter from the optical perception node #2. The message #2 corresponds to the fifth message described above.

[0387] It can be understood that the message #2 and the message #1 can be the same or different, such as the content of the message #2 and the message #1 being the same and the information elements being different.

[0388] S1606, the optical perception node #2 calculates the scaling parameter #2.

[0389] The specific implementation of the optical perception node #2 calculating the scaling parameter according to the spatial range of each of the N point cloud data and the normal line of each of the N point cloud data can refer to the related description in the foregoing embodiment shown in FIG. 11, and will not be described here again.

[0390] S1607, the optical perception node #2 sends the scaling parameter #2 to the central processing node. Correspondingly, the central processing node receives the scaling parameter #2 from the optical perception node #2.

[0391] S1608, the central processing node determines the real point cloud data in the N point cloud data based on the scaling parameter #1 and the scaling parameter #2.

[0392] That is, the central processing node can determine the neighborhood matching degree of each of the N point cloud data based on the scaling parameter #1 and the scaling parameter #2, N being an integer greater than 1, and filter the noise in the N point cloud data based on the neighborhood matching degree of each of the N point cloud data to obtain the real point cloud data, which can refer to the foregoing related description and will not be described here again.

[0393] S1609, the central processing node performs multi-modal fusion perception based on the real point cloud data.

[0394] The specific implementation principle of S1609 can refer to the related description of S14010 described above, and will not be described here again.

[0395] It can be understood that S1601-S1609 can refer to the related description in the foregoing FIG. 11, and details are not described herein again. It can also be understood that the order of each step of the interaction between the center processing node and the optical perception node #1 and the order of each step of the interaction between the center processing node and the optical perception node #2 are not limited in the embodiments of the present application. For example, the center processing node can first interact with the optical perception node #1 and then interact with the optical perception node #2, that is, first perform S1602-S1604 and then perform S1605-S1607. For another example, the order of S1602, S1605, S1603, S1606, S1604, and S1607 is performed.

[0396] It can be understood that for the embodiments shown in FIGS. 14-16, the center processing node can pre-agree with the optical perception node #1 and the optical perception node #2 on the capabilities that need to be reported. At this time, the optical perception node #1 and the optical perception node #2 can indicate the state of each capability corresponding to the center processing node according to the pre-agreed capabilities, that is, at this time, each capability parameter is a parameter for indicating the state of each capability corresponding to the optical perception node #1 or the optical perception node #2.

[0397] For example, the center processing node can pre-agree with the optical perception node #1 and the optical perception node #2 on the reporting of the internal and external parameter transmission capability and the scaling factor calculation transmission capability. The internal and external parameter transmission capability is used to indicate whether to support providing related parameters and images for two-dimensional perception of an object. The scaling factor calculation and transmission capability is used to indicate whether to support providing a scaling factor. For the embodiment shown in FIG. 14, the capability parameter #1 and the capability parameter #2 can be as shown in Table 4. For the embodiment shown in FIG. 15, the capability parameter #1 and the capability parameter #2 can be as shown in Table 5. For the embodiment shown in FIG. 16, the capability parameter #1 and the capability parameter #2 can be as shown in Table 6.

[0398] Table 4

[0399] Table 5

[0400] Table 6

[0401] It can also be understood that in various embodiments of the present application, the names of each message, each information, and each parameter are only an example of a description manner, and each message, each information, and each parameter can also be replaced by any possible description manner, which is not limited. In addition, the embodiments shown in FIGS. 14-16 are only an example, and in different cases, each step in the embodiments shown in FIGS. 14-16 can also be changed accordingly, which is not limited.

[0402] The communication method provided by the embodiments of the present application is described in detail above in combination with FIG. 11. The communication apparatus for performing the communication method provided by the embodiments of the present application is described in detail below in combination with FIGS. 17-18.

[0403] FIG. 17 is a structural schematic diagram of a communication apparatus provided by the embodiments of the present application. As an example, as shown in FIG. 17, the communication apparatus 1700 includes a transceiver module 1701 and a processing module 1702. For the convenience of description, FIG. 17 only shows the main components of the communication apparatus.

[0404] The transceiver module 1701 is configured to perform the transceiving functions of the method shown in FIG. 11, and the processing module 1702 is configured to perform the functions other than the transceiving functions of the method shown in FIG. 11.

[0405] Optionally, the transceiver module 1701 can include a sending module (not shown in FIG. 17) and a receiving module (not shown in FIG. 17). The sending module is configured to implement the sending functions of the communication apparatus 1700, and the receiving module is configured to implement the receiving functions of the communication apparatus 1700.

[0406] Optionally, the communication apparatus 1700 can further include a storage module (not shown in FIG. 17), which stores programs or instructions. When the processing module 1702 executes the programs or instructions, the communication apparatus 1700 can perform the functions of the terminal device or the network device (such as the first apparatus, the second apparatus, etc.) in the method shown in FIG. 11.

[0407] It can be understood that the communication apparatus 1700 can be a terminal device or a network device, or a chip (system) or other components or assemblies that can be arranged in the terminal device or the network device, or an apparatus containing the terminal device or the network device, and the present application does not limit this.

[0408] In addition, the technical effects of the communication apparatus 1700 can refer to the technical effects of the communication method shown in FIG. 11, which will not be described here again.

[0409] FIG. 18 is a structural schematic diagram of a communication apparatus provided by the embodiments of the present application. As an example, the communication apparatus can be a terminal device or a network device, or a chip (system) or other components or assemblies that can be arranged in the terminal device or the network device. As shown in FIG. 18, the communication apparatus 1800 can include a processor 1801. Optionally, the communication apparatus 1800 can further include a memory 1802 and / or a transceiver 1803. The processor 1801 is coupled with the memory 1802 and the transceiver 1803, for example, through a communication bus.

[0410] The components of the communication apparatus 1800 are described in detail below in combination with FIG. 18.

[0411] The processor 1801 is a control center of the communication device 1800, which can be one processor or collectively refer to a plurality of processing elements. For example, the processor 1801 is one or more central processing units (CPUs), application specific integrated circuits (ASICs), or one or more integrated circuits configured to perform the functions of the embodiments of the present application, such as one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0412] Optionally, the processor 1801 can perform various functions of the communication device 1800 by running or executing software programs or instructions stored in the memory 1802, and calling data stored in the memory 1802, such as performing the above communication method.

[0413] In a specific implementation, as an embodiment, the processor 1801 can include one or more CPUs, such as CPU0 and CPU1 shown in FIG. 18.

[0414] In a specific implementation, as an embodiment, the communication device 1800 can also include a plurality of processors, such as the processor 1801 and the processor 1804 shown in FIG. 18. Each of the processors can be a single-CPU or a multi-CPU. The processor here can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0415] The memory 1802 is configured to store software programs or instructions for performing the schemes of the present application, and the processor 1801 is configured to control the execution. The specific implementation can refer to the above method embodiments, and will not be described here.

[0416] Optionally, the memory 1802 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited to this. The memory 1802 can be integrated with the processor 1801 or exist independently and be coupled to the processor 1801 through the interface circuit (not shown in FIG. 18) of the communication apparatus 1800, and the embodiments of the present application are not limited in this regard.

[0417] The transceiver 1803 is configured to communicate with other communication apparatuses. For example, the communication apparatus 1800 is a terminal, and the transceiver 1803 can be configured to communicate with a network apparatus or another terminal. For another example, the communication apparatus 1800 is a network apparatus, and the transceiver 1803 can be configured to communicate with a terminal or another network apparatus.

[0418] Optionally, the transceiver 1803 can include a receiver and a transmitter (not shown separately in FIG. 18). The receiver is configured to implement the receiving function, and the transmitter is configured to implement the transmitting function.

[0419] Optionally, the transceiver 1803 can be integrated with the processor 1801 or exist independently and be coupled to the processor 1801 through the interface circuit (not shown in FIG. 18) of the communication apparatus 1800, and the embodiments of the present application are not limited in this regard.

[0420] It can be understood that the structure of the communication apparatus 1800 shown in FIG. 18 does not constitute a limitation on the communication apparatus, and an actual communication apparatus can include more or fewer components than those shown, or combine certain components, or have different arrangement of components.

[0421] In addition, the technical effects of the communication apparatus 1800 can refer to the technical effects of the methods described in the above method embodiments, which will not be described here.

[0422] It should be appreciated that a processor in the embodiments of the present application can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0423] It should also be understood that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM) or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).

[0424] The above-described embodiments can be implemented in whole or in part by software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (e.g., infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.

[0425] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects, but it can also represent an "and / or" relationship, which can be understood in the context.

[0426] In this application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0427] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined by their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0428] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0429] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.

[0430] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0431] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0432] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.

[0433] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0434] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A communication method characterized by comprising: The method comprises: According to N point cloud data, obtain the first neighborhood i and the second neighborhood i, the N point cloud data is the three-dimensional point cloud data obtained by perceiving the target object, N is an integer greater than 1, i is an integer traversing 1 to N, the first neighborhood i is the area where the i-th point cloud data in the N point cloud data is located after being projected onto the first image, and the second neighborhood i is the area where the i-th point cloud data is located after being projected onto the second image, the first image and the second image both include the perception result of the target object, the imaging plane of the first image and the imaging plane of the second image are parallel, and the first image and the second image are different two-dimensional images; According to the first neighborhood i and the second neighborhood i, multi-modal fusion perception is performed.

2. The method of claim 1, wherein, According to the N point cloud data, the first neighborhood i is obtained, comprising: Sending a first message to a second device, the first message being used to request the area where each point cloud data in the N point cloud data is located after being projected onto a target image, the target image being an image obtained by the second device performing two-dimensional perception on the target object and then being converted according to a first parameter, and the first parameter being an internal and external parameter associated with the target image; Receiving the first neighborhood i from the second device.

3. The method of claim 2, wherein, The first message comprises at least one of the following: the spatial range of each of the N point cloud data, the normal of each of the N point cloud data, or the first parameter, the spatial range of each of the N point cloud data being used to indicate the size and / or shape of the space where each of the N point cloud data is located, and the normal of each of the N point cloud data being used to determine the rotation matrix corresponding to each of the N point cloud data.

4. The method according to claim 2 or 3, characterized in that, The first neighborhood i is the area where the i-th point cloud data is located after being projected onto the first image, comprising: the first neighborhood i is the area where the first space i where the i-th point cloud data is located is projected onto the first image; the second neighborhood i is the area where the i-th point cloud data is located after being projected onto the second image, comprising: the second neighborhood i is the area where the first space i is projected onto the second image; the first message is further used to request a scaling parameter, the scaling parameter being the scaling ratio between the object performing two-dimensional perception and the imaging result of the object; and the receiving of the first neighborhood i from the second device comprises: Receiving the first neighborhood i and a first scaling parameter from the second device, the first scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the first image; The method further comprises: According to the first neighborhood i, the first scaling parameter and a second scaling parameter, determining a third neighborhood i, the second scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the second image, and the third neighborhood i comprising the same number of pixel points as the second neighborhood i; The multi-modal fusion perception according to the first neighborhood i and the second neighborhood i comprises: According to the third neighborhood i and the second neighborhood i, multi-modal fusion perception is performed.

5. The method of claim 4, wherein, Before the first message is sent to the second device, the method further includes: receiving a first capability parameter from the second device, the first capability parameter being used to indicate that the second device supports providing the scaling parameter; the first message sent to the second device includes: According to the first capability parameter, the first message is sent to the second device.

6. The method according to any one of claims 2-5, characterized in that, The first message is also used to indicate an angle range, the pointing of the second device belonging to the angle range, and the two-dimensional image obtained by the device belonging to the angle range being capable of being used to determine the two-dimensional image for the N point cloud data projection.

7. The method according to any one of claims 1-3, characterized in that, The first neighborhood i is the area where the i-th point cloud data is located after being projected onto the first image, including: the first neighborhood i is the area where the first space i is projected onto the first image; the second neighborhood i is the area where the i-th point cloud data is located after being projected onto the second image, including: the second neighborhood i is the area where the first space i is projected onto the second image; the method further includes: According to the first neighborhood i, the first scaling parameter and the second scaling parameter, a third neighborhood i is determined, the first scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the first image, the second scaling parameter being used to indicate the scaling ratio between the target object and the imaging result of the target object in the second image, the third neighborhood i including the same number of pixel points as the second neighborhood i; According to the first neighborhood i and the second neighborhood i, multi-modal fusion perception is performed, including: According to the third neighborhood i and the second neighborhood i, multi-modal fusion perception is performed.

8. The method of claim 7, wherein, The method further includes: sending a second message to the second device, the second message being used to request a scaling parameter, the scaling parameter being the scaling ratio between the object for two-dimensional perception and the imaging result of the object; receiving the first scaling parameter from the second device.

9. The method of claim 8, wherein, The second message includes at least one of the following: the spatial range of each of the N point cloud data, or the normal line of each of the N point cloud data, the spatial range of each of the N point cloud data being used to indicate the size and / or shape of the space where each point cloud data in the N point cloud data is located, and the normal line of each of the N point cloud data being used to determine the rotation matrix corresponding to each of the N point cloud data.

10. The method according to claim 8 or 9, characterized in that, Before the second message is sent to the second device, the method further includes: receiving a second capability parameter from the second device, the second capability parameter being used to indicate that the second device supports providing the scaling parameter; the second message sent to the second device includes: According to the second capability parameter, the second message is sent to the second device.

11. The method according to any one of claims 1-10, characterized in that, The first neighborhood is a region where the i-th point cloud data is located after being projected onto the first image, and the first neighborhood i is a region where the first space i where the i-th point cloud data is located is projected onto the first image; the second neighborhood i is a region where the i-th point cloud data is located after being projected onto the second image, and the second neighborhood is a region where the first space i is projected onto the second image. The first space i where the i-th point cloud data is located belongs to a first effective perception space, the first effective perception space is an intersection of a viewing pyramid of a second virtual device and a viewing pyramid of a second device, the second virtual device is associated with the first image, the first image is determined according to an image obtained by two-dimensional perception of the second device, and the second virtual device has the same position as the second device; and / or, The first space i where the i-th point cloud data is located belongs to a second effective perception space, the second effective perception space is an intersection of a viewing pyramid of a third virtual device and a viewing pyramid of a third device, the third virtual device is associated with the second image, the second image is determined according to an image obtained by two-dimensional perception of the third device, and the third virtual device has the same position as the third device.

12. A communication method, comprising: The method comprises: receiving a first message from a first device, the first message being used to request a region where each point cloud data in N point cloud data is located after being projected onto a target image, the N point cloud data being three-dimensional point cloud data obtained by perceiving a target object, the target image being an image obtained by converting an image obtained by two-dimensional perception of the target object according to a first parameter, the first parameter being internal and external parameters of the target image, and N being an integer greater than 1; sending a first neighborhood i to the first device according to the first message, the first neighborhood i being a region where i-th point cloud data in the N point cloud data is located after being projected onto a first image, the first image comprising a perception result of the target object, and i being an integer traversing 1 to N.

13. The method of claim 12, wherein, The first message comprises at least one of the following: a spatial range of each of the N point cloud data, a normal line of each of the N point cloud data, or the first parameter, the spatial range of each of the N point cloud data being used to indicate a size and / or shape of a space where each of the N point cloud data is located, and the normal line of each of the N point cloud data being used to determine a rotation matrix corresponding to each of the N point cloud data.

14. The method according to claim 12 or 13, characterized in that, The first neighborhood i is a region where the i-th point cloud data is located after being projected onto the first image, and the first neighborhood i is a region where the first space i where the i-th point cloud data is located is projected onto the first image, and the first message is further used to request a scaling parameter, the scaling parameter being a scaling ratio between an object performing two-dimensional perception and an imaging result of the object; and the sending of the first neighborhood i to the first device comprises: sending the first neighborhood i and a first scaling parameter to the first device, the first scaling parameter being used to indicate a scaling ratio between the target object and an imaging result of the target object in the first image.

15. The method according to any one of claims 12-14, characterized in that, Before the receiving the first message from the first device, the method further comprises: sending a first capability parameter to the first device, the first capability parameter being used to indicate that providing the scaling parameter is supported.

16. The method according to any one of claims 12-15, characterized in that, The first message is further used to indicate an angle range, a two-dimensional image obtained by a device belonging to the angle range and sensing the target object can be used to determine a two-dimensional image for the N point cloud data projection.

17. A method of communication, comprising: The method comprises: receiving a second message from a first device, the second message being used to request a scaling parameter, the scaling parameter being a scaling ratio between an object for two-dimensional sensing and an imaging result of the object; sending a first scaling parameter to the first device according to the second message, the first scaling parameter being used to indicate a scaling ratio between a target object and an imaging result of the target object in a first image, the first image comprising a sensing result of the target object.

18. The method of claim 17, wherein, The second message comprises at least one of the following: a spatial range of each of the N point cloud data, or a normal line of each of the N point cloud data, the spatial range of each of the N point cloud data being used to indicate a size and / or shape of a space where each of the N point cloud data is located, and the normal line of each of the N point cloud data being used to determine a rotation matrix corresponding to each of the N point cloud data.

19. The method of claim 17 or 18, wherein, Before the receiving the second message from the first device, the method further comprises: sending a second capability parameter to the first device, the second capability parameter being used to indicate that providing the scaling parameter is supported.

20. A communications device, characterized by The communication device is configured to perform the communication method according to any one of claims 1-19.

21. A communications device, characterized by comprises: a processor and a memory; The memory is configured to store a computer program or instructions, when the processor executes the program or instructions, so as to make the communication device perform the communication method according to any one of claims 1-19.

22. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a computer program or instructions, when the computer program or instructions are run on the communication device, so as to make the communication device perform the method according to any one of claims 1-19.

23. A computer program product, characterised in that, The computer program product comprises a computer program or instructions, when the computer program or instructions are run on the communication device, so as to make the method according to any one of claims 1-19 be performed.

Citation Information

Patent Citations

  • Camera, camera set, its control method, apparatus and system

    CN101404725A

  • Multi-sensor data fusion method, device and system

    CN115908578A

  • Fusion sensing method and system based on multi-class sensor information

    CN116245961A

  • Target detection method based on fusion of vision, lidar, and millimeter wave radar

    US20220277557A1