Sensing method and communication apparatus
By combining information from 2D and 3D perception, and utilizing key point back projection and surface information, the problems of lack of depth information in 2D perception and susceptibility to noise interference in 3D perception are solved, achieving more efficient perception performance and accuracy.
Patent Information
- Application Number
- PCT/CN2025/111241
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-19
- Filing Date
- 2025-07-29
- Publication Date
- 2026-02-26
AI Technical Summary
Existing 2D perception methods can only generate planar images and lack depth information, while 3D perception is susceptible to noise and interference, leading to a decline in perception performance.
By combining information from 2D and 3D perception, key points from the 2D perception results are back-projected into 3D space, and combined with information from the surface where the target is located, the shape and position of the target are determined, thereby improving perception performance.
It improved the perception coverage and accuracy, restored the depth information and two-dimensional shape of the target, and optimized the perception results.
Smart Images

Figure CN2025111241_26022026_PF_FP_ABST
Abstract
Description
Method and communication apparatus for perception
[0001] This application claims priority to the Chinese Patent Application No. 202411142888.2, filed on August 19, 2024, and entitled "Method and communication apparatus for perception", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of communication, and more particularly, to a method and communication apparatus for perception. BACKGROUND
[0003] Perception can be implemented in a variety of different ways, for example, it can be divided into two categories: two-dimensional (2D) perception and three-dimensional (3D) perception. The 2D perception process is relatively simple, but based on 2D perception, only a planar image can be generated, i.e., only width and height information is included, without depth information. While 3D perception can provide depth information, it is susceptible to noise and interference, thereby affecting the performance of perception. Therefore, how to improve the performance of perception needs further study. SUMMARY
[0004] The present application provides a method and communication apparatus for perception, which can improve the performance of perception and optimize the result of perception.
[0005] In a first aspect, a method for perception is provided. The method is applied to a first device, or the method is performed by the first device. In the absence of special description, "the first device" in the present application can refer to the first device itself (e.g., a 2D perception node, a 3D perception node), a component (e.g., a processor, an apparatus, a circuit, a chip, or a chip system, etc.) in the first device, or a logic module or software capable of realizing all or part of the functions of the first device.
[0006] The method comprises: obtaining first information, the first information being obtained based on a 3D perception node perceiving a target, the first information indicating information of a surface where the target is located; obtaining second information, the second information being obtained based on a 2D perception node perceiving the target, the second information indicating information of a projection straight line corresponding to each key point in a plurality of key points of a 2D perception result of the target, the key point being obtained by inversely projecting the key point to the 3D space; and determining information of the target based on the first information and the second information.
[0007] In some implementations of the first aspect, the information of the target comprises a shape of the target and / or a position of the target.
[0008] In the present application, the "target" can also be referred to as the "perceived target", and the two terms can be used interchangeably.
[0009] In the technical solution, the intersection information of the projection straight line and the surface where the perception target is located can be obtained based on the projection straight line of the key point in the 2D perception result of the perception target and the surface where the perception target is located, so that the related information of the perception target, such as the shape of the perception target and the position of the perception target, is obtained. The method combines the perception information of 2D perception and the perception information of 3D perception, and compared with 2D perception or 3D perception alone, the perception performance can be improved and the perception result can be optimized. For example, compared with 2D perception alone, the depth information of the perception target can be recovered, and compared with 3D perception alone, the contour and the two-dimensional shape of the perception target can be obtained based on the 2D perception result, so that the 3D perception coverage rate and the perception accuracy are improved.
[0010] In some implementations of the first aspect, determining the information of the target based on the first information and the second information comprises: determining the intersection of the projection straight line corresponding to the plurality of key points and the surface where the perception target is located based on the first information and the second information; and determining the information of the target based on the intersection.
[0011] In some implementations of the first aspect, the projection straight line corresponding to the first key point in the plurality of key points comprises a first point and a second point, the first point is any point obtained by inversely projecting the first key point into the 3D space, and the second point is a point where the 2D perception node is located.
[0012] In the technical solution, an implementation of determining the projection straight line corresponding to the first key point is given. Two points can determine a straight line. Therefore, the projection straight line corresponding to the first key point can be determined based on the coordinates of the first point in the 3D space and the coordinates of the 2D perception node in the 3D space.
[0013] In some implementations of the first aspect, the first information comprises at least one of the following: an equation of the surface where the target is located, a normal set corresponding to a discrete plane of the surface where the target is located, or a 3D perception result of the target.
[0014] It can be understood that the first information can directly or indirectly indicate the information of the surface where the target is located. The first information indicating the equation or the normal of the surface where the perception target is located can be regarded as direct indication, and the first information indicating the 3D perception result can be regarded as indirect indication.
[0015] In some implementations of the first aspect, the second information comprises a 2D perception result of the target, and / or an equation of the projection straight line corresponding to the plurality of key points.
[0016] It can be understood that the second information can directly or indirectly indicate the information of the surface where the target is located. The second information indicating the equation of the projection straight line corresponding to the key point can be regarded as direct indication, and the second information indicating the 2D perception result can be regarded as indirect indication.
[0017] In some implementations of the first aspect, the obtaining the second information comprises: receiving the second information from N1 2D perception nodes, N1 being a positive integer.
[0018] In some implementations of the first aspect, the N1 2D perception nodes are 2D perception nodes satisfying a first condition, the first condition being a predefined condition or an indicated condition.
[0019] It can be understood that the perception quality of the 2D perception node affects the final joint perception result, and therefore, the 2D perception node with better perception quality can be selected based on the first condition to perceive the perception target, so as to further optimize the joint perception result.
[0020] In some implementations of the first aspect, the first condition is that the perception error of each 2D perception node participating in the perception is less than a first threshold, or the first condition is that the 2D perception nodes participating in the perception are N1 nodes with the best ranking in the candidate 2D perception nodes in a descending order of the perception error, wherein the perception error of each 2D perception node is determined based on at least one parameter in the perception parameter of each 2D perception node.
[0021] In some implementations of the first aspect, the perception parameter of the 2D perception node comprises at least one of the following parameters: a focal length of the camera, a field of view angle of the camera, a resolution of the camera, a position of the camera, or a pointing direction of the camera.
[0022] In some implementations of the first aspect, the perception error error satisfies the following formula:
[0023] wherein FoV is the field of view angle of the camera, a is an included angle between the pointing direction of the optical axis of the camera and the plane of the target, r is a physical distance between the intersection of the plane of the target and the pointing direction of the optical axis of the camera and the position of the camera, resolution is the resolution of the camera, and a is determined based on the first information.
[0024] In some implementations of the first aspect, the method further comprises: sending third information to at least one candidate 2D perception node, the third information indicating that if the corresponding candidate 2D perception node satisfies the first condition, the corresponding second information is reported.
[0025] In the above technical solution, the first device can instruct the 2D perception node to determine whether the first condition is satisfied, and if so, report the corresponding second information.
[0026] In some implementations of the first aspect, the method further includes: determining a candidate 2D perception node that satisfies the first condition from the at least one candidate 2D perception node; and sending a first request message to the candidate 2D perception node that satisfies the first condition, the first request message being used to request reporting of the corresponding second information.
[0027] In the technical solution described above, the first device itself determines whether the 2D perception node satisfies the first condition, and if so, instructs the 2D perception node to report the corresponding second information.
[0028] In some implementations of the first aspect, the method further includes: sending a second request message to the L1 perception nodes, the second request message being used to request the perception nodes to report corresponding perception capabilities, L1 being a positive integer; and receiving a second request response message from at least one perception node of the L1 perception nodes, the second request response message including fourth information, the fourth information being used to indicate the perception capability of the corresponding perception node, wherein the N1 2D perception nodes are at least one of the at least one perception node that has a 2D perception capability.
[0029] In the technical solution described above, the first device can request the perception nodes to report capabilities, so that the first device can select appropriate perception nodes for joint perception based on its own needs.
[0030] In some implementations of the first aspect, the fourth information includes specific information of at least one of the following of the corresponding perception node: supported perception dimension, supported perception modality, and supported perception parameter.
[0031] In some implementations of the first aspect, the fourth information does not include the required perception capability information, and the method further includes: sending a third request message to the at least one candidate 2D perception node, the third request message being used to request the 2D perception node to report the required perception capability information, the at least one candidate 2D perception node being at least one of the at least one perception node that has a 2D perception capability, and the N1 2D perception nodes being at least one of the at least one candidate 2D perception node; and receiving a third request response message from the at least one candidate 2D perception node, the third request response message including the required perception capability information of the corresponding 2D perception node.
[0032] In some implementations of the first aspect, the first information is obtained by: receiving the first information from the N2 3D perception nodes, N2 being a positive integer.
[0033] In some implementations of the first aspect, the N2 3D perception nodes are 3D perception nodes that satisfy a second condition, the second condition being a predefined condition or an indicated condition.
[0034] In some implementations of the first aspect, the second condition is that a surface estimation error of each 3D perception node of the participating 3D perception nodes is less than a second threshold, or the second condition is that the participating 3D perception nodes are top N2 nodes in the candidate 3D perception nodes in an ascending order of the surface estimation error.
[0035] In some implementations of the first aspect, the method further includes: sending, to the at least one candidate 3D perception node, fifth information indicating that the corresponding candidate 3D perception node reports the corresponding first information if the second condition is met.
[0036] In some implementations of the first aspect, the method further includes: determining the 3D perception nodes of the at least one candidate 3D perception node that meet the second condition; and sending, to the 3D perception nodes that meet the second condition, a fourth request message for requesting the 3D perception nodes to report the corresponding first information.
[0037] In some implementations of the first aspect, the method further includes: sending, to the L2 perception nodes respectively, a fifth request message for requesting the perception nodes to report corresponding perception capabilities, L2 being a positive integer; and receiving a fifth request response message from at least one perception node of the L2 perception nodes, the fifth request response message including sixth information for indicating the perception capabilities of the perception node, wherein the N2 3D perception nodes are at least one of the at least one perception node that has a 3D perception capability.
[0038] In some implementations of the first aspect, the sixth information includes specific information of at least one of the following of the corresponding perception node: supported perception dimension, supported perception modality, and supported perception parameter.
[0039] In some implementations of the first aspect, the sixth information does not include the required perception capability information for calculation, and the method further includes: sending, to the at least one candidate 3D perception node, a sixth request message for requesting the 3D perception node to report the required perception capability information, the at least one candidate 3D perception node being at least one of the at least one perception node that has a 3D perception capability, and the N2 3D perception nodes being at least one of the at least one candidate 3D perception node; and receiving a sixth request response message from the at least one candidate 3D perception node, the sixth request response message including the required perception capability information of the 3D perception node.
[0040] In some implementations of the first aspect, the perception dimension includes 2D perception and / or 3D perception, the perception modality of the 2D perception includes optical perception and / or radar perception, and the perception modality of the 3D perception includes radio frequency perception and / or computer tomography.
[0041] In a second aspect, a method for sensing is provided. The method can be applied to, or performed by, a first device. In the absence of specific description, the first device in the present application can refer to the first device itself (e.g., a 2D sensing node, a 3D sensing node), a component (e.g., a processor, a device, a circuit, a chip, or a chip system, etc.) in the first device, or a logic module or software capable of realizing all or part of the functions of the first device.
[0042] The method comprises: sending a first request message to L sensing nodes, the first request message being used to request the sensing nodes to report corresponding sensing capabilities, L being a positive integer; and receiving a first request response message from at least one of the L sensing nodes, the first request response message comprising third information, the third information being used to indicate the sensing capabilities of the sensing nodes.
[0043] In the above technical solution, the first device can request the sensing nodes to report capabilities, so that the first device can select appropriate sensing nodes for joint sensing based on its own needs.
[0044] For example, the third information comprises specific information of at least one of the following for the corresponding sensing node: supported sensing dimension, supported sensing modality, or supported sensing parameter.
[0045] In some implementations of the second aspect, the third information does not comprise the sensing capability information required for calculation, and the method further comprises: sending a second request message to at least one candidate sensing node, the second request message being used to request the sensing nodes to report the required sensing capability information, the at least one candidate sensing node being at least one of the at least one sensing node; and receiving a second request response message from the at least one candidate sensing node, the second request response message comprising the required sensing capability information of the corresponding sensing node.
[0046] It can be understood that if the third information does not comprise the capability information of the corresponding sensing node required for subsequent calculation by the first device, for example, in the method for sensing provided in the present application, the first device can calculate the sensing error based on the 2D sensing node sensing parameter to determine whether the 2D sensing node meets the first condition. If the third information only reports the sensing dimension of the corresponding sensing node, the first device cannot determine whether the 2D sensing node meets the first condition. Then, the first device can obtain the required related information of the sensing capability of the corresponding sensing node again based on the method.
[0047] In some implementations of the second aspect, the sensing dimension comprises 2D sensing and / or 3D sensing, the sensing modality of the 2D sensing comprises optical sensing and / or radar sensing, and the sensing modality of the 3D sensing comprises radio frequency sensing and / or computed tomography.
[0048] In some implementations of the second aspect, the perception parameters of the 2D perception node include at least one of: a focal length of the camera, a field of view angle, a resolution, a position, or a pointing direction.
[0049] In some implementations of the second aspect, the perception parameters of the 3D perception node include a surface estimation error.
[0050] It can be understood that the capability reporting scheme described in the second aspect can be executed alone or in combination with the joint perception method described in the first aspect. For example, if combined, in some implementations of the second aspect, the method further includes: obtaining first information, the first information being obtained based on a three-dimensional (3D) perception node perceiving the target, the first information indicating information of a surface on which the target is located; obtaining second information, the second information being obtained based on a two-dimensional (2D) perception node perceiving the target, the second information indicating information of a projection straight line corresponding to each key point in a 2D perception result of the target when the key point is inversely projected into a 3D space; and determining information of the target based on the first information and the second information.
[0051] In some implementations of the second aspect, the method further includes: obtaining first information, the first information being obtained based on a 3D perception node perceiving the target, the first information indicating information of a surface on which the target is located; receiving second information from N1 2D perception nodes, N1 being a positive integer, the second information being obtained based on the 2D perception nodes perceiving the target, the second information indicating information of a projection straight line corresponding to each key point in a 2D perception result of the target in a 3D space, the N1 2D perception nodes being at least one node in at least one candidate perception node, the at least one candidate perception node being at least one node in at least one perception node having a 2D perception capability; and determining information of the target based on the first information and the second information.
[0052] In some implementations of the second aspect, the information of the target includes a shape of the target and / or a position of the target.
[0053] For beneficial effects of the second aspect, refer to the description of the first aspect, which will not be repeated here.
[0054] In some implementations of the second aspect, the N1 2D perception nodes are 2D perception nodes satisfying a first condition, the first condition being a predefined condition or an indicated condition.
[0055] In some implementations of the second aspect, the first condition is that a perception error of each 2D perception node participating in the perception is less than a first threshold, or the first condition is that the 2D perception nodes participating in the perception are a top N1 number of nodes in the candidate 2D perception nodes in an ascending order of the perception error, where the perception error of each 2D perception node is determined based on at least one of the perception parameters of each 2D perception node.
[0056] In some implementations of the second aspect, the perception error error satisfies the following formula:
[0057] wherein FoV is a field of view of the camera, a is an included angle between a pointing direction of an optical axis of the camera and a plane of the target, r is a physical distance between an intersection of the plane of the target and the pointing direction of the optical axis of the camera and a position of the camera, resolution is a resolution of the camera, and a is determined based on the first information.
[0058] In some implementations of the second aspect, the method further includes:
[0059] receiving first information from N2 3D perception nodes, the first information being obtained based on the 3D perception nodes perceiving the target, the first information indicating information of a surface on which the target is located, N2 being a positive integer, and the N2 3D perception nodes being at least one node in at least one candidate perception node, the at least one candidate perception node being at least one node in at least one perception node having a 3D perception capability;
[0060] determining second information, the second information being obtained based on the 2D perception node perceiving the target, the second information indicating information of a projection straight line corresponding to each key point in a 2D perception result of the target in the 3D space, and determining the information of the target based on the first information and the second information.
[0061] In some implementations of the second aspect, the N2 3D perception nodes are 3D perception nodes satisfying a second condition, the second condition being a predefined condition or an indicated condition.
[0062] In some implementations of the second aspect, the second condition is that a surface estimation error of each 3D perception node participating in the perception is less than a second threshold, or the second condition is that the 3D perception nodes participating in the perception are a top N2 number of nodes in the candidate 3D perception nodes in an ascending order of the surface estimation error.
[0063] In some implementations of the second aspect, the projection straight line corresponding to the first key point in the plurality of key points includes a first point and a second point, the first point being any point obtained by inversely projecting the first key point into the 3D space, and the second point being a point at which the 2D perception node is located.
[0064] In some implementations of the second aspect, determining the information of the target based on the first information and the second information comprises: determining an intersection of a straight line on which the plurality of key points are located and a surface on which the target is located based on the first information and the second information; and determining the information of the target based on the intersection.
[0065] In some implementations of the second aspect, the information of the target comprises a shape and a position of the target.
[0066] In some implementations of the second aspect, the first information comprises at least one of: an equation of a surface on which the target is located, a set of normal vectors corresponding to discrete planes of the surface on which the target is located, or a 3D perception result of the target.
[0067] In some implementations of the second aspect, the second information comprises a 2D perception result of the target, and / or an equation of a projection straight line corresponding to the plurality of key points.
[0068] In a third aspect, a communication apparatus is provided. The apparatus is configured to perform the method in any of the above aspects and / or implementations.
[0069] In one implementation, the apparatus is a 2D perception node or a 3D perception node or an intermediate node. When the apparatus is a 2D perception node or a 3D perception node or an intermediate node, the transceiver unit can be a transceiver, or an input / output interface, or a communication interface; and the processing unit can be at least one processor. Optionally, the transceiver is a transceiver circuit. Optionally, the input / output interface is an input / output circuit.
[0070] In another implementation, the apparatus is a chip, chip system or circuit for a 2D perception node or a 3D perception node or an intermediate node. When the apparatus is a chip, chip system or circuit for a 2D perception node or a 3D perception node or an intermediate node, the transceiver unit can be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin or related circuitry, etc. on the chip, chip system or circuit; and the processing unit can be at least one processor, a processing circuit or a logic circuit, etc.
[0071] In a fourth aspect, a communication apparatus is provided. The apparatus comprises: a memory configured to store a program; and at least one processor configured to execute the computer program or instructions stored in the memory to perform the method in any of the above aspects and / or implementations.
[0072] In an implementation, the apparatus is a 2D or 3D perception node or an intermediate node. For example, the intermediate node can be a base station or a terminal, or the intermediate node can be a management unit in a core network, such as a perception management unit.
[0073] In another implementation, the apparatus is a chip, chip system or circuit for a 2D or 3D perception node or an intermediate node.
[0074] In a fifth aspect, a communication apparatus is provided, which includes at least one processor and a communication interface, the at least one processor being configured to acquire a computer program or instructions stored in a memory via the communication interface, so as to execute the method provided in any one of the aspects or the implementations thereof. The communication interface can be implemented by hardware or software.
[0075] In an implementation, the apparatus further includes the memory.
[0076] In a sixth aspect, a processor is provided, which is configured to execute the method provided in the aspects.
[0077] For the sending and acquiring / receiving operations of the processor, if no special description is provided, or if it does not conflict with the actual role or inherent logic in the related description, it can be understood as the output and receiving, input operations of the processor, or the sending and receiving operations performed by the radio frequency circuit and the antenna, which are not limited in the present application.
[0078] In a seventh aspect, a computer readable storage medium is provided, which stores program codes for execution by an apparatus, and the program codes include codes for executing the method provided in any one of the aspects or the implementations thereof.
[0079] In an eighth aspect, a computer program product is provided, which includes instructions, and when the computer program product is run on a computer, the computer is caused to execute the method provided in any one of the aspects or the implementations thereof.
[0080] In a ninth aspect, a chip is provided, which includes a processor and a communication interface, the processor being configured to read instructions stored in a memory via the communication interface, and execute the method provided in any one of the aspects or the implementations thereof. The communication interface can be implemented by hardware or software.
[0081] Optionally, as an implementation, the chip further includes the memory, and the memory stores computer programs or instructions, and the processor is configured to execute the computer programs or instructions stored in the memory, and when the computer programs or instructions are executed, the processor is configured to execute the method provided in any one of the aspects or the implementations thereof.
[0082] When the method provided in the present application is executed by a chip, the present application does not limit the number of chips for implementing the method of the present application, for example, the method can be executed by one chip, or two or more chips. Moreover, when the number of chips for implementing the method of the present application is two or more, the chip manufacturers are not limited, and can be the same manufacturer or different manufacturers.
[0083] In a tenth aspect, a computer program is provided, which, when running on a computer, causes the method provided in any one of the above aspects or the implementation manner thereof to be executed.
[0084] In an eleventh aspect, a communication system is provided, which includes at least one of the 2D sensing node, the 3D sensing node, or the intermediate node described above. BRIEF DESCRIPTION OF DRAWINGS
[0085] FIG. 1 is a schematic diagram of a communication system to which embodiments of the present application are applicable.
[0086] FIG. 2 is a schematic diagram of four coordinate systems in an optical sensing scenario.
[0087] FIG. 3 is a schematic diagram of projection of a target from a camera coordinate system to a 2D imaging plane.
[0088] FIG. 4 is a schematic diagram of a possible sensing scenario to which embodiments of the present application are applicable.
[0089] FIG. 5 is a schematic diagram of a 3D point cloud sensed by a radio frequency sensing node.
[0090] FIG. 6 is a schematic flowchart of a sensing method 600 provided in the present application.
[0091] FIG. 7 is a schematic diagram of parameters related to sensing errors.
[0092] FIG. 8 is a schematic flowchart of sensing capability reporting provided in the present application.
[0093] FIGS. 9 to 13 are schematic flowcharts of possible sensing methods provided in the present application.
[0094] FIGS. 14 and 15 are schematic block diagrams of communication apparatuses provided in embodiments of the present application. DETAILED DESCRIPTION
[0095] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0096] In order to facilitate understanding of the above embodiments provided in the present application, the following points are explained:
[0097] 1) In the present application, the terms and / or descriptions among different embodiments are consistent and can be referred to each other if there is no special description and logical conflict. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0098] 2) In the present application, the terms "system" and "network" can be used interchangeably. "At least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following cases: A exists alone, A and B exist together, B exists alone, where A and B can be singular or plural. In the literal description of the present application, the character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single item or multiple items. For example, at least one of a, b and c can represent: a, or b, or c, or a and b, or a and c, or b and c, or a, b and c. Where a, b and c can be single or multiple. "Comma" in A, B, C or D means or.
[0099] 3) The ordinal numbers such as "first", "second" and the like mentioned in the embodiments of the present application are used to distinguish a plurality of objects, and are not used to limit the size, content, order, time sequence, priority or importance of the plurality of objects. For example, the first indication information and the second indication information can be the same information or different information, and such names do not mean that the contents, sizes, application scenarios, sending / receiving ends, priorities or importance of the two messages are different. In addition, the numbering of steps in each embodiment introduced in the present application is only to distinguish different steps, and is not used to limit the order between the steps.
[0100] 4) In the present application, the descriptions such as "when", "in the case of" and "if" all mean that the device will make corresponding processing under certain objective circumstances, not limited to time, and also does not require the device to have a judgment action when implemented, nor does it mean that there are other limitations.
[0101] 5) In the present application, "indication" or "for indicating" can include direct indication and indirect indication. When describing that certain indication information is used to indicate A, it can include that the indication information directly indicates A or indirectly indicates A, and it does not mean that A must be carried in the indication information.
[0102] The indication manner involved in the embodiments of the present application should be understood as covering various methods that can enable the to-be-indicated party to know the to-be-indicated information. The to-be-indicated information can be sent together as a whole, or can be sent separately into multiple sub-information, and the sending period and / or sending occasion of the sub-information can be the same or different, and the present application does not limit the sending method.
[0103] The "indication information" in the embodiments of the present application can be explicit indication, i.e., directly indicated through signaling, or obtained according to the indicated parameters, combined with other rules or combined with other parameters or through derivation. Or it can be implicit indication, i.e., obtained according to rules or relationships, or according to other parameters, or through derivation. The present application does not make specific limitations on this.
[0104] 6) The "protocol" involved in the present application can refer to a standard protocol in the communication field, which can include, for example, a fourth generation (4th generation, 4G) network, a fifth generation (5th generation, 5G) network protocol, an NR protocol, a 5.5G network protocol, and a related protocol applied to a future communication network, and the present application does not limit this.
[0105] 7) In the present application, "communication" can also be described as "data transmission", "information transmission", "data processing", etc. "Transmission" includes "sending" and "receiving".
[0106] 8) In the present application, "sending information" in the present application can be understood as a device sending information to another device, or can also be understood as a logical module in a device sending information to another logical module. For example, "3D sensing node sending information" can be understood as the 3D sensing node sending information to another device (such as a 2D sensing node), or can be understood as a logical module 1 in the 3D sensing node sending information to a logical module 2 in the 2D sensing node.
[0107] "Receiving information" in the present application can be understood as a device receiving information from another device, or can also be understood as a logical module in a device receiving information from another logical module. For example, "3D sensing node receiving information" can be understood as the 3D sensing node receiving information from another device (such as a 2D sensing node), or can be understood as a logical module 1 in the 3D sensing node receiving information from a logical module 2 in the 3D sensing node.
[0108] In this application, "sending information to" or related expressions in the figures can be understood as the destination of the information is the 2D perception node. It can include directly or indirectly sending information to the 2D perception node. "Receiving information from" or "receiving information from" or "receiving information sent by" or related expressions in the figures can be understood as the source of the information is the 2D perception node, which can include directly or indirectly receiving information from the 2D perception node. The information between the source and the destination of the information sending may be necessary processing, such as format change, etc., but the destination can understand the effective information from the source. Similar expressions in this application can be similarly understood, and will not be repeated here.
[0109] 9) The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices.
[0110] 10) The blocks or arrows shown by dashed lines in the schematic diagrams in the drawing part of the specification of this application represent optional steps or optional modules.
[0111] The technical solutions of the embodiments of the present application can be applied to various communication systems, for example: long term evolution (LTE) system, LTE frequency division duplex (FDD) system, LTE time division duplex (TDD), 5th generation (5G) system or new radio (NR), and future communication systems, vehicle-to-X (V2X), which can include vehicle to network (V2N), vehicle to vehicle (V2V), vehicle to infrastructure (V2I), vehicle to pedestrian (V2P), etc., long term evolution-vehicle (LTE-V), Internet of Vehicles, machine type communication (MTC), Internet of Things (IoT), long term evolution-machine (LTE-M), machine to machine (M2M), etc.
[0112] FIG. 1 is a schematic diagram of a communication system to which embodiments of the present application are applicable. It can be appreciated that FIG. 1 is a possible, non-restrictive, schematic diagram of a system. As shown in FIG. 1, the communication system 10 includes a radio access network (RAN) 100 and a core network (CN) 200. The RAN 100 includes at least one RAN node (e.g., 110a and 110b in FIG. 1, collectively referred to as 110) and at least one terminal (e.g., 120a-120j in FIG. 1, collectively referred to as 120). The RAN 100 can further include other RAN nodes, such as a wireless relay device and / or a wireless backhaul device (not shown in FIG. 1), etc. The terminal 120 is connected to the RAN node 110 in a wireless manner. The RAN node 110 is connected to the core network 200 in a wireless or wired manner. The core network device in the core network 200 and the RAN node 110 in the RAN 100 can be different physical devices respectively, or can be the same physical device integrated with the logical functions of the core network and the radio access network.
[0113] The RAN 100 can be a 3rd generation partnership project (3GPP) related cellular system, e.g., a 4G, 5G mobile communication system, or a future evolvement of the system. The RAN 100 can also be an open radio access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. The RAN 100 can also be a communication system that combines two or more of the above systems.
[0114] The RAN nodes 110, which can also be referred to as access network devices, RAN entities or access nodes or network devices, etc., form part of the communication system 100 and are responsible for enabling wireless access to the communication system 100 for terminals 120. The RAN nodes 110 in the communication system 100 can be the same type of nodes or different types of nodes. In the present application, the RAN nodes and network devices can be replaced by each other unless specifically stated otherwise. In some scenarios, the roles of the RAN nodes 110 and the terminals 120 are relative, e.g., the network element 120i in Figure 1 can be a helicopter or a drone, which can be configured as a mobile base station. For those terminals 120j that access the RAN 100 through the network element 120i, the network element 120i is a base station; but for the base station 110a, the network element 120i is a terminal. The RAN nodes 110 and the terminals 120 are sometimes referred to as communication apparatuses, e.g., the network elements 110a and 110b in Figure 1 can be understood as communication apparatuses with base station functions, and the network elements 120a-120j can be understood as communication apparatuses with terminal functions.
[0115] In a possible scenario, the RAN node can be a base station, an evolved NodeB (eNodeB), an access point (AP), a transmission reception point (TRP), a next generation NodeB (gNB), a base station in a future mobile communication system, or an access node in a WiFi system, etc. The RAN node can be a macro base station (such as 110a in FIG. 1), a micro base station or an indoor station (such as 110b in FIG. 1), a relay node or a donor node, or a wireless controller in a CRAN scenario. Optionally, the RAN node can also be a server, a wearable device, a vehicle or a vehicle-mounted device, etc. For example, the access network device in vehicle to everything (V2X) technology can be a road side unit (RSU). All or part of the functions of the RAN node in this application can also be implemented by software functions running on hardware, or by virtualized functions instantiated on a platform (such as a cloud platform). The RAN node can also be provided with a communication module, circuit or chip for performing corresponding communication functions, and program instructions for performing corresponding communication functions. The RAN node in this application can also be a logical node, a logical module or software that can implement all or part of the functions of the RAN node.
[0116] In another possible scenario, multiple RAN nodes cooperate to assist a terminal to implement wireless access, and different RAN nodes respectively implement part of the functions of a base station. For example, the RAN node can be a central unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU), etc. The CU and the DU can be separately arranged, or can be included in the same network element, such as a baseband unit (BBU). The RU can be included in a radio frequency device or a radio frequency unit, such as a remote radio unit (RRU), an active antenna processing unit (AAU), or a remote radio head (RRH).
[0117] In different systems, the CU (or CU-CP and CU-UP), DU or RU can also have different names, but those skilled in the art can understand their meanings. For example, in the ORAN system, the CU can also be referred to as O-CU (open CU), the DU can also be referred to as O-DU, the CU-CP can also be referred to as O-CU-CP, the CU-UP can also be referred to as O-CU-UP, and the RU can also be referred to as O-RU. For the convenience of description, the CU, CU-CP, CU-UP, DU and RU are taken as examples for description in this application. Any one of the CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0118] A terminal can access the above-mentioned communication system and has corresponding communication functions. The terminal can also be referred to as a terminal, user equipment (UE), mobile station, mobile terminal, etc. The terminal can be widely used in various scenarios, such as device-to-device (D2D) communication, vehicle-to-everything (V2X) communication, machine-type communication (MTC), internet of things (IOT), virtual reality, augmented reality, industrial control, autonomous driving, remote medical treatment, smart power grid, smart furniture, smart office, smart wear, smart transportation, smart city, etc. The terminal can be a mobile phone, tablet computer, computer with wireless transceiver function, wearable device, vehicle, unmanned aerial vehicle, helicopter, airplane, ship, robot, mechanical arm, smart home device, transport vehicle with wireless communication function, communication module, etc. Embodiments of the present application do not limit the device form of the terminal. The terminal is usually provided with a communication module, circuit or chip for executing corresponding communication functions. The terminal is also configured with program instructions for executing corresponding communication functions.
[0119] The technical problems to be solved and the technical solutions adopted by the present application are described below.
[0120] Perception can be achieved by a variety of different methods and can be divided into two categories: 2D perception and 3D perception. 2D perception generates a planar image, which only contains width and height information and has no depth information. 3D perception captures the shape and structure of a scene in three-dimensional space and generates a stereoscopic image, which contains width, height and depth information. Common 2D perception schemes include optical perception and radar imaging. Common 3D perception schemes include radio frequency perception and computer tomography. 2D optical perception and 3D radio frequency perception are taken as examples to introduce 2D perception and 3D perception.
[0121] (One) Radio frequency sensing
[0122] In this sensing scheme, the sensing of the target is achieved by receiving the echo signal of the sensed target (also referred to as the sensing target), where the echo signal can be a reflected signal, a diffracted signal, or a scattered signal, etc.
[0123] (Two) Optical sensing
[0124] In this sensing scheme, the light wave is detected by the image sensor of the camera and an image is generated. The parameters of the optical sensing node are divided into camera intrinsics and camera extrinsics, where the intrinsics are parameters related to the characteristics of the camera itself, such as the focal length f of the camera, the field of view, the resolution, etc.; the extrinsics are parameters in the world coordinate system, such as the position, pointing direction, and rotation direction of the camera. The imaging process involves four coordinate systems, which are the world coordinate system, the camera coordinate system, the image coordinate system, and the pixel coordinate system. The following describes these four coordinate systems in detail in conjunction with FIG. 2.
[0125] ① World coordinate system: a coordinate system of a three-dimensional world defined by a user, introduced to describe the position of the sensing target in the real world, and is a three-dimensional coordinate system of the objective world. As shown in FIG. 2, O w , X w , Y w , and Z w are the origin, X-axis, Y-axis, and Z-axis of the world coordinate system, respectively. For example, there is a point P in the real world, and the coordinates of the P point in the world coordinate system can be represented as (X w , Y w , Z w ).
[0126] ② Camera coordinate system: a three-dimensional coordinate system established with the optical sensing node (such as the optical center of the camera) as the coordinate origin and the optical axis of the camera as the Z-axis. As shown in FIG. 2, O c , X c , Y c , and Z c are the origin, X-axis, Y-axis, and Z-axis of the camera coordinate system, respectively. It can be understood that the camera coordinate system is obtained from the world coordinate system through rotation and translation operations, and correspondingly, the world coordinate system can also be obtained from the camera coordinate system through rotation and translation operations. For example, the coordinates of the P point in the camera coordinate system can be represented as (X c , Y c , Z c ).
[0127] ③Image coordinate system: a 2D coordinate system introduced to describe the projection transmission relationship of the perceived target from the camera coordinate system to the image coordinate system in the imaging process. Among them, the image coordinate system takes the center of the image sensor as the coordinate origin, and the X-axis and Y-axis are parallel to the two vertical edges of the image sensor, as shown in FIG. 2, o, x, and y are the origin, X-axis and Y-axis of the image coordinate system, respectively.
[0128] ④Pixel coordinate system: a 2D coordinate system introduced to describe the coordinates of the image point on the digital image (photo) after the imaging of the object, which is the coordinate system where the user actually reads the information from the camera. The image coordinate system takes the top left corner of the image sensor as the origin, and the X-axis and Y-axis of the image coordinate system are parallel to the X-axis and Y-axis of the image coordinate system, as shown in FIG. 2, u and v are the X-axis and Y-axis of the image coordinate system, respectively. Each image contains multiple elements, each element is a pixel, and the pixel coordinate system is in units of pixels.
[0129] Among them, the Z-axis of the camera coordinate system is perpendicular to the image coordinate system plane and passes through the origin of the image coordinate system, and the origin O c of the camera coordinate system is the optical center of the camera. The distance between the camera coordinate system origin O c and the image coordinate system origin o is the focal length f (i.e. the distance from the camera focus to the optical center). The pixel coordinate system plane u-v coincides with the image coordinate system plane x-y, but the pixel coordinate system origin is located at the top left corner of the figure. The reason for such definition is to start reading and writing from the first address of the stored information.
[0130] The conversion process of the four coordinate systems is introduced below.
[0131] (1) World coordinate system -> camera coordinate system conversion (3D -> 3D projection)
[0132] First, the external matrix is determined by the camera external parameters (such as the position and direction of the camera), and the coordinates of the perceived target are transformed from the world coordinate system to the camera coordinate system through rotation and translation based on the external matrix. The camera position refers to the position of the camera at the origin O c of the world coordinate system, and the camera direction refers to the direction of the three axes of the camera coordinate system in the world coordinate system [X axis , Y axis , Z axis ]. When any two coordinate axes in the camera direction are determined, the direction of the other axis can be calculated, for example, Y axis = Z axis *X axis . The rotation matrix from the world coordinate system to the camera coordinate system is a 3*3 orthogonal matrix, represented as where ||.|| represents the modulus, and the translation vector is a 3*1 vector, represented as t = -R*O c Therefore, the external matrix of the camera can be represented as
[0133] (2) Camera coordinate system -> image coordinate system conversion (3D -> 2D projection)
[0134] Fig. 3 is a schematic diagram of point projection in camera coordinate system to image coordinate system plane. Wherein, P, B, A points are points in camera coordinate system, p, c, o are points of P, B, A projected to image coordinate system plane, f is the focal length of the camera, wherein, the coordinates of P point are (X c ,Y c ,Z c ), the coordinates of p point are (x, y). The transformation from camera coordinate system to image coordinate system can be calculated according to the principle of similar triangles.
[0135] Since △ABO c ≈△oCO c , △PBO c ≈△pCO c , ∽ represents similar, then
[0136] Therefore, The transformation process is represented in matrix form as follows:
[0137] It can be seen that in this projection process, the depth information (Z c ) is lost. For example, as shown in Fig. 3, other points on the ray Oc-P will also be projected to the p point on the image coordinate system plane, that is, according to the projection result, the depth information cannot be recovered.
[0138] (3) Image coordinate system -> pixel coordinate system conversion (2D -> 2D projection)
[0139] The pixel coordinate system and the image coordinate system are both on the imaging plane of the optical perception node, but the origins and the units of measurement are different. The image coordinate system generally takes the center of the image sensor as the coordinate origin, and the unit is a physical unit, which is taken as an example below. Millimeter (mm). The pixel coordinate system generally takes the top left corner of the image sensor as the origin, and the unit is pixel. The number of pixels determines the resolution of the camera. For example, in Fig. 4, the number of pixels of the camera is M*N, which means that the number of pixels in each row is M and the number of pixels in each column is N, that is, the corresponding resolution is M*N.
[0140] Wherein, the relationship conversion between the pixel coordinate system coordinate point (u, v) and the image coordinate system coordinate point (x, y) is: Wherein, dx and dy represent how many mm each column of pixels and each row of pixels represent, u m and v mThe pixel point coordinates corresponding to the center of the image sensor are equal to half of the horizontal and vertical resolutions, respectively, i.e. The matrix is expressed as follows:
[0141] In combination with (2) (camera->image coordinate system conversion) and (3) (image->pixel coordinate system conversion), the internal matrix of the camera can be defined as:
[0142] wherein, f represents the number of pixels corresponding to the focal length in the horizontal direction, f represents the number of pixels corresponding to the focal length in the vertical direction. Given the field of view angle in the horizontal direction as θ u , there is Similarly, θ v represents the field of view angle in the vertical direction. Therefore, when the focal length, the resolution, or the field of view angle, the resolution is known, the internal matrix of the camera can be determined.
[0143] FIG. 4 is a schematic diagram of a possible perception scene to which the embodiments of the present application are applicable. As shown in FIG. 4, the perception scene includes 3D perception node A, 3D perception node B, 2D perception node A, and 2D perception node B. These perception nodes can perceive perception targets (buildings, cars, etc.) in the perception area. In FIG. 4, the perception nodes can perceive the perception targets in the perception area (the solid lines in the figure represent that perception can be performed), and there is a data link between the perception nodes (the dashed solid lines in the figure represent the data link). It can be understood that FIG. 4 only exemplarily shows two 2D perception nodes and two 3D perception nodes, and FIG. 4 can further include more 2D perception nodes and 3D perception nodes, and the perception modalities (i.e., perception schemes) of the perception nodes in FIG. 4 are not specifically limited by the present application.
[0144] At present, the 2D perception process is relatively simple, and can quickly capture and display images, and optical imaging has high resolution, which can provide clear texture, color, and other information, but the 2D perception cannot provide depth information. Although the 3D perception provides depth information, it is easily affected by noise and interference, which affects the perception performance. In addition, in the 3D perception scheme such as radio frequency perception, when the beam of the array element is not omnidirectional, the coverage is limited, as shown in FIG. 5, and only a small amount of 3D point cloud can be obtained based on the radio frequency perception scheme, which further affects the perception performance.
[0145] Therefore, the present application provides a perception method, which can effectively solve the above technical problems. The method provided by the present application will be described in detail below.
[0146] FIG. 6 is a schematic flowchart of a perception method 600 provided by the present application. The method includes the following steps.
[0147] S601, the first device obtains first information, the first information being obtained based on a 3D perception node perceiving a perception target, the first information indicating information of a surface on which the perception target is located.
[0148] It can be understood that the first device is taken as an example for illustration of the execution subject in the present application, but the present application does not limit the execution subject for illustration. For example, the method executed by the first device in the present application can also be implemented by a component (such as a circuit, a processor, a chip or a chip system, etc.) in the first device, or a logic node, a logic module or software capable of realizing all or part of the function of the first device.
[0149] It can also be understood that the 3D perception node refers to a node with a 3D perception function, and the above-mentioned 3D perception node can also be replaced by “a module with a 3D perception function” or “an entity with a 3D perception function” or “a node with a 3D perception function” or “a device with a 3D perception function” or “an apparatus with a 3D perception function”. For example, the 3D perception function can be realized by electronic hardware, computer software, or a combination of computer software and electronic hardware, and the present application does not limit this.
[0150] The first information can be direct indication or indirect indication. The first information is illustrated below.
[0151] Example one, the first information includes an equation F(x, y, z) = 0 of the surface on which the perception target is located. For example, the surface on which the perception target is located is a plane, and the corresponding equation can be Ax + By + Cz + D = 0. For another example, the surface on which the target is located is a spherical surface, and the corresponding equation can be (x-a) 2 (y-b) 2 +(z-c) 2 =r 2 .
[0152] Example two, the first information includes a normal of the surface on which the perception target is located or a set of normals of discrete subplanes of the surface on which the perception target is located. For example, the equation of the surface on which the target is located is Ax + By + Cz + D = 0, and the corresponding normal is (A, B, C). For another example, if the surface on which the target is located is a non-planar surface, the surface on which the target is located can be discretized to obtain a plurality of discrete subplanes, and the first information includes the normals corresponding to the plurality of discrete subplanes.
[0153] Example three, the first information includes a 3D perception result of the perception target. For example, if the 3D perception node is a radio frequency perception node, the 3D perception result is a 3D point cloud.
[0154] It can be understood that the above example one and example two are direct indication, and example three is indirect indication. It can also be understood that the above three examples can be used alone or in combination. For example, the first information can include the equation of the surface where the perception target is located and / or the 3D perception result of the perception target.
[0155] S602, the first device obtains second information, the second information is obtained based on the 2D perception node perceiving the perception target, and the second information indicates information of each key point in a plurality of key points of a 2D perception result of the perception target reversely projecting to a corresponding projection straight line in a 3D space.
[0156] It can be understood that the 2D perception node refers to a node with 2D perception function, and the above 2D perception node can be replaced by "a module with 2D perception function" or "an entity with 2D perception function" or "a node with 2D perception function" or "a device with 2D perception function" or "an apparatus with 2D perception function". For example, the 2D perception function can be realized by electronic hardware, computer software, or a combination of computer software and electronic hardware, which is not limited in the present application.
[0157] For example, if the 2D perception node is an optical perception node, the 3D space can be a space corresponding to a camera coordinate system of the optical perception node.
[0158] For example, if the 2D perception node is an optical perception node, the 2D perception result is a 2D perception picture. For example, the key points of the 2D perception picture can be the boundary pixel points of the 2D perception picture.
[0159] The second information can be indirect indication or direct indication. The second information is exemplified as follows.
[0160] Example one, the second information includes equations of the plurality of key points reversely projecting to the corresponding projection straight lines in the 3D space.
[0161] Example two, the second information includes the 2D perception result of the perception target.
[0162] Similarly, the above example one is direct indication, and example two is indirect indication. It can also be understood that the above two examples can be used alone or in combination. For example, the second information includes the equation of the plurality of key points reversely projecting to the corresponding projection straight lines in the 3D space and / or the 2D perception result of the perception target.
[0163] It can be understood that the projection straight line corresponding to any of the plurality of key points (taking the first key point as an example) includes a first point and a second point, the first point is any point obtained by inversely projecting the first key point into the 3D space, and the second point is the point where the 2D perception node is located. In order to facilitate the understanding of the inverse projection, how to determine the projection straight line corresponding to the first key point is illustrated by taking FIG. 4 as an example. For example, the first key point is the p point in the 2D imaging plane in FIG. 4, and the 3D space is the space corresponding to the camera coordinate system in FIG. 4. The projection straight line corresponding to the p point in the camera coordinate system is the straight line where the Oc-P point is located. It can be understood that the Oc point is a known point, and then, based on the internal matrix of the camera, the coordinates (u, v) of the p point in the pixel coordinate system, and any Z c1 (Z c1 ≠ 0), the P' point in the camera coordinate system is inversely deduced, the coordinates of the P' point are (X cP′ ,Y cP′ ,Z cP′ )(Z cP′ = Z c1 ), so that the projection straight line corresponding to the p point can be determined based on the Oc point (i.e., the second point) and the P' point (i.e., the first point).
[0164] S603, the first device determines the information of the perception target based on the first information and the second information.
[0165] For example, the information of the perception target includes the shape and position of the perception target.
[0166] It can be understood that the first device determines the information of the perception target based on the first information and the second information, including: the first device determines the intersection point of the projection straight line corresponding to each key point of the 2D perception result of the perception target and the surface of the perception target based on the first information and the second information; and the first device determines the information of the perception target based on the obtained intersection point. For example, the first information can indicate the information of the surface of the perception target in the 3D space, and the second information can indicate the information of the projection straight line corresponding to the plurality of key points in the 3D space, so as to obtain the intersection information of the surface of the perception target and the projection straight line in the 3D space, and then determine the information of the perception target based on the intersection information.
[0167] Taking the first information including the 3D point cloud of the perception target and the second information including the 2D perception picture of the perception target as an example, how to determine the information of the perception target is illustrated below.
[0168] (1) The first device selects a plurality of key points (for example, q) in the 2D perception picture of the perception target based on the second information, and then obtains the equation of the projection straight line corresponding to the inverse projection of the q key points into the camera coordinate system based on the internal matrix K of the camera and the coordinates of the q key points in the pixel coordinates. The equations of the projection straight lines corresponding to the q key points are R E1, R E2 , …, R Eq .
[0169] (2) The first information includes coordinates of 3D point cloud of the perception target (for example, the 3D point cloud includes s 3D perception points) in a world coordinate system, respectively (X w1 , Y w1 , Z w1 ), (X w2 , Y w2 , Z w2 ), …, (X ws , Y ws , Z ws ). The first device can obtain coordinates of the s 3D perception points in a camera coordinate system based on an extrinsic matrix E of the camera, respectively (X c1 , Y c1 , Z c1 ), (X c2 , Y c2 , Z c2 ), …, (X cs , Y cs , Z cs ), and then determine an equation F(x, y, z) = 0 of a surface on which the perception target is located based on the coordinates of the s 3D perception points in the camera coordinate system.
[0170] (3) The first device brings equations R E1 , R E2 , …, R Eq of the projection straight lines corresponding to the q key points into the equation F(x, y, z) = 0 of the surface on which the perception target is located, to obtain coordinates of the intersection points corresponding to each key point.
[0171] It can be understood that, in order to obtain correct intersection point coordinates, the equations of the projection straight lines of the multiple key points and the equation of the surface on which the perception target is located need to take the same coordinate system as a reference coordinate system. In the above (1) and (2), the equations of the projection straight lines of the multiple key points and the equation of the surface on which the perception target is located are only exemplarily expressed in the camera coordinate system, and the equations of the projection straight lines of the multiple key points and the equation of the surface on which the perception target is located can also be expressed in the world coordinate system, and the reference coordinate system is not limited in the present application.
[0172] Exemplarily, the q key points include a p point in FIG. 3, an equation of a projection straight line (a straight line on which the p point is located) corresponding to the p point is R E1 , and a surface on which a perception target is located is shown as a 3D perception plane in FIG. 3, and a corresponding equation is F(x, y, z) = 0. An intersection point of the equation R E1 and the equation F(x, y, z) = 0 is the P point in FIG. 4.
[0173] (4) The first device determines information of the sensing target based on the intersection coordinates corresponding to the q key points.
[0174] It can be understood that the Z-axis size in the intersection coordinates is the depth information recovered by the back projection.
[0175] The technical solution combines the sensing information of 2D sensing with the sensing information of 3D sensing. Compared with 2D sensing or 3D sensing alone, the sensing performance can be improved, and the sensing result can be optimized. For example, compared with 2D sensing alone, the depth information of the sensing target can be recovered. Compared with 3D sensing alone, the 2D sensing result can be used to obtain the contour and two-dimensional shape information of the sensing target, thereby improving the coverage rate and sensing accuracy of 3D sensing.
[0176] In a possible implementation, the first device can be a 2D sensing node, and the first device performs joint sensing based on the second information obtained by itself and the first information obtained by the 3D sensing node; or, the first device can be a 3D sensing node, and the first device performs joint sensing based on the first information obtained by itself and the second information obtained by the 2D sensing node. For example, the first device can be the RAN node 110 or the terminal 120 in FIG. 1.
[0177] In another possible implementation, the first device is a sensing node that has both 2D sensing function and 3D sensing function. For example, the first device can perform joint sensing based on the first information obtained by itself and the second information obtained by itself. For example, the first device can be the RAN node 110 or the terminal 120 in FIG. 1.
[0178] In yet another possible implementation, the first device can be an intermediate node that performs joint sensing based on the first information and the second information. For example, the intermediate node can be the RAN node 110 or the terminal 120 of the RAN 100 in FIG. 1, or the intermediate node can be a management unit such as a sensing management unit of the core network 200 in FIG. 1. The node type of the intermediate node is not limited in the present application.
[0179] The specific examples of the first device will be described in detail hereinafter, which will not be described here.
[0180] In the joint sensing scheme described above, the sensing quality of the 2D sensing node and the 3D sensing node will affect the final joint sensing result. Therefore, it is necessary to select a sensing node with good sensing quality to sense the sensing target, so as to further optimize the joint sensing result. The selection of the 2D sensing node and the 3D sensing node will be described below.
[0181] (1) Selection of 2D sensing node
[0182] In this application, the 2D perception node participating in joint perception can be a node satisfying a first condition, and the first condition is a predefined condition or an indicated condition.
[0183] For example, the index affecting the perception quality of the 2D perception node includes a perception error. The following possible first conditions are given by taking the perception error as an example.
[0184] For example, the first condition is that the perception error of each 2D perception node participating in perception is less than a first threshold, or the first condition is that the 2D perception node participating in perception is the top N1 nodes in the candidate 2D perception nodes in the order of the perception error from small to large, and N1 is a positive integer.
[0185] For example, the standard can stipulate the evaluation method of the perception error, for example, the perception error is a function determined based on at least one of the perception parameters of the 2D perception node.
[0186] The perception error of the 2D perception node is specifically described below as an optical perception node. It can be understood that the perception error of the optical perception node is related to the perception parameters of the camera. The perception parameters of the camera include at least one of the following parameters: focal length, field of view, resolution, position, or pointing direction.
[0187] In one possible implementation, the perception error error of the optical perception node satisfies the following formula:
[0188] Wherein, FoV is the field of view of the camera, a is the angle between the pointing direction of the optical axis of the camera (i.e. the Z axis in the camera coordinate system) and the plane of the perception target, and a is determined based on the first information, r is the physical distance between the intersection of the plane of the perception target and the pointing direction of the optical axis of the camera and the position of the camera, resolution is the resolution of the camera, and parameters such as a and r are shown in FIG. 7.
[0189] (2) Selection of 3D perception node
[0190] In this application, the 3D perception node participating in joint perception can be a node satisfying a second condition, and the second condition is a predefined condition or an indicated condition.
[0191] For example, the index affecting the perception quality of the 3D perception node includes a surface estimation error. The following possible second conditions are given by taking the surface estimation error as an example.
[0192] For example, the second condition is that the perception error of each 3D perception node participating in perception is less than a first threshold, or the second condition is that the 3D perception node participating in perception is the top N2 nodes in the candidate 3D perception nodes in the order of the perception error from small to large, and N2 is a positive integer.
[0193] It can be understood that in the joint perception scheme of the present application or other possible joint perception schemes, the first device can also request the at least one perception device to report its corresponding perception capability, so that the first device can select appropriate perception nodes to participate in subsequent perception based on the perception capability of the perception nodes. The flow of reporting the perception capability is described below.
[0194] FIG. 8 is a schematic flowchart of a perception capability reporting method according to an embodiment of the present application. The method comprises the following steps.
[0195] S801, the first device sends a request message #1 to L perception nodes respectively, the request message #1 being used to request the corresponding perception nodes to report the perception capability, L being a positive integer. Correspondingly, the L perception nodes receive the request message #1 from the first device respectively.
[0196] Optionally, the perception capability of a perception node is associated with the perception dimension supported by the perception node, the perception modality supported by the perception node, and the perception parameter supported by the perception node.
[0197] For example, the perception dimension supported by a perception node includes 2D perception and 3D perception.
[0198] For example, the 2D perception modality supported by a perception node includes optical perception and / or radar perception, and the 3D perception modality supported by a perception node includes radio frequency perception and / or computer tomography.
[0199] For example, the perception parameter of a 2D perception node (such as an optical perception node) includes at least one of the following parameters: focal length, field of view, resolution, position, or pointing direction; and the perception parameter of a 3D perception node includes surface estimation error.
[0200] Optionally, the reporting content and format of the capability can be agreed by the capability reporting requester (i.e. the first device), the reporting content can include reporting all the information related to the perception capability at one time, or reporting only part of the information related to the perception capability, and the reporting format can be in the form of a table or a list.
[0201] For example, if it is agreed to report all the capability information at one time and the reporting format can be in the form of a table, one possible reporting format is shown in Table 1.
[0202] Table 1
[0203] For example, if it is agreed to report all the capability information at one time and the reporting format can be in the form of a list, one possible reporting format is {supported perception dimension}, {supported perception modality}, and {supported perception parameter}.
[0204] S802, at least one of the L sensing nodes respectively sends a request response message #1 to the first device, the request response message #1 comprises information #1, the information #1 is used to indicate the sensing capability of the sensing node. Correspondingly, the first device respectively receives the request response message #1 from at least one of the L sensing nodes.
[0205] It can be understood that, if a node of the L sensing nodes is willing to participate in joint sensing, it can carry its corresponding sensing capability in the corresponding request response message according to the agreed content; if it is unwilling to participate in sensing, it can inform the first device that it does not participate in joint sensing through the corresponding request response message (for example, carrying "Fail" (or failure indication) in the request response), or it can not send the corresponding request response message, and the first device defaults that the sensing node does not participate in joint sensing, or it can set all the contents of the request report to "empty".
[0206] Optionally, when the first device does not participate in joint sensing, it can further indicate the reason for not participating in sensing. For example, the reason includes at least one of the following: not supporting joint sensing, or not enough resources to participate in joint sensing, where the resources can be at least one of time domain, frequency domain or spatial domain resources.
[0207] For example, if it is agreed to report all the capability information at one time and the reporting format can be in table form, the capability reported by the sensing node willing to participate in sensing can be as shown in Tables 2 to 4.
[0208] Table 2
[0209] Table 3
[0210] Table 4
[0211] In one possible implementation, the sensing node can also actively report its own sensing capability to the first device, that is, the first device does not need to request the sensing node to report the capability information in S801. For example, the first device is a base station, and the sensing nodes served by the base station can actively report their own sensing capability to the base station.
[0212] Optionally, the method further comprises S803 and S804.
[0213] S803, the first device sends information #2 to at least one candidate sensing node, the information #2 is used to indicate the sensing scene. Correspondingly, at least one candidate sensing node receives the information #2 from the first device.
[0214] It can also be understood that the at least one candidate sensing node is at least one of the nodes of the L sensing nodes willing to participate in sensing in S802.
[0215] For example, the first device can determine the candidate sensing nodes based on the capability information of the sensing nodes willing to participate in sensing in S802 and the self-demand of the first device. For example, the first device is a 3D sensing node, and the 3D sensing node needs to obtain 2D sensing information, the first device can send information #2 to the node with 2D sensing function in the nodes willing to participate in sensing in S802.
[0216] For example, the information #2 includes sensing time and / or spatial range of sensing.
[0217] S804, at least one candidate sensing node sends information #3 to the first device, the information #3 indicates that the specified sensing scene can be sensed or the specified sensing scene cannot be sensed.
[0218] For example, if the sensing can be performed, the information #3 includes “ACK”.
[0219] For example, if the sensing cannot be performed, the information #3 includes “Fail” (or failure indication), or the sensing node can also not feedback the information #3, and the first device defaults that the sensing node cannot sense the specified scene, and the specific feedback mode can be agreed in the standard.
[0220] Optionally, when the first device indicates that the specified sensing scene cannot be sensed, the first device can further indicate the reason for the inability to sense. For example, the reason for the inability to sense includes at least one of the following: unable to sense within a specified time, unable to sense in a specified sensing space, or insufficient resources to participate in joint sensing, where the resources can be at least one of time domain, frequency domain or spatial domain resources.
[0221] Optionally, if the information #1 of S802 does not include the capability information of the corresponding sensing node required by the first device for subsequent calculation, for example, in the joint sensing method proposed in the present application, the first device can calculate the sensing error based on the 2D sensing node sensing parameter to determine whether the 2D sensing node meets the first condition, if the information #1 only reports the sensing dimension of the corresponding sensing node, the first device cannot determine whether the 2D sensing node meets the first condition. Therefore, optionally, the method further includes S805 and S806.
[0222] S805, the first device sends a request message #2 to at least one node (i.e. a new candidate sensing node) capable of sensing the specified scene, the request message #2 is used to request the sensing node to report the required sensing capability information. Correspondingly, the new candidate sensing node receives the request message #2 from the first device.
[0223] S806, the new candidate perception node sends a request response message #2 to the first device, the request response message #2 including the required perception capability information corresponding to the perception node. The first device receives the request response message #2 from the new candidate perception node.
[0224] It can be understood that the perception nodes receiving the information sent by the first device in different steps of the method can be different, and in FIG. 8, the perception nodes participating in all steps are taken as examples for description for the sake of drawing. The perception nodes actually participating in each step can be understood with reference to the detailed description.
[0225] It can also be understood that the capability reporting scheme shown in FIG. 8 can be executed alone or in combination with the joint perception method shown in FIG. 6. Based on the joint perception method and the capability reporting method proposed in the present application, a possible complete joint perception process is given taking the first device as a 3D perception node, a 2D perception node, and an intermediate node as examples. In FIGS. 9 and 10, the first device is taken as a 3D perception node for description, in FIGS. 11 and 12, the first device is taken as a 2D perception node for description, and in FIG. 13, the first device is taken as an intermediate node for description.
[0226] FIG. 9 is a schematic diagram of a possible perception method proposed in the present application. In the method, the joint perception is initiated by a 3D perception node, and information fusion is performed. The method includes the following steps.
[0227] S901, the 3D perception node sends a request message #1 to L1 perception nodes, the request message #1 being used to request the perception nodes to report the corresponding perception capability, L1 being a positive integer. Correspondingly, the L1 perception nodes receive the request message #1 from the 3D perception node.
[0228] For the detailed description of S901, refer to S801, which will not be repeated here.
[0229] S902, the 3D perception node receives a request response message #1 from at least one of the L1 perception nodes, the request response message #1 including information #1, the information #1 being used to indicate the perception capability of the corresponding perception node. Correspondingly, the 3D perception node receives the request response message #1 from at least one of the L perception nodes.
[0230] For S902, refer to the description in S802, which will not be repeated here.
[0231] S903, the 3D perception node sends information #2 to at least one candidate 2D perception node, the information #2 being used to indicate the perception scene. Correspondingly, the at least one candidate 2D perception node receives the information #2 from the first device.
[0232] The at least one candidate 2D sensing node in S903 and S904 is at least one node with 2D sensing capability in the at least one sensing node of the L1 sensing nodes.
[0233] It can be understood that in the joint sensing method, the joint sensing can be regarded as being initiated by the 3D sensing node, and the 3D sensing node hopes to perform joint sensing with the 2D sensing node (for example, the optical sensing node). In this case, the node with optical sensing capability in the at least one sensing node of the L1 sensing nodes can be sent information #2.
[0234] For other descriptions of S903, refer to S803, which will not be repeated here.
[0235] S904, the at least one candidate 2D sensing node sends information #3 to the 3D sensing node, and the information #3 indicates that the specified sensing scene can be sensed or cannot be sensed.
[0236] For S904, refer to the description in S804, which will not be repeated here.
[0237] S905, the 3D sensing node senses the sensing target to obtain first information, and the first information indicates information of a surface on which the sensing target is located.
[0238] For the first information, refer to the description in S601, which will not be repeated here.
[0239] S906, the 3D sensing node sends message #1 to the at least one candidate 2D sensing node, and the message #1 includes the first information and information #4. The information #4 is used to indicate that if the corresponding candidate sensing node meets a first condition, the corresponding second information is reported.
[0240] The at least one candidate 2D sensing node in S906 is at least one node in the nodes that can sense the specified sensing scene in S904.
[0241] For example, the first condition can be that the sensing error of each 2D sensing node participating in sensing is less than a first threshold value, and the specific value of the first threshold value is included in the information #3. For example, the formula for calculating the sensing error is a predefined formula, or the information #3 can further include the formula for calculating the sensing error.
[0242] S907, the at least one candidate 2D sensing node calculates its own sensing error and judges whether it meets the first condition.
[0243] If the first condition is met, S908 is performed; if the first condition is not met, S908 does not need to be performed.
[0244] S908, the 2D perception nodes satisfying the first condition respectively send respective corresponding second information to the 3D perception node, the second information indicating information that each key point in the plurality of key points of the 2D perception result of the perception target is reversely projected to a corresponding projection straight line in the 3D space.
[0245] For the second information, refer to the description in S602, which will not be repeated here.
[0246] S909, the 3D perception node determines the information of the perception target based on the first information and the second information.
[0247] For S909, refer to the description in S603, which will not be repeated here.
[0248] It can be understood that the perception nodes receiving the respective information sent by the 3D perception node in different steps of the method can be different. In FIG. 9, the 2D perception nodes participating in all steps are taken as an example for illustrative description. The perception nodes actually participating in each step are described in detail below. In FIG. 10, the 2D perception nodes participating in all steps are also taken as an example for illustrative description, which will not be repeated hereinafter.
[0249] FIG. 10 is a schematic diagram of a possible perception method proposed in the present application. The difference between the method shown in FIG. 10 and the method shown in FIG. 9 is that, in the method shown in FIG. 9, the 3D perception node instructs the 2D perception node to determine whether the first condition is satisfied by itself, and if the first condition is satisfied, the corresponding second information is reported. In the method, the 3D perception node itself determines whether the candidate 2D perception node satisfies the first condition, and then instructs the 2D perception node satisfying the first condition to report the corresponding second information. The method comprises the following steps.
[0250] S1001-S1004 are the same as S901-S904, which will not be repeated here.
[0251] Optionally, if the perception node only reports part of the capability information when reporting the capability, the part of the capability information does not include part or all of the perception capability information for determining whether the perception node satisfies the first condition, the method further comprises S1005 and S1006.
[0252] S1005, the 3D perception node sends a request message #2 to at least one candidate 2D perception node, the request message #2 being used to request the perception node to report the required perception capability information. Correspondingly, the at least one candidate perception node receives the request message #2 from the 3D perception node.
[0253] The at least one candidate 2D perception node in S1005 (corresponding to the “new candidate perception node” in S805) is at least one perception node in the at least one candidate 2D perception node in S1004 that can perform perception on the specified perception scene.
[0254] S1006, at least one candidate 2D perception node sends a request response message #2 to the 3D perception node, the request response message #2 comprising the required perception capability information of the perception node. Correspondingly, the 3D perception node receives the request response message #2 from at least one candidate perception node.
[0255] S1007, the 3D perception node determines the node meeting the first condition from the at least one candidate 2D perception node. The first condition is described above and will not be repeated here.
[0256] S1008, the 3D perception node sends a request message #3 to the candidate 2D perception node meeting the first condition, the request message #3 being used to request reporting the corresponding second information. Correspondingly, the candidate 2D perception node meeting the first condition receives the request message #3 from the 3D perception node.
[0257] S1009, the 2D perception node meeting the first condition respectively sends the corresponding second information to the 3D perception node, the second information indicating the information of each key point of the 2D perception result of the perception target reversely projecting to the corresponding projection straight line in the 3D space. Correspondingly, the 3D perception node receives the corresponding second information from the 2D perception node meeting the first condition.
[0258] S1010, the 3D perception node perceives the perception target to obtain the first information, the first information indicating the information of the surface where the perception target is located.
[0259] S1011, the 3D perception node determines the information of the perception target based on the first information and the second information.
[0260] FIG. 11 is a schematic diagram of a possible perception method proposed in the present application. The method is similar to the method shown in FIG. 9, except that the joint perception and information fusion are initiated by the 2D perception node, and the 3D perception node meeting the second condition is instructed by the 2D perception node to report the corresponding first information. The method comprises the following steps.
[0261] S1101, the 2D perception node sends a request message #1 to L2 perception nodes, the request message #1 being used to request the perception nodes to report the corresponding perception capability, L2 being a positive integer. Correspondingly, the L2 perception nodes receive the request message #1 from the 2D perception node.
[0262] The specific description of S1101 can be referred to S801, which will not be repeated here.
[0263] S1102, the 2D perception node receives a request response message #1 from at least one of the L2 perception nodes, the request response message #1 comprising information #1, the information #1 being used to indicate the perception capability of the corresponding perception node. Correspondingly, the 2D perception node receives the request response message #1 from at least one of the L perception nodes respectively.
[0264] For S1102, refer to the description in S802, which will not be repeated here.
[0265] S1103, the 2D perception node sends information #2 to at least one candidate 3D perception node, the information #2 being used to indicate the perception scene. Correspondingly, the at least one candidate 3D perception node receives the information #2 from the first device.
[0266] Among them, at least one candidate 3D perception node in S1103 and S1104 is at least one of the nodes with 3D perception capability in the at least one of the L2 perception nodes.
[0267] It can be understood that in the joint perception method, the 2D perception node initiates joint perception, and the 2D perception node hopes to perform joint perception with the 3D perception node (such as the radio frequency perception node), and then the information #2 can be sent to at least one of the nodes with radio frequency perception capability in the at least one of the L2 perception nodes.
[0268] For other descriptions of S1103, refer to S803, which will not be repeated here.
[0269] S1104, the at least one candidate 3D perception node sends information #3 to the 2D perception node, the information #3 indicating that the specified perception scene can be perceived, or the specified perception scene cannot be perceived.
[0270] For S1104, refer to the description in S804, which will not be repeated here.
[0271] S1105, the 2D perception node perceives the perception target to obtain second information, the second information indicating information of each key point in a plurality of key points of a 2D perception result of the perception target, the information being used to indicate that the corresponding projection straight line in the 3D space of each key point is projected reversely.
[0272] For the second information, refer to the description in S601, which will not be repeated here.
[0273] S1106, the 2D perception node sends information #4 to at least one candidate 3D perception node, the information #4 being used to indicate that if the corresponding candidate perception node satisfies a second condition, the corresponding first information is reported.
[0274] The at least one candidate 3D perception node in S1106 is at least one of the nodes that can perceive the designated perception scene in S1104.
[0275] For example, the second condition can be that the surface estimation error of each 3D perception node participating in the perception is less than a second threshold value, and the specific value of the second threshold value is included in information #3. For example, the formula for calculating the surface estimation error is a predefined formula, or information #3 can also include the formula for calculating the surface estimation error.
[0276] S1107, the at least one candidate 3D perception node calculates its own perception error and determines whether it satisfies the second condition.
[0277] If the first condition is satisfied, S908 is executed; if the first condition is not satisfied, S908 does not need to be executed.
[0278] S1108, the 3D perception nodes that satisfy the second condition respectively send their respective first information to the 2D perception node, and the first information indicates the information of the surface where the perception target is located.
[0279] S1109, the 2D perception node determines the information of the perception target based on the first information and the second information.
[0280] For S1109, refer to the description in S603, which will not be repeated here.
[0281] It can be understood that the perception nodes that receive the corresponding information sent by the 2D perception node in different steps of the method can be different. In FIG. 11, the 3D perception nodes participating in all steps are taken as an example for illustrative description. The perception nodes actually participating in each step are described in detail below. In FIG. 12, the 2D perception node participating in all steps is taken as an example for illustrative description, which will not be repeated here.
[0282] FIG. 12 is a schematic diagram of a possible perception method proposed in the present application. The difference between this method and the method shown in FIG. 11 is that in the method shown in FIG. 11, the 2D perception node instructs the 3D perception node to determine whether it satisfies the second condition by itself, and if it satisfies the second condition, it reports the corresponding first information. In this method, the 2D perception node itself determines whether the candidate 3D perception node satisfies the second condition, and then instructs the 3D perception node that satisfies the second condition to report the corresponding first information. The method includes the following steps.
[0283] S1201-S1204 are the same as S1101-S1104, which will not be repeated here.
[0284] Optionally, if the sensing node only reports part of the capability information when reporting the capability, the part of the capability information does not include part or all of the sensing capability information used to determine whether the sensing node meets the second condition, the method further includes S1205 and S1206.
[0285] S1205, the 2D sensing node sends a request message #2 to the at least one candidate 3D sensing node, the request message #2 is used to request the sensing node to report the required sensing capability information. Correspondingly, the at least one candidate sensing node receives the request message #2 from the 2D sensing node.
[0286] In S1205, the at least one 3D candidate sensing node (corresponding to the "new candidate sensing node" in S805) is at least one sensing node in the at least one candidate 3D sensing node in S1204 that can perform sensing on the specified sensing scene.
[0287] S1206, the at least one candidate 3D sensing node sends a request response message #2 to the 2D sensing node, the request response message #2 includes the required sensing capability information corresponding to the sensing node. Correspondingly, the 2D sensing node receives the request response message #2 from the at least one candidate sensing node.
[0288] S1207, the 2D sensing node determines the node in the at least one candidate 3D sensing node that meets the second condition. The second condition is described above and will not be repeated here.
[0289] S1208, the 2D sensing node sends a request message #3 to the candidate 3D sensing node that meets the second condition, the request message #3 is used to request to report the corresponding first information. Correspondingly, the candidate 3D sensing node that meets the second condition receives the request message #3 from the 2D sensing node.
[0290] S1209, the 3D sensing node that meets the second condition respectively sends the respective corresponding first information to the 2D sensing node, the first information indicates the information of the surface where the sensing target is located.
[0291] S1210, the 2D sensing node performs sensing on the sensing target to obtain second information, the second information indicates the information of the projection straight line in the 3D space corresponding to each key point in the plurality of key points of the 2D sensing result of the sensing target.
[0292] S1211, the 2D sensing node determines the information of the sensing target based on the first information and the second information.
[0293] FIG. 13 is a schematic diagram of a possible sensing method proposed in the present application. The method is similar to the method shown in FIG. 9, except that joint sensing and information fusion are initiated by a 2D sensing node, and the 3D sensing nodes satisfying the second condition are instructed by the 2D sensing node to report the corresponding first information.
[0294] It can be understood that in the method, the intermediate node can request multiple sensing nodes to report capabilities, and the flow is the same as the flow in which the first device requests multiple sensing nodes to report sensing capabilities in FIG. 8. The intermediate node can determine candidate 3D sensing nodes and candidate 2D sensing nodes participating in sensing based on the reported sensing capability information, and finally, the information of the sensing target can be obtained based on the second information of the sensing nodes satisfying the first condition in the candidate 2D sensing nodes and the first information of the sensing nodes satisfying the second condition in the candidate 3D sensing nodes. Among them, the candidate 3D sensing nodes and the candidate 2D sensing nodes can judge whether they satisfy the corresponding conditions based on the indication of the intermediate node, and if so, report the corresponding first information or second information, or the intermediate node can also determine whether the candidate sensing nodes satisfy the corresponding conditions based on the sensing parameters of the candidate 3D sensing nodes and the candidate 2D sensing nodes, and if so, instruct the candidate sensing nodes to report the corresponding first information or second information. For the flow of capability reporting and node selection, refer to the description in FIGS. 9-12, which will not be repeated here. Here, only the corresponding flow when the first device in FIG. 6 is an intermediate node is simply described.
[0295] S1301, at least one 3D sensing node sends first information to the intermediate node, the first information indicating information of a plane where the target is located. Correspondingly, the intermediate node receives the first information from the at least one 3D sensing node.
[0296] S1302, at least one 2D sensing node sends second information to the intermediate node, the second information indicating information that each key point in a plurality of key points of a 2D sensing result of the sensing target is inversely projected to a corresponding projection straight line in a 3D space. Correspondingly, the intermediate node receives the second information from the at least one 2D sensing node.
[0297] S1303, the intermediate node determines information of the sensing target based on the first information and the second information.
[0298] It should be understood that the size of the serial number of each process does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0299] It should also be understood that, in various embodiments of the present application, the terms and / or descriptions between different embodiments are consistent and can be referred to each other if there is no special description and logical conflict, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationship.
[0300] It should also be understood that, in some of the above embodiments, the devices in the existing network architecture are mainly exemplarily described, and it should be understood that the specific forms of the devices are not limited in the embodiments of the present application. For example, devices that can realize the same functions in the future are also applicable to the embodiments of the present application.
[0301] It can be understood that, in each of the above method embodiments, the method and operation implemented by the first device (such as the 2D perception node, the 3D perception node, the intermediate node, etc. described above) can also be implemented by a component (such as a chip or a circuit or a processor or a chip system) of the first device.
[0302] The above, in combination with FIGS. 1 to 13, details the method provided by the embodiments of the present application. The above method is mainly introduced from the perspective of the interaction between the first device and the perception node. It can be understood that the first device contains the corresponding hardware structure and / or software module for executing each function in order to realize the above functions.
[0303] Those skilled in the art should realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is realized in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0304] In the following, the communication apparatus provided by the embodiments of the present application is described in detail in combination with FIGS. 14 and 15. It should be understood that the description of the apparatus embodiments corresponds to the description of the method embodiments, therefore, the content not described in detail can be referred to the above method embodiments, and for brevity, some content will not be described again. The embodiments of the present application can divide the functional modules of the first device according to the above method examples, for example, each functional module can be divided according to each function, or two or more functions can be integrated in one processing module. The above integrated module can be realized in the form of hardware or in the form of software functional module. It should be noted that the division of the modules in the embodiments of the present application is illustrative, and is only a logical function division, and another division mode can be used in actual implementation. In the following, taking the division of each functional module according to each function as an example for description.
[0305] The perception method provided in the present application is described in detail above, and the communication device provided in the present application is introduced below. In a possible implementation, the device is used to implement the steps or processes corresponding to the first device in the method embodiments.
[0306] FIG. 14 is a schematic block diagram of a communication device 1400 provided in an embodiment of the present application. As shown in FIG. 14, the communication device 1400 can include modules or units for implementing the steps or processes corresponding to the method embodiments. In a possible design, the communication device 1400 includes a processing unit 1402 and a communication unit 1403. Optionally, the communication device 1400 can further include a storage unit 1401 configured to store device program code and / or data.
[0307] The communication device 1400 can be a device on the first device side in the above-described embodiments, for example, a communication module in the first device or a circuit or chip responsible for the communication function in the terminal.
[0308] For example, in an embodiment, the processing unit 1402 is configured to obtain first information, the first information being obtained based on a 3D perception node perceiving a target, the first information indicating information of a surface where the target is located; the processing unit 1402 is further configured to obtain second information, the second information being obtained based on a 2D perception node perceiving the target, the second information indicating information of a projection straight line corresponding to each key point in a 2D perception result of the target, the key point being obtained by inversely projecting the key point to a 3D space; and the processing unit 1402 is further configured to determine information of the target based on the first information and the second information.
[0309] In a possible design, the processing unit 1402 is specifically configured to determine, based on the first information and the second information, an intersection of the projection straight line corresponding to each key point and the surface where the target is located; and determine the information of the target based on the intersection.
[0310] In a possible design, the projection straight line corresponding to the first key point in the plurality of key points includes a first point and a second point, the first point being any point obtained by inversely projecting the first key point to the 3D space, and the second point being a point where the 2D perception node is located.
[0311] In a possible design, the first information includes at least one of the following: an equation of the surface where the target is located, a normal line set corresponding to a discrete plane of the surface where the target is located, or a 3D perception result of the target.
[0312] In a possible design, the second information includes the 2D perception result of the target, and / or an equation of the projection straight line corresponding to each key point in the 3D space.
[0313] In a possible design, the second information is acquired by receiving the second information corresponding to the N1 2D perception nodes, where N1 is a positive integer.
[0314] In a possible design, the N1 2D perception nodes are 2D perception nodes that satisfy a first condition, where the first condition is a predefined condition or an indicated condition.
[0315] In a possible design, the first condition is that a perception error of each 2D perception node participating in perception is less than a first threshold, or the first condition is that the 2D perception nodes participating in perception are N1 nodes that are ranked in front in a sequence of the candidate 2D perception nodes in ascending order of the perception error, where the perception error of each 2D perception node is determined based on at least one of the perception parameters of each 2D perception node.
[0316] In a possible design, the perception parameters of the 2D perception node include at least one of the following: a focal length of the camera, a field of view angle of the camera, a resolution of the camera, a position of the camera, or a pointing direction of the camera.
[0317] In a possible design, the perception error error satisfies the following formula:
[0318] where FoV is the field of view angle of the camera, a is an included angle between a pointing direction of the optical axis of the camera and a plane of the target, r is a physical distance between the intersection of the plane of the target and the pointing direction of the optical axis of the camera and the position of the camera, resolution is the resolution of the camera, and a is determined based on the first information.
[0319] In a possible design, the communication unit 1403 is configured to: send, to at least one candidate 2D perception node, third information indicating that, if the corresponding candidate 2D perception node satisfies a first condition, the corresponding second information is reported.
[0320] In a possible design, the processing unit 1402 is configured to: determine a candidate 2D perception node that satisfies the first condition from the at least one candidate 2D perception node; and the communication unit 1403 is configured to: send, to the candidate 2D perception node that satisfies the first condition, a first request message, where the first request message is used to request reporting of the corresponding second information.
[0321] In a possible design, the communication unit 1403 is configured to send a second request message to L1 perception nodes, where L1 is a positive integer, the second request message being used to request the perception nodes to report corresponding perception capabilities; and the communication unit 1403 is configured to receive a second request response message from at least one of the L1 perception nodes, the second request response message including fourth information, the fourth information being used to indicate the perception capability of the corresponding perception node, where the N1 2D perception nodes are at least one of the nodes having 2D perception capability in the at least one perception node.
[0322] In a possible design, the fourth information includes specific information of at least one of the following of the corresponding perception node: supported perception dimension, supported perception mode, and supported perception parameter.
[0323] In a possible design, the fourth information does not include the required perception capability information, and the communication unit 1403 is configured to send a third request message to at least one candidate 2D perception node, the third request message being used to request the 2D perception node to report the required perception capability information, the at least one candidate 2D perception node being at least one of the nodes having 2D perception capability in the at least one perception node, and the N1 2D perception nodes being at least one of the at least one candidate 2D perception node; and the communication unit 1403 is configured to receive a third request response message from the at least one candidate 2D perception node, the third request response message including the required perception capability information of the corresponding 2D perception node.
[0324] In a possible design, the communication unit 1403 is specifically configured to receive the corresponding first information from the N2 3D perception nodes, where N2 is a positive integer.
[0325] In a possible design, the N2 3D perception nodes are 3D perception nodes satisfying a second condition, where the second condition is a predefined condition or an indicated condition.
[0326] In a possible design, the second condition is that the surface estimation error of each 3D perception node participating in perception is less than a second threshold, or the second condition is that the 3D perception nodes participating in perception are the top N2 nodes in the candidate 3D perception nodes in a surface estimation error ascending order.
[0327] In a possible design, the communication unit 1403 is configured to send fifth information to the at least one candidate 3D perception node, the fifth information indicating that the corresponding candidate 3D perception node reports the corresponding first information if the second condition is satisfied.
[0328] In a possible design, the processing unit 1402 is configured to: determine a 3D perception node that meets the second condition from the at least one candidate 3D perception node; and the communication unit 1403 is configured to: send, to the 3D perception node that meets the second condition, a fourth request message, where the fourth request message is used to request reporting of the corresponding first information.
[0329] In a possible design, the communication unit 1403 is configured to: send, to the L2 perception nodes respectively, a fifth request message, where the fifth request message is used to request the perception nodes to report corresponding perception capabilities, and L2 is a positive integer; and the communication unit 1403 is configured to: receive a fifth request response message from at least one of the L2 perception nodes, where the fifth request response message includes sixth information, and the sixth information is used to indicate the perception capability of the perception node, and the N2 3D perception nodes are at least one of the at least one perception node that has a 3D perception capability.
[0330] In a possible design, the sixth information includes specific information of at least one of the following of the corresponding perception node: supported perception dimension, supported perception modality, and supported perception parameter.
[0331] In a possible design, the sixth information does not include the required perception capability information, and the communication unit 1403 is configured to: send, to the at least one candidate 3D perception node, a sixth request message, where the sixth request message is used to request the 3D perception node to report the required perception capability information, the at least one candidate 3D perception node is at least one of the at least one perception node that has a 3D perception capability, and the N2 3D perception nodes are at least one of the at least one candidate 3D perception node; and the communication unit 1403 is configured to: receive a sixth request response message from the at least one candidate 3D perception node, where the sixth request response message includes the required perception capability information of the 3D perception node.
[0332] In a possible design, the perception dimension includes 2D perception and 3D perception, the perception modality of the 2D perception includes optical perception and radar perception, and the perception modality of the 3D perception includes radio frequency perception and computer tomography.
[0333] In a possible design, when the communication apparatus 1400 is a first device or a communication module in the first device, the function of the processing unit 1402 can be implemented by one or more processors. Specifically, the processor can include a modem chip, or a system on chip (SoC) chip or a system in package (SIP) chip that includes a modem core. The function of the communication unit 1403 can be implemented by a transceiver circuit.
[0334] In a possible design, when the communication apparatus 1400 is a circuit or a chip responsible for communication functions in the first device, such as a modem chip or a system on chip (SoC) chip or a system in package (SIP) chip including a modem core, the function of the processing unit 1402 can be implemented by circuitry including one or more processors or processor cores in the chip. The function of the communication unit 1403 can be implemented by interface circuitry or data transceiver circuitry on the chip.
[0335] In a possible design, when the communication apparatus 1400 is the first device or a processing module in the first device, the function of the processing unit 1402 can be implemented by one or more processors. Specifically, the processor can include a graphics processing unit (GPU), or a system on chip (SoC) chip or a system in package (SIP) chip including a GPU. The function of the communication unit 1403 can be implemented by transceiver circuitry.
[0336] In a possible design, when the communication apparatus 1400 is a circuit or a chip responsible for processing functions in the first device, such as a GPU or a system on chip (SoC) chip or a system in package (SIP) chip including a GPU, the function of the processing unit 1402 can be implemented by circuitry including one or more processors or processor cores in the chip. The function of the communication unit 1403 can be implemented by interface circuitry or data transceiver circuitry on the chip.
[0337] It can be understood that the division of the units in the apparatus is merely a logical division of functions, and one function can be implemented by one functional unit, or two or more functions can be integrated into one functional unit. In actual implementation, all or part of the units can be integrated into one physical entity, or distributed on different physical entities. In addition, the functional units can be implemented in the form of hardware, software, or a combination of hardware and software. Whether a function is implemented in the form of hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can implement the described functions by using different methods for specific applications, but such implementation should not be considered beyond the scope of the present application.
[0338] In one example, the functional units in any of the above apparatuses can be one or more integrated circuits configured to implement the above methods, for example: one or more application specific integrated circuits (ASICs), or, one or more central processing units (CPUs), one or more microcontroller units (MCUs), one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0339] In one example, the storage unit 1401 can include random access memory, flash memory, read only memory, programmable read only memory, electrically erasable programmable memory, and / or registers, etc.
[0340] FIG. 15 is a schematic block diagram of a communication apparatus 1500 according to an embodiment of the present application. The apparatus 1500 includes a processor 1510 and a transceiver 1520. The processor 1510 and the transceiver 1520 communicate with each other through an internal connection path. The processor 1510 is configured to execute instructions to control the transceiver 1520 to transmit and / or receive signals.
[0341] Optionally, the apparatus 1500 can further include a memory 15150, which communicates with the processor 1510 and the transceiver 1520 through an internal connection path. The memory 15150 is configured to store instructions, and the processor 1510 can execute the instructions stored in the memory 15150. In a possible implementation, the apparatus 1500 is configured to implement the respective procedures and steps of the first device in the above method embodiments.
[0342] It should be understood that the apparatus 1500 can be specifically the first device in the above-described embodiments, or can be a chip or a chip system. Correspondingly, the transceiver 1520 can be a transceiver circuit of the chip, which is not limited herein. Specifically, the apparatus 1500 can be configured to perform each step and / or process in the above-described method embodiments corresponding to the first device. Optionally, the memory 15150 can include read-only memory and random access memory, and provide instructions and data for the processor. Part of the memory can also include non-volatile random access memory. For example, the memory can also store device type information. The processor 1510 can be configured to execute instructions stored in the memory, and when the processor 1510 executes the instructions stored in the memory, the processor 1510 is configured to perform each step and / or process in the above-described method embodiments corresponding to the first device.
[0343] In the implementation process, each step of the above method can be completed by integrated logic circuit of hardware in the processor or instruction in the form of software. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as hardware processor execution completion, or executed by hardware and software module combination in the processor. The software module can be located in the mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0344] It should be noted that the processor in the embodiments of the present application can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method embodiment can be completed by integrated logic circuit of hardware in the processor or instruction in the form of software. The above processor can be a general processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component. The processor in the embodiments of the present application can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as hardware decoding processor execution completion, or executed by hardware and software module combination in the decoding processor. The software module can be located in the mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory, and combines the hardware to complete the steps of the above method.
[0345] It is to be understood that the memory in the embodiments of the present application can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. Among them, the nonvolatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM) used as an external cache. By way of example, and not limitation, many forms of RAM can be used, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the system and method described herein is intended to include, but not be limited to, these and any other suitable types of memory.
[0346] It should be noted that when the processor is a general processor, DSP, ASIC, FPGA or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, the memory (storage module) can be integrated in the processor.
[0347] In addition, the present application also provides a computer readable storage medium, the computer readable storage medium stores computer instructions, when the computer instructions run on the computer, the operations and / or processes performed by the first device in the method embodiments of the present application are executed.
[0348] The present application also provides a computer program product, the computer program product includes computer program code or instructions, when the computer program code or instructions run on the computer, the operations and / or processes performed by the first device in the method embodiments of the present application are executed.
[0349] Further, the chip can further include a communication interface. The communication interface can be an input / output interface, an interface circuit, or the like. Further, the chip can further include a memory.
[0350] Further, the chip can further include a communication interface. The communication interface can be an input / output interface, an interface circuit, or the like. Further, the chip can further include a memory.
[0351] It should be further noted that the memory described herein is intended to include, but not be limited to, these and any other suitable type of memory.
[0352] Those skilled in the art can understand that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized by electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to implement the described functions for each specific application. However, the implementation should not be considered beyond the scope of the present application. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be described here. In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can be another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms. The units described as separate components can be or can not be physically separate, and the components shown as units can be or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0353] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk, and various other program code storage media.
[0354] It should be understood that the "embodiments" mentioned throughout the specification mean that the specific features, structures or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.
[0355] It should also be understood that in the various embodiments of the present application, "A corresponds to B" means that B is associated with A and can be determined according to A. However, it should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information.
[0356] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of perception, the method comprising: The method comprises: obtaining first information, the first information being obtained based on a three-dimensional (3D) perception node perceiving a target, the first information indicating information of a surface on which the target is located; obtaining second information, the second information being obtained based on a two-dimensional (2D) perception node perceiving the target, the second information indicating information of a projection straight line corresponding to each key point in a 2D perception result of the target, the key point being obtained by inversely projecting the key point into a 3D space; determining information of the target based on the first information and the second information.
2. The method of claim 1, wherein, A projection straight line corresponding to a first key point in the key points comprises a first point and a second point, the first point being any point obtained by inversely projecting the first key point into the 3D space, and the second point being a point at which the 2D perception node is located.
3. The method according to claim 1 or 2, characterized in that, The determining of the information of the target based on the first information and the second information comprises: determining an intersection of the projection straight line corresponding to the key points and the surface on which the target is located based on the first information and the second information; and determining the information of the target based on the intersection.
4. The method according to any one of claims 1 to 3, characterized in that, The information of the target comprises a shape and a position of the target.
5. The method according to any one of claims 1 to 4, characterized in that, The first information comprises at least one of the following: an equation of the surface on which the target is located, a normal line set corresponding to a discrete plane of the surface on which the target is located, or a 3D perception result of the target.
6. The method according to any one of claims 1 to 5, characterized in that, The second information comprises the 2D perception result of the target and / or an equation of the projection straight line corresponding to the key points.
7. The method of any one of claims 1 to 6, wherein The obtaining of the second information comprises: receiving the second information corresponding to N1 2D perception nodes, N1 being a positive integer.
8. The method of claim 7, wherein, The N1 2D perception nodes are 2D perception nodes satisfying a first condition, the first condition being a predefined condition or an indicated condition.
9. The method of claim 8, wherein The first condition is that a perception error of each 2D perception node participating in perception is less than a first threshold, or The first condition is that the 2D perception nodes participating in perception are N1 nodes that are ranked in front in a sequence of candidate 2D perception nodes in ascending order of perception errors, wherein the perception error of each 2D perception node is determined based on at least one parameter in a perception parameter of the 2D perception node.
10. The method of claim 9, wherein, The perception parameter of the 2D perception node comprises at least one of the following parameters: a focal length of a camera, a field of view angle of the camera, a resolution of the camera, a position of the camera, or a pointing direction of the camera.
11. The method of claim 10, wherein, The perception error error satisfies the following equation: wherein FoV is the field of view angle of the camera, a is an included angle between a pointing direction of an optical axis of the camera and a plane of the target, r is a physical distance between an intersection of the plane of the target and the pointing direction of the optical axis of the camera and the position of the camera, and resolution is the resolution of the camera, and the a is determined based on the first information.
12. The method according to any one of claims 8 to 11, characterized in that, The method further comprises: sending, to at least one candidate 2D perception node, third information indicating that, if the corresponding candidate 2D perception node satisfies the first condition, the corresponding second information is reported.
13. The method according to any one of claims 8 to 11, characterized in that, The method further comprises: determining a candidate 2D perception node that meets the first condition from the at least one candidate 2D perception node; sending a first request message to the candidate 2D perception node that meets the first condition, the first request message being used to request reporting of the corresponding second information.
14. The method according to any one of claims 7 to 13, characterized in that, The method further comprises: sending a second request message to L1 perception nodes, the second request message being used to request the perception nodes to report corresponding perception capabilities, L1 being a positive integer; receiving a second request response message from at least one of the L1 perception nodes, the second request response message including fourth information, the fourth information being used to indicate the perception capabilities of the corresponding perception node, wherein the N1 2D perception nodes are at least one of the nodes with 2D perception capabilities among the at least one perception node.
15. The method of claim 14, wherein, The fourth information does not include the required perception capability information for calculation, and the method further comprises: sending a third request message to at least one candidate 2D perception node, the third request message being used to request the 2D perception nodes to report the required perception capability information, the at least one candidate 2D perception node being at least one of the nodes with 2D perception capabilities among the at least one perception node, and the N1 2D perception nodes being at least one of the at least one candidate 2D perception node; receiving a third request response message from the at least one candidate 2D perception node, the third request response message including the required perception capability information of the corresponding 2D perception node.
16. The method of claim 15, wherein the fourth information includes specific information of at least one of the following of the corresponding perception node: supported perception dimension, supported perception modality, or supported perception parameter.
17. The method of any one of claims 1 to 16, wherein the obtaining the first information comprises: receiving the first information from N2 3D perception nodes, N2 being a positive integer.
18. The method of claim 17, wherein, The N2 3D perception nodes are 3D perception nodes that meet a second condition, the second condition being a predefined condition or an indicated condition.
19. The method of claim 18, wherein the second condition is that a surface estimation error of each 3D perception node of the 3D perception nodes participating in perception is less than a second threshold value, or the second condition is that the 3D perception nodes participating in perception are the top N2 nodes in the candidate 3D perception nodes in an ascending order of surface estimation error.
20. The method of claim 18 or 19, wherein, The method further comprises: sending fifth information to at least one candidate 3D perception node, the fifth information indicating that the corresponding candidate 3D perception node reports the corresponding first information if the second condition is met.
21. The method of claim 18 or 19, wherein, The method further comprises: determining a 3D perception node that meets the second condition from the at least one candidate 3D perception node; sending a fourth request message to the 3D perception node that meets the second condition, the fourth request message being used to request reporting of the corresponding first information.
22. The method of any one of claims 17-21, wherein, The method further comprises: transmitting a fifth request message to L2 perception nodes respectively, the fifth request message being used for requesting the perception nodes to report corresponding perception capabilities, L2 being a positive integer; receiving a fifth request response message from at least one of the L2 perception nodes, the fifth request response message including sixth information, the sixth information being used for indicating the perception capability of the perception node, wherein, the N2 3D perception nodes are at least one of the nodes having 3D perception capability in the at least one perception node.
23. The method of claim 22, wherein, the sixth information does not include the required perception capability information for calculation, and the method further includes: transmitting a sixth request message to at least one candidate 3D perception node, the sixth request message being used for requesting the 3D perception nodes to report the required perception capability information, the at least one candidate 3D perception node being at least one of the nodes having 3D perception capability in the at least one perception node, and the N2 3D perception nodes being at least one of the at least one candidate 3D perception node; receiving a sixth request response message from the at least one candidate 3D perception node, the sixth request response message including the required perception capability information corresponding to the 3D perception node.
24. The method of claim 22 or 23, wherein the sixth information includes specific information of at least one of the following of the corresponding perception node: supported perception dimension, supported perception modality, and supported perception parameter.
25. The method of claim 16 or 24, wherein the perception dimension includes 2D perception and / or 3D perception, the perception modality of the 2D perception includes optical perception and / or radar perception, and the perception modality of the 3D perception includes radio frequency perception and / or computer tomography.
26. A communications device, characterized by The communication device includes at least one processor coupled with at least one memory, the at least one processor being configured to execute computer programs or instructions stored in the at least one memory to cause the communication device to perform the method of any one of claims 1 to 25.
27. A communications device, characterized by The computer readable storage medium stores computer instructions, when the computer instructions are run on a computer, the method of any one of claims 1 to 25 is performed.
28. A computer-readable storage medium, characterized in that, The computer program product includes computer program codes, when the computer program codes are run on a computer, the method of any one of claims 1 to 25 is performed.
29. A computer program product, characterised in that,
Citation Information
Patent Citations
Context aware measurement
US20230141372A1
Joint 2d and 3D object tracking for autonomous systems and applications
US20230360231A1
Perception of 3D objects in sensor data
WO2023006836A1
KR20210081223A