A cooperative perception method, apparatus, device and medium
Patent Information
- Application Number
- CN202311711567.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2043-12-13
AI Technical Summary
[0004]本申请提供了一种协同感知方法、装置、设备及介质,可以解决协同感知结果的精确度低的问题
[0069]在本申请的实施例中,通过获取目标区域的多个智能体的传感器数据,然后分别针对每个智能体,基于智能体的传感器数据获取多个候选区域,并从所有智能体的所有候选区域中获取智能体的至少一个前景区域,再获取每个前景区域的多个兴趣点和多个关键点,并将每个前景区域划分为多个网格区域,并将网格区域的中心作为网格区域的特权点,然后基于每个前景区域中的所有关键点的坐标和兴趣点的坐标,获取前景区域中每个特权点的几何特征,再分别针对每个前景区域中的每个特权点,基于特权点的坐标和前景区域中每个关键点之间的坐标,获取特权点的偏移特征,最后基于每个智能体对应的所有特权点的几何特征和偏移特征,获取目标区域的协同感知结果。其中,从所有智能体的所有候选区域中获取智能体的至少一个前景区域,使得每个前景区域对应至少一个智能体,进而使得前景区域的数据来自至少一个智能体,提高前景区域的数据全面性和准确性,基于数据准确的前景区域的所有关键点和兴趣点获取特权点的几何特征,能够提高几何特征的准确度,基于特权点的坐标和前景区域中每个关键点的坐标,获取特权点的偏移特征,能够对特权点发生偏移的相关信息进行准确描述,基于准确的几何特征和偏移特征获取协同感知结果,能够提高协同感知结果的精确度。
Smart Images

Figure CN117522989B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a collaborative sensing method, apparatus, device, and medium. Background Technology
[0002] Due to limitations in the performance of their individual sensors, single-agent systems struggle to achieve efficient perception when obstacles are obstructed or located at considerable distances. Collaborative perception is an effective technique to address these shortcomings. By leveraging advanced network communication technologies, agents participating in collaborative perception can share their perceived environmental information, promoting more comprehensive perception and fundamentally overcoming the unavoidable problems of single-agent perception, such as occlusion and long-distance obstacles. This improves the overall accuracy of perception tasks and has significant application value in fields such as autonomous driving and drone swarms.
[0003] Existing collaborative sensing methods are mostly used to study collaborative sensing tasks under ideal conditions, and they have high communication overhead and low efficiency. However, in real communication environments, factors such as communication delay and positioning offset are unavoidable, resulting in low accuracy of collaborative sensing results. Summary of the Invention
[0004] This application provides a collaborative sensing method, apparatus, device, and medium that can solve the problem of low accuracy in collaborative sensing results.
[0005] In a first aspect, embodiments of this application provide a collaborative sensing method, which includes:
[0006] Acquire sensor data from multiple agents in the target area;
[0007] For each agent, multiple candidate regions are constructed based on the agent's sensor data, and at least one foreground region of the agent is obtained from all candidate regions of all agents. The candidate region is the area occupied by other objects within the coverage area of the agent, and the foreground region is either the candidate region of the agent or the candidate region of another agent.
[0008] Obtain multiple points of interest and multiple key points for each foreground region, and divide each foreground region into multiple grid regions, with the center point of each grid region serving as the privileged point of that grid region; points of interest are points at random locations within the foreground region, and key points are corner points or the center point of the foreground region;
[0009] Based on the coordinates of all key points and interest points in each foreground region, obtain the geometric features of each privileged point in the foreground region;
[0010] For each privileged point in each foreground region, the offset features of the privileged point are obtained based on the coordinates of the privileged point and the coordinates of each key point in the foreground region; the offset features are used to describe the positional offset of the privileged point.
[0011] Based on the geometric and offset features of all privileged points corresponding to each agent, the collaborative perception results of the target area are obtained.
[0012] Optionally, at least one foreground region of an agent is obtained from all candidate regions of all agents, including:
[0013] Through the formula:
[0014]
[0015] Obtain the set F of the foreground regions of the i-th agent. i ;
[0016] in, This represents the first foreground region of the i-th agent. This represents the second foreground region of the i-th agent. Represents the i-th intelligent agent. R A foreground area, B i This represents the set of candidate regions for the i-th agent. This represents the first candidate region for the i-th agent. This represents the second candidate region for the i-th agent. Represents the Nth intelligent agent of the i-th agent. i Candidate regions, ζ i B represents the information of the i-th agent. j This represents the set of candidate regions for the j-th agent. This represents the first candidate region for the j-th agent. This represents the second candidate region for the j-th agent. Represents the Nth intelligent agent of the j-th agent. j Candidate regions, ζ j This represents the information of the j-th agent, where i,j = 1, 2, ..., N. agent N agent This represents the total number of agents, Transform() represents coordinate transformation, and Association() represents the foreground region matching algorithm.
[0017] Foreground region matching algorithms include:
[0018] For each candidate region, iterate through each other candidate region and perform the following steps:
[0019] Perform intersection-union (IoU) matching between candidate regions and other candidate regions;
[0020] If a match is successful, it is determined whether other candidate regions are used as foreground regions. If other candidate regions are used as foreground regions, then other candidate regions are used as foreground regions of the agent corresponding to the candidate regions.
[0021] If all other candidate regions fail to match the candidate region, or if multiple other candidate regions that successfully match the candidate region are not used as foreground regions, then the candidate region is used as a foreground region for each agent.
[0022] Optionally, based on the coordinates of all keypoints and points of interest in each foreground region, the geometric features of each privileged point in the foreground region are obtained, including:
[0023] For each point of interest in each foreground region, the geometric code of the point of interest is obtained based on the coordinates of the point of interest and the coordinates of all key points in the foreground region corresponding to the point of interest.
[0024] For each privileged point in each foreground region, the geometric codes of all interest points located in the same grid region as the privileged point are aggregated to obtain the geometric features of the privileged point.
[0025] Optionally, based on the coordinates of the point of interest and the coordinates of all keypoints in the foreground region corresponding to the point of interest, the geometric encoding of the point of interest is obtained, including:
[0026] Through the formula:
[0027]
[0028] Calculate the geometric encoding of the k-th interest point in the t-th foreground region of the i-th agent.
[0029] Here, Concat() represents the concatenation operation, MLP() represents the multilayer perceptron, and ξ() represents the function that converts the three-dimensional Cartesian coordinate system to the spherical coordinate system. This represents the k-th point of interest in the r-th foreground region of the i-th agent. This represents the coordinates of the k-th point of interest in the r-th foreground region of the i-th agent. This represents the v-th keypoint in the r-th foreground region of the i-th agent. This represents the coordinates of the v-th keypoint in the r-th foreground region of the i-th agent, where x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the vertical axis, i = 1, 2, ..., N. agent N agentLet r represent the total number of agents, r = 1, 2, ..., i R i R Let k represent the total number of foreground regions for the i-th agent, where k = 1, 2, ..., r. K r K Let v = 1, 2, ..., r, representing the total number of interest points in the r-th foreground region of the i-th agent. v r v This represents the total number of key points in the r-th foreground region of the i-th agent.
[0030] Optionally, the geometric codes of all interest points located in the same grid region as the privileged point are aggregated to obtain the geometric features of the privileged point, including:
[0031] Through the formula:
[0032]
[0033] Calculate the geometric features of the m-th privileged point in the r-th foreground region of the i-th agent.
[0034] Where SA() represents the SA operation. This represents the m-th privileged point in the r-th foreground region of the i-th agent, where m = 1, 2, ..., r. m r m This represents the total number of privileged points in the r-th foreground region of the i-th agent.
[0035] Optionally, based on the coordinates of the privileged point and the coordinates of each keypoint in the foreground region, the offset features of the privileged point are obtained, including:
[0036] Through the formula:
[0037]
[0038] Obtain the offset feature of the m-th privileged point in the r-th foreground region of the i-th agent.
[0039] Here, Concat() represents the concatenation operation, and diff() represents the operation of obtaining the relative deviation between the privileged point and the key point. This represents the m-th privileged point in the r-th foreground region of the i-th agent. This represents the v-th keypoint in the r-th foreground region of the i-th agent. This represents the coordinates of the m-th privileged point in the r-th foreground region of the i-th agent. This represents the coordinates of the v-th keypoint in the r-th foreground region of the i-th agent, where x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the vertical axis, i = 1, 2, ..., N.agent N agent Let r represent the total number of agents, r = 1, 2, ..., i R i R Let v represent the total number of foreground regions for the i-th agent, where v = 1, 2, ..., r. v r v Let m represent the total number of keypoints in the r-th foreground region of the i-th agent, where m = 1, 2, ..., r. m r m This represents the total number of privileged points in the r-th foreground region of the i-th agent.
[0040] Optionally, based on the geometric and offset features of all privileged points corresponding to each agent, the collaborative perception results of the target region are obtained, including:
[0041] For each agent, collective perception information is obtained based on the geometric and offset features of all privileged points corresponding to the agent; the collective perception information is used to describe the information of all foreground regions of the agent.
[0042] By using a multilayer perceptron to perform feature interaction on the collective perceptual information of all agents, the initial comprehensive features of the target region are obtained.
[0043] A multi-head self-attention mechanism is used to extract features from the initial integrated features to obtain the final integrated features of the target region;
[0044] The final integrated features are input into the detection head to generate collaborative perception results for the target area.
[0045] Optionally, based on the geometric and offset features of all privileged points corresponding to the agent, the collective perception information of the agent is obtained, including:
[0046] Through the formula:
[0047]
[0048] Calculate the collective perception information C of the i-th agent. i ;
[0049] Where Concat() represents the concatenation operation, e i,r g represents the set of offset features of the privileged points in the r-th foreground region of the i-th agent. i,r The set of geometric features representing the privileged points of the r-th foreground region of the i-th agent, where r = 1, 2, ..., i R i R This represents the total number of foreground regions for the i-th agent, where i = 1, 2, ..., N. agent N agentIndicates the total number of intelligent agents;
[0050] A multi-head self-attention mechanism is used to extract features from the initial integrated features to obtain the final integrated features of the target region, including:
[0051] Through the formula:
[0052] H' = MSA(Q(H+E),K(H+E),V(H))
[0053] Obtain the comprehensive features H' of the target region;
[0054] Where MSE() represents the multi-head self-attention method, Q() represents the query linear layer, K() represents the key linear layer, V() represents the value linear layer, H represents the initial integrated features of the target region, and E represents the position embedding of all agents.
[0055] Through the formula:
[0056] O=MSA(Q(q),K(H'+E),V(H'))
[0057] Obtain the final comprehensive feature O of the target region;
[0058] Where q represents the query vector.
[0059] Secondly, embodiments of this application provide a collaborative sensing device, comprising:
[0060] The first acquisition module acquires sensor data from multiple agents in the target area;
[0061] The construction module constructs multiple candidate regions for each agent based on the agent's sensor data, and obtains at least one foreground region for the agent from all candidate regions of all agents. The candidate region is the area occupied by other objects within the agent's coverage area, and the foreground region is either the agent's candidate region or the candidate region of another agent.
[0062] The module divides the foreground region into multiple interest points and multiple key points, and divides each foreground region into multiple grid regions. The center point of each grid region is designated as the privileged point of the grid region. Interest points are points at random locations within the foreground region, and key points are corner points or center points of the foreground region.
[0063] The geometric feature acquisition module acquires the geometric features of each privileged point in the foreground region based on the coordinates of all key points and interest points in each foreground region.
[0064] The offset feature acquisition module acquires the offset features of each privileged point in each foreground region based on the coordinates of the privileged point and the coordinates of each key point in the foreground region; the offset features are used to describe the positional offset of the privileged point.
[0065] The second acquisition module acquires the collaborative perception results of the target area based on the geometric features and offset features of all privileged points corresponding to each agent.
[0066] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aforementioned collaborative perception method.
[0067] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned collaborative sensing method.
[0068] The above-mentioned solution in this application has the following beneficial effects:
[0069] In the embodiments of this application, sensor data of multiple agents in the target area are acquired. Then, for each agent, multiple candidate areas are acquired based on the agent's sensor data. At least one foreground area of the agent is acquired from all candidate areas of all agents. Then, multiple points of interest and multiple key points are acquired for each foreground area. Each foreground area is divided into multiple grid areas, and the center of the grid area is taken as the privileged point of the grid area. Then, based on the coordinates of all key points and points of interest in each foreground area, the geometric features of each privileged point in the foreground area are acquired. Then, for each privileged point in each foreground area, the offset features of the privileged point are acquired based on the coordinates of the privileged point and the coordinates between each key point in the foreground area. Finally, based on the geometric features and offset features of all privileged points corresponding to each agent, the collaborative perception result of the target area is acquired. Specifically, at least one foreground region of an agent is obtained from all candidate regions of all agents, such that each foreground region corresponds to at least one agent, thereby ensuring that the data of the foreground region comes from at least one agent, improving the comprehensiveness and accuracy of the foreground region data. Based on all key points and interest points of the foreground region with accurate data, the geometric features of the privileged point are obtained, which can improve the accuracy of the geometric features. Based on the coordinates of the privileged point and the coordinates of each key point in the foreground region, the offset features of the privileged point are obtained, which can accurately describe the relevant information of the privileged point offset. Based on the accurate geometric features and offset features, the collaborative perception results are obtained, which can improve the accuracy of the collaborative perception results.
[0070] Furthermore, the privileged point is the center point of the grid region, which can accurately represent the shape of the foreground region while reducing the amount of data, improving the efficiency of obtaining collaborative perception results, and reducing communication overhead.
[0071] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description
[0072] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0073] Figure 1 A flowchart of a collaborative sensing method provided in an embodiment of this application;
[0074] Figure 2 A schematic diagram illustrating the acquisition of geometric features according to an embodiment of this application;
[0075] Figure 3 This is a schematic diagram illustrating the acquisition of offset features according to an embodiment of this application;
[0076] Figure 4 This is a schematic diagram of the structure of a collaborative sensing device provided in an embodiment of this application;
[0077] Figure 5 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0078] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0079] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0080] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0081] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0082] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0083] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0084] To address the issue of low accuracy in existing collaborative perception results, this application provides a collaborative perception method. This method acquires sensor data from multiple agents in a target area. Then, for each agent, multiple candidate regions are obtained based on the agent's sensor data. At least one foreground region for each agent is then acquired from all candidate regions. Multiple points of interest and multiple key points are acquired for each foreground region. Each foreground region is divided into multiple grid regions, with the center of each grid region designated as a privileged point. Based on the coordinates of all key points and points of interest in each foreground region, the geometric features of each privileged point in the foreground region are obtained. For each privileged point in each foreground region, the offset features of the privileged point are obtained based on the coordinates of the privileged point and the coordinates between each key point in the foreground region. Finally, based on the geometric features and offset features of all privileged points corresponding to each agent, the collaborative perception result for the target area is obtained. Specifically, at least one foreground region of an agent is obtained from all candidate regions of all agents, such that each foreground region corresponds to at least one agent, thereby ensuring that the data of the foreground region comes from at least one agent, improving the comprehensiveness and accuracy of the foreground region data. Based on all key points and interest points of the foreground region with accurate data, the geometric features of the privileged point are obtained, which can improve the accuracy of the geometric features. Based on the coordinates of the privileged point and the coordinates of each key point in the foreground region, the offset features of the privileged point are obtained, which can accurately describe the relevant information of the privileged point offset. Based on the accurate geometric features and offset features, the collaborative perception results are obtained, which can improve the accuracy of the collaborative perception results.
[0085] Furthermore, the privileged point is the center point of the grid region, which can accurately represent the shape of the foreground region while reducing the amount of data, improving the efficiency of obtaining collaborative perception results, and reducing communication overhead.
[0086] The collaborative sensing method provided in this application will be illustrated below.
[0087] like Figure 1 As shown, the collaborative sensing method provided in this application includes the following steps:
[0088] Step 11: Obtain sensor data from multiple agents in the target area.
[0089] The aforementioned target area refers to the area requiring collaborative sensing, and the aforementioned intelligent agent refers to devices, systems, etc., that can perform sensing.
[0090] For example, the target area can be the area around a drone or an autonomous vehicle, and the intelligent agent can be a device or system with a sensing module, such as a drone, an autonomous vehicle, or a traffic facility. Sensor data of the intelligent agent can be obtained by accessing the intelligent agent's system.
[0091] Step 12: For each agent, construct multiple candidate regions based on the agent's sensor data, and obtain at least one foreground region of the agent from all candidate regions of all agents.
[0092] The aforementioned candidate region is the area occupied by other objects within the coverage range of the intelligent agent. The coverage range of the intelligent agent is the perception range of the intelligent agent's sensors. For example, if the intelligent agent is an autonomous vehicle, the candidate region is the area occupied by other vehicles, obstacles, and other objects within the perception range of the autonomous vehicle's sensors. For example, for a 4×4×6 meter cuboid obstacle, the area it occupies is a 4×4×6 meter cuboid area. The foreground region is the candidate region of the intelligent agent or the candidate region of other intelligent agents.
[0093] It should be noted that single-stage 3D target detectors and other devices can be used to process sensor data and construct multiple candidate regions.
[0094] In some embodiments of this application, the step of obtaining at least one foreground region of an agent from all candidate regions of all agents specifically involves:
[0095] Through the formula:
[0096]
[0097] Obtain the set F of the foreground regions of the i-th agent. i .
[0098] in, This represents the first foreground region of the i-th agent. This represents the second foreground region of the i-th agent. Represents the i-th intelligent agent. R A foreground area, B i This represents the set of candidate regions for the i-th agent. This represents the first candidate region for the i-th agent. This represents the second candidate region for the i-th agent. Represents the Nth intelligent agent of the i-th agent. i Candidate regions, ζ i B represents the information of the i-th agent. j This represents the set of candidate regions for the j-th agent. This represents the first candidate region for the j-th agent. This represents the second candidate region for the j-th agent. Represents the Nth intelligent agent of the j-th agent. jCandidate regions, ζ j This represents the information of the j-th agent, where i,j = 1, 2, ..., N. agent N agent This represents the total number of agents, Transform() represents coordinate transformation, and Association() represents the foreground region matching algorithm.
[0099] The foreground region matching algorithm mentioned above includes:
[0100] For each candidate region, iterate through each other candidate region and perform the following steps:
[0101] Perform intersection-union matching on the candidate regions and other candidate regions.
[0102] If a match is successful, it is determined whether other candidate regions are used as foreground regions. If other candidate regions are used as foreground regions, then the other candidate regions are used as foreground regions for each agent.
[0103] If all other candidate regions fail to match the candidate region, or if multiple other candidate regions that successfully match the candidate region are not used as foreground regions, then the candidate region is used as a foreground region of the agent corresponding to the candidate region.
[0104] For example, for the first candidate region of the first agent, the only candidate region that successfully matches the first candidate region after cross-intersection over union (CUI) matching is the fifth candidate region of the third agent, and the fifth candidate region of the third agent is not a foreground region. In this case, the first candidate region of the first agent is taken as a foreground region of each agent. When obtaining the foreground region of the third agent, the fifth candidate region of the third agent successfully matches the first candidate region of the first agent, and at this time, the first candidate region of the first agent is the foreground region of the third agent. In this case, no processing is performed, and cross-intersection over union (CUI) matching is performed on the next candidate region. For the third candidate region of the second agent, the candidate region that successfully matches the third candidate region after cross-intersection over union (CUI) matching is the first candidate region of the third candidate region, and the first candidate region of the third candidate region is a foreground region. In this case, the first candidate region of the third candidate region is also taken as a foreground region of the second agent.
[0105] It should be noted that the information ζ of the i-th agent mentioned above... i This includes sensor data for the i-th agent and the location of each candidate region. If the foreground region is not within the sensor range of the agent, the relevant data for that foreground region is empty.
[0106] It is worth mentioning that candidate regions constructed using sensor data from a single agent may become offset or distorted due to factors such as occlusion and data loss. By acquiring the foreground region, multiple candidate regions that have successfully matched among all candidate regions can be retained as a single foreground region, and the foreground region corresponds to one or more agents. This means that the information of the foreground region comes from one or more agents, thereby improving the accuracy and comprehensiveness of the foreground region information.
[0107] Step 13: Obtain multiple points of interest and multiple key points for each foreground region, and divide each foreground region into multiple grid regions, using the center point of each grid region as the privileged point of that grid region.
[0108] The points of interest mentioned above are points at random locations in the foreground region, while key points are corner points or center points of the foreground region.
[0109] For example, for a 4×4×6 meter cuboid foreground region, multiple points of interest are obtained by randomly selecting points in the foreground region. The eight corner points and one center point of the foreground region are all used as key points, resulting in a total of nine key points. The foreground region is divided into four 2×2×3 meter cuboid grid regions, and the center point of each grid region is used as a privileged point, resulting in a total of four privileged points.
[0110] It is worth mentioning that privileged points are the center points of the grid area, and the combination of all privileged points can accurately represent the spatial shape of the foreground area. Key points are corner points or center points, which can describe the size and shape of the foreground area. Key points are random points in the foreground area, which facilitates subsequent processing of the foreground area information.
[0111] Step 14: Based on the coordinates of all key points and points of interest in each foreground region, obtain the geometric features of each privileged point in the foreground region.
[0112] The coordinates of the key points and points of interest mentioned above are all in the world coordinate system.
[0113] In some embodiments of this application, the step of obtaining the geometric features of each privileged point in the foreground region based on the coordinates of all key points and points of interest in each foreground region specifically includes:
[0114] The first step is to obtain the geometric code of each point of interest in each foreground region, based on the coordinates of the point of interest and the coordinates of all key points in the foreground region corresponding to the point of interest.
[0115] Specifically, through the formula:
[0116]
[0117] Calculate the geometric encoding of the k-th interest point in the r-th foreground region of the i-th agent.
[0118] Here, Concat() represents the concatenation operation, MLP() represents the multilayer perceptron, and ξ() represents the function that converts the three-dimensional Cartesian coordinate system to the spherical coordinate system. This represents the k-th point of interest in the r-th foreground region of the i-th agent. This represents the coordinates of the k-th point of interest in the r-th foreground region of the i-th agent. This represents the v-th keypoint in the r-th foreground region of the i-th agent. This represents the coordinates of the v-th keypoint in the r-th foreground region of the i-th agent, where x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the vertical axis, i = 1, 2, ..., N. agent N agent Let r represent the total number of agents, r = 1, 2, ..., i R i R Let k represent the total number of foreground regions for the i-th agent, where k = 1, 2, ..., r. K r K Let v = 1, 2, ..., r, representing the total number of interest points in the r-th foreground region of the i-th agent. v r v This represents the total number of key points in the r-th foreground region of the i-th agent.
[0119] The second step involves aggregating the geometric codes of all interest points located in the same grid region as the privileged point for each privileged point in each foreground region, thereby obtaining the geometric features of the privileged point.
[0120] Specifically, through the formula:
[0121]
[0122] Calculate the geometric features of the m-th privileged point in the r-th foreground region of the i-th agent.
[0123] Where SA() represents the SA operation. This represents the m-th privileged point in the r-th foreground region of the i-th agent, where m = 1, 2, ..., r. m r m This represents the total number of privileged points in the r-th foreground region of the i-th agent.
[0124] It should be noted that the geometric encoding described above is used to represent the distance relationship between points of interest and all corresponding keypoints. The SetAbstraction (SA) operation consists of three parts: a sampling layer, which selects feature points as centroids; a grouping layer, which determines the scale, finds the neighboring points of the centroid, and constructs local regions; and a PointNet layer, which extracts features from each local region. Specifically, it transforms the coordinates of all neighboring points in a local region into coordinates relative to the centroid, and then uses the PointNet to extract features, capturing the point-to-point relationships within the local region through relative coordinates.
[0125] For example, the foreground region is divided into four 2×2×3 meter cuboid grid regions. The first grid region has three points of interest. The geometric codes of these three points of interest are aggregated to obtain the geometric features of the privileged points in the first grid region.
[0126] It is worth mentioning that by aggregating the geometric encoding of points of interest, the amount of data can be reduced and the processing efficiency can be improved. Furthermore, by obtaining the geometric features of privileged points based on all key points and points of interest, the accuracy of geometric features can be improved.
[0127] The steps for obtaining the geometric features of privileged points are illustrated below with a specific example.
[0128] like Figure 2 As shown in Figure a, Figure a is a schematic diagram of the process of obtaining the geometric codes of interest points, and Figure b is a schematic diagram of the process of aggregating the geometric codes of interest points. In the figures, black circles represent key points, black dots represent interest points, white circles represent privileged points, black boxes represent the top view of the foreground area, straight lines with arrows represent the distance between interest points and key points, and dashed lines represent the aggregation of the geometric codes of interest points to privileged points.
[0129] Step 15: For each privileged point in each foreground region, obtain the offset features of the privileged point based on the coordinates of the privileged point and the coordinates of each key point in the foreground region.
[0130] The coordinates of the privileged points and key points mentioned above are all in the world coordinate system.
[0131] The aforementioned offset features are used to describe the positional offset of privileged points.
[0132] Specifically, through the formula:
[0133]
[0134] Obtain the offset feature of the m-th privileged point in the r-th foreground region of the i-th agent.
[0135] Here, Concat() represents the concatenation operation, and diff() represents the operation of obtaining the relative deviation between the privileged point and the key point. This represents the m-th privileged point in the r-th foreground region of the i-th agent. This represents the v-th keypoint in the r-th foreground region of the i-th agent. This represents the coordinates of the m-th privileged point in the r-th foreground region of the i-th agent. This represents the coordinates of the v-th keypoint in the r-th foreground region of the i-th agent, where x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the vertical axis, i = 1, 2, ..., N. agent N agent Let r represent the total number of agents, r = 1, 2, ..., i R i R Let v represent the total number of foreground regions for the i-th agent, where v = 1, 2, ..., r. v r v Let m represent the total number of keypoints in the r-th foreground region of the i-th agent, where m = 1, 2, ..., r. m r m This represents the total number of privileged points in the r-th foreground region of the i-th agent.
[0136] For example, the above formula can be run using computer software for data computation such as Matlab and Mathematica to obtain the offset characteristics of each privileged point.
[0137] It is worth mentioning that by obtaining the offset features of the privileged point based on the coordinates of the privileged point and the coordinates of each key point in the foreground region, the relevant information about the offset of the privileged point can be accurately described.
[0138] The steps for obtaining the offset features of privileged points are illustrated below with a specific example.
[0139] A schematic diagram of obtaining the offset characteristics of privileged points is shown below. Figure 3 As shown in the figure, black circles represent key points, white circles represent privilege points, two black boxes represent top views of the foreground area, and straight lines with arrows represent the distance between privilege points and key points.
[0140] Step 16: Based on the geometric features and offset features of all privileged points corresponding to each agent, obtain the collaborative perception results of the target area.
[0141] The aforementioned collaborative perception results include the positions of multiple objects in the target area, as well as information such as the size and shape of each object. For example, if the target area is a certain range around an autonomous vehicle, then the collaborative perception results for that area include the size, shape, and position of other vehicles, obstacles, and other objects within a certain range around the autonomous vehicle.
[0142] In some embodiments of this application, the steps for obtaining the collaborative perception results of the target region based on the geometric features and offset features of all privileged points corresponding to each agent specifically include:
[0143] The first step is to obtain the collective perception information of each agent based on the geometric features and offset features of all privileged points corresponding to the agent.
[0144] The aforementioned collective perception information is used to describe information about all foreground regions of the agent.
[0145] Specifically, through the formula:
[0146]
[0147] Calculate the collective perception information C of the i-th agent. i .
[0148] Where Concat() represents the concatenation operation, e i,r g represents the set of offset features of the privileged points in the r-th foreground region of the i-th agent. i,r The set of geometric features representing the privileged points of the r-th foreground region of the i-th agent, where r = 1, 2, ..., i R i R This represents the total number of foreground regions for the i-th agent, where i = 1, 2, ..., N. agent N agent This represents the total number of intelligent agents.
[0149] The second step involves using a multilayer perceptron to perform feature interaction on the collective perception information of all agents, thereby obtaining the initial comprehensive features of the target region.
[0150] It should be noted that the aforementioned multilayer perceptron can be a hybrid multilayer perceptron (MLP Mixer). When using the multilayer perceptron to perform feature interaction on the collective perceptual information of all agents, the collective perceptual information is embedded in the 3D scene, including the three dimensions of x, y, and z. During feature interaction, token and channel dimension mixing is performed in the order of x, y, and z, respectively. This makes it possible to independently fuse spatial and channel information, thereby extracting local and spatial information from all collective perceptual information to obtain the initial comprehensive features of the target region.
[0151] The third step is to use a multi-head self-attention mechanism to extract features from the initial integrated features to obtain the final integrated features of the target region.
[0152] Specifically, through the formula:
[0153] H' = MSA(Q(H+E),K(H+E),V(H))
[0154] Obtain the comprehensive features H' of the target region.
[0155] In this context, MSA() represents the multi-head self-attention method, Q() represents the query linear layer, K() represents the key linear layer, V() represents the value linear layer, H represents the initial integrated features of the target region, and E represents the position embeddings of all agents.
[0156] Through the formula:
[0157] O=MSA(Q(q),K(H'+E),V(H'))
[0158] Obtain the final comprehensive feature O of the target region.
[0159] Where q represents the query vector.
[0160] It should be noted that the aforementioned position embedding E is obtained by processing the agent's position information through a multilayer perceptron and is a learnable parameter used to represent the agent's position.
[0161] The fourth step is to input the final integrated features into the detection head to generate collaborative perception results for the target area.
[0162] For example, the detection head described above could be a query-based Transformer decoder.
[0163] In some embodiments of this application, the above steps can be run using computer software such as PreScan and PTVVissim to obtain the collaborative perception results of the target area.
[0164] It is worth mentioning that at least one foreground region of an agent is obtained from all candidate regions of all agents, so that each foreground region corresponds to at least one agent. This ensures that the data of the foreground region comes from at least one agent, improving the comprehensiveness and accuracy of the foreground region data. Based on all key points and interest points of the foreground region with accurate data, the geometric features of the privileged point are obtained, which can improve the accuracy of the geometric features. Based on the coordinates of the privileged point and the coordinates of each key point in the foreground region, the offset features of the privileged point are obtained, which can accurately describe the relevant information of the offset of the privileged point. Based on the accurate geometric features and offset features, the collaborative perception results are obtained, which can improve the accuracy of the collaborative perception results.
[0165] Furthermore, the privileged point is the center point of the grid region, which can accurately represent the shape of the foreground region while reducing the amount of data, improving the efficiency of obtaining collaborative perception results, and reducing communication overhead.
[0166] The collaborative sensing device provided in this application is described below as an example.
[0167] like Figure 4 As shown, this application embodiment provides a cooperative sensing device, the cooperative sensing device 400 including:
[0168] The first acquisition module 401 acquires sensor data from multiple intelligent agents in the target area;
[0169] Module 402 constructs multiple candidate regions for each agent based on the agent's sensor data, and obtains at least one foreground region of the agent from all candidate regions of all agents; the candidate region is the area occupied by other objects within the coverage area of the agent; the foreground region is either the candidate region of the agent or the candidate region of another agent.
[0170] The segmentation module 403 obtains multiple points of interest and multiple key points for each foreground region, and divides each foreground region into multiple grid regions, with the center point of each grid region serving as the privileged point of that grid region; points of interest are points at random locations within the foreground region, and key points are corner points or the center point of the foreground region;
[0171] The geometric feature acquisition module 404 acquires the geometric features of each privileged point in the foreground region based on the coordinates of all key points and interest points in each foreground region;
[0172] The offset feature acquisition module 405 acquires the offset feature of each privileged point in each foreground region based on the coordinates of the privileged point and the coordinates of each key point in the foreground region; the offset feature is used to describe the positional offset of the privileged point.
[0173] The second acquisition module 406 acquires the collaborative perception results of the target area based on the geometric features and offset features of all privileged points corresponding to each agent.
[0174] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0175] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0176] like Figure 5 As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 5 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.
[0177] Specifically, when the processor D100 executes the computer program D102, it acquires sensor data from multiple agents in the target area, then acquires multiple candidate areas for each agent based on the agent's sensor data, and acquires at least one foreground area for each agent from all candidate areas of all agents. It then acquires multiple points of interest and multiple key points for each foreground area, divides each foreground area into multiple grid areas, and uses the center of each grid area as the privileged point. Based on the coordinates of all key points and points of interest in each foreground area, it acquires the geometric features of each privileged point in the foreground area. Then, for each privileged point in each foreground area, based on the coordinates of the privileged point and the coordinates between each key point in the foreground area, it acquires the offset features of the privileged point. Finally, based on the geometric features and offset features of all privileged points corresponding to each agent, it acquires the collaborative perception result of the target area. Specifically, at least one foreground region of an agent is obtained from all candidate regions of all agents, such that each foreground region corresponds to at least one agent, thereby ensuring that the data of the foreground region comes from at least one agent, improving the comprehensiveness and accuracy of the foreground region data. Based on all key points and interest points of the foreground region with accurate data, the geometric features of the privileged point are obtained, which can improve the accuracy of the geometric features. Based on the coordinates of the privileged point and the coordinates of each key point in the foreground region, the offset features of the privileged point are obtained, which can accurately describe the relevant information of the privileged point offset. Based on the accurate geometric features and offset features, the collaborative perception results are obtained, which can improve the accuracy of the collaborative perception results.
[0178] Furthermore, the privileged point is the center point of the grid region, which can accurately represent the shape of the foreground region while reducing the amount of data, improving the efficiency of obtaining collaborative perception results, and reducing communication overhead.
[0179] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0180] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.
[0181] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0182] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.
[0183] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the collaborative sensing method device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0184] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0185] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0186] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A collaborative sensing method, characterized in that, include: Acquire sensor data from multiple agents in the target area; For each of the intelligent agents, multiple candidate regions are constructed based on the sensor data of the intelligent agent, and at least one foreground region of the intelligent agent is obtained from all candidate regions of all intelligent agents; the candidate region is the area occupied by other objects within the coverage area of the intelligent agent, and the foreground region is the candidate region of the intelligent agent or the candidate region of other intelligent agents. Multiple points of interest and multiple key points are obtained for each foreground region, and each foreground region is divided into multiple grid regions, with the center point of each grid region being used as the privileged point of that grid region. The point of interest is a point at a random location in the foreground region, and the key point is a corner point or the center point of the foreground region; Based on the coordinates of all key points and points of interest in each foreground region, the geometric features of each privileged point in the foreground region are obtained; For each privileged point in each foreground region, an offset feature is obtained based on the coordinates of the privileged point and the coordinates of each key point in the foreground region; the offset feature is used to describe the positional offset of the privileged point. Based on the geometric and offset features of all privileged points corresponding to each agent, the collaborative perception results of the target region are obtained.
2. The collaborative sensing method according to claim 1, characterized in that, The step of obtaining at least one foreground region of the agent from all candidate regions of all agents includes: Through the formula: Obtain the set F of the foreground regions of the i-th agent. i ; in, This represents the first foreground region of the i-th agent. This represents the second foreground region of the i-th agent. The i-th intelligent agent represents the i-th intelligent agent. R A foreground area, B i This represents the set of candidate regions for the i-th agent. This represents the first candidate region of the i-th agent. This represents the second candidate region of the i-th agent. The Nth term of the i-th agent is represented by i Candidate regions, ζ i B represents the information of the i-th agent. j This represents the set of candidate regions for the j-th agent. This represents the first candidate region of the j-th agent. This represents the second candidate region of the j-th agent. The Nth term of the j-th agent is represented by... j Candidate regions, ζ j This represents the information of the j-th agent, where i,j = 1, 2, ..., N. agent N agent This represents the total number of the intelligent agents, Transform() represents coordinate transformation, and Association() represents the foreground region matching algorithm. The foreground region matching algorithm includes: For each candidate region, iterate through each other candidate region and perform the following steps: Perform intersection-union (IUU) matching on the candidate regions and the other candidate regions; If a match is successful, it is determined whether the other candidate regions are used as foreground regions. If the other candidate regions are used as foreground regions, the other candidate regions are used as foreground regions of the agent corresponding to the candidate regions. If all other candidate regions fail to match the candidate region, or if multiple other candidate regions that successfully match the candidate region are not used as foreground regions, then the candidate region is used as a foreground region for each agent.
3. The collaborative sensing method according to claim 1, characterized in that, The process of obtaining the geometric features of each privileged point in the foreground region based on the coordinates of all key points and points of interest in each foreground region includes: For each point of interest in each foreground region, the geometric code of the point of interest is obtained based on the coordinates of the point of interest and the coordinates of all key points in the foreground region corresponding to the point of interest. For each privileged point in each foreground region, the geometric codes of all interest points located in the same grid region as the privileged point are aggregated to obtain the geometric features of the privileged point.
4. The collaborative sensing method according to claim 3, characterized in that, The process of obtaining the geometric code of the point of interest based on its coordinates and the coordinates of all key points in the foreground region corresponding to the point of interest includes: Through the formula: Calculate the geometric encoding of the k-th interest point in the r-th foreground region of the i-th agent. Here, Concat() represents the concatenation operation, MLP() represents the multilayer perceptron, and ξ() represents the function that converts the three-dimensional Cartesian coordinate system to the spherical coordinate system. This represents the k-th point of interest in the r-th foreground region of the i-th agent. This represents the coordinates of the k-th point of interest in the r-th foreground region of the i-th agent. This represents the v-th key point in the r-th foreground region of the i-th agent. This represents the coordinates of the v-th keypoint in the r-th foreground region of the i-th agent, where x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the vertical axis, i = 1, 2, ..., N. agent N agent Let r represent the total number of agents, r = 1, 2, ..., i R i R This represents the total number of foreground regions for the i-th agent, k = 1, 2, ..., r K r K This represents the total number of interest points in the r-th foreground region of the i-th agent, where v = 1, 2, ..., r v r v This represents the total number of key points in the r-th foreground region of the i-th agent.
5. The collaborative sensing method according to claim 4, characterized in that, The aggregation of the geometric codes of all interest points located in the same grid region as the privileged point to obtain the geometric features of the privileged point includes: Through the formula: Calculate the geometric features of the m-th privileged point in the r-th foreground region of the i-th agent. Where SA() represents the SA operation. This represents the m-th privileged point in the r-th foreground region of the i-th agent, where m = 1, 2, ..., r. m r m This represents the total number of privileged points in the r-th foreground region of the i-th agent.
6. The collaborative sensing method according to claim 1, characterized in that, The process of obtaining the offset features of the privileged point based on its coordinates and the coordinates of each key point in the foreground region includes: Through the formula: Obtain the offset feature of the m-th privileged point in the r-th foreground region of the i-th agent. Here, Concat() represents the concatenation operation, and diff() represents the operation of obtaining the relative deviation between the privileged point and the key point. This represents the m-th privileged point in the r-th foreground region of the i-th agent. This represents the v-th key point in the r-th foreground region of the i-th agent. This represents the coordinates of the m-th privileged point in the r-th foreground region of the i-th agent. This represents the coordinates of the v-th keypoint in the r-th foreground region of the i-th agent, where x represents the horizontal coordinate, y represents the vertical coordinate, and z represents the vertical axis, i = 1, 2, ..., N. agent N agent Let r represent the total number of agents, r = 1, 2, ..., i R i R Let v represent the total number of foreground regions for the i-th agent, where v = 1, 2, ..., r. v r v This represents the total number of keypoints in the r-th foreground region of the i-th agent, where m = 1, 2, ..., r. m r m This represents the total number of privileged points in the t-th foreground region of the i-th agent.
7. The collaborative sensing method according to claim 1, characterized in that, The method of obtaining the collaborative perception result of the target region based on the geometric features and offset features of all privileged points corresponding to each agent includes: For each agent, collective perception information of the agent is obtained based on the geometric features and offset features of all privileged points corresponding to the agent; the collective perception information is used to describe the information of all foreground regions of the agent. By using a multilayer perceptron to perform feature interaction on the collective perceptual information of all agents, the initial comprehensive features of the target region are obtained. The initial integrated features are extracted using a multi-head self-attention mechanism to obtain the final integrated features of the target region; The final integrated features are input into the detection head to generate the collaborative perception results of the target region.
8. The collaborative sensing method according to claim 7, characterized in that, The process of obtaining collective perception information of the agent based on the geometric features and offset features of all privileged points corresponding to the agent includes: Through the formula: Calculate the collective perception information C of the i-th agent. i ; Where Concat() represents the concatenation operation, e i,r g represents the set of offset features of the privileged point in the r-th foreground region of the i-th agent. i,r The set of geometric features representing the privileged points of the r-th foreground region of the i-th agent, where r = 1, 2, ..., i R i R This represents the total number of foreground regions for the i-th agent, where i = 1, 2, ..., N. agent N agent This represents the total number of the intelligent agents; The step of using a multi-head self-attention mechanism to extract features from the initial integrated features to obtain the final integrated features of the target region includes: Through the formula: G'=MSA(Q(H+E),K(H+E),V(H)) Obtain the comprehensive feature H' of the target region; Where MSA() represents the multi-head self-attention method, Q() represents the query linear layer, K() represents the key linear layer, V() represents the value linear layer, H represents the initial integrated features of the target region, and E represents the position embedding of all agents. Through the formula: O=MSA(Q(q),K(H'+E),V(H')) Obtain the final comprehensive feature O of the target region; Where q represents the query vector.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the collaborative perception method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the collaborative perception method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method and system for realizing combined semantic hierarchical connection model based on panoramic area scene perception
CN110533048A
Travelable area generation method and device, electronic equipment and storage medium
CN116311114A