Methods for determining object information and storage media
By generating and analyzing point cloud data of traffic roads, and extracting spatial and temporal features, the problem of low obstacle recognition efficiency in autonomous driving is solved, and efficient determination and safe control of obstacle information are achieved.
Patent Information
- Application Number
- CN202210680775.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-06-16
AI Technical Summary
In the process of autonomous driving, the existing technology of using LiDAR to identify obstacles is inefficient and cannot effectively determine the movement information of obstacles.
By detecting traffic roads to generate point cloud data, spatial and temporal features within the grid area are extracted, and the speed of obstacles is analyzed to improve the efficiency of obstacle information determination.
It enables efficient determination of obstacle information, accurately responds to various types of obstacles, and improves the safety and efficiency of autonomous driving.
Smart Images

Figure CN115116031B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a method for determining information about an object and a storage medium. Background Technology
[0002] Currently, acquiring motion information of all dynamic obstacles is a prerequisite for ensuring the safety of autonomous driving decision-making and control during the autonomous driving process.
[0003] In related technologies, lidar is typically used as a high-precision three-dimensional ranging device in autonomous driving perception systems to identify and determine obstacles and other situations in road scenes. However, obstacle categories exhibit a wide variety of motion behaviors, resulting in the technical problem of low efficiency in determining objects using the aforementioned methods.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This invention provides a method and storage medium for determining object information, thereby at least solving the technical problem of low efficiency in determining objects.
[0006] According to one aspect of the present invention, a method for determining object information is provided, comprising: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data located within at least one grid range in a target view of the traffic roads; extracting spatial features and temporal features from the point cloud data within any one grid range, wherein the spatial features are used to characterize the spatial information of the grid range and the temporal features are used to characterize the temporal information of the grid range; and analyzing the speed of obstacles located within the grid range based on the spatial features and the temporal features.
[0007] According to another aspect of the present invention, another method for determining object information is also provided, comprising: determining the traffic road on which the vehicle is traveling; retrieving point cloud data of the traffic road, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one grid area in a target view of the traffic road; extracting spatial features and temporal features from the point cloud data within any one grid area, wherein the spatial features are used to characterize the spatial information of the grid area and the temporal features are used to characterize the temporal information of the grid area; analyzing the speed of obstacles located within the grid area based on the spatial features and temporal features; and controlling the vehicle to avoid obstacles based on the speed of the obstacles.
[0008] According to another aspect of the present invention, another method for determining object information is also provided, comprising: responding to a data input instruction acting on an operating interface, displaying point cloud data of a traffic road on the operating interface, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one grid range in a target view of the traffic road; and responding to an information generation instruction acting on the operating interface, displaying the speed of obstacles within the grid range on the operating interface, wherein the speed of the obstacles is obtained based on spatial and temporal features within any one grid range of the point cloud data, the spatial features being used to characterize the spatial information of the grid range and the temporal features being used to characterize the temporal information of the grid range.
[0009] According to another aspect of the present invention, another method for determining object information is also provided, comprising: displaying point cloud data of traffic roads on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid range in a target view of the traffic roads; extracting spatial features and temporal features from the point cloud data within any one grid range, wherein the spatial features are used to characterize the spatial information of the grid range and the temporal features are used to characterize the temporal information of the grid range; analyzing the speed of obstacles located within the grid range based on the spatial features and temporal features; and driving the VR device or AR device to display the speed of the obstacles.
[0010] According to one aspect of the present invention, an object information determination apparatus is provided, comprising: a generation unit for generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data located within at least one grid area in a target view of the traffic roads; a first extraction unit for extracting spatial features and temporal features from the point cloud data within any one grid area, wherein the spatial features are used to characterize the spatial information of the grid area and the temporal features are used to characterize the temporal information of the grid area; and a first analysis unit for analyzing the speed of obstacles located within the grid area based on the spatial features and the temporal features.
[0011] According to another aspect of the present invention, another object information determination apparatus is also provided, comprising: a determination unit, configured to determine a traffic road in which a vehicle is traveling; retrieve point cloud data of the traffic road, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one grid area in a target view of the traffic road; a second extraction unit, configured to extract spatial features and temporal features from the point cloud data within any one grid area, wherein the spatial features characterize the spatial information of the grid area and the temporal features characterize the temporal information of the grid area; a second analysis unit, configured to analyze the speed of obstacles located within the grid area based on the spatial features and temporal features; and control the vehicle to avoid obstacles based on the speed of the obstacles.
[0012] According to another aspect of the present invention, another object information determination apparatus is also provided, comprising: a first display unit, configured to display point cloud data of a traffic road on the operation interface in response to a data input command applied to an operation interface, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one grid range in a target view of the traffic road; and a second display unit, configured to display the speed of obstacles within the grid range on the operation interface in response to an information generation command applied to the operation interface, wherein the speed of the obstacles is obtained based on spatial and temporal features within any one grid range of the point cloud data, the spatial features being used to characterize the spatial information of the grid range, and the temporal features being used to characterize the temporal information of the grid range.
[0013] According to another aspect of the present invention, another object information determination apparatus is also provided, comprising: a presentation unit for displaying point cloud data of a traffic road on a presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one grid range in a target view of the traffic road; a third extraction unit for extracting spatial features and temporal features from the point cloud data within any one grid range, wherein the spatial features characterize the spatial information of the grid range and the temporal features characterize the temporal information of the grid range; a third analysis unit for analyzing the speed of obstacles located within the grid range based on the spatial features and temporal features; and a driving unit for driving the VR device or AR device to display the speed of the obstacles.
[0014] According to another aspect of the present invention, an object information confirmation system is also provided, comprising: a processor; and a memory connected to the processor for providing the processor with instructions to process the following steps: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data located within at least one grid range in a target view of the traffic roads; extracting spatial features and temporal features from the point cloud data within any one grid range, wherein the spatial features are used to characterize the spatial information of the grid range and the temporal features are used to characterize the temporal information of the grid range; and analyzing the speed of obstacles located within the grid range based on the spatial features and temporal features.
[0015] According to another aspect of the present invention, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, the device where the storage medium is located executes the method for determining information of any of the above-mentioned objects.
[0016] According to another aspect of the present invention, a processor is also provided, which is used to run a program, wherein the method for determining the information of any of the above-mentioned objects is executed during program execution.
[0017] In this embodiment of the invention, point cloud data is generated by detecting traffic roads. The point cloud data represents object data within at least one grid area in the target view of the traffic road. Spatial and temporal features are extracted from the point cloud data within any grid area. The spatial features represent the spatial information of the grid area, and the temporal features represent the temporal information of the grid area. Based on the spatial and temporal features, the speed of obstacles located within the grid area is analyzed. In other words, in this embodiment of the invention, point cloud data generated by detecting traffic roads is used as input. Effective temporal and spatial features are extracted from it, and then the information of objects on the grid occupied by the point cloud data is accurately determined based on the extracted temporal and spatial features. This effectively addresses all types of obstacles, thereby improving the efficiency of determining object information and solving the technical problem of low efficiency in object determination. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0019] Figure 1 This is a hardware structure block diagram of a virtual reality device according to an embodiment of the present invention, which describes a method for determining information about an object.
[0020] Figure 2This is a flowchart of an object information determination method according to an embodiment of the present invention;
[0021] Figure 3 This is a flowchart of another method for determining object information according to an embodiment of the present invention;
[0022] Figure 4 This is a flowchart of another method for determining object information according to an embodiment of the present invention;
[0023] Figure 5 This is a flowchart of another method for determining object information according to an embodiment of the present invention;
[0024] Figure 6 This is a schematic diagram illustrating the result of determining the information of an object according to an embodiment of the present invention;
[0025] Figure 7 This is a flowchart of a laser point cloud full-scene velocity estimation method based on a spatiotemporal bidirectional augmentation network according to an embodiment of the present invention;
[0026] Figure 8 This is a schematic diagram of a spatiotemporal bidirectional augmentation network framework according to an embodiment of the present invention;
[0027] Figure 9 This is a schematic diagram of a time-enhanced spatial module according to an embodiment of the present invention;
[0028] Figure 10 This is a schematic diagram of a space-enhanced time module according to an embodiment of the present invention;
[0029] Figure 11 This is a schematic diagram illustrating the effect of determining object information according to an embodiment of the present invention;
[0030] Figure 12 This is a schematic diagram of an object information determination device according to an embodiment of the present invention;
[0031] Figure 13 This is a schematic diagram of another object information determination device according to an embodiment of the present invention;
[0032] Figure 14 This is a schematic diagram of another object information determination device according to an embodiment of the present invention;
[0033] Figure 15 This is a schematic diagram of an object information determination device according to an embodiment of the present invention;
[0034] Figure 16 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Detailed Implementation
[0035] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0037] Example 1
[0038] According to an embodiment of the present invention, an embodiment of a method for determining information of an object is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0039] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the present invention, which describes a method for determining object information. Figure 1 As shown, the virtual reality device 104 is connected to the terminal 106, and the terminal 106 is connected to the server 102 via a network. The virtual reality device 104 is not limited to virtual reality headsets, virtual reality glasses, virtual reality all-in-one machines, etc. The terminal 106 is not limited to PCs, mobile phones, tablets, etc. The server 102 can be a server corresponding to a media file operator. The network includes, but is not limited to, wide area networks, metropolitan area networks, or local area networks.
[0040] Optionally, the virtual reality device 104 in this embodiment includes a memory, a processor, and a transmission device. The memory stores an application program that can be used to perform: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data within at least one grid area in a target view of the traffic roads; extracting spatial and temporal features from the point cloud data within any one grid area, wherein the spatial features characterize the spatial information of the grid area, and the temporal features characterize the temporal information of the grid area; and analyzing the speed of obstacles located within the grid area based on the spatial and temporal features, thereby solving the technical problem of low efficiency in object detection and achieving the goal of improving the efficiency of object detection.
[0041] The terminal in this embodiment can be used to display point cloud data of traffic segments on the display screen of a virtual reality (VR) device or an augmented reality (AR) device. The point cloud data is generated by detecting traffic roads and is used to represent object data within at least one grid area in the target view of the traffic road. Spatial and temporal features are extracted from the point cloud data within any one grid area, where the spatial features represent the spatial information of the grid area and the temporal features represent the temporal information of the grid area. Based on the spatial and temporal features, the speed of obstacles located within the grid area is analyzed. The VR or AR device is driven to display the speed of the obstacles. Information is output to the virtual reality device 104, which, upon receiving the prompt information, displays the speed of the obstacles at the target projection location.
[0042] Optionally, the virtual reality device 104 in this embodiment includes an eye-tracking head-mounted display (HMD) and an eye-tracking module that function the same as in the embodiments described above. That is, the screen in the HMD displays real-time images, and the eye-tracking module in the HMD acquires the real-time movement path of the user's eyes. In this embodiment, the terminal acquires the user's position and movement information in real three-dimensional space through a tracking system, and calculates the three-dimensional coordinates of the user's head in virtual three-dimensional space, as well as the user's field of vision orientation in virtual three-dimensional space.
[0043] Figure 1 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned AR / VR device (or mobile device), but also as an exemplary block diagram of the aforementioned server. Under the operating environment described above, this invention provides... Figure 2 The method for determining object information shown is illustrated. It should be noted that the object information determination method in this embodiment can be derived from... Figure 1The mobile terminal in the illustrated embodiment is executed.
[0044] Figure 2 This is a flowchart of an object information determination method according to an embodiment of the present invention, such as... Figure 2 As shown, the method may include the following steps:
[0045] Step S202: Point cloud data is generated by detecting traffic roads, wherein the point cloud data is used to characterize object data within at least one raster range in the target view of the traffic roads.
[0046] In the technical solution provided by step S202 of the present invention, traffic roads are detected, and point cloud data generated by the detected traffic roads is obtained. The point cloud data can be used to characterize the data of objects located within at least one grid range in the target view of the traffic road. It can be a set of points in three-dimensional space, such as discretized laser point cloud data, lidar point cloud, lidar point cloud, etc. The type of point cloud data is not specifically limited here. The target view can be a view of a specific object, scene, or object, such as a view of a selected location in the traffic road, or a top view of a specific scene or object. The grid can be a grid obtained by discretizing the point cloud data, such as a top view grid.
[0047] Optionally, traffic roads can be detected to determine the point cloud data of the target view. This can be achieved by discretizing the point cloud data into a grid under the top view, thereby generating object data within at least one grid range in the target view of the traffic roads.
[0048] Optionally, point cloud data is generated by detecting traffic roads. The acquired time-series point cloud data is used as input, and the point cloud data can be sorted according to time order to obtain time-series point cloud data. The detected point cloud data is then discretized to achieve the purpose of discretizing the point cloud data into a raster under the top view.
[0049] For example, a lidar device can be installed in a vehicle to scan the surrounding environment (e.g., traffic sections) in real time to obtain laser scanning data, identify target views in the surrounding environment, and discretize the data to obtain object data within at least one grid area.
[0050] Step S204: Extract spatial and temporal features from any raster range in the point cloud data. The spatial features are used to characterize the spatial information of the raster range, and the temporal features are used to characterize the temporal information of the raster range.
[0051] In the technical solution provided by step S204 of the present invention, a grid to be processed is arbitrarily selected in the point cloud data, and spatial features and temporal features within the grid range are extracted from it. The spatial features can be used to characterize the spatial information of the grid range, such as the spatial location information of the selected object; the temporal features can be used to characterize the temporal information of the grid range, such as the point cloud data of each frame in the grid range.
[0052] Optionally, a spatiotemporal bidirectional augmentation network can be used to extract temporal and spatial features from point cloud data, thereby achieving the goal of extracting spatial and temporal features within any grid range from the point cloud data. It should be noted that no specific restrictions are placed on the methods for obtaining spatial and temporal features here, the purpose being that all methods for obtaining spatial and temporal features should be within the protection scope of this invention.
[0053] Step S206: Analyze the velocity of obstacles located within the grid area based on spatial and temporal features.
[0054] In the technical solution provided by step S206 of the present invention, object data within at least one grid range in the target view is obtained. Based on the object data within at least one grid range in the target view, spatial and temporal information within any grid range is extracted from the obtained point cloud data. The velocity of the object within the grid range is analyzed and determined. The object can be an obstacle.
[0055] Optionally, a time-series point cloud sequence can be used as input to extract temporal and spatial features from the point cloud data. The extracted temporal and spatial features are then input into the velocity estimation module to analyze the velocity of obstacles within the grid area, thereby estimating the velocity of objects occupying the grid in the top view.
[0056] Through steps S202 to S206 of the present invention, point cloud data is generated by detecting traffic roads. The point cloud data is used to represent object data within at least one grid area in the target view of the traffic road. Spatial and temporal features are extracted from the point cloud data within any grid area. The spatial features represent the spatial information of the grid area, and the temporal features represent the temporal information of the grid area. Based on the spatial and temporal features, the speed of obstacles located within the grid area is analyzed. In other words, in this embodiment of the invention, point cloud data generated by detecting traffic roads is used as input. Effective temporal and spatial features are extracted from it, and then the information of objects on the grid occupied by the point cloud data is accurately determined based on the extracted temporal and spatial features. This effectively addresses all types of obstacles, thereby improving the efficiency of determining object information and solving the technical problem of low efficiency in object determination.
[0057] The method described in this embodiment will be further described below.
[0058] As an optional implementation, step S204, extracting spatial and temporal features from any grid range of point cloud data, includes: extracting spatial and temporal features from multiple frames of point cloud data, wherein the point cloud data includes multiple frames of point cloud data, and the multiple frames of point cloud data are sorted according to time.
[0059] In this embodiment, considering that the motion of objects in the target view is a change of point cloud scene in point cloud data over time, the point cloud data may include multiple frames of point cloud data. Spatial and temporal features can be extracted from the multiple frames of point cloud data. The multiple frames of point cloud data are sorted in chronological order, and the multiple frames of point cloud data can also be called a temporal point cloud sequence.
[0060] Optionally, by detecting traffic roads and generating point cloud data, the point cloud data can be sorted in chronological order to obtain temporal point cloud data. The temporal point cloud data can then be processed to extract spatial and temporal features within any grid area from the point cloud data.
[0061] As an optional implementation, step S204, extracting spatial features from multi-frame point cloud data, includes: extracting common features from multi-frame point cloud data, wherein the common features are features with the same attributes in multi-frame point cloud data; and determining spatial features based on the common features.
[0062] In this embodiment, features with the same attributes can be extracted from multi-frame point cloud data in the time dimension, thereby achieving the purpose of extracting common features from multi-frame point cloud data, and determining spatial features based on the extracted common features.
[0063] Optionally, to address the difficulty of spatial semantic understanding caused by the sparsity of single-frame point clouds, this embodiment of the invention provides a temporal augmentation spatial module, which extracts common features from multi-frame point cloud data in the temporal dimension and determines spatial features based on the extracted common features. The extraction of common features from multi-frame point cloud data can be performed on point cloud data of adjacent frames or on point cloud data of non-adjacent frames, and no specific limitation is made here.
[0064] In this embodiment of the invention, the spatial semantic features of a single frame of point cloud data are enhanced by utilizing the similarity of scenes in multi-frame point cloud data, thereby achieving better spatial semantic understanding.
[0065] As an optional implementation, determining spatial features based on common features includes: fusing subspace semantic features and common features of each frame of point cloud data to obtain a first fused feature corresponding to each frame of point cloud data, wherein the common features are used to enhance the subspace semantic features in the first fused feature, and the subspace semantic features are used to characterize the semantic information of each frame of point cloud data in space; and determining spatial features based on the first fused feature corresponding to each frame of point cloud data.
[0066] In this embodiment of the invention, features with the same attributes can be extracted from multi-frame point cloud data in the time dimension, thereby achieving the purpose of extracting common features from multi-frame point cloud data. Spatial features are determined based on the extracted common features, and then the subspace semantic features and common features of each frame of point cloud data are fused to obtain the first fused feature corresponding to each frame of point cloud data. Spatial features can be determined based on the first fused feature corresponding to each frame of point cloud data. The subspace semantic features can be used to characterize the semantic information of each frame of point cloud data in space, for example, it can be a single-frame spatial semantic feature.
[0067] Optionally, the common feature map can be stitched and fused with the single-frame point cloud feature map to enhance the spatial semantic features of a single frame by utilizing the similarity of the scene in the point cloud data of adjacent frames.
[0068] For example, to address the difficulty of spatial semantic understanding caused by the sparsity of single-frame point clouds, a temporal augmentation spatial module is provided. This module extracts common features from multi-frame point cloud data in the temporal dimension, determines spatial features based on the extracted common features, and then stitches and fuses the extracted common feature maps with the single-frame point cloud feature maps respectively.
[0069] For example, common features (T”) can be extracted from the semantic features of the first frame subspace (T1), the second frame subspace (T2), the third frame subspace (T3), the fourth frame subspace (T4), and the fifth frame subspace (T5). T1, T2, T3, T4, T5, and T” are then concatenated and fused to obtain the first fusion features (T1T”, T2T”, T3T”, T4T”, and T5T”) corresponding to each frame of point cloud data. Spatial features can be determined based on the first fusion features corresponding to each frame of point cloud data, thereby enhancing the spatial semantic features of a single frame by utilizing the similarity of point cloud scenes in adjacent frames, and achieving better spatial semantic understanding.
[0070] As an optional implementation, spatial features are determined based on the first fusion feature corresponding to each frame of point cloud data, including: decoding the first fusion feature corresponding to each frame of point cloud data based on the semantic decoder to obtain spatial semantic features, wherein the semantic decoder is trained based on semantic segmentation task data, the semantic segmentation task data is used to represent the semantic segmentation task, and the spatial features include spatial semantic features, which are used to represent the semantic information of the raster range in space.
[0071] In this embodiment, the first fusion feature corresponding to each frame of point cloud data is decoded based on the semantic decoder to determine the spatial semantic features. The semantic decoder can be trained based on semantic segmentation task data. The semantic segmentation data can be used to represent the semantic segmentation task, for example, it can be a pixel-level object classification task. The spatial feature information can include spatial semantic features, which can be used to represent the semantic information of the grid range in space.
[0072] Optionally, the point cloud data is transmitted to a spatiotemporal bidirectional augmentation network, and a semantic decoder is introduced between the backbone network encoder and the velocity decoder, thereby completing the extraction of spatial and temporal semantic features of the augmentation network and enhancing the network's ability to extract semantic features in space.
[0073] The spatiotemporal bidirectional augmentation network framework in this embodiment of the invention introduces a semantic decoder between the backbone network encoder and the velocity decoder, based on the traditional velocity estimation encoder-decoder framework. The training of this decoder is supervised by a semantic segmentation task, thereby enhancing the network's ability to extract spatial semantic features. The extracted spatial semantic features are then input into subsequent modules, thereby achieving the goal of using semantic information to assist the velocity estimation task and improve velocity estimation.
[0074] As an optional implementation, common features are extracted from multi-frame point cloud data, including: extracting common features from adjacent frame point cloud data in multi-frame point cloud data.
[0075] In this embodiment, in order to address the difficulty of spatial semantic understanding caused by the sparsity of single-frame point clouds, common features are extracted from the point cloud data of adjacent frames in the time dimension. This enhances the spatial semantic features of a single frame by utilizing the similarity of the point cloud scenes of adjacent frames, thereby achieving better spatial semantic understanding.
[0076] As an optional implementation, step S204, extracting temporal features from multi-frame point cloud data, includes: extracting difference features from multi-frame point cloud data, wherein the difference features are features with different attributes in multi-frame point cloud data; and determining temporal features based on the difference features.
[0077] In this embodiment, difference features are extracted from multi-frame point cloud data, and time characteristics are determined based on the determined difference features. The difference features can be features with different attributes in multi-frame point cloud data.
[0078] Optionally, considering that the motion of an object is represented by point cloud data and the scene changes over time, this embodiment of the invention provides a spatially enhanced temporal module that extracts differential features from multiple frames of point cloud data in the time dimension, thereby enabling the capture of dynamic motion cues of an object using the differential features of non-adjacent frame point cloud data.
[0079] As an optional implementation, determining the temporal feature based on the difference features includes: fusing multiple difference features to obtain a second fused feature, wherein the second fused feature is used to characterize the state of multi-frame point cloud data changing over time; and determining the temporal feature corresponding to the second fused feature.
[0080] In this embodiment, multiple differential features are fused to obtain a second fused feature, and a time feature corresponding to the second fused feature is determined. The second fused feature can be used to characterize the state of multiple frames of point cloud data changing over time, such as the state of scene data in point cloud data changing over time.
[0081] Optionally, differential features are extracted from non-adjacent frames in the time dimension, and then the extracted differential features are fused to obtain a second fused feature. The time feature corresponding to the second fused feature is determined to obtain the time feature of the point cloud data.
[0082] For example, differential features are extracted from T1 and T5, T2 and T4, and T3, and the extracted features are fused to obtain a second fused feature. The temporal feature corresponding to the second fused feature is then determined to obtain the temporal feature of the point cloud data.
[0083] As an optional implementation, extracting difference features from multi-frame point cloud data includes: extracting difference features from non-adjacent frame point cloud data in multi-frame point cloud data.
[0084] In this embodiment, considering that the motion of an object is a point cloud scene that changes over time, a spatial augmentation time module is provided. This module extracts differentiated features from non-adjacent frames in the time dimension and then fuses the extracted features. This achieves the goal of capturing motion cues of dynamic targets by utilizing the differences in point cloud scenes of non-adjacent frames, thereby achieving a more accurate estimation of the velocity of objects in point cloud data.
[0085] As an optional implementation, step S204 involves extracting spatial and temporal features from any grid area of the point cloud data, including: extracting spatial and temporal features from the point cloud data based on a spatiotemporal feature extraction model, wherein the spatiotemporal feature extraction model is used to enhance the spatial and temporal features in the point cloud data.
[0086] In this embodiment, point cloud data is processed based on a spatiotemporal feature extraction model to extract spatial and temporal features from the point cloud data. The spatiotemporal model can be used to enhance the temporal and spatial features in the point cloud data. The spatiotemporal feature extraction model can also be called a spatiotemporal bidirectional enhancement network model.
[0087] In this embodiment of the invention, in order to enhance the spatial and temporal features in point cloud data, a spatiotemporal bidirectional enhancement network model is proposed. The spatiotemporal bidirectional enhancement network model may include: a backbone network encoder, a semantic decoder, a velocity decoder, a temporal-enhanced spatial encoder (TeSE) and a spatial-enhanced temporal encoder (SeTE).
[0088] Optionally, the point cloud data is transmitted to a spatiotemporal bidirectional augmentation network. A semantic decoder is introduced between the backbone network encoder and the velocity decoder to extract the spatial and temporal semantic features of the augmentation network. The extracted spatiotemporal features are then input into the velocity estimation module to achieve velocity estimation of the grid occupied by the top view. Specifically, the temporal augmentation spatial module is used to enhance spatial features, and the spatial augmentation temporal module is used to enhance temporal features, thereby extracting more effective spatiotemporal features and achieving more accurate velocity estimation.
[0089] As an optional implementation, step S206, based on spatial and temporal features, analyzes the speed of obstacles located within the grid range, including: analyzing the motion information of obstacles located within the grid range based on spatial and temporal features, wherein the motion information is at least used to characterize the motion behavior of the obstacles; and predicting the speed of any obstacle located within the grid range based on the motion information of the obstacles.
[0090] In this embodiment, temporal point cloud data is input into a spatiotemporal bidirectional augmentation network. The spatiotemporal bidirectional augmentation network extracts the temporal and spatial features of the point cloud data. Based on the extracted temporal and spatial features, the motion information of obstacles within the grid area is analyzed. Based on the motion information of the obstacles, the velocity of any obstacle located within the grid area is predicted. The motion information can be used to characterize the motion behavior of obstacles within the grid area, such as being stationary, moving forward rapidly, or decelerating backward.
[0091] For example, point cloud data is transmitted to a spatiotemporal bidirectional augmentation network. A semantic decoder is introduced between the backbone network encoder and the velocity decoder to extract the spatial and temporal semantic features of the augmentation network. The extracted spatiotemporal features are then input into the velocity estimation module, which analyzes the motion information of obstacles within the grid area to obtain information about any obstacle located within the grid area.
[0092] In this embodiment of the invention, point cloud data is generated by detecting traffic roads. The point cloud data represents object data within at least one grid area in the target view of the traffic road. Spatial and temporal features are extracted from the point cloud data within any grid area. The spatial features represent the spatial information of the grid area, and the temporal features represent the temporal information of the grid area. Based on the spatial and temporal features, the speed of obstacles located within the grid area is analyzed. In other words, this embodiment of the invention uses point cloud data generated by detecting traffic roads as input, extracts effective temporal and spatial features from it, and then accurately determines the information of objects on the grid occupied by the point cloud data based on the extracted temporal and spatial features. This effectively addresses all types of obstacles, thereby improving the efficiency of object information determination and solving the technical problem of low efficiency in object determination.
[0093] This invention also provides another method for determining object information, which can be applied in open road scenarios.
[0094] Figure 3 This is a flowchart of another method for determining object information according to an embodiment of the present invention, such as... Figure 3 As shown, the method may include the following steps:
[0095] Step S302: Determine the traffic road on which the vehicle is traveling.
[0096] In the technical solution provided by step S302 of the present invention, the traffic road in which the vehicle is traveling on an open road is determined, wherein the vehicle can be a vehicle being driven, for example, an autonomous vehicle.
[0097] Step S304: Retrieve point cloud data of traffic roads, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one raster range in the target view of the traffic roads.
[0098] In the technical solution provided by step S304 of the present invention, point cloud data of the traffic road on which the vehicle travels is retrieved. The point cloud data can be object data generated by detecting the traffic road and can be used to characterize the traffic road within at least one grid range in the target view. The object data can be data of obstacles in the traffic road, such as information of other moving or stationary vehicles on the traffic road.
[0099] Step S306: Extract spatial and temporal features from any raster range in the point cloud data. The spatial features are used to characterize the spatial information of the raster range, and the temporal features are used to characterize the temporal information of the raster range.
[0100] In the technical solution provided by step S306 of the present invention, spatial and temporal features within any grid range are extracted from point cloud data. For example, spatial and temporal features within the grid range corresponding to the current road segment can be extracted from point cloud data. Spatial features can be used to characterize the spatial information of the grid range, such as the position of other objects within the grid range, such as up, down, left, right, etc. Here, spatial features are only given as examples and are not specifically limited. Temporal features can be used to characterize the temporal information of the grid range, such as the time taken for an object within the grid to move from one place to another. Here, temporal features are only given as examples and are not specifically limited.
[0101] Step S308: Analyze the velocity of obstacles located within the grid area based on spatial and temporal features.
[0102] In the technical solution provided by step S308 of the present invention, the speed of the obstacle located on the grid range is determined by utilizing the spatial and temporal characteristics of the obstacle on the grid range. The obstacle can be an object on the traffic road, such as other moving vehicles on the traffic road.
[0103] Step S310: Control the vehicle to avoid obstacles based on the speed of the obstacles.
[0104] In the technical solution provided by step S310 of the present invention, the vehicle is controlled based on the speed of other obstacles to achieve the purpose of allowing the vehicle to avoid other obstacles on the road during driving.
[0105] For example, by obtaining the speed of obstacles on the road, when the time it takes for an obstacle to reach a certain point is the same as the time it takes for a vehicle to travel, the vehicle's route can be replanned, thereby enabling the vehicle to avoid other obstacles on the road while traveling.
[0106] Through steps S302 to S310 of the present invention, the traffic road on which the vehicle is traveling is determined; point cloud data of the traffic road is retrieved, wherein the point cloud data is generated by detecting the traffic road and is used to represent object data within at least one grid range in the target view of the traffic road; spatial features and temporal features within any grid range are extracted from the point cloud data, wherein the spatial features are used to represent the spatial information of the grid range and the temporal features are used to represent the temporal information of the grid range; based on the spatial features and temporal features, the speed of obstacles located within the grid range is analyzed; the vehicle is controlled to avoid obstacles based on the speed of the obstacles, thereby achieving the technical effect of rationally planning the vehicle's driving route and speed, and solving the technical problem of unreasonable vehicle driving route and speed planning.
[0107] The present invention also provides another method for determining the information of an object.
[0108] Figure 4 This is a flowchart of another method for determining object information according to an embodiment of the present invention. Figure 4 As shown, the method may include the following steps.
[0109] Step S402: In response to a data input command applied to the operation interface, point cloud data of traffic roads is displayed on the operation interface, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid range in the target view of the traffic roads.
[0110] In the technical solution provided by step S402 of the present invention, the input operation command can be triggered by the user to display the point cloud data of traffic roads on the operation interface. Thus, this embodiment responds to the input operation command acting on the interactive interface to display the point cloud data of traffic roads. The point cloud data can be generated by detecting traffic roads and can be used to characterize object data located within at least one grid range in the target view of the traffic roads.
[0111] Step S404: In response to the information generation command applied to the operation interface, display the speed of obstacles in the grid range on the operation interface.
[0112] In the technical solution provided by step S404 of the present invention, the speed of obstacles within the grid range is obtained in response to the information generation command on the operation interface acting on the interactive interface. The speed of the obstacles is obtained based on the analysis of spatial and temporal features within any grid range of the point cloud data. The spatial features are used to characterize the spatial information of the grid range, and the temporal features are used to characterize the temporal information of the grid range.
[0113] Optionally, the spatial and temporal features within any grid area of the point cloud data are analyzed to obtain the velocity of the obstacle, and the velocity of the obstacle within the grid area is displayed on the operation interface. The spatial features can be used to characterize the spatial information of the grid area, and the temporal features can be used to characterize the temporal information of the grid area.
[0114] This invention also provides a method for determining the information of objects in virtual reality scenarios such as VR devices and AR devices.
[0115] Figure 5 This is a flowchart of a method for determining object information according to an embodiment of the present invention. Figure 5 As shown, the method may include the following steps.
[0116] Step S502: Display point cloud data of traffic roads on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid range in the target view of the traffic roads.
[0117] In the technical solution provided in step S502 of the present invention, there are detection devices for capturing traffic roads around the vehicle, such as cameras, etc., to obtain traffic road information captured by the detection devices, to detect the captured traffic road segments, to generate point cloud data of the traffic roads, and to display the point cloud data of the traffic road segments on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device. The point cloud data can be used to characterize object data within at least one grid range in the target view of the traffic road.
[0118] Step S504: Extract spatial and temporal features from any raster range in the point cloud data. The spatial features are used to characterize the spatial information of the raster range, and the temporal features are used to characterize the temporal information of the raster range.
[0119] In the technical solution provided by step S504 of the present invention, the point cloud data is processed to obtain the spatial features and temporal features within any grid range of the point cloud data. The spatial features can be used to characterize the spatial information of the grid range, and the temporal features can be used to characterize the temporal information of the grid range.
[0120] Step S506: Analyze the velocity of obstacles located within the grid area based on spatial and temporal features.
[0121] Step S508: Drive the VR or AR device to display the speed of the obstacle.
[0122] In the technical solution provided by step S508 of the present invention, a VR device or an AR device is driven to display the speed of the obstacle on the VR device or AR device.
[0123] Optionally, in this embodiment, the method for determining the information of the object described above can be applied to a hardware environment consisting of a server and a virtual reality device. The server, displayed on the screen of the virtual reality device or augmented reality device, can be a server corresponding to a media file operator. The aforementioned network includes, but is not limited to, a wide area network (WAN), a metropolitan area network (MAN), or a local area network (LAN). The aforementioned virtual reality device is not limited to, for example, a virtual reality headset, virtual reality glasses, or a standalone virtual reality device.
[0124] Optionally, the virtual reality device includes: a memory, a processor, and a transmission device. The memory stores an application that can be used to execute: displaying point cloud data of traffic roads on the rendering screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid area in a target view of the traffic roads; extracting spatial and temporal features from the point cloud data within any one grid area, wherein the spatial features characterize the spatial information of the grid area and the temporal features characterize the temporal information of the grid area; analyzing the speed of obstacles located within the grid area based on the spatial and temporal features; and driving the VR or AR device to display the speed of the obstacles.
[0125] It should be noted that the above-described method for verifying object information in VR or AR devices in this embodiment may include... Figure 2 The method of the illustrated embodiment is used to achieve the purpose of driving VR or AR devices to output prompt information based on obstacle speed.
[0126] Optionally, the processor in this embodiment can invoke the application stored in the memory via the transmission device to perform the above steps. The transmission device can receive media files sent by the server via a network, and can also be used for data transmission between the processor and the memory.
[0127] Optionally, in a virtual reality device, there is a head-mounted display with eye tracking. The screen in the HMD is used to display the video footage. The eye-tracking module in the HMD is used to acquire the real-time movement path of the user's eyes. The tracking system is used to track the user's position and movement information in real three-dimensional space. The computing and processing unit is used to acquire the user's real-time position and movement information from the tracking system and calculate the three-dimensional coordinates of the user's head in the virtual three-dimensional space, as well as the user's field of vision orientation in the virtual three-dimensional space.
[0128] In this embodiment of the invention, the virtual reality device can be connected to a terminal, and the terminal and the server are connected through a network. The virtual reality device is not limited to: virtual reality helmet, virtual reality glasses, virtual reality all-in-one machine, etc. The terminal is not limited to PC, mobile phone, tablet computer, etc. The server can be the server corresponding to the media file operator. The network includes, but is not limited to: wide area network, metropolitan area network or local area network.
[0129] Figure 6 This is a schematic diagram illustrating the result of determining the information of an object according to an embodiment of the present invention, such as... Figure 6 As shown, on the display screen of a virtual reality (VR) device or an augmented reality (AR) device, the prompt information is used to display the speed of an obstacle located on any grid, such as the speed of obstacle A, obstacle B, and obstacle C respectively. The pointing of the obstacle can be used to indicate the direction of the obstacle.
[0130] This invention uses point cloud data generated by detecting traffic roads as input, extracts effective temporal and spatial features from it, and then accurately determines the information of objects on the grid occupied by the point cloud data based on the extracted temporal and spatial features. This effectively addresses all types of obstacles, thereby improving the efficiency of determining object information and solving the technical problem of low efficiency in object determination.
[0131] Example 2
[0132] The preferred implementation of the above method of this embodiment will be further described below, specifically using a method for estimating the velocity of a target in a laser point cloud across the entire scene based on a spatiotemporal bidirectional augmentation network.
[0133] LiDAR, as a high-precision three-dimensional ranging device, is widely used in the development of autonomous driving perception systems.
[0134] In related technologies, velocity estimation tasks are often modeled as object tracking problems. This method includes cascaded object detection and object tracking modules. However, object tracking methods can only process objects detected by object detection. Therefore, existing object detection methods rely on neural network models trained on datasets and require pre-setting of object detection categories, which makes it difficult to predict obstacles of unknown categories.
[0135] Based on this, we propose a method to model the velocity estimation task as a scene flow estimation problem. This method directly predicts the dense motion field from the raw point cloud data of adjacent frames, thus freeing it from the restriction of target category. However, since existing scene flow estimation methods take hundreds of milliseconds, they are difficult to apply in real time.
[0136] Building upon this, a method for velocity estimation under a top-view grid representation is proposed. This method discretizes the point cloud into a grid under a top-view representation and jointly estimates the semantics and velocity of each grid. Although this method can handle all types of obstacles and has a computation speed on the order of tens of milliseconds, the framework of this method has a trade-off between semantic segmentation and velocity estimation performance, which makes it difficult to meet the predetermined conditions on a single task, resulting in inaccurate velocity estimation.
[0137] To address the aforementioned issues, this invention proposes a method for estimating the velocity of targets across all scenes using laser point clouds based on a spatiotemporal bidirectional augmentation network. This method applies environmental perception technology from the field of autonomous driving and is a laser point cloud-based method for estimating the velocity of targets across all scenes. By estimating the velocity of all grids occupied by laser point clouds within the entire scene based on the top-view grid representation, it can effectively address the long-tail problem of target categories in autonomous driving environmental perception, thereby accurately predicting the velocity of objects in open road scenarios and improving the robustness and safety of autonomous driving systems.
[0138] The method described in this embodiment will be further described below.
[0139] Figure 7 This is a flowchart of a laser point cloud full-scene velocity estimation method based on a spatiotemporal bidirectional augmentation network according to an embodiment of the present invention, as follows: Figure 7 As shown, the method may include the following steps:
[0140] Step S702: Take time-series point cloud data as input.
[0141] In this embodiment, point cloud data is generated by detecting traffic roads. The point cloud data can be sorted in chronological order to obtain temporal point cloud data, which is then used as input.
[0142] Optionally, in an autonomous driving system, a temporal point cloud sequence is used as input.
[0143] Step S704: Discretize the point cloud data.
[0144] In this embodiment, the detected point cloud data is discretized to achieve the purpose of discretizing the point cloud data into a grid under the top view.
[0145] Step S706: Process the data using a spatiotemporal bidirectional augmentation network.
[0146] In this embodiment, a spatiotemporal bidirectional augmentation network is used to extract temporal and spatial features from point cloud data.
[0147] Figure 8 This is a schematic diagram of a spatiotemporal bidirectional augmentation network framework according to an embodiment of the present invention, as shown below. Figure 8 As shown, point cloud data is transmitted to a spatiotemporal bidirectional augmentation network. A semantic decoder is introduced between the backbone network encoder and the velocity decoder to extract the spatial and temporal semantic features of the augmentation network. The extracted spatiotemporal features are then input into the velocity estimation module to achieve velocity estimation of the grid occupied by the top view. Specifically, the temporal augmentation spatial module is used to enhance spatial features, and the spatial augmentation temporal module is used to enhance temporal features, thereby extracting more effective spatiotemporal features and achieving more accurate velocity estimation.
[0148] The spatiotemporal bidirectional augmentation network framework in this embodiment of the invention introduces a semantic decoder between the backbone network encoder and the velocity decoder, based on the traditional velocity estimation encoder-decoder framework. The training of this decoder is supervised by a semantic segmentation task, thereby enhancing the network's ability to extract spatial semantic features. The extracted spatial semantic features are then input into the subsequent SeTE and motion decoder modules, thereby achieving the goal of using semantic information to assist the velocity estimation task and improve velocity estimation.
[0149] Optionally, to address the difficulties in spatial semantic understanding caused by the sparsity of single-frame point clouds, embodiments of the present invention provide a temporal augmentation spatial module. Figure 9 This is a schematic diagram of a time-enhanced spatial module according to an embodiment of the present invention, as shown below. Figure 9 As shown, common features are extracted from adjacent frame point cloud data in the time dimension. Then, the extracted common feature maps are concatenated and fused with the single frame point cloud feature maps. For example, the same T” is extracted from T1, T2, T3, T4, and T5. T1, T2, T3, T4, T5 and T” are concatenated and fused to obtain T1T”, T2T”, T3T”, T4T”, and T5T”. In this way, the similarity of adjacent frame point cloud scenes is used to enhance the spatial semantic features of a single frame and achieve better spatial semantic understanding.
[0150] Optionally, considering that the motion of objects manifests as changes in the point cloud scene over time, this embodiment of the invention provides a spatial augmentation time module. Figure 10 This is a schematic diagram of a spatial augmentation time module according to an embodiment of the present invention, as shown below. Figure 10 As shown, differential features are extracted from non-adjacent frames in the time dimension, and then the extracted features are fused. For example, differential features are extracted between T1 and T5, T2 and T4, and T3, and then the extracted features are fused. This achieves the goal of capturing motion cues of dynamic targets by utilizing the differences in point cloud scenes of non-adjacent frames, and thus achieves more accurate velocity estimation.
[0151] Step S708: Determine the top view of the grid speed.
[0152] In this embodiment, Figure 11 This is a schematic diagram illustrating the effect of determining object information according to an embodiment of the present invention, such as... Figure 11 As shown, the length and direction of the arrows in each top-view grid represent the magnitude and direction of the grid's velocity. This embodiment of the invention achieves the effect of accurately predicting the velocity of an object.
[0153] This invention presents a spatiotemporally enhanced velocity estimation network. By using a discretized laser point cloud occupancy grid as input, it outputs the velocity of each occupied grid in the top view end-to-end, resulting in a more stable and accurate velocity estimation result.
[0154] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the method for determining object information according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0156] Example 3
[0157] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 2 The method for determining the information of an object shown is a device for determining the information of an object.
[0158] Figure 12 This is a schematic diagram of an apparatus for determining object information according to an embodiment of the present invention. Figure 12 As shown, the device 1200 for determining the information of the object may include: a generation unit 1202, a first extraction unit 1204, and a first analysis unit 1206.
[0159] The generation unit 1202 is used to generate point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data located within at least one raster range in a target view of the traffic roads.
[0160] The first extraction unit 1204 is used to extract spatial and temporal features from point cloud data within any raster range. The spatial features are used to characterize the spatial information of the raster range, and the temporal features are used to characterize the temporal information of the raster range.
[0161] The first analysis unit 1206 is used to analyze the velocity of obstacles located within the grid range based on spatial and temporal features.
[0162] It should be noted that the above-mentioned generation unit 1202, first extraction unit 1204, and first analysis unit 1206 correspond to steps S202 to S206 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above-mentioned units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0163] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 3 The method for determining the information of an object shown is a device for determining the information of an object.
[0164] Figure 13 This is a schematic diagram of an apparatus for determining information about an object according to an embodiment of the present invention, which is applied in an open road scenario. Figure 13 As shown, the device 1300 for determining the information of the object may include: a determining unit 1302, a second extraction unit 1304, and a second analysis unit 1306.
[0165] The determining unit 1302 is used to determine the traffic road on which the vehicle is traveling; and to retrieve point cloud data of the traffic road, wherein the point cloud data is generated by detecting the traffic road and is used to represent object data located within at least one grid range in the target view of the traffic road.
[0166] The second extraction unit 1304 is used to extract spatial and temporal features from point cloud data within any raster range. The spatial features are used to characterize the spatial information of the raster range, and the temporal features are used to characterize the temporal information of the raster range.
[0167] The second analysis unit 1306 is used to analyze the speed of obstacles located within the grid range based on spatial and temporal features; and to control the vehicle to avoid obstacles based on the speed of the obstacles.
[0168] It should be noted that the aforementioned determining unit 1302, the second extraction unit 1304, and the second analysis unit 1306 correspond to steps S302 to S306 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0169] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 4 The method for determining the information of an object shown is a device for determining the information of an object.
[0170] Figure 14 This is a schematic diagram of an apparatus for determining information about an object according to another embodiment of the present invention. Figure 14 As shown, the device 1400 for determining the information of the object may include: a first display unit 1402 and a second display unit 1404.
[0171] The first display unit 1402 is configured to respond to a data input command applied to the operation interface and display point cloud data of traffic roads on the operation interface, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid range in the target view of the traffic roads.
[0172] The second display unit 1404 is used to respond to the information generation command applied to the operation interface and display the speed of obstacles in the grid range on the operation interface. The speed of the obstacles is obtained based on the analysis of spatial and temporal features within any grid range in the point cloud data. The spatial features are used to characterize the spatial information of the grid range, and the temporal features are used to characterize the temporal information of the grid range.
[0173] It should be noted that the first display unit 1402 and the second display unit 1404 mentioned above correspond to steps S402 to S404 in Embodiment 1. The two units and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0174] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 5 The device for determining the information of the object shown.
[0175] Figure 15 This is a schematic diagram of an object information determination device according to an embodiment of the present invention. Figure 15 As shown, the object information determination device 1500 may include: a presentation unit 1502, a third extraction unit 1504, a third analysis unit 1506, and a driving unit 1508.
[0176] The presentation unit 1502 is used to display point cloud data of traffic roads on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid range in the target view of the traffic roads.
[0177] The third extraction unit 1504 is used to extract spatial and temporal features from point cloud data within any raster range. The spatial features are used to characterize the spatial information of the raster range, and the temporal features are used to characterize the temporal information of the raster range.
[0178] The third analysis unit 1506 is used to analyze the velocity of obstacles located within the grid range based on spatial and temporal features.
[0179] The drive unit 1508 is used to drive the speed at which VR or AR devices display obstacles.
[0180] It should be noted that the aforementioned presentation unit 1502, third extraction unit 1504, third analysis unit 1506, and driving unit 1508 correspond to steps S502 to S508 in Embodiment 1. The four units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0181] In the object information determination apparatus of this embodiment, point cloud data generated by detecting traffic roads is used as input, effective temporal and spatial features are extracted from it, and then the object information on the grid occupied by the point cloud data is accurately determined based on the extracted temporal and spatial features, thereby effectively dealing with all types of obstacles, thereby achieving the technical effect of improving the efficiency of object information determination and solving the technical problem of low efficiency in object determination.
[0182] Example 4
[0183] Embodiments of the present invention can provide an object information verification system, which may include a server and a client, and the AR / VR device may be any AR / VR device in a group of AR / VR devices. Optionally, the object information verification device includes: a processor; and a memory connected to the processor, used to provide the processor with instructions to process the following processing steps: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data located within at least one grid area in a target view of the traffic roads; extracting spatial features and temporal features from the point cloud data within any one grid area, wherein the spatial features are used to characterize the spatial information of the grid area, and the temporal features are used to characterize the temporal information of the grid area; and analyzing the speed of obstacles located within the grid area based on the spatial features and temporal features.
[0184] In this embodiment of the invention, by using a discretized laser point cloud occupancy grid as input, the velocity of each occupancy grid in the top view is output end-to-end, resulting in a more stable and accurate velocity estimation result. This achieves the technical effect of improving the efficiency of determining object information and solves the technical problem of low efficiency in determining objects.
[0185] Example 5
[0186] Embodiments of the present invention may provide a processor for a method of determining object information. This processor may include a computer terminal, which can be any computer terminal device from a group of computer terminals. Optionally, in this embodiment, the computer terminal may be replaced with a mobile terminal or other terminal device.
[0187] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0188] In this embodiment, the computer terminal described above can execute the program code for the following steps in the method for determining the information of objects in an application: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data within at least one grid range in the target view of the traffic roads; extracting spatial and temporal features from the point cloud data within any one grid range, wherein the spatial features are used to characterize the spatial information of the grid range and the temporal features are used to characterize the temporal information of the grid range; and analyzing the speed of obstacles located on the grid range based on the spatial and temporal features.
[0189] Optionally, Figure 16 This is a structural block diagram of a computer terminal according to an embodiment of the present invention. Figure 16 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1602, memory 1604, and transmission devices 1606.
[0190] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the object information determination method and apparatus in this embodiment of the invention. The processor executes various functional applications and predictions by running the software programs and modules stored in the memory, thereby realizing the aforementioned object information determination method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to computer terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0191] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data within at least one raster range in a target view of the traffic roads; extracting spatial and temporal features from the point cloud data within any raster range, wherein the spatial features are used to characterize the spatial information of the raster range and the temporal features are used to characterize the temporal information of the raster range; and analyzing the speed of obstacles located within the raster range based on the spatial and temporal features.
[0192] Optionally, the processor may also execute program code that extracts spatial and temporal features from multi-frame point cloud data, wherein the point cloud data includes multi-frame point cloud data and the multi-frame point cloud data is sorted according to time.
[0193] Optionally, the processor may also execute program code that performs the following steps: extracting common features from multi-frame point cloud data, wherein the common features are features with the same attributes in the multi-frame point cloud data; and determining spatial features based on the common features.
[0194] Optionally, the processor may also execute program code that performs the following steps: fusing the subspace semantic features and common features of each frame of point cloud data to obtain a first fused feature corresponding to each frame of point cloud data, wherein the common features are used to enhance the subspace semantic features in the first fused feature, and the subspace semantic features are used to characterize the semantic information of each frame of point cloud data in space; and determining spatial features based on the first fused feature corresponding to each frame of point cloud data.
[0195] Optionally, the processor may also execute program code that performs the following steps: decodes the first fusion feature corresponding to each frame of point cloud data based on the semantic decoder to obtain spatial semantic features, wherein the semantic decoder is trained based on semantic segmentation task data, the semantic segmentation task data is used to characterize the semantic segmentation task, and the spatial features include spatial semantic features, which are used to characterize the semantic information of the raster range in space.
[0196] Optionally, the processor may also execute program code that extracts common features from adjacent frame point cloud data in multi-frame point cloud data.
[0197] Optionally, the processor may also execute program code that performs the following steps: extracting difference features from multi-frame point cloud data, wherein the difference features are features with different attributes in the multi-frame point cloud data; and determining time features based on the difference features.
[0198] Optionally, the processor may also execute program code that performs the following steps: fusing multiple differential features to obtain a second fused feature, wherein the second fused feature is used to characterize the state of multi-frame point cloud data changing over time; and determining the time feature corresponding to the second fused feature.
[0199] Optionally, the processor may also execute program code that extracts differential features from non-adjacent frame point cloud data in multi-frame point cloud data.
[0200] Optionally, the processor may also execute program code that extracts spatial and temporal features from point cloud data based on a spatiotemporal feature extraction model, wherein the spatiotemporal feature extraction model is used to enhance the spatial and temporal features in the point cloud data.
[0201] Optionally, the processor may also execute program code that performs the following steps: analyzing motion information of obstacles located within the grid area based on spatial and temporal features, wherein the motion information is used at least to characterize the motion behavior of the obstacles; and predicting the velocity of any obstacle located within the grid area based on the motion information of the obstacles.
[0202] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: determining the traffic road on which the vehicle is traveling; retrieving point cloud data of the traffic road, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data within at least one raster range in the target view of the traffic road; extracting spatial and temporal features from the point cloud data within any raster range, wherein the spatial features characterize the spatial information of the raster range and the temporal features characterize the temporal information of the raster range; analyzing the speed of obstacles located within the raster range based on the spatial and temporal features; and controlling the vehicle to avoid obstacles based on the speed of the obstacles.
[0203] As an alternative example, the processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: in response to a data input command applied to the operating interface, displaying point cloud data of a traffic road on the operating interface, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one grid area in the target view of the traffic road; in response to an information generation command applied to the operating interface, displaying the speed of obstacles within the grid area on the operating interface, wherein the speed of the obstacles is obtained based on spatial and temporal features within any grid area of the point cloud data, the spatial features being used to characterize the spatial information of the grid area, and the temporal features being used to characterize the temporal information of the grid area.
[0204] As an alternative example, the processor can invoke information and applications stored in memory via a transmission device to perform the following steps: displaying point cloud data of traffic roads on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting traffic roads and is used to characterize object data located within at least one grid area in the target view of the traffic roads; extracting spatial and temporal features from the point cloud data within any one grid area, wherein the spatial features characterize the spatial information of the grid area and the temporal features characterize the temporal information of the grid area; analyzing the speed of obstacles located within the grid area based on the spatial and temporal features; and driving the VR or AR device to display the speed of the obstacles.
[0205] This invention provides a method for determining object information. By using point cloud data generated from detecting traffic roads as input, effective temporal and spatial features are extracted from the data. Based on the extracted temporal and spatial features, the information of the object on the grid occupied by the point cloud data is accurately determined, thereby effectively dealing with all types of obstacles. This achieves the technical effect of improving the efficiency of determining object information and solves the technical problem of low efficiency in object determination.
[0206] Those skilled in the art will understand that Figure 16 The structure shown is for illustrative purposes only. Computer terminal A can also be a smartphone (such as an Android phone, iOS phone, etc.), tablet computer, mobile internet device (MID), PAD and other terminal devices. Figure 16 This does not limit the structure of the aforementioned computer terminal A. For example, computer terminal A may also include components that are more complex than those described above. Figure 16 Showing more or fewer components (such as network interfaces, display devices, etc.), or having the same Figure 16 The different configurations shown.
[0207] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0208] Example 6
[0209] Embodiments of the present invention also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method for determining the information of the object provided in Embodiment 1.
[0210] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0211] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: generating point cloud data by detecting traffic roads, wherein the point cloud data is used to characterize object data within at least one grid range in a target view of the traffic roads; extracting spatial and temporal features from the point cloud data within any one grid range, wherein the spatial features are used to characterize the spatial information of the grid range and the temporal features are used to characterize the temporal information of the grid range; and analyzing the speed of obstacles located on the grid range based on the spatial and temporal features.
[0212] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: extracting spatial and temporal features from multi-frame point cloud data, wherein the point cloud data includes multi-frame point cloud data, and the multi-frame point cloud data is sorted according to time.
[0213] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: fusing the subspace semantic features and common features of each frame of point cloud data to obtain a first fused feature corresponding to each frame of point cloud data, wherein the common features are used to enhance the subspace semantic features in the first fused feature, and the subspace semantic features are used to characterize the semantic information of each frame of point cloud data in space; and determining spatial features based on the first fused feature corresponding to each frame of point cloud data.
[0214] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: decoding the first fusion feature corresponding to each frame of point cloud data based on the semantic decoder to obtain spatial semantic features, wherein the semantic decoder is trained based on semantic segmentation task data, the semantic segmentation task data is used to characterize the semantic segmentation task, and the spatial features include spatial semantic features, which are used to characterize the semantic information of the raster range in space.
[0215] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: extracting common features from adjacent frame point cloud data in multi-frame point cloud data.
[0216] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: extracting difference features from multi-frame point cloud data, wherein the difference features are features with different attributes in the multi-frame point cloud data; and determining time features based on the difference features.
[0217] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: fusing multiple differential features to obtain a second fused feature, wherein the second fused feature is used to characterize the state of multi-frame point cloud data changing over time; and determining the time feature corresponding to the second fused feature.
[0218] Optionally, the computer-readable storage medium may also execute program code that performs the following steps: extracting difference features from non-adjacent frame point cloud data in multi-frame point cloud data.
[0219] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: extracting spatial and temporal features from point cloud data based on a spatiotemporal feature extraction model, wherein the spatiotemporal feature extraction model is used to enhance the spatial and temporal features in the point cloud data.
[0220] Optionally, the aforementioned computer-readable storage medium may also execute program code that performs the following steps: analyzing motion information of obstacles located within a grid area based on spatial and temporal features, wherein the motion information is at least used to characterize the motion behavior of the obstacles; and predicting the velocity of any obstacle located within the grid area based on the motion information of the obstacles.
[0221] As an alternative example, a computer-readable storage medium is configured to store program code for performing the following steps: determining the traffic road in which the vehicle is traveling; retrieving point cloud data of the traffic road, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one raster range in a target view of the traffic road; extracting spatial and temporal features from the point cloud data within any raster range, wherein the spatial features characterize the spatial information of the raster range and the temporal features characterize the temporal information of the raster range; analyzing the speed of obstacles located within the raster range based on the spatial and temporal features; and controlling the vehicle to avoid obstacles based on the speed of the obstacles.
[0222] As an alternative example, a computer-readable storage medium is configured to store program code for performing the following steps: in response to a data input instruction applied to an operating interface, displaying point cloud data of a traffic road on the operating interface, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one raster range in a target view of the traffic road; in response to an information generation instruction applied to the operating interface, displaying the speed of obstacles within the raster range on the operating interface, wherein the speed of the obstacles is obtained based on spatial and temporal features within any raster range of the point cloud data, the spatial features being used to characterize the spatial information of the raster range and the temporal features being used to characterize the temporal information of the raster range.
[0223] As an optional example, a computer-readable storage medium is configured to store program code for performing the following steps: displaying point cloud data of a traffic road on a rendering screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one raster range in a target view of the traffic road; extracting spatial and temporal features from the point cloud data within any raster range, wherein the spatial features characterize the spatial information of the raster range and the temporal features characterize the temporal information of the raster range; analyzing the velocity of obstacles located within the raster range based on the spatial and temporal features; and driving the VR or AR device to display the velocity of the obstacles.
[0224] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0225] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0226] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0227] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0228] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0229] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0230] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for determining information about an object, characterized in that, include: Point cloud data is generated by detecting traffic roads, wherein the point cloud data is used to characterize object data located within at least one raster range in a target view of the traffic roads; Spatial and temporal features are extracted from any one of the grid ranges from the point cloud data, wherein the spatial features are used to characterize the spatial information of the grid range, and the temporal features are used to characterize the temporal information of the grid range; Based on the spatial and temporal characteristics, the velocity of obstacles located within the grid area is analyzed; The point cloud data includes multiple frames of point cloud data; Extracting the spatial features within any one of the grid cells from the point cloud data includes: Common features are extracted from the multi-frame point cloud data; the subspace semantic features of each frame of point cloud data and the common features are fused to obtain a first fused feature corresponding to each frame of point cloud data, wherein the subspace semantic features are used to characterize the spatial semantic information of each frame of point cloud data, and the common features are used to enhance the subspace semantic features in the first fused feature; based on the first fused feature, the spatial features within any one of the grid ranges are determined. Extracting the temporal feature from any one of the grid ranges in the point cloud data includes: Multiple difference features are extracted from the multi-frame point cloud data; the multiple difference features are fused to obtain a second fused feature, and the time feature corresponding to the second fused feature is determined.
2. The method according to claim 1, characterized in that, The multi-frame point cloud data is sorted according to time.
3. The method according to claim 2, characterized in that, The common features extracted from the multi-frame point cloud data include: The common features are extracted from the multi-frame point cloud data, wherein the common features are features with the same attributes in the multi-frame point cloud data.
4. The method according to claim 3, characterized in that, Based on the first fusion feature, the spatial features within any one of the grid areas are determined, including: The spatial features are determined based on the first fusion feature corresponding to the point cloud data in each frame.
5. The method according to claim 4, characterized in that, Determining the spatial features based on the first fusion feature corresponding to the point cloud data in each frame includes: Based on the semantic decoder, the first fusion feature corresponding to each frame of point cloud data is decoded to obtain spatial semantic features. The semantic decoder is trained based on semantic segmentation task data, which is used to represent the semantic segmentation task. The spatial features include the spatial semantic features, which are used to represent the semantic information of the raster range in space.
6. The method according to claim 3, characterized in that, The common features extracted from the multi-frame point cloud data include: The common features are extracted from adjacent frame point cloud data in the multi-frame point cloud data.
7. The method according to claim 2, characterized in that, The difference features are the features with different attributes in the multi-frame point cloud data.
8. The method according to claim 7, characterized in that, The second fusion feature is used to characterize the state of the multi-frame point cloud data as it changes over time.
9. The method according to claim 7, characterized in that, The differential features extracted from the multi-frame point cloud data include: The difference features are extracted from the non-adjacent frame point cloud data in the multi-frame point cloud data.
10. The method according to claim 1, characterized in that, Extracting spatial and temporal features from any given grid area from the point cloud data includes: The spatial features and temporal features are extracted from the point cloud data based on the spatiotemporal feature extraction model, wherein the spatiotemporal feature extraction model is used to enhance the spatial features and temporal features in the point cloud data.
11. The method according to any one of claims 1 to 10, characterized in that, Based on the spatial and temporal features, the velocity of obstacles located within the grid area is analyzed, including: Based on the spatial features and the temporal features, the motion information of obstacles located within the grid range is analyzed, wherein the motion information is used at least to characterize the motion behavior of the obstacles; Based on the motion information of the obstacles, the velocity of any obstacle located within the grid range is predicted.
12. A method for determining information about an object, characterized in that, include: Determine the traffic route the vehicle is traveling on; Retrieve point cloud data of the traffic road, wherein the point cloud data is generated by detecting the traffic road and is used to characterize object data located within at least one raster range in the target view of the traffic road; Spatial and temporal features are extracted from any one of the grid ranges from the point cloud data, wherein the spatial features are used to characterize the spatial information of the grid range, and the temporal features are used to characterize the temporal information of the grid range; Based on the spatial and temporal characteristics, the velocity of obstacles located within the grid area is analyzed; The vehicle is controlled to avoid the obstacle based on the obstacle's speed; The point cloud data includes multiple frames of point cloud data; Extracting the spatial features within any one of the grid cells from the point cloud data includes: Common features are extracted from the multi-frame point cloud data; the subspace semantic features of each frame of point cloud data and the common features are fused to obtain a first fused feature corresponding to each frame of point cloud data, wherein the subspace semantic features are used to characterize the spatial semantic information of each frame of point cloud data, and the common features are used to enhance the subspace semantic features in the first fused feature; based on the first fused feature, the spatial features within any one of the grid ranges are determined. Extracting the temporal feature from any one of the grid ranges in the point cloud data includes: Multiple difference features are extracted from the multi-frame point cloud data; the multiple difference features are fused to obtain a second fused feature, and the time feature corresponding to the second fused feature is determined.
13. A method for determining information about an object, characterized in that, include: In response to a data input command applied to the operation interface, point cloud data of traffic roads is displayed on the operation interface, wherein the point cloud data is generated by detecting the traffic roads and is used to characterize object data located within at least one grid area in the target view of the traffic roads. In response to an information generation command applied to the operation interface, the speed of obstacles within the grid range is displayed on the operation interface. The speed of the obstacles is obtained by analyzing the spatial and temporal features within any one of the grid ranges in the point cloud data. The spatial features are used to characterize the spatial information of the grid range, and the temporal features are used to characterize the temporal information of the grid range. The point cloud data includes multiple frames of point cloud data; The spatial features within any of the grid areas are determined based on a first fusion feature. The first fusion feature is obtained by fusing the subspace semantic features and common features of each frame of point cloud data. The common features are obtained based on multiple frames of point cloud data, and the common features are used to enhance the subspace semantic features in the first fusion feature. The subspace semantic features are used to characterize the spatial semantic information of each frame of point cloud data. The temporal feature within any of the grid ranges corresponds to the second fusion feature, which is obtained by fusing multiple difference features, and the difference features are extracted from multiple frames of point cloud data.
14. A method for determining information about an object, characterized in that, include: Displaying point cloud data of traffic roads on the presentation screen of a virtual reality (VR) device or an augmented reality (AR) device, wherein the point cloud data is generated by detecting the traffic roads and is used to characterize object data located within at least one grid area in a target view of the traffic roads; Spatial and temporal features are extracted from any one of the grid ranges from the point cloud data, wherein the spatial features are used to characterize the spatial information of the grid range, and the temporal features are used to characterize the temporal information of the grid range; Based on the spatial and temporal characteristics, the velocity of obstacles located within the grid area is analyzed; The speed at which the VR device or AR device displays the obstacle; The point cloud data includes multiple frames of point cloud data; Extracting the spatial features within any one of the grid cells from the point cloud data includes: Common features are extracted from the multi-frame point cloud data; the subspace semantic features of each frame of point cloud data and the common features are fused to obtain a first fused feature corresponding to each frame of point cloud data, wherein the subspace semantic features are used to characterize the spatial semantic information of each frame of point cloud data, and the common features are used to enhance the subspace semantic features in the first fused feature; based on the first fused feature, the spatial features within any one of the grid ranges are determined. Extracting the temporal feature from any one of the grid ranges in the point cloud data includes: Multiple difference features are extracted from the multi-frame point cloud data; the multiple difference features are fused to obtain a second fused feature, and the time feature corresponding to the second fused feature is determined.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program is run by a processor, it controls the device in which the computer-readable storage medium resides to perform the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Obstacle detection method and device for automatic driving scene of port
CN110764108A
Obstacle recognition model training method, obstacle recognition method, device and system
CN112347999A