Autonomous driving equipment and control method thereof

By using a unified neural network model for multi-task prediction in autonomous driving equipment, the problem of information transmission faults between different tasks is solved, and fast, accurate scene response and efficient task processing are achieved.

CN120792859AActive Publication Date: 2025-10-17浙江人形机器人创新中心有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510908035.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing technologies have gaps and conflicts in the information transmission between different tasks in autonomous driving equipment, resulting in complex and inefficient post-processing fusion, making it difficult to respond quickly and accurately to scene changes.

Method used

A unified neural network model is used for multi-task prediction. Through image feature extraction, feature space conversion and time series information fusion, the synchronous output and high consistency of multiple task results are achieved, reducing the amount of calculation.

Benefits of technology

It improves the response speed and accuracy of autonomous driving equipment in scene changes, reduces the amount of post-processing calculations, and improves task processing efficiency and perception consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention is suitable for the technical field of automatic driving perception, and provides autonomous driving equipment and a control method thereof, and the method comprises the steps: continuously collecting images during the driving of the autonomous driving equipment; according to a first image collected at a first moment T, a neural network model is utilized to complete one-time forward propagation so as to synchronously output a plurality of task prediction results, and the neural network model comprises an image feature extraction module, a feature space conversion module, a time sequence information fusion module and a plurality of prediction modules; and adjusting the driving state of the autonomous driving equipment according to the plurality of task prediction results. Through the above method, during the driving period of the autonomous driving equipment, a plurality of task prediction results are reasoned at the same time by using unified features, so that the reasoned results have high consistency, the post-processing calculation amount is reduced, the task processing efficiency is significantly improved, and then the autonomous driving equipment can quickly and accurately respond to scene changes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of autonomous driving perception technology, and particularly relates to an autonomous driving device and a control method thereof. BACKGROUND

[0002] The space occupancy (Occupancy) perception technology originates from the occupancy grid mapping method in the field of robots. The method discretizes the environment into two-dimensional or three-dimensional grid cells, and represents the obstacle distribution through binary states (occupied / free), providing directly usable perception information for path planning. In this method, a neural network model is often used to process the collected initial data to obtain perception information. When different models are used to process multiple tasks, there is a fault in the information transmission between different tasks, not only complex post-processing fusion is needed, but there may also be conflicting results. SUMMARY

[0003] Therefore, the present application provides an autonomous driving device and a control method thereof, which can use a unified feature to infer multiple task prediction results during the driving of the autonomous driving device, so that each result inferred has high consistency, and the post-processing calculation amount is reduced, the task processing efficiency is significantly improved, and further, the autonomous driving device can quickly and accurately respond to scene changes.

[0004] In a first aspect, the present application provides an autonomous driving device, comprising: The vehicle body; a plurality of drive wheels for driving the vehicle body to move; an image acquisition unit arranged on the vehicle body; a memory, wherein a neural network model is stored, the neural network model comprising an image feature extraction module, a feature space conversion module, a time sequence information fusion module and a plurality of prediction modules; a processor configured to: control the autonomous driving device to drive; continuously acquire images using the image acquisition unit during driving of the autonomous driving device; at a first time T, complete a forward propagation using the neural network model according to a first image acquired at the first time T, to synchronously output a plurality of task prediction results; and adjust the driving state of the autonomous driving device according to the plurality of task prediction results; wherein, in the process of forward propagation, the processor extracts two-dimensional image features from the first image through the image feature extraction module, projects the two-dimensional image features into a three-dimensional space through the feature space conversion module to obtain three-dimensional features, receives a fusion request through the time sequence information fusion module, fuses the three-dimensional features with historical three-dimensional features according to the fusion request to obtain fused features, calculates time sequence information according to the fused features, and calculates the three-dimensional features and the time sequence information through the plurality of prediction modules respectively to output the plurality of task prediction results.

[0005] In some embodiments, after projecting the two-dimensional image features into the three-dimensional space, a scattered feature point cloud is obtained, and the processor performs aggregation processing on the feature point cloud to obtain the three-dimensional features, which are used to represent objects in a scene where the autonomous driving device is located.

[0006] In some embodiments, the aggregation processing is pooling processing.

[0007] In some embodiments, the processor is further configured to: maintain a queue through the time sequence information fusion module, store the three-dimensional features output by the feature space conversion module into the queue each time, remove the three-dimensional features stored in the queue earliest from the queue when the queue reaches a preset length; and record the three-dimensional features in the queue as historical three-dimensional features.

[0008] In some embodiments, the processor extracts semantic information from the two-dimensional image features through the feature space conversion module, the semantic information comprising target categories and scene geometric structures, predicts depth values of each pixel in the first image according to the semantic information, and projects the two-dimensional image features into the three-dimensional space according to the depth values.

[0009] In some embodiments, the plurality of prediction modules include a semantic segmentation module, an instance segmentation module, and a velocity estimation module, and the plurality of task prediction results include a geometric structure of a scene in which the autonomous driving device is located, a division result of semantic information of different regions in the scene, differentiation information of different individuals of a same target category in the scene, and a velocity of each voxel in the scene.

[0010] In some embodiments, the neural network model further includes an image feature enhancement module, and the processor is further configured to: process the two-dimensional image features through the image feature enhancement module to obtain the two-dimensional image features of multiple scales, and input the two-dimensional image features of multiple scales into a feature space conversion module.

[0011] In a second aspect, the present application provides a control method of an autonomous driving device, comprising: controlling the autonomous driving device to drive; continuously collecting images using an image collection unit during driving of the autonomous driving device; at a first time T, completing one forward propagation using a neural network model according to a first image collected at the first time T to synchronously output a plurality of task prediction results, the neural network model including an image feature extraction module, a feature space conversion module, a time sequence information fusion module, and a plurality of prediction modules; and adjusting a driving state of the autonomous driving device according to the plurality of task prediction results; wherein, in the process of the forward propagation, the image feature extraction module extracts two-dimensional image features from the first image, the feature space conversion module projects the two-dimensional image features into a three-dimensional space to obtain three-dimensional features, the time sequence information fusion module receives a fusion request, fuses the three-dimensional features with historical three-dimensional features according to the fusion request to obtain fused features, and calculates time sequence information according to the fused features, and the plurality of prediction modules respectively calculate the three-dimensional features and the time sequence information to output the plurality of task prediction results.

[0012] In some embodiments, the adjusting the driving state of the autonomous driving device according to the plurality of task prediction results includes: constructing a map of a scene in which the autonomous driving device is located according to the plurality of task prediction results; and adjusting the driving state of the autonomous driving device according to the map.

[0013] In some embodiments, the image feature extraction module, the feature space conversion module, the time sequence information fusion module, and the plurality of prediction modules interact through standardized interfaces.

[0014] In a third aspect, the present application provides an electronic device, comprising a memory and a processor, the memory being coupled to the processor; The memory stores program instructions which, when executed by the processor, cause the electronic device to perform the control method provided in the second aspect.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program which, when running on a computer or a processor, causes the computer or the processor to perform the control method provided in the second aspect.

[0016] In a fifth aspect, the present application provides a computer program product which, when running on an electronic device, causes the electronic device to perform the control method provided in the second aspect.

[0017] The autonomous driving device provided in the present application continuously collects images using an image collection unit during the driving of the autonomous driving device, completes one forward propagation using a neural network model according to a first image collected at a first time T to synchronously output a plurality of task prediction results, the neural network model includes an image feature extraction module, a feature space conversion module, a time sequence information fusion module, and a plurality of prediction modules; and adjusts the driving state of the autonomous driving device according to the plurality of task prediction results. Through the above method, during the driving of the autonomous driving device, a plurality of task prediction results are inferred using unified features, so that each result inferred has high consistency, and multi-task prediction is completed in a single forward propagation, which can reduce the amount of calculation compared with a serial architecture. In addition, projecting two-dimensional image features into a three-dimensional space helps the plurality of prediction modules to understand the space and provides a basis for subsequent time sequence fusion, and the time sequence information fusion module combines the three-dimensional features extracted at the current time with the three-dimensional features (historical three-dimensional features) at the past time, which can capture the dynamic changes of the scene and improve the continuity of perception and prediction ability.

[0018] It can be understood that the beneficial effects of the second aspect to the fifth aspect can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0020] Figure 1 is a network structure schematic diagram of the neural network model provided in the embodiments of the present application.

[0021] Figure 2FIG. 1 is a flowchart of a control method of an autonomous mobile device according to an embodiment of the present application.

[0022] Figure 3 FIG. 2 is a structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular sequences of acts, techniques, etc., in order to provide a thorough understanding of the present embodiments. However, it will be apparent to those skilled in the art that the present embodiments can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, and circuits are omitted so as not to obscure the description of the present embodiments.

[0024] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or "comprising", when used in this specification and in the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0025] It is also to be understood that the terminology "and / or" when used in this specification and in the following claims, refers to at least one of the items, or any combination of the items, and includes all possible combinations of the items.

[0026] As used in this specification and in the claims, the term "if" can be interpreted as meaning "when", or "once", or "in response to determining", or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be interpreted as meaning "once it is determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]", depending on the context.

[0027] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", etc. are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0028] Reference within the specification to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places within specified The embodiment of the present application provides an autonomous driving device, which comprises a vehicle body, a plurality of driving wheels for driving the vehicle body to move, an image acquisition unit arranged on the vehicle body, and a memory. The plurality of driving wheels can be installed on the vehicle body and can drive the vehicle body to move when the driving wheels rotate. The image acquisition unit can be a camera, for example, a binocular camera, a trinocular camera, etc. The camera can be installed on the front of the vehicle body, and can continuously acquire images in front of the autonomous driving device during driving of the autonomous driving device.

[0029] The memory stores a neural network model, and the neural network model comprises an image feature extraction module, a feature space conversion module, a time sequence information fusion module and a plurality of prediction modules. The processor is configured to control the autonomous driving device to drive, continuously acquire images by using the image acquisition unit during driving of the autonomous driving device, complete one forward propagation by using the neural network model according to a first image acquired at a first time T to synchronously output a plurality of task prediction results at the first time T, and adjust a driving state of the autonomous driving device according to the plurality of task prediction results. During the forward propagation, the processor extracts two-dimensional image features from the first image by using the image feature extraction module, projects the two-dimensional image features into a three-dimensional space by using the feature space conversion module to obtain three-dimensional features, receives a fusion request by using the time sequence information fusion module, fuses the three-dimensional features with historical three-dimensional features according to the fusion request to obtain fused features, calculates time sequence information according to the fused features, and respectively calculates the three-dimensional features and the time sequence information by using the plurality of prediction modules to output the plurality of task prediction results.

[0030] Figure 1It is the structure of a neural network model. Each of the image feature extraction module, feature space conversion module, time series information fusion module, and multiple prediction modules is part of the neural network model. In a forward propagation, the image is input from the image feature extraction module. After calculation by each module, the prediction module outputs the task prediction result. The autonomous driving device can be a car or a robot. While the autonomous driving device is driving, an image acquisition unit installed on the autonomous driving device can be used to continuously capture images. For example, the image acquisition unit captures images in front of the autonomous driving device in real time at a preset frequency. Each image captured by the image acquisition unit can be input into the image feature extraction module. The first moment T represents any moment during the driving of the autonomous driving device. At the first moment T, the image captured by the image acquisition unit can be recorded as the first image. After the first image is captured, the first image can be immediately input into the image feature extraction module to meet the real-time requirements of autonomous driving. The first image can be an RGB image. The image feature extraction module is used to extract 2D image features from the first image, and the feature space conversion module is used to project the 2D image features into 3D space, thereby obtaining a large number of 3D features. 2D image feature extraction extracts key information from the image for subsequent processing and analysis. For example, deep learning-based feature extraction methods can be used: Convolutional Neural Networks (CNNs) can automatically learn features within an image. For example, when using the AlexNet network architecture for image feature extraction, multiple layers of convolution, pooling, and fully connected operations are performed to extract a high-dimensional feature vector from the input image. If the feature space conversion module detects a moving object based on 3D features, it can initiate a fusion request to the time series information fusion module. In response to the fusion request, the time series information fusion module fuses the currently received 3D features with historical 3D features to obtain fused features and calculates time series information based on the fused features. This time series information helps the prediction module infer the object's motion, such as its speed. The time series information fusion module calculates the time series information only when it receives a fusion request (for example, when a moving object is detected), rather than calculating it for each frame of data, thereby reducing the average computing load. For example, the feature space conversion module initiates a fusion request only when a moving object is detected. Among them, the prediction module can be expanded according to actual conditions, and the embodiment of the present application does not limit the number of prediction modules. The task prediction result is used to represent relevant information of the scene where the autonomous driving device is located, such as the category of objects in the scene and the movement speed of the objects. The processor can control the behavior of the autonomous driving device based on multiple task prediction results, thereby adjusting the driving state of the autonomous driving device. The adjustment of the driving state can be deceleration, steering, parking, etc. For example, if the processor determines that there is an animal in front of the autonomous driving device based on the task prediction results, it controls the autonomous driving device to stop moving.

[0031] In some embodiments, after projecting the two-dimensional image features into the three-dimensional space, a dispersed feature point cloud is obtained, and the processor aggregates the feature point cloud to obtain the three-dimensional features, which are used to represent objects in a scene in which the autonomous driving device is located. Projecting two-dimensional image features into a three-dimensional space is an important step in realizing two-dimensional image to three-dimensional scene understanding. For example, two-dimensional image features can be mapped to a three-dimensional space using depth information. The feature point cloud is a set of discrete points obtained after projecting two-dimensional image features into a three-dimensional space, each point containing three-dimensional coordinate information and attributes related to the original two-dimensional image features. The position of each feature point in the three-dimensional space is represented by its x, y, z coordinates. For example, in a projection based on depth information, the depth value directly determines the position of the feature point in the z-axis direction, and the image pixel coordinates determine the positions in the x and y-axis directions. The aggregation of the feature point cloud is the process of integrating the dispersed feature point cloud data into three-dimensional features that can effectively represent objects in a three-dimensional scene. This process is crucial for subsequent scene understanding, object recognition, and decision-making of the autonomous driving device.

[0032] In some embodiments, the processor extracts semantic information from the two-dimensional image features through the feature space conversion module, the semantic information including target categories and scene geometry, predicts the depth value of each pixel in the first image according to the semantic information, and projects the two-dimensional image features into the three-dimensional space according to the depth value. Illustratively, the feature space conversion module includes a monocular depth estimation network, which mainly functions to convert two-dimensional image features into depth information, thereby providing a basis for subsequent 3D projection. The monocular depth estimation network adopts an encoder-decoder structure. The encoder part abstracts the input two-dimensional image features layer by layer to extract high-level semantic features. These features contain rich target category and structure information, such as human height ratio, vehicle size, etc. The decoder part gradually recovers the depth information of the image based on the semantic features extracted by the encoder to generate a depth map corresponding to the size of the input image. The feature space conversion module also includes a feature projection part, which projects the two-dimensional image features into a three-dimensional space using the estimated depth information to obtain a large number of dispersed feature point clouds.

[0033] In some embodiments, the aggregation processing is pooling processing. When processing the feature point cloud, pooling can aggregate the dispersed feature point cloud into a more compact three-dimensional feature representation. For example, the feature space conversion module has a pooling layer that processes the feature point cloud according to a set aggregation rule (such as maximum pooling, average pooling, etc.).

[0034] In some embodiments, the processor is further configured to: maintain, by the temporal information fusion module, a queue, store the three-dimensional features output by the feature space conversion module into the queue each time, remove the three-dimensional features stored in the queue earliest when the queue reaches a preset length, and record the three-dimensional features in the queue as historical three-dimensional features. The main function of the temporal information fusion module is to manage the historical three-dimensional features so as to utilize the historical information in subsequent processing. The queue stores the three-dimensional features output by the feature space conversion module, and by maintaining the historical three-dimensional features, the motion trajectory of the object can be tracked more accurately, and the future position of the object can be predicted. For example, the maximum length L of the queue is set, and each time the feature space conversion module outputs a new three-dimensional feature, the new three-dimensional feature is stored into the queue. When the length of the queue reaches L, the three-dimensional feature stored earliest is removed, so that the most recent and most timely three-dimensional features are always stored in the queue, and the dynamic adjustment mechanism of the queue avoids high memory occupation and improves the robustness of the system. When the feature space conversion module initiates a request for fusion of historical feature information, the temporal information fusion module receives the current three-dimensional feature. The current three-dimensional feature is spliced (fused) with the historical three-dimensional features in the queue, and then the spliced features (fused features) are sent to the fusion network for processing. The fusion network adopts a convolutional neural network architecture, performs layer-by-layer calculation on the input fused features, and generates more abundant and accurate temporal information for target detection, semantic segmentation, motion prediction and other tasks. The queue is a linear data structure that follows the first-in first-out (FIFO) principle. It only allows elements to be inserted at one end (tail) and deleted at the other end (head). Basic operations of the queue include enqueue, dequeue, front, etc. For example, the queue can be implemented by using an array.

[0035] In some embodiments, the temporal information fusion module can adopt a convolutional neural network (CNN) or a long short-term memory network (LSTM).

[0036] In some embodiments, the plurality of prediction modules include a semantic segmentation module, an instance segmentation module, and a speed estimation module, and the plurality of task prediction results include a geometric structure of a scene where the autonomous driving device is located, a division result of semantic information of different regions in the scene, differentiation information of different individuals of the same target category in the scene, and a speed of each voxel in the scene. In the embodiments of the present application, the autonomous driving device focuses on three tasks of semantic segmentation, instance segmentation and speed estimation, and the combination of the three tasks can provide the autonomous driving device with more comprehensive environmental perception capability.

[0037] In some embodiments, the semantic segmentation module can employ FCN networks (Fully Convolutional Networks), U-Net, DeepLab networks, etc. The instance segmentation module can employ Mask R-CNN networks or YOLO networks. The speed estimation module can employ FlowNet or PWC-Net.

[0038] The function of the semantic segmentation module is geometric structure estimation and semantic information division. Geometric structure estimation refers to estimating the geometric structure of the scene, which can involve predicting the depth information, 3D shape, etc. of the scene. For example, by predicting the depth value of each pixel or voxel, the three-dimensional structure of the scene can be constructed. Semantic information division refers to dividing the semantic information of different regions in the scene, in other words, classifying each pixel or voxel in the scene into a predefined semantic category, such as road, building, pedestrian, etc.

[0039] The instance segmentation module is responsible for distinguishing different individuals belonging to the same semantic category. For example, different vehicles, pedestrians, etc. are identified in the scene, and each individual is assigned a unique identifier.

[0040] The speed estimation module voxelizes the scene, i.e. divides the continuous scene space into discrete voxel units, each of which can contain part of one or more objects in the scene. Then the speed of each voxel is estimated, which can provide information about the motion of objects in the scene for the autonomous driving device, helping to predict the future position of the objects and thus achieve safer navigation.

[0041] In some embodiments, the neural network model further comprises an image feature enhancement module, and the processor is further configured to process the two-dimensional image features through the image feature enhancement module to obtain image features of multiple scales, and input the image features of multiple scales to the feature space conversion module. By processing the two-dimensional image features through the image feature enhancement module, image features of multiple scales can be obtained. This is because features of different scales can capture information of different sizes and details in the image. For example, in the target detection task, large-scale features are more suitable for detecting large objects, while small-scale features are more advantageous for detecting small objects. The feature space conversion module can fuse and integrate the two-dimensional image features of multiple scales to generate more representative and discriminative feature representations. For example, the image feature enhancement module can employ a spatial pyramid pooling (SPP) algorithm or a feature pyramid network (FPN) algorithm.

[0042] During driving of the autonomous driving device, a unified feature is used to infer multiple task prediction results simultaneously, so that each result inferred has high consistency, and multi-task prediction is completed in a single forward propagation, which can reduce the amount of calculation compared with a serial architecture. In addition, projecting a two-dimensional image feature into a three-dimensional space helps multiple prediction modules to understand space and provides a basis for subsequent temporal fusion. The temporal information fusion module combines the three-dimensional feature extracted at the current time with the three-dimensional feature (historical three-dimensional feature) at the past time, which can capture the dynamic changes of the scene and improve the continuity of perception and prediction ability. Figure 2 A flowchart of a control method of an autonomous driving device provided by an embodiment of the present application is shown, and the details are as follows: Step 201, control the autonomous driving device to drive.

[0043] Step 202, during driving of the autonomous driving device, continuously acquire images using an image acquisition unit.

[0044] Step 203, at a first time T, according to a first image acquired at the first time T, complete a forward propagation once using a neural network model to synchronously output multiple task prediction results.

[0045] Step 204, adjust the driving state of the autonomous driving device according to the multiple task prediction results.

[0046] In an embodiment of the present application, the neural network model includes an image feature extraction module, a feature space conversion module, a temporal information fusion module, and multiple prediction modules. In the process of the forward propagation, the image feature extraction module extracts a two-dimensional image feature from the first image, the feature space conversion module projects the two-dimensional image feature into a three-dimensional space to obtain a three-dimensional feature, the temporal information fusion module receives a fusion request, fuses the three-dimensional feature and a historical three-dimensional feature according to the fusion request to obtain a fusion feature, and calculates temporal information according to the fusion feature, and the multiple prediction modules respectively calculate the three-dimensional feature and the temporal information to output the multiple task prediction results. The control method provided by the present application can be specifically referred to the related description of the above device embodiments, which will not be repeated here.

[0047] In some embodiments, the adjusting the driving state of the autonomous driving device according to the multiple task prediction results comprises: constructing a map of a scene where the autonomous driving device is located according to the multiple task prediction results; and adjusting the driving state of the autonomous driving device according to the map. The multiple task prediction results include semantic segmentation results, instance segmentation results, and speed estimation results. The semantic segmentation results provide semantic information of different regions in the scene, such as roads, buildings, pedestrians, vehicles, etc. These information can help to construct the basic framework of the map, and clarify the function and type of different regions. The instance segmentation results can distinguish different individuals in the same semantic category, such as marking different vehicles and pedestrians. These information can further refine the map, and assign a unique identifier to each individual, so as to realize accurate representation of each object in the scene. The speed estimation results provide speed information of each voxel in the scene. These speed information can be used for the construction of dynamic map, and help to predict the future position of the object, so as to realize more accurate real-time map updating. By combining these multiple task prediction results, a high-precision scene map containing static and dynamic information can be constructed. Based on the constructed map, the autonomous driving device can plan an optimal path. Since the map is constructed based on dynamic information, the device can continuously adjust the driving state according to the real-time updated map information. For example, if a new obstacle is detected in front of the device, the device can adjust the path or speed in time to avoid collision.

[0048] In some embodiments, the image feature extraction module, the feature space conversion module, the time sequence information fusion module, and the multiple prediction modules interact through standardized interfaces. Through the standardized interfaces, each module (such as the image feature extraction module, the feature space conversion module, the time sequence information fusion module, and the multiple prediction modules) can be independently developed and tested. The standardized interfaces make the system have good scalability. When new functions or modules need to be added, only the new module needs to conform to the interface specification. At the same time, it is more convenient to maintain and update the modules, because the coupling degree between each module is low.

[0049] As can be seen from the above, in the scheme of the present application, during the driving of the autonomous driving device, a unified feature is used to infer multiple task prediction results, so that each result inferred has high consistency, and the multi-task prediction is completed in a single forward propagation, which can reduce the amount of calculation compared with a serial architecture. In addition, projecting the two-dimensional image feature into a three-dimensional space helps the multiple prediction modules to understand the space and provides a basis for subsequent time sequence fusion. The time sequence information fusion module combines the three-dimensional features extracted at the current time with the three-dimensional features (historical three-dimensional features) at past times, which can capture the dynamic changes of the scene and improve the coherence of perception and prediction ability.

[0050] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. Figure 3 The structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 3. As shown in FIG. 3, the electronic device 3 of the embodiment includes at least one processor 30 (only one is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, the electronic device performs the steps of the control method embodiment. Figure 3 Figure 3 The processor 30 executes the computer program 32, so that the electronic device performs the steps of the control method embodiment.

[0051] The electronic device 3 can include, but is not limited to, the processor 30 and the memory 31. Those skilled in the art can understand that the electronic device 3 is only an example of the electronic device 3 and does not constitute a limitation on the electronic device 3, and can include more or fewer components than those shown in the figure, or combine certain components, or different components, for example, can also include input / output devices, network access devices, etc. Figure 3 The processor 30 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or can also be any conventional processor.

[0052]

[0053] ​​The memory 31 can be an internal storage unit of the electronic device 3 in some embodiments, such as a hard disk or a memory of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3 in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Further, the memory 31 can include both an internal storage unit and an external storage device of the electronic device 3. The memory 31 is used to store an operating system, an application program, a boot loader, data, and other programs, such as program codes of the computer program. The memory 31 can also be used to temporarily store data that has been output or is to be output. It should be noted that the information interaction and execution process between the above devices / units are based on the same concept as the method embodiments of the present application, and the specific functions and technical effects brought about can be referred to the method embodiments part. Therefore, no further description is given here.

[0054] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example for illustration. In actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for mutual distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, and will not be described here.

[0055] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0056] The embodiments of the present application provide a computer program product, which, when running on an electronic device, causes the electronic device to perform the steps in the above method embodiments.

[0057] The chip system provided in the embodiments of the present application comprises a processor, and the processor is configured to call and run a computer program from a memory, so that an electronic device installed with the chip system performs the steps in each of the method embodiments.

[0058] The integrated units described above, if realized in the form of software function units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the method embodiments described above can be completed by a computer program instructing relevant hardware. The computer program described above can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of each of the method embodiments described above can be implemented. The computer program described above includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer-readable medium described above can at least include any entity or device capable of carrying the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunications signal and a software distribution medium. For example, a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunications signal.

[0059] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0060] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0061] In the embodiments provided by the present application, it should be understood that the disclosed apparatus / network device and method can be implemented in other manners. For example, the embodiments of the apparatus / network device described above are merely schematic, and the division of the modules or units is merely logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.

[0062] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0063] The above embodiments are merely used to describe the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An autonomous driving device, characterized in that: include: body; A plurality of driving wheels, wherein the driving wheels are used to drive the vehicle body to move; An image acquisition unit, arranged on the vehicle body; A memory storing a neural network model, wherein the neural network model includes an image feature extraction module, a feature space conversion module, a time series information fusion module, and multiple prediction modules; The processor is configured to: control the autonomous driving device to travel; while the autonomous driving device is traveling, continuously collect images using the image acquisition unit; at a first time T, perform a forward propagation using the neural network model based on a first image collected at the first time T to simultaneously output multiple task prediction results; and adjust the driving state of the autonomous driving device based on the multiple task prediction results; In which, during the forward propagation process, the processor extracts two-dimensional image features from the first image through the image feature extraction module, and projects the two-dimensional image features into three-dimensional space through the feature space conversion module to obtain three-dimensional features. The processor also receives a fusion request through the time series information fusion module, fuses the three-dimensional features with historical three-dimensional features according to the fusion request to obtain a fusion feature, and calculates the time series information based on the fusion feature. The processor also calculates the three-dimensional features and the time series information respectively through the multiple prediction modules to output the multiple task prediction results.

2. The autonomous driving device according to claim 1, wherein: After projecting the two-dimensional image features into the three-dimensional space, a dispersed feature point cloud is obtained. The processor aggregates the feature point cloud to obtain the three-dimensional features, which are used to represent objects in the scene where the autonomous driving device is located.

3. The autonomous driving device according to claim 2, wherein: The aggregation process is a pooling process.

4. The autonomous driving device according to claim 1, wherein: The processor is further configured to: maintain a queue through the temporal information fusion module, store the three-dimensional features output by the feature space conversion module each time into the queue, and when the queue reaches a preset length, remove the three-dimensional features stored earliest in the queue from the queue; and record the three-dimensional features in the queue as historical three-dimensional features.

5. The autonomous driving device according to claim 1, wherein: The processor extracts semantic information from the two-dimensional image features through the feature space conversion module, where the semantic information includes target categories and scene geometry, predicts the depth value of each pixel in the first image based on the semantic information, and projects the two-dimensional image features into the three-dimensional space based on the depth value.

6. The autonomous driving device according to claim 5, wherein: The multiple prediction modules include a semantic segmentation module, an instance segmentation module and a speed estimation module. The multiple task prediction results include the geometric structure of the scene where the autonomous driving device is located, the division results of semantic information of different areas in the scene, the distinction information of different individuals of the same target category in the scene, and the speed of each voxel in the scene.

7. The autonomous driving device according to claim 1, wherein: The neural network model also includes an image feature enhancement module, and the processor is further configured to: process the two-dimensional image features through the image feature enhancement module to obtain the two-dimensional image features of multiple scales, and input the two-dimensional image features of multiple scales into the feature space conversion module.

8. A control method for an autonomous driving device, characterized in that: include: controlling the autonomous driving device to travel; During the driving of the autonomous driving device, continuously capturing images using an image capture unit; At a first time point T, a neural network model is used to complete a forward propagation based on a first image collected at the first time point T to simultaneously output multiple task prediction results. The neural network model includes an image feature extraction module, a feature space conversion module, a time series information fusion module, and multiple prediction modules; adjusting the driving state of the autonomous driving device according to the plurality of task prediction results; Among them, during the forward propagation process, the image feature extraction module extracts two-dimensional image features from the first image, the feature space conversion module projects the two-dimensional image features into three-dimensional space to obtain three-dimensional features, the time series information fusion module receives a fusion request, fuses the three-dimensional features with historical three-dimensional features according to the fusion request to obtain fused features, and calculates time series information based on the fused features, and the multiple prediction modules respectively calculate the three-dimensional features and the time series information to output the multiple task prediction results.

9. The control method according to claim 8, wherein: The adjusting the driving state of the autonomous driving device according to the plurality of task prediction results includes: Constructing a map of the scene where the autonomous driving device is located based on the multiple task prediction results; The driving state of the autonomous driving device is adjusted according to the map.

10. The control method according to claim 8, wherein: The image feature extraction module, the feature space conversion module, the temporal information fusion module and the multiple prediction modules interact with each other through a standardized interface.

Citation Information

Patent Citations

  • Automatic driving 3D target detection method and related device

    CN116259043A

  • Fusion tracking method and device based on multi-sensor time sequence sensing result post-fusion

    CN116433714A

  • Automatic driving BEV task learning method and related device

    CN116469079A

  • Roadway infrastructure monitoring based on aggregated mobile vehicle communication parameters

    US20160011124A1

  • Spatio-temporal pose / object database

    US20210101614A1