Speed estimation method and device, electronic equipment and storage medium

By employing a multimodal fusion method, which utilizes feature extraction and interactive fusion of image, radar, and driving trajectory data, the problem of inaccurate speed estimation in map navigation applications is solved, enabling more accurate speed and travel time predictions and improving the user's travel experience.

CN116504057BActive Publication Date: 2026-04-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2023-04-18
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing map navigation mobile applications do not accurately estimate speed based on trajectory data, resulting in inaccurate overall travel time estimates.

Method used

A multimodal fusion method is adopted to acquire driving data of various modalities (such as images, radar, and driving trajectory data). Through feature extraction and interactive fusion, speed prediction is performed using a pre-trained large model and a Siamese network model, and accurate prediction is achieved by combining information from different modalities.

Benefits of technology

It achieves more accurate speed and overall travel time prediction, reduces the timeliness issues of congestion gathering and dissipation, ensures users' rationality in road selection, saves users' travel time, and improves the perceived experience of travel time prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116504057B_ABST
    Figure CN116504057B_ABST
Patent Text Reader

Abstract

The present disclosure provides a speed estimation method and device, electronic equipment and storage medium, relates to the technical field of artificial intelligence, in particular to data processing, deep learning and intelligent transportation technology. The method comprises: obtaining driving data of multiple modalities; performing feature extraction on the driving data of multiple modalities to obtain driving feature data of multiple modalities; performing interactive fusion on the driving feature data of different modalities to obtain driving feature fusion data of multiple fusion modalities; and performing speed estimation according to the driving feature fusion data of multiple fusion modalities to obtain a target speed estimation value. The present disclosure adopts a multi-modal fusion manner, fuses driving feature data of different modalities, and interacts driving feature data of different modalities, so as to obtain more representative driving feature fusion data, and then more accurate speed estimation and overall travel time estimation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, specifically to data processing, deep learning and intelligent transportation technologies, and particularly to a speed prediction method, apparatus, electronic device and storage medium. Background Technology

[0002] Currently, map navigation mobile applications typically estimate vehicle speed in real time based on vehicle trajectory data. However, in practice, speed calculations based on trajectory data are limited by the quality and quantity of the trajectory data, resulting in inaccurate speed estimates, which in turn lead to inaccurate estimates of overall travel time. Summary of the Invention

[0003] A method, apparatus, electronic device, and storage medium for speed estimation are provided.

[0004] According to the first aspect, a speed prediction method is provided, comprising: acquiring driving data of multiple modes; extracting features from the driving data of multiple modes to obtain driving feature data of multiple modes; interactively fusing the driving feature data of different modes to obtain driving feature fusion data of multiple fused modes; and predicting speed based on the driving feature fusion data of multiple fused modes to obtain a target speed prediction value.

[0005] According to a second aspect, a speed prediction device is provided, comprising: an acquisition module for acquiring driving data of multiple modes; an extraction module for extracting features from the driving data of multiple modes to obtain driving feature data of multiple modes; a fusion module for interactively fusing the driving feature data of different modes to obtain driving feature fusion data of multiple fusion modes; and a prediction module for predicting speed based on the driving feature fusion data of multiple fusion modes to obtain a target speed prediction value.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to said at least one processor; wherein the memory stores instructions executable by said at least one processor, said instructions being executed by said at least one processor to enable said at least one processor to perform the speed estimation method described in the first aspect of this disclosure.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a speed estimation method according to a first aspect of this disclosure.

[0008] According to a fifth aspect, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the speed estimation method according to the first aspect of this disclosure.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0011] Figure 1 This is a schematic flowchart of a speed estimation method according to the first embodiment of this disclosure;

[0012] Figure 2 This is a schematic flowchart of a speed estimation method according to a second embodiment of the present disclosure;

[0013] Figure 3 This is a schematic flowchart of a speed estimation method according to a third embodiment of the present disclosure;

[0014] Figure 4 This is a schematic diagram of the overall process of the speed estimation method according to the fourth embodiment of this disclosure;

[0015] Figure 5 This is a block diagram of a speed estimation device according to a first embodiment of the present disclosure;

[0016] Figure 6 This is a block diagram of a speed estimation device according to a second embodiment of the present disclosure;

[0017] Figure 7 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation

[0018] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0019] Artificial intelligence (AI) is a technical science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. Currently, AI technology has the advantages of high automation, high accuracy, and low cost, and has been widely applied.

[0020] Data processing (DP) is the acquisition, storage, retrieval, processing, transformation, and transmission of data. The fundamental purpose of data processing is to extract and derive valuable and meaningful data from large amounts of potentially chaotic and incomprehensible data. Data processing is a fundamental component of systems engineering and automatic control. It permeates all areas of social production and life. The development of data processing technology and the breadth and depth of its applications have profoundly influenced the progress of human society.

[0021] Deep learning (DL) is a new research direction in the field of machine learning (ML). It learns the inherent patterns and hierarchical representations of sample data. The information gained during this learning process greatly aids in interpreting data such as text, images, and sound. Its ultimate goal is to enable machines to possess analytical and learning capabilities like humans, allowing them to recognize text, images, and sound. Specific research areas mainly include convolutional neural networks (CNNs) based on convolutional operations; autoencoder neural networks based on multi-layered neurons; and deep belief networks that are pre-trained using multi-layered autoencoder neural networks and then further optimized by incorporating discriminative information. Deep learning has achieved significant results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech recognition, recommendation and personalization technologies, and other related fields. Deep learning enables machines to mimic human activities such as sight, hearing, and thinking, solving many complex pattern recognition problems and leading to significant advancements in artificial intelligence technologies.

[0022] Intelligent Traffic System (ITS), also known as Intelligent Transportation System, effectively integrates advanced science and technology (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing. It strengthens the connection between vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, improves the environment, and saves energy.

[0023] The following description, in conjunction with the accompanying drawings, describes a speed estimation method, apparatus, electronic device, and storage medium according to embodiments of the present disclosure.

[0024] Figure 1 This is a schematic flowchart of a speed estimation method according to the first embodiment of this disclosure.

[0025] like Figure 1 As shown, the speed estimation method of this disclosure embodiment may specifically include the following steps:

[0026] S101 acquires driving data in multiple modes.

[0027] Specifically, the execution entity of the speed estimation method in this embodiment of the disclosure can be the speed estimation device provided in this embodiment of the disclosure. This speed estimation device can be a hardware device with data processing capabilities and / or the necessary software to drive the hardware device. Optionally, the execution entity may include a workstation, server, computer, user terminal, and other devices. The user terminal includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and vehicle terminals.

[0028] In this embodiment of the disclosure, driving data refers to data related to vehicle movement. A modality is any source or form of information. Multiple modalities refer to information with at least two modalities, including text, images, video, and audio. This step is used to acquire driving data in multiple modalities. Specifically, the driving data in multiple modalities may include, but is not limited to, at least two of the following modalities: image data, radar data, and driving trajectory data.

[0029] The image data consists of images captured by the vehicle's camera while the vehicle is in motion. It's important to note that each image returned includes a timestamp, because in the final layer of feature representation using the image data, the image features are embedded into each time slice.

[0030] Radar data is information from vehicle-mounted radar, mainly including the distance and relative speed between vehicles in the same lane and adjacent lanes. The relative speed can effectively reflect the congestion situation, and we can obtain a wealth of road condition information from radar data.

[0031] Driving trajectory data consists of information transmitted by users of navigation products during driving and while driving. It mainly depicts the speed information of vehicles in the real world, as well as the speed information of various vehicles passing through the same road segment. From the driving trajectory data, we can obtain speed information at the road segment level.

[0032] S102, extract features from driving data of multiple modes to obtain driving feature data of multiple modes.

[0033] In this embodiment of the disclosure, features can be extracted from driving data of different modes to obtain driving feature data of different modes, such as driving feature data of image mode, driving feature data of radar mode, and driving feature data of driving trajectory mode.

[0034] S103, interactively fuse driving feature data from different modalities to obtain fused driving feature data with multiple fused modalities.

[0035] In this embodiment, driving feature data from different modalities can be fused pairwise to obtain driving feature fusion data of different fusion modalities, such as driving feature fusion data of image-radar fusion modality, driving feature fusion data of image-trajectory fusion modality, and driving feature fusion data of radar-trajectory fusion modality. By employing a multimodal fusion approach, driving feature data from images, radar, and driving trajectories are fused, i.e., feature representations. The feature representations of images, radar signals, and driving trajectory data are interacted to obtain more representative feature representations, i.e., driving feature fusion data.

[0036] S104. Speed ​​is estimated based on the fusion data of driving characteristics from multiple fusion modes to obtain the target speed estimate.

[0037] In this embodiment of the disclosure, speed prediction is performed based on the driving feature fusion data of multiple fusion modes obtained in step S103, thereby obtaining the final speed prediction value, i.e., the target speed prediction value.

[0038] In summary, the speed estimation method of this disclosure extracts features from driving data of multiple modalities to obtain driving feature data of multiple modalities. It then inter-fuses these driving feature data of different modalities to obtain fused driving feature data of multiple fused modalities. Based on this fused driving feature data of multiple fused modalities, speed estimation is performed to obtain a target speed estimate. This disclosure adopts a multimodal fusion approach, fusing driving feature data of different modalities and inter-fusing them to obtain more representative driving feature fusion data. This enables more accurate speed estimation and overall travel time estimation, while reducing the timeliness issues of congestion accumulation and dissipation, ensuring the rationality of users' road selection, scientifically guiding users' travel, saving users' travel time, and continuously improving users' perception of overall travel time estimation.

[0039] Figure 2 This is a schematic flowchart of a speed estimation method according to a second embodiment of the present disclosure.

[0040] like Figure 2 As shown, in Figure 1Based on the illustrated embodiments, the speed estimation method of this disclosure specifically includes the following steps:

[0041] S201 acquires driving data in multiple modes.

[0042] In this embodiment, step S201 is the same as step S101 in the above embodiment, and will not be described again here.

[0043] The step S102 in the above embodiment, "extracting features from driving data of multiple modes to obtain driving feature data of multiple modes", may specifically include the following step S202.

[0044] S202, a feature extraction model is used to extract features from the driving data to obtain driving feature data.

[0045] In this embodiment of the disclosure, a pre-trained large model can be used as a feature extraction model to extract features from driving data, thereby obtaining driving feature data. The pre-trained large model is a model that has been trained in advance with massive amounts of data. For example, five years of accumulated driving data from three sources—images, radar, and driving trajectories—can be used for pre-training of the model. Through learning from massive amounts of multi-source data, the model can learn general traffic pattern features and traffic trend features.

[0046] For image data, image feature extraction models such as the ResNet 152 model can be used to extract features from the image data to obtain driving feature data of the image modality. The ResNet 152 model is a FAST-RNN model that identifies the distance of the current vehicle and the three vehicles ahead in the same lane based on the image data, and uses this as the feature representation of the image data, i.e., the driving feature data of the image modality.

[0047] For radar data and driving trajectory data, time series models such as the Transformer model can be used to extract features from the radar data and driving trajectory data to obtain driving feature data of radar mode and driving feature data of driving trajectory mode.

[0048] The step S103 in the above embodiment, "interactively fusing driving feature data of different modes to obtain driving feature fusion data of multiple fusion modes", may specifically include the following step S203.

[0049] S203 uses a twin network model to interactively fuse driving feature data from different modalities, resulting in fused driving feature data with multiple fusion modalities.

[0050] In this embodiment, the twin network model, also known as the dual-tower network model, receives driving feature data from two different modalities. The twin network model then interactively fuses the driving feature data from the two modalities to obtain fused driving feature data for the fused modality. For example, when encoding driving feature data from the image modality, the twin network model can fuse driving feature data from the radar modality or driving feature data from the driving trajectory modality.

[0051] Step S104 in the above embodiment, "Estimating speed based on driving feature fusion data of multiple fusion modes to obtain target speed estimate", may specifically include the following step S204.

[0052] S204: Input the fusion data of driving characteristics from multiple fusion modes into the same speed prediction model to predict the speed and obtain the target speed prediction value.

[0053] In this embodiment, the fused driving feature data from multiple fusion modalities obtained in step S203 can be input into the same speed prediction model. The speed prediction model performs speed prediction based on the input data and outputs a target speed prediction value. Specifically, the speed prediction model can be a time series model, such as the Transformer model. For driving feature data from different sources, a multimodal fusion approach is adopted, combining intuitive road condition depictions from image data, more accurate vehicle distance and relative speed depictions from radar data, and speed depictions from driving trajectory data. By combining these modalities, the speed prediction model can learn information from different sources, achieving more accurate real-time speed prediction.

[0054] It should be noted here that while the correlation between driving data from different modalities is high, there is a problem of synchronization issues between the multimodal data sources. To address this issue, the data from different modalities can be used for separate speed prediction, and then the final speed prediction value can be obtained through integration. Specifically, step S104 in the above embodiment, "Speed ​​prediction is performed based on the fusion data of driving features from multiple fusion modalities to obtain the target speed prediction value," can include the following steps S301-S302:

[0055] S301, the driving feature fusion data of multiple fusion modes are input into different speed prediction models according to the fusion mode to predict the speed and obtain the initial speed prediction value of multiple fusion modes.

[0056] In this embodiment of the disclosure, the driving feature fusion data of multiple fusion modes obtained in step S203 can be input into the same speed prediction model according to different fusion modes. The driving feature fusion data of the same fusion mode can be input into different speed prediction models. The speed prediction model performs speed prediction based on the input data and outputs the initial speed prediction value of the corresponding fusion mode.

[0057] S302, determine the target velocity estimate based on the initial velocity estimates of multiple fusion modes.

[0058] In this embodiment of the disclosure, the initial velocity estimates of the various fusion modes obtained in step S301 can be fused using rules to obtain the target velocity estimate. Specifically, the rule fusion may include, but is not limited to, any of the following fusion methods: maximum value fusion and average value fusion, etc.

[0059] In summary, the speed prediction method of this disclosure adopts a multimodal fusion approach, integrating driving feature data from different modalities. This involves the interaction of driving feature data from different modalities to obtain more representative driving feature fusion data. This results in more accurate speed and overall travel time predictions, while also reducing the timeliness issues of congestion accumulation and dissipation, ensuring the rationality of users' road choices, scientifically guiding users' travel, saving users' travel time, and continuously improving users' perception of overall travel time prediction. Furthermore, the use of a pre-trained large model for feature extraction, through learning from massive amounts of multi-source data, allows the pre-trained large model to learn general traffic pattern features and traffic trend features, improving the representativeness of the extracted driving feature data and thus achieving more accurate speed and overall travel time predictions. For driving feature data from different sources, a multimodal fusion approach is adopted, combining intuitive road condition descriptions from image data, more accurate vehicle distance and relative speed descriptions from radar data, and speed descriptions from driving trajectory data. By combining these modalities, the speed prediction model can learn information from different sources and achieve more accurate real-time speed prediction.

[0060] To clearly illustrate the speed estimation method of the embodiments of this disclosure, it is now combined with Figure 4 Provide a detailed description. Figure 4 This is a schematic diagram of the overall process of a speed estimation method according to an embodiment of the present disclosure.

[0061] like Figure 4As shown, taking image-mode driving data, radar-mode driving data, and driving trajectory-mode driving data as examples, the ResNet 152 model is used to extract features from the image data to obtain image-mode driving feature data. The Transformer model is used to extract features from the radar data and driving trajectory data to obtain radar-mode driving feature data and driving trajectory-mode driving feature data. The Siamese network model is used to perform pairwise interaction and fusion of the three modes of driving feature data to obtain fused driving feature data of the three modes. This fused data is then input into the speed prediction model to predict the target speed.

[0062] Figure 5 This is a block diagram of a speed estimation device according to a first embodiment of the present disclosure.

[0063] like Figure 5 As shown, the speed estimation device 500 of this embodiment includes: an acquisition module 501, an extraction module 502, a fusion module 503, and an estimation module 504.

[0064] The acquisition module 501 is used to acquire driving data in multiple modes.

[0065] The extraction module 502 is used to extract features from driving data of multiple modes to obtain driving feature data of multiple modes.

[0066] The fusion module 503 is used to interactively fuse driving feature data from different modalities to obtain fused driving feature data with multiple fusion modalities.

[0067] The prediction module 504 is used to predict the speed based on the fusion data of driving characteristics from multiple fusion modes, and obtain the target speed prediction value.

[0068] It should be noted that the explanation of the above-described embodiments of the speed estimation method also applies to the speed estimation device of the embodiments of this disclosure, and the specific process will not be repeated here.

[0069] In summary, the speed estimation device of this disclosure extracts features from driving data of multiple modalities to obtain driving feature data of multiple modalities. It then inter-fuses these driving feature data of different modalities to obtain fused driving feature data of multiple fused modalities. Based on this fused driving feature data of multiple fused modalities, it estimates the target speed. This disclosure adopts a multimodal fusion approach, fusing driving feature data of different modalities and inter-fusing them to obtain more representative fused driving feature data. This enables more accurate speed estimation and overall travel time estimation, while reducing the timeliness issues of congestion accumulation and dissipation, ensuring the rationality of users' road selection, scientifically guiding users' travel, saving users' travel time, and continuously improving users' perception of overall travel time estimation.

[0070] Figure 6 This is a block diagram of a speed estimation device according to a second embodiment of the present disclosure.

[0071] like Figure 6 As shown, the speed estimation device 600 of this embodiment includes: an acquisition module 601, an extraction module 602, a fusion module 603, and an estimation module 604.

[0072] The acquisition module 601 has the same structure and function as the acquisition module 501 in the previous embodiment, the extraction module 602 has the same structure and function as the extraction module 502 in the previous embodiment, the fusion module 603 has the same structure and function as the fusion module 503 in the previous embodiment, and the prediction module 604 has the same structure and function as the prediction module 504 in the previous embodiment.

[0073] Furthermore, the multimodal driving data includes at least two of the following modalities: image data, radar data, and driving trajectory data.

[0074] Furthermore, the fusion module 603 includes a fusion unit 6031, which is used to interactively fuse driving feature data of different modalities using a twin network model to obtain driving feature fusion data of multiple fusion modalities.

[0075] Furthermore, the prediction module 604 is further used to: input the fusion data of driving characteristics from multiple fusion modes into the same speed prediction model to predict the speed and obtain the target speed prediction value.

[0076] Furthermore, the prediction module 604 is further used to: input the driving feature fusion data of multiple fusion modes into different speed prediction models according to the fusion mode to perform speed prediction, and obtain the initial speed prediction value of multiple fusion modes; and determine the target speed prediction value based on the initial speed prediction value of multiple fusion modes.

[0077] Furthermore, the prediction module 604 is further used to: perform rule fusion on the initial velocity prediction values ​​of multiple fusion modes to obtain the target velocity prediction value. The rule fusion includes any one of the following fusion methods: maximum value fusion and average value fusion.

[0078] Furthermore, the extraction module 602 is further used to: extract features from the driving data using a feature extraction model to obtain driving feature data.

[0079] Furthermore, the extraction module 602 is further used to: extract features from the image data using an image feature extraction model to obtain driving feature data.

[0080] Furthermore, the extraction module 602 is further used to: extract features from radar data or driving trajectory data using a time series model to obtain driving feature data.

[0081] It should be noted that the explanation of the above-described embodiments of the speed estimation method also applies to the speed estimation device of the embodiments of this disclosure, and the specific process will not be repeated here.

[0082] In summary, the speed prediction device of this disclosure adopts a multimodal fusion approach, integrating driving feature data from different modalities. By interacting with the driving feature data from different modalities, more representative driving feature fusion data is obtained, thereby achieving more accurate speed and overall travel time prediction. Simultaneously, it reduces the timeliness issues of congestion gathering and dissipation, ensuring the rationality of users' road choices, scientifically guiding users' travel, saving users' travel time, and continuously improving users' perception of overall travel time prediction. Furthermore, by employing a pre-trained large model for feature extraction, and through learning from massive amounts of multi-source data, the pre-trained large model can learn general traffic pattern features and traffic trend features, improving the representativeness of the extracted driving feature data, thus achieving more accurate speed and overall travel time prediction. For driving feature data from different sources, a multimodal fusion approach is adopted, combining intuitive road condition descriptions from image data, more accurate vehicle distance and relative speed descriptions from radar data, and speed descriptions from driving trajectory data. By combining these modalities, the speed prediction model can learn information from different sources and achieve more accurate real-time speed prediction.

[0083] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0084] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0085] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0086] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0087] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0088] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as... Figures 1 to 4The speed estimation method is illustrated. For example, in some embodiments, the speed estimation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by computing unit 701, one or more steps of the semantic parsing method described above may be performed. Alternatively, in other embodiments, computing unit 701 may be configured to execute the speed estimation method by any other suitable means (e.g., by means of firmware).

[0089] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0090] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0091] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0092] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0093] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0094] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the management difficulties and weak business scalability inherent in traditional physical hosts and VPS (Virtual Private Server) services. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0095] According to embodiments of this disclosure, this disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the steps of the speed estimation method shown in the above embodiments of this disclosure.

[0096] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0097] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for predicting speed, comprising: Acquire driving data in multiple modalities; Feature extraction is performed on the driving data of the various modalities to obtain driving feature data of the various modalities; The driving feature data of different modalities are interactively fused to obtain driving feature fusion data of multiple fusion modalities; as well as Speed ​​prediction is performed based on the fusion data of driving characteristics from the multiple fusion modes to obtain the target speed prediction value; The process involves interactively fusing driving feature data from different modalities to obtain fused driving feature data with multiple fused modalities, including: The driving feature data of different modalities are interactively fused using a twin network model to obtain driving feature fusion data of multiple fusion modalities. The driving feature fusion data of multiple fusion modalities includes driving feature fusion data of image radar fusion modality, driving feature fusion data of image driving trajectory fusion modality, and driving feature fusion data of radar driving trajectory fusion modality. The step of estimating the target speed based on the fusion data of driving features from the multiple fusion modalities to obtain the target speed estimate includes: The fusion data of the same fusion mode from the fusion data of the multiple fusion modes are input into the same speed prediction model, and the fusion data of different fusion modes are input into different speed prediction models, outputting the initial speed prediction values ​​of the multiple fusion modes; and The initial velocity estimates of the multiple fusion modes are fused according to rules to obtain the target velocity estimate. The rule fusion includes any one of the following fusion methods: maximum value fusion and average value fusion.

2. The estimation method according to claim 1, wherein, The driving data of the multiple modes includes at least two of the following modes: Image data, radar data, and driving trajectory data.

3. The estimation method according to claim 1, wherein, The step of estimating the target speed based on the fusion data of driving features from the multiple fusion modalities to obtain the target speed estimate also includes: The driving feature fusion data of the multiple fusion modes are input into the same speed prediction model to predict the speed, and the target speed prediction value is obtained.

4. The estimation method according to claim 2, wherein, The feature extraction of the driving data of the multiple modalities yields driving feature data of the multiple modalities, including: The driving data is subjected to feature extraction using a feature extraction model to obtain the driving feature data.

5. The estimation method according to claim 4, wherein, The step of using a feature extraction model to extract features from the driving data to obtain the driving feature data includes: The image data is processed using an image feature extraction model to extract features, thereby obtaining the driving feature data.

6. The estimation method according to claim 4, wherein, The step of using a feature extraction model to extract features from the driving data to obtain the driving feature data includes: The radar data or driving trajectory data are used to extract features using a time series model to obtain the driving feature data.

7. A speed prediction device, comprising: The acquisition module is used to acquire driving data in multiple modes; The extraction module is used to extract features from the driving data of the multiple modes to obtain driving feature data of the multiple modes; The fusion module is used to interactively fuse the driving feature data of different modalities to obtain driving feature fusion data of multiple fusion modalities; as well as The prediction module is used to predict the speed based on the driving feature fusion data of the multiple fusion modes, and obtain the target speed prediction value; The fusion module includes: The fusion unit is used to interactively fuse the driving feature data of different modalities using a Siamese network model to obtain driving feature fusion data of multiple fusion modalities, wherein the driving feature fusion data of multiple fusion modalities includes driving feature fusion data of image radar fusion modality, driving feature fusion data of image driving trajectory fusion modality, and driving feature fusion data of radar driving trajectory fusion modality; The prediction module is further used for: The fusion data of the same fusion mode from the fusion data of the multiple fusion modes are input into the same speed prediction model, and the fusion data of different fusion modes are input into different speed prediction models, outputting the initial speed prediction values ​​of the multiple fusion modes; and The target velocity estimate is determined based on the initial velocity estimate of the multiple fusion modes; The prediction module is further used for: The initial velocity estimates of the multiple fusion modes are fused according to rules to obtain the target velocity estimate. The rule fusion includes any one of the following fusion methods: maximum value fusion and average value fusion.

8. The prediction device according to claim 7, wherein, The driving data of the multiple modes includes at least two of the following modes: Image data, radar data, and driving trajectory data.

9. The prediction device according to claim 7, wherein, The prediction module is further used for: The driving feature fusion data of the multiple fusion modes are input into the same speed prediction model to predict the speed, and the target speed prediction value is obtained.

10. The prediction device according to claim 8, wherein, The extraction module is further used for: The driving data is subjected to feature extraction using a feature extraction model to obtain the driving feature data.

11. The prediction device according to claim 10, wherein, The extraction module is further used for: The image data is processed using an image feature extraction model to extract features, thereby obtaining the driving feature data.

12. The prediction device according to claim 10, wherein, The extraction module is further used for: The radar data or driving trajectory data are used to extract features using a time series model to obtain the driving feature data.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

15. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method and system based on cross-modal attention and hierarchical fusion

    CN115063709A

  • Multi-target detection and tracking method fusing millimeter wave radar and depth vision

    CN115308732A