Driving assistance information processing method and device and electronic equipment

The encoder network processed dialogue query data and environmental images, and used the collaborative cooperation of multimodal model and image segmentation model to generate target mask images to provide dialogue response data, solving the problem of poor traffic sign recognition effect in the prior art and improving the reliability and user experience of driving assistance.

CN120220113APending Publication Date: 2025-06-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510323698.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing technology has poor identification of traffic signs and cannot provide users with more reliable driving assistance, affecting driving safety.

Method used

The dialogue query data and the vehicle's environmental image are processed through the encoder network, and the question-asked semantic features, environmental semantic features and environmental image features are output. Combined with the collaborative cooperation of the multimodal model and the image segmentation model, the target mask image is generated to generate dialogue response data.

Benefits of technology

Improves the accuracy of the target mask image, enhances the user's driving experience during driving the vehicle, and provides more reliable driving assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220113A_ABST
    Figure CN120220113A_ABST
Patent Text Reader

Abstract

The invention provides a driving assistance information processing method, a driving assistance information processing device and electronic equipment, which can be applied to the technical field of intelligent driving and data processing. The method comprises the following steps: acquiring dialogue query data and an environment image of a vehicle, wherein the environment image comprises at least one traffic marker; the dialogue query data and the environment image are input into an encoder network, question semantic features, environment semantic features and at least one environment image feature corresponding to the dialogue query data are output, and the encoder network comprises an encoder of a multi-modal model and an image encoder of an image segmentation model; inputting the question semantic features and the environment semantic features into a target language model, and outputting text potential features, wherein the target decoding network comprises a decoder of an image segmentation model; and inputting the text potential feature and the at least one environment image feature into a target decoding network, and outputting a target mask image so as to generate dialogue response data for the dialogue query data according to the target mask image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of intelligent driving and data processing, and particularly to a method for processing driving assistance information, a device for processing driving assistance information, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] During the driving process of a vehicle, the accurate recognition of traffic signs is a key link to ensure the safe driving of the vehicle and compliance with traffic rules. With the continuous development of intelligent in-vehicle devices, the accurate recognition and understanding of traffic signs have become one of the important conditions for realizing the intelligence of the driving system.

[0003] However, the related technology has a poor recognition effect on traffic signs, thus unable to provide relatively reliable help for users, which affects driving safety. Summary of the Invention

[0004] In view of the above problems, the present application provides a method for processing driving assistance information, a device for processing driving assistance information, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to the first aspect of the present application, there is provided a method for processing driving assistance information, including:

[0006] Obtaining dialogue query data and an environmental image of the vehicle, the environmental image including at least one traffic marker;

[0007] Inputting the dialogue query data and the environmental image into an encoder network, and outputting a question semantic feature corresponding to the dialogue query data, an environmental semantic feature corresponding to the environmental image, and at least one environmental image feature, the encoder network including an encoder of a multimodal model and an image encoder of an image segmentation model;

[0008] Inputting the question semantic feature and the environmental semantic feature into a target language model, and outputting a text latent feature, the text latent feature being used to indicate the attention relationship of the target decoding network to different traffic markers, the target decoding network including a decoder of an image segmentation model;

[0009] Inputting the text latent feature and at least one environmental image feature into the target decoding network, and outputting a target mask image, so as to generate dialogue response data for the dialogue query data according to the target mask image.

[0010] The second aspect of the present application provides a device for processing driving assistance information, including:

[0011] An obtaining module, configured to obtain dialogue query data and an environmental image of the vehicle, the environmental image including at least one traffic marker;

[0012] An encoding module, configured to input the dialogue query data and the environmental image into an encoder network, and output a question semantic feature corresponding to the dialogue query data, an environmental semantic feature corresponding to the environmental image, and at least one environmental image feature, where the encoder network includes an encoder of a multimodal model and an image encoder of an image segmentation model;

[0013] An extraction module, configured to input the question semantic feature and the environmental semantic feature into a target language model, and output a text latent feature, where the text latent feature is used to indicate the attention relationship of the target decoding network to different traffic markers, and the target decoding network includes a decoder of an image segmentation model;

[0014] An output module, configured to input the text latent feature and at least one environmental image feature into the target decoding network, and output a target mask image, so as to generate dialogue response data for the dialogue query data according to the target mask image.

[0015] A third aspect of the present application provides an electronic device, including: one or more processors; a memory, configured to store one or more computer programs, where the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0016] A fourth aspect of the present application further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0017] A fifth aspect of the present application further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0018] According to the embodiments of the present application, the dialogue query data and the environmental image of the vehicle are processed by an encoder network to obtain a question semantic feature, an environmental semantic feature, and at least one environmental image feature. The question semantic feature and the environmental semantic feature are input into a target language model to output a text latent feature, and the text latent feature and at least one environmental image feature are input into a target decoding network to output a target mask image. Since the encoder network in this embodiment includes an encoder of a multimodal model and an image encoder of an image segmentation model, the accurate recognition of the dialogue query data and the environmental image is completed through the collaborative cooperation of the two models, thereby improving the accuracy of the target mask image and further improving the driving experience of the user during driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Through the following description of the embodiments of the present application with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present application will become clearer. In the drawings:

[0020] Figure 1 The figure shows an application scenario diagram of a method for processing driving assistance information according to an embodiment of the present application;

[0021] Figure 2 The figure shows a flowchart of a method for processing driving assistance information according to an embodiment of the present application;

[0022] Figure 3 The figure shows a flowchart of a method for processing driving assistance information according to another embodiment of the present application;

[0023] Figure 4 The figure shows a flowchart of a method for processing driving assistance information according to still another embodiment of the present application;

[0024] Figure 5 The figure shows a schematic diagram of an application scenario of a method for processing driving assistance information according to the first embodiment of the present application;

[0025] Figure 6 The figure shows a schematic diagram of an application scenario of a method for processing driving assistance information according to the second embodiment of the present application;

[0026] Figure 7 The figure shows a schematic diagram of an application scenario of a method for processing driving assistance information according to the third embodiment of the present application;

[0027] Figure 8 The figure shows a structural block diagram of a device for processing driving assistance information according to an embodiment of the present application; and

[0028] Figure 9 The figure shows a block diagram of an electronic device suitable for implementing a method for processing driving assistance information according to an embodiment of the present application. Detailed implementation manners

[0029] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present application. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present application. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present application.

[0030] The terms used herein are merely for describing specific embodiments and are not intended to limit the present application. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0032] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning that those of ordinary skill in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0033] With the rapid development of intelligent in-vehicle devices, traffic sign recognition, as an important part of vehicle environment perception, is increasingly becoming the focus of research. Traffic sign recognition requires precise detection, segmentation, and recognition of traffic signs in complex scenarios to ensure that vehicles can accurately understand road rules and guarantee driving safety. However, due to the diversity and complexity of the driving environment, the characteristics of traffic signs such as shape, color, size, position, and lighting conditions vary greatly, posing a severe challenge to assisted driving.

[0034] An image segmentation model can generate high-quality segmentation masks from the input image. The working mechanism of the image segmentation model is that after inputting the image into the model, it uses the powerful visual feature extraction ability of the image segmentation model to generate a segmentation result with a mask. However, this image segmentation model only supports processing images and cannot process text information, lacking semantic understanding ability. Therefore, it is difficult for the image segmentation model to provide effective assistance for assisted driving.

[0035] A multimodal model is a model that integrates visual and language capabilities. Through joint modeling of images and text, it can achieve multimodal tasks such as visual question answering (VQA) and image generation of text descriptions. Although the multimodal model has cross-modal capabilities and powerful semantic reasoning capabilities, the segmentation ability of the multimodal model is weak, and the ability to process details in images is weak. Therefore, it is difficult for the multimodal model to provide relatively reliable assistance for assisted driving.

[0036] In view of this, embodiments of the present application provide a method and apparatus for processing driving assistance information, and an electronic device. The method includes obtaining conversation query data and an environmental image of a vehicle, where the environmental image includes at least one traffic sign; inputting the conversation query data and the environmental image into an encoder network to output a question semantic feature corresponding to the conversation query data, an environmental semantic feature corresponding to the environmental image, and at least one environmental image feature, where the encoder network includes an encoder of a multimodal model and an image encoder of an image segmentation model; inputting the question semantic feature and the environmental semantic feature into a target language model to output a text latent feature, where the text latent feature is used to indicate an attention focus relationship of a target decoding network to different traffic signs, and the target decoding network includes a decoder of an image segmentation model; and inputting the text latent feature and at least one environmental image feature into the target decoding network to output a target mask image, so as to generate conversation response data for the conversation query data according to the target mask image.

[0037] In the technical solution of the present application, the involved user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure, and application, complies with relevant laws, regulations, and standards, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or reject.

[0038] Figure 1 FIG. shows an application scenario diagram of the method for processing driving assistance information according to an embodiment of the present application.

[0039] As Figure 1 shown, the application scenario 100 according to this embodiment may include a vehicle driving on a highway to select a highway exit. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0040] The user may use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103.

[0041] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and intelligent central control computers in vehicles, etc.

[0042] The server 105 may be a server providing various services, such as a background management server (only for example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests, etc.) to the terminal device.

[0043] It should be noted that the method for processing driving assistance information provided in the embodiments of the present application can generally be executed by the server 105. Correspondingly, the device for processing driving assistance information provided in the embodiments of the present application can generally be set in the server 105. The method for processing driving assistance information provided in the embodiments of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the device for processing driving assistance information provided in the embodiments of the present application can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.

[0044] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0045] are merely illustrative. According to actual needs, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 7 scenario described below, the method for processing driving assistance information of the disclosed embodiments will be described in detail through

[0046] Figure 2 FIG. shows a flowchart of the method for processing driving assistance information according to an embodiment of the present application.

[0047] As Figure 2 shown, the method for processing driving assistance information of this embodiment includes operations S210 to S240, and this transaction processing method can be executed by an electronic device.

[0048] In operation S210, dialogue query data and an environmental image of the vehicle are obtained, and the environmental image includes at least one traffic sign.

[0049] According to an embodiment of the present application, the dialogue query data may include data in voice format and / or text format. For example, the dialogue query data in voice format can be obtained through an audio acquisition device such as a microphone on an electronic device such as a mobile phone, a computer, and an in-vehicle computer, and the dialogue query data in text format can be obtained through a text input box on the electronic device.

[0050] According to an embodiment of the present application, the environmental image of the vehicle may refer to an image that can reflect the driving situation of the vehicle captured by an auxiliary device such as a drone, an electronic device such as an in-vehicle computer, and an image acquisition device such as a camera on the vehicle. For example, traffic markers for indicating traffic are set on both sides of the road before the vehicle reaches a fork. The traffic markers may include various types of markers such as markers indicating the destination ahead of the road, left (or right) driving signs, speed limit signs, and lane direction signs.

[0051] In an embodiment of the present application, before obtaining the user's dialogue query data, the consent or authorization of the user may be obtained. For example, before operation S210, a request for obtaining user information may be sent to the user. When the user consents or authorizes to obtain user information, the above operation S210 is executed.

[0052] In an embodiment of the present application, after the user enters the vehicle, the user's dialogue query data and the environmental image of the vehicle may be obtained during the process of driving the vehicle, or the user's dialogue query data and the environmental image of the vehicle may be obtained when the vehicle is in a started state but not moving.

[0053] In operation S220, the dialogue query data and the environmental image are input into the encoder network, and a question semantic feature corresponding to the dialogue query data, an environmental semantic feature corresponding to the environmental image, and at least one environmental image feature are output.

[0054] According to an embodiment of the present application, the encoder network includes an encoder of a multimodal model and an image encoder of an image segmentation model. For example, the encoder network may include a text encoder, an image encoder of a Multimodal Large Language Model (MLLM), and an image encoder of a Segment Anything Model (SAM).

[0055] According to an embodiment of the present application, the MLLM model is a model that combines the advantages of a large language model (LLM) and a large vision model (LVM). The LLM model performs excellently in language understanding and reasoning capabilities, but has limitations in processing visual information; while the LVM model performs outstandingly in visual tasks, but its reasoning ability needs to be improved. The MLLM model can demonstrate stronger capabilities when processing multi-modal tasks by integrating the advantages of these two models. However, multi-modal models such as the MLLM model have weak image segmentation capabilities and fine-grained visual processing capabilities, and it is difficult for multi-modal models to provide assistance in driving a vehicle during the driving process.

[0056] ‌According to an embodiment of the present application, the SAM model is a general image segmentation model, and its core ability is to generate high-quality segmentation masks from the input image. The working mechanism of the SAM model is to input the image into the model and then use the model's powerful visual feature extraction ability to generate segmentation results with masks. However, the SAM model cannot process text descriptions or semantic information related to pictures, and it cannot accurately understand the specific semantics of traffic markers during vehicle driving, thus providing limited assistance in driving assistance.

[0057] It should be noted that in the above embodiment, the MLLM model is used as an example of a multi-modal model, and it can also be other types of multi-modal models that can achieve the same function. Similarly, the SAM model can be replaced by other image segmentation models.

[0058] According to an embodiment of the present application, during the driving process of a vehicle, the obtained dialogue query data and environmental images are input into the encoder network, and a question semantic feature corresponding to the dialogue query data, an environmental semantic feature corresponding to the environmental image, and at least one environmental image feature are output. Since the encoder network includes the encoder of the multi-modal model and the image encoder of the image segmentation model, the question semantic feature can accurately reflect the dialogue query data. Similarly, the environmental semantic feature can accurately reflect the environmental image, and each environmental image feature marks a traffic marker in the environmental image.

[0059] In a specific embodiment, for example, when multiple traffic markers respectively correspond to the upcoming cities A, B, and C corresponding to different roads, the encoder network can output three environmental image features. The first environmental image feature uses a mask to mark the upcoming city A, the second environmental image feature uses a mask to mark the upcoming city B, and the third environmental image feature uses a mask to mark the upcoming city C.

[0060] In operation S230, the question semantic feature and the environmental semantic feature are input into the target language model, and a text latent feature is output. The text latent feature is used to indicate the attention relationship of the target decoding network to different traffic markers. The target decoding network includes the decoder of the image segmentation model.

[0061] According to an embodiment of the present application, the target language model may be a large language model such as an MLLM model.

[0062] According to an embodiment of the present application, the question semantic feature and the environmental semantic feature output by the encoder network are input into the target language model, and the target language model may output a text latent feature indicating the attention relationship of the target decoding network to different traffic markers in the next operation.

[0063] In a specific embodiment, when the dialogue query data representation asks about the direction to the target city A, and multiple traffic markers respectively correspond to the cities A, B, and C ahead corresponding to different roads, the text latent feature output by the target language model may include the attention relationship between the user's target city A and the cities A, B, and C ahead.

[0064] In operation S240, the text latent feature and at least one environmental image feature are input into the target decoding network, and a target mask image is output, so as to generate dialogue response data for the dialogue query data based on the target mask image.

[0065] According to an embodiment of the present application, the target decoding network may be an image decoder of an image segmentation model such as a SAM model.

[0066] According to an embodiment of the present application, the text latent feature output by the target language model and at least one environmental image feature output by the encoder network are input into the target decoding network. The target decoding network generates a target mask image based on the attention relationship to different traffic markers in the text latent feature according to the text latent feature and at least one environmental image feature.

[0067] In a specific embodiment, the target decoding network determines the target mask image from multiple environmental image features based on the attention relationship between the user's target city A and the cities A, B, and C ahead in the text latent feature. A mask is marked on the target traffic marker of the city A ahead in the target mask image, and the mask may be represented by color or a wireframe.

[0068] According to an embodiment of the present application, after obtaining the target mask image, dialogue response data for dialogue query data can be generated based on the target mask image and the question semantic features output by the encoder network. For example, when the dialogue query data represents an inquiry about the direction to the target city A, the dialogue response data can be "Drive along the left road to the target city A".

[0069] According to an embodiment of the present application, the encoder network processes the dialogue query data input by the user during vehicle driving and the environmental image of the vehicle to obtain question semantic features, environmental semantic features, and at least one environmental image feature. The question semantic features and environmental semantic features are input into the target language model to output text latent features, and the text latent features and at least one environmental image feature are input into the target decoding network to output the target mask image. Since the encoder network in this embodiment includes an encoder of a multimodal model and an image encoder of an image segmentation model, the accurate recognition of the dialogue query data and the environmental image is completed through the collaborative cooperation of the two models, thereby improving the accuracy of the target mask image and further improving the driving experience of the user during vehicle driving.

[0070] Figure 3 The flowchart of the processing method of driving assistance information according to another embodiment of the present application is shown.

[0071] As Figure 3 shown, the dialogue query data and the environmental image are input into the encoder network, and the question semantic features corresponding to the dialogue query data, the environmental semantic features corresponding to the environmental image, and at least one environmental image feature are output, including operations S311 to S312.

[0072] In operation S311, the dialogue query data is input into the text encoder for feature convolution processing to output question semantic features.

[0073] In operation S312, the environmental image is input into the image encoding model, and feature convolution and image segmentation processing are performed on the environmental image to output environmental semantic features and at least one environmental image feature.

[0074] According to an embodiment of the present application, the text encoder may include an encoder of a multimodal model, for example, it may be the TXT Encoder of the MLLM model.

[0075] According to an embodiment of the present application, when the dialogue query data is in text format, in operation S311, the text encoder can perform feature convolution on the dialogue query data to split the dialogue query data into individual tokens (basic data processing units), thereby obtaining question semantic features.

[0076] According to an embodiment of the present application, during the process of the text encoder processing the dialogue query data, in operation S312, the image encoding model will simultaneously perform feature convolution and image segmentation on the environmental image, so as to obtain environmental semantic features and at least one environmental image feature. Each environmental image feature is marked with a traffic sign, so as to facilitate the determination of the target mask image related to the dialogue query data.

[0077] According to an embodiment of the present application, refer to Figure 3 , input the environmental image into the image encoding model, perform feature convolution and image segmentation processing on the environmental image, and output environmental semantic features and at least one environmental image feature, including operations S321 to S322.

[0078] In operation S321, input the environmental image into the first image encoder to output environmental semantic features.

[0079] In operation S322, input the environmental image into the second image encoder to perform marking processing on the traffic signs in the environmental image, and output at least one environmental image feature, where the environmental image feature is marked with a traffic sign.

[0080] According to an embodiment of the present application, the first image encoder includes the image encoder of the multimodal model, for example, it can be the Image Encoder of the MLLM model. The second image encoder includes the image encoder of the image segmentation model, for example, the Image Encoder of the SAM model.

[0081] According to an embodiment of the present application, in operation S321, the convolution features obtained by the first image encoder performing feature convolution on the pixel points in the environmental image can be regarded as feature tensors, and continuously perform convolution on the feature tensors so that the tensor obtained by the last convolution is output, that is, the environmental semantic features.

[0082] According to an embodiment of the present application, in operation S322, the second image encoder can mark each traffic sign in the environmental image, and thus the second image encoder can output the environmental image features corresponding to each traffic sign.

[0083] According to an embodiment of the present application, by utilizing the superiority of the image encoder of the multimodal model in language understanding and the superiority of the image encoder of the image segmentation model in visual feature extraction, more accurate environmental semantic features and environmental image features can be output, and thus the accuracy of the target mask image can be improved.

[0084] According to an embodiment of the present application, refer to Figure 3, input the text latent features and at least one environmental image feature into the target decoding network to output the target mask image, including operations S331 to S332.

[0085] In operation S331, perform a splicing process on the text latent features and at least one environmental image feature to obtain the target spliced feature.

[0086] In operation S332, input the target spliced feature into the target decoding network, and perform a screening process on at least one environmental image feature based on the attention relationship between the text latent features and different traffic markers, and output the target mask image related to the dialogue query data.

[0087] According to an embodiment of the present application, the target decoding network may be the decoder of an image segmentation model, for example, the decoder (Decoder) of the SAM model.

[0088] According to an embodiment of the present application, in operation S331, the text latent features and at least one environmental image feature are spliced according to a certain splicing rule to obtain the spliced target spliced feature, where the splicing rule refers to connecting multiple features together in a certain order to form a new feature, for example, the head and tail connection of features.

[0089] According to an embodiment of the present application, in operation S332, the spliced target spliced feature is input into the target decoding network, and the target decoding network performs a screening process on at least one environmental image feature based on the attention relationship between the text latent features and different traffic markers, so that the target decoding network belongs to the target mask image related to the dialogue query data, and the target mask image includes the target traffic marker.

[0090] In a specific embodiment, the target decoding network determines the target mask image from multiple environmental image features based on the attention relationship between the user's target city A and the upcoming cities A, B, and C, and the target traffic marker of the upcoming city A in the target mask image is marked with a mask.

[0091] According to an embodiment of the present application, before the splicing process, the following operations are further included:

[0092] Perform a reshaping process on the feature dimensions of the initial features to obtain the target features, where the initial features include text latent features and / or environmental image features, and the feature dimensions of the target features are the same as those of the target decoding network.

[0093] According to an embodiment of the present application, in the reshaping of feature dimensions, the text latent features are preferentially reshaped. The purpose of the reshaping is to make the dimensions of the reshaped target features the same as those of the target decoding network, thereby facilitating the target decoding network to perform decoding processing on the reshaped target features.

[0094] In a specific embodiment, if the dimension of the text latent features is 9×3, the dimension of the environmental image features is 3×3×3, and the dimension of the target decoding network is 3×3×3, at this time, the text latent features with a dimension of 9×3 can be reshaped into 3×3×3, and thus a feature splicing operation is performed on the reshaped text latent features and the environmental image feature dimensions.

[0095] According to an embodiment of the present application, by reshaping the feature dimensions, the reshaped target features have the same dimensions as the target decoding network, thereby further improving the processing ability of the target decoding network for the input features, and further improving the output accuracy of the target mask image.

[0096] Figure 4 The flowchart of the processing method of driving assistance information according to another embodiment of the present application is shown.

[0097] According to an embodiment of the present application, as Figure 4 shown, the dialogue response data for the dialogue query data is generated according to the target mask image, including operation S411 to operation S412.

[0098] In operation S411, the target mask image and the question semantic features are input into the language encoding network, and the question semantic features are enhanced through feature fusion, and the response latent features are output.

[0099] In operation S412, the response latent features are input into the text decoder, and the relationship between the dialogue query data and the target traffic signs is decoded, and the dialogue response data is output.

[0100] According to an embodiment of the present application, the language encoding network may include the target language model and the image encoder of the multimodal model. Among them, the target language model may be a large language model such as the MLLM model. The text decoder is the decoder of the image segmentation model, such as the TXT Decoder of the MLLM model.

[0101] According to an embodiment of the present application, the language encoding network fuses the features on the target mask image into the question semantic features through feature fusion, thereby obtaining the response latent features, and the response latent features describe the relationship between the dialogue query data and the target traffic signs in the target mask image.

[0102] According to an embodiment of the present application, the response latent features output by the language encoding network are input into the text decoder, and the text encoder can decode the relationship between the dialogue query data and the target traffic sign in the response latent features, so as to output the dialogue response data of the dialogue query data.

[0103] According to an embodiment of the present application, the features on the target mask image are fused into the question semantic features in a feature fusion manner to enhance the question semantic features, thereby improving the accuracy of the dialogue response data obtained by inputting the resulting response latent features into the text decoder, and thus improving the user experience.

[0104] According to an embodiment of the present application, referring to Figure 4 , the target mask image and the question semantic features are input into the language encoding network, and the question semantic features are enhanced by a feature fusion method, and response latent features are output, including operations S421 to S422.

[0105] In operation S421, the target mask image is input into the third image encoder, and feature convolution processing is performed on the target traffic sign to output image mask features.

[0106] In operation S422, the question semantic features and the image mask features are fused and input into the target language model, and feature association is performed on the dialogue query data and the target traffic sign to output response latent features.

[0107] According to an embodiment of the present application, the third image encoder is an image encoder of a multimodal model, for example, it can be the Image Encoder of the MLLM model.

[0108] According to an embodiment of the present application, before fusing the target mask image into the question semantic features, the target mask image needs to be converted into data that can be processed by the target language model. Thus, the third image encoder performs feature convolution on each pixel point in the target mask image, and thus image mask features related to the target traffic sign can be obtained.

[0109] According to an embodiment of the present application, the image mask features output by the third image encoder are fused into the question semantic features for feature enhancement, and then the fused features are input into the target language model. The target language model performs feature association based on the dialogue query data and the target traffic sign to generate response latent features. Thus, based on the response latent features, dialogue response data for answering the dialogue query data can be generated.

[0110] Figure 5 The schematic diagram of the application scenario of the processing method of driving assistance information according to the first embodiment of the present application is shown. Figure 6The figure shows a schematic diagram of an application scenario of a method for processing driving assistance information according to the second embodiment of the present application. Figure 7 The figure shows a schematic diagram of an application scenario of a method for processing driving assistance information according to the third embodiment of the present application.

[0111] According to an embodiment of the present application, the dialogue query data includes data in voice format and / or text format. When the dialogue query data represents an inquiry about the vehicle driving direction related to the destination, the dialogue response data represents the recommended driving direction to the destination. When the dialogue query data represents an inquiry about the road section parameters related to the driving section, the dialogue response data represents the parameter values of the driving section, where the road section parameters include at least one of a speed limit parameter, a no-entry parameter, and a traffic flow density parameter.

[0112] According to an embodiment of the present application, the dialogue response data can be data in voice format and / or text format, and can be displayed through a speaker and / or a display screen on an electronic device. Preferably, the vehicle's speaker and in-vehicle computer are used to display the dialogue response data.

[0113] In the first specific embodiment, refer to Figure 5 , the user inputs dialogue query data in voice format through a microphone, such as "I need to turn left at the intersection ahead. Which lane should I take?", and the dialogue query data can be format-converted to obtain dialogue query data in text format. Through the method for processing driving assistance information according to the embodiment of the present application, the target mask image and dialogue response data of the dialogue query data can be obtained, such as "Please take the leftmost lane".

[0114] In the second specific embodiment, refer to Figure 6 , the user inputs dialogue query data in voice format through a microphone, such as "What is the speed limit on the current road?", and the dialogue query data can be format-converted to obtain dialogue query data in text format. Through the method for processing driving assistance information according to the embodiment of the present application, the target mask image and dialogue response data of the dialogue query data can be obtained, such as "The speed limit on the current lane is 40 KM / h".

[0115] In the third specific embodiment, refer to Figure 7 , the user inputs dialogue query data in voice format through a microphone, such as "The destination is City A. Which exit should I take to get off the highway?", and the dialogue query data can be format-converted to obtain dialogue query data in text format. Through the method for processing driving assistance information according to the embodiment of the present application, the dialogue response data of the dialogue query data can be obtained, such as "Destination City A. Get off the highway at Exit 23B".

[0116] Based on the above method for processing driving assistance information, the present application further provides a device for processing driving assistance information. The following will be combined with Figure 8 to describe this device in detail.

[0117] Figure 8 Fig. 6 shows a structural block diagram of a device for processing driving assistance information according to an embodiment of the present application.

[0118] As Figure 8 shown, the device 800 for processing driving assistance information in this embodiment includes an acquisition module 810, an encoding module 820, an extraction module 830, and an output module 840.

[0119] The acquisition module 810 is configured to acquire dialogue query data and an environmental image of the vehicle, where the environmental image includes at least one traffic sign.

[0120] The encoding module 820 is configured to input the dialogue query data and the environmental image into an encoder network, and output a question semantic feature corresponding to the dialogue query data, an environmental semantic feature corresponding to the environmental image, and at least one environmental image feature. The encoder network includes an encoder of a multimodal model and an image encoder of an image segmentation model.

[0121] The extraction module 830 is configured to input the question semantic feature and the environmental semantic feature into a target language model, and output a text latent feature, where the text latent feature is used to indicate the attention focus relationship of the target decoding network to different traffic signs. The target decoding network includes a decoder of an image segmentation model.

[0122] The output module 840 is configured to input the text latent feature and at least one environmental image feature into the target decoding network, and output a target mask image, so as to generate dialogue response data for the dialogue query data according to the target mask image.

[0123] According to an embodiment of the present application, the dialogue query data and the environmental image of the vehicle are processed by the encoder network to obtain a question semantic feature, an environmental semantic feature, and at least one environmental image feature. The question semantic feature and the environmental semantic feature are input into the target language model to output a text latent feature. The text latent feature and at least one environmental image feature are input into the target decoding network to output a target mask image. Since the encoder network in this embodiment includes an encoder of a multimodal model and an image encoder of an image segmentation model, the accurate recognition of the dialogue query data and the environmental image is completed through the collaborative cooperation of the two models, thereby improving the accuracy of the target mask image and further improving the driving experience of the user during driving the vehicle.

[0124] According to an embodiment of the present application, the encoding module 820 includes a text encoding unit and an image processing unit.

[0125] The text encoding unit is used to input the dialogue query data into the text encoder for feature convolution processing, and output the question semantic features. The text encoder includes the encoder of the multimodal model.

[0126] The image processing unit is used to input the environmental image into the image encoding model, perform feature convolution and image segmentation processing on the environmental image, and output the environmental semantic features and at least one environmental image feature.

[0127] According to an embodiment of the present application, the image processing unit includes an image encoding subunit and an image marking subunit.

[0128] The image encoding subunit is used to input the environmental image into the first image encoder and output the environmental semantic features. Among them, the first image encoder includes the image encoder of the multimodal model.

[0129] The image marking subunit is used to input the environmental image into the second image encoder, perform marking processing on the traffic signs in the environmental image, and output at least one environmental image feature. Among them, the environmental image feature is marked with a traffic sign, and the second image encoder includes the image encoder of the image segmentation model.

[0130] According to an embodiment of the present application, the output module 840 includes a feature splicing unit and a feature decoding unit.

[0131] The feature splicing unit is used to perform splicing processing on the text latent features and at least one environmental image feature to obtain the target splicing features.

[0132] The feature decoding unit is used to input the target splicing features into the target decoding network, perform screening processing on at least one environmental image feature based on the attention relationship between the text latent features and different traffic signs, and output the target mask image related to the dialogue query data. The target mask image includes the target traffic signs related to the dialogue query data among at least one traffic sign.

[0133] According to an embodiment of the present application, the output module 840 further includes a feature reshaping unit.

[0134] The feature reshaping unit is used to perform reshaping processing on the feature dimensions of the initial features to obtain the target features. Among them, the initial features include text latent features and / or environmental image features, and the feature dimensions of the target features are the same as those of the target decoding network.

[0135] According to an embodiment of the present application, the output module 840 includes a feature fusion unit and a text decoding unit.

[0136] The feature fusion unit is used to input the target mask image and the question semantic features into the language encoding network, enhance the question semantic features through feature fusion, and output response latent features, where the response latent features are used to describe the relationship between the dialogue query data and the target traffic signs in the target mask image.

[0137] The text decoding unit is used to input the response latent features into the text decoder, decode the relationship between the dialogue query data and the target traffic signs, and output the dialogue response data, where the text decoder is the decoder of the image segmentation model.

[0138] According to an embodiment of the present application, the feature fusion unit includes a feature convolution unit and a feature association unit.

[0139] The feature convolution unit is used to input the target mask image into the third image encoder, perform feature convolution processing on the target traffic signs, and output image mask features, where the third image encoder is the image encoder of the multimodal model.

[0140] The feature association unit is used to fuse the question semantic features and the image mask features and input them into the target language model, perform feature association on the dialogue query data and the target traffic signs, and output response latent features.

[0141] According to an embodiment of the present application, the dialogue query data includes data in voice format and / or text format.

[0142] In the case where the dialogue query data represents an inquiry about the driving direction of a vehicle related to the destination, the dialogue response data represents the recommended driving direction to the destination. In the case where the dialogue query data represents an inquiry about the road segment parameters related to the driving segment, the dialogue response data represents the parameter values of the driving segment, where the road segment parameters include at least one of a speed limit parameter, a no-entry parameter, and a traffic flow density parameter.

[0143] According to an embodiment of the present application, any combination of the acquisition module 810, the encoding module 820, the extraction module 830, and the output module 840 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present application, at least one of the acquisition module 810, the encoding module 820, the extraction module 830, and the output module 840 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the acquisition module 810, the encoding module 820, the extraction module 830, and the output module 840 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0144] Figure 9 A block diagram of an electronic device suitable for implementing a method for processing driving assistance information according to an embodiment of the present application is shown.

[0145] As Figure 9 shown, the electronic device 900 according to an embodiment of the present application includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 901 may also include on-board memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present application.

[0146] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs can also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in one or more memories.

[0147] According to an embodiment of the present application, the electronic device 900 may further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the input / output (I / O) interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage portion 908 as needed.

[0148] The present application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present application is implemented.

[0149] According to an embodiment of the present application, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, the computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than the ROM 902 and RAM 903.

[0150] An embodiment of the present application further includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present application.

[0151] When the computer program is executed by the processor 901, it executes the above functions defined in the system / apparatus of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0152] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the removable medium 911. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0153] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or be installed from the removable medium 911. When the computer program is executed by the processor 901, it executes the above functions defined in the system of the embodiment of the present application. According to an embodiment of the present application, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0154] In accordance with embodiments of the present application, program code for executing the computer programs provided by the embodiments of the present application may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0156] Those skilled in the art can understand that the features described in the various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application.

[0157] The above describes the embodiments of the present application. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present application.

Claims

1. A method for processing driving assistance information, characterized in that: The method comprises: Acquire dialogue query data and an environmental image of the vehicle, wherein the environmental image includes at least one traffic sign; Inputting the dialogue query data and the environment image into an encoder network, outputting a question semantic feature corresponding to the dialogue query data, an environment semantic feature corresponding to the environment image, and at least one environment image feature, the encoder network comprising an encoder of a multimodal model and an image encoder of an image segmentation model; Inputting the question semantic features and the environment semantic features into a target language model, and outputting textual latent features, wherein the textual latent features are used to indicate the attention relationship of a target decoding network to different traffic signs, wherein the target decoding network includes a decoder of the image segmentation model; The text latent features and at least one of the environmental image features are input into the target decoding network, and a target mask image is output, so as to generate dialogue response data for the dialogue query data according to the target mask image.

2. The method according to claim 1, characterized in that Inputting the dialogue query data and the environment image into an encoder network, outputting a question semantic feature corresponding to the dialogue query data, an environment semantic feature corresponding to the environment image, and at least one environment image feature, including: Inputting the dialogue query data into a text encoder for feature convolution processing and outputting the question semantic features, wherein the text encoder includes an encoder of the multimodal model; The environment image is input into an image coding model, feature convolution and image segmentation processing are performed on the environment image, and the environment semantic feature and at least one environment image feature are output.

3. The method according to claim 2, characterized in that Inputting the environment image into an image coding model, performing feature convolution and image segmentation processing on the environment image, and outputting the environment semantic feature and at least one environment image feature, including: Inputting the environment image into a first image encoder and outputting the environment semantic features, wherein the first image encoder includes an image encoder of the multimodal model; The environment image is input to a second image encoder, the traffic signs in the environment image are marked, and at least one feature of the environment image is output, wherein the environment image feature is marked with the traffic sign, and the second image encoder includes an image encoder of the image segmentation model.

4. The method according to claim 1, characterized in that: Inputting the text potential feature and at least one of the environment image features into the target decoding network, and outputting a target mask image, comprising: Performing splicing processing on the text potential feature and at least one of the environmental image features to obtain a target splicing feature; The target splicing feature is input into the target decoding network, at least one environmental image feature is screened based on the attention relationship between the text potential feature and different traffic signs, and the target mask image is output, wherein the target mask image includes at least one target traffic sign related to the dialogue query data among the traffic signs.

5. The method according to claim 4, characterized in that Before the splicing process, it also includes: The initial features are reshaped in feature dimension to obtain target features, wherein the initial features include the text latent features and / or the environmental image features, and wherein the feature dimension of the target features is the same as the dimension of the target decoding network.

6. The method according to claim 1, characterized in that Generating dialogue response data for the dialogue query data according to the target mask image includes: Inputting the target mask image and the question semantic features into a language encoding network, enhancing the question semantic features by feature fusion, and outputting a response latent feature, wherein the response latent feature is used to describe the relationship between the dialogue query data and the target traffic sign in the target mask image; The response potential features are input into a text decoder, the relationship between the dialogue query data and the target traffic sign is decoded, and the dialogue response data is output, wherein the text decoder includes a decoder of the image segmentation model.

7. The method according to claim 6, characterized in that The target mask image and the question semantic features are input into the language encoding network, the question semantic features are enhanced by feature fusion, and the answer potential features are output, including: Inputting the target mask image into a third image encoder, performing feature convolution processing on the target traffic sign, and outputting image mask features, wherein the third image encoder includes an image encoder of the multimodal model; The question semantic features and the image mask features are fused and input into the target language model, the dialogue query data and the target traffic sign are feature associated, and the response potential features are output.

8. The method according to claim 7, characterized in that The dialogue query data includes data in voice format and / or text format; In the case where the dialogue query data represents an inquiry about vehicle driving directions associated with a destination, the dialogue response data represents a suggested driving direction to the destination; In the case where the dialogue query data represents an inquiry into a section parameter related to a driving section, the dialogue response data represents a parameter value of the driving section, wherein the section parameter includes at least one of a speed limit parameter, a prohibited driving parameter, and a traffic density parameter.

9. A driving assistance information processing device, characterized in that: include: An acquisition module, used to acquire dialogue query data and an environmental image of the vehicle, wherein the environmental image includes at least one traffic sign; an encoding module, configured to input the dialogue query data and the environment image into an encoder network, and output a question semantic feature corresponding to the dialogue query data, an environment semantic feature corresponding to the environment image, and at least one environment image feature, wherein the encoder network includes an encoder of a multimodal model and an image encoder of an image segmentation model; An extraction module, used for inputting the question semantic features and the environment semantic features into a target language model, and outputting textual latent features, wherein the textual latent features are used for indicating the attention relationship of a target decoding network to different traffic signs, wherein the target decoding network includes a decoder of the image segmentation model; An output module is used to input the text potential features and at least one of the environmental image features into the target decoding network, and output a target mask image so as to generate dialogue response data for the dialogue query data according to the target mask image.

10. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.