Multi-mode vehicle machine voice assistant interaction method, device and equipment

By using a multimodal in-vehicle voice assistant interaction method that combines user voice, geographic information, and image information, the problem of limited interaction capabilities of existing in-vehicle voice assistants is solved, achieving more efficient vehicle control and interaction.

CN120892008APending Publication Date: 2025-11-04CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510975532.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing in-vehicle voice assistants can only recognize text information in the user's voice and cannot combine it with geographical and image information, resulting in limited interactive capabilities.

Method used

By using a multimodal vehicle-mounted voice assistant interaction method, user voice commands are obtained, and when they are determined to be related to geographical location, point of interest tags are added. Combined with the vehicle's real-time location and street view images, a multimodal geolinguistic model is used for information recommendation and vehicle control.

Benefits of technology

It enables multi-dimensional interaction of user voice, geographic information, and image information, improving the efficiency of vehicle-machine interaction and vehicle control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892008A_ABST
    Figure CN120892008A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode in-vehicle voice assistant interaction method, device and equipment, belongs to the technical field of intelligent in-vehicle inforams.When it is judged that a user voice instruction is related to a geographic position, an interest point label can be added to the user voice instruction, and therefore the interest point label related to the geographic position in the user voice instruction is extracted; subsequently, recommendation information can be pushed to the user according to the interest point label, and the vehicle is further controlled according to the recommendation information, so that text information in the voice of the user is combined with information of other dimensions, such as geographic information and image information, and the vehicle machine can interact with the user and control the vehicle in combination with multi-dimensional information. And the efficiency of vehicle-machine interaction and vehicle control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent car machine, and particularly relates to a multi-modal car machine voice assistant interaction method, device and equipment. BACKGROUND

[0002] With the development of the automobile manufacturing industry and the improvement of people's living standards, the popularity rate of private cars has been increasingly high, and consumers' demand for vehicles has not been limited to the function of transportation. Consumers have begun to focus on the intelligent level of the car when choosing a vehicle.

[0003] Thanks to the progress of chip manufacturing technology and the use of AI large model technology, the intelligent car machine voice assistant has been equipped on more and more vehicle models. The car machine voice assistant can realize some simple functions such as adjusting the air conditioner and opening the window according to the user's voice instruction.

[0004] However, the recognition model used by the existing car machine voice assistant is generally a single-modal model, which can only recognize the text information in the user's voice and cannot combine with other dimensional information such as geographic information and image information to interact with the user. SUMMARY

[0005] Therefore, the present application provides a multi-modal car machine voice assistant interaction method, which enables the car machine to interact with the user and control the vehicle by combining multi-dimensional information. The method comprises:

[0006] In one aspect, the present application provides a multi-modal car machine voice assistant interaction method, which comprises:

[0007] Obtaining a user voice instruction.

[0008] Determining whether the user voice instruction is related to a geographic location.

[0009] When it is determined that the user voice instruction is related to the geographic location, adding a point of interest label to the user voice instruction.

[0010] Obtaining recommended information according to the user voice instruction with the added point of interest label.

[0011] Controlling the vehicle according to the recommended information.

[0012] Optionally, when it is determined that the user voice instruction is related to the geographic location, adding a point of interest label to the user voice instruction comprises:

[0013] When it is determined that the user voice instruction is related to the geographic location, calling a geographic enhanced voice recognition model to add a point of interest label to the user voice instruction, the method further comprises:

[0014] When it is judged that the user voice instruction is not related to the geographical location, a standard voice recognition model is called.

[0015] Alternatively, the method further comprises:

[0016] The real-time position of the vehicle and the real-time street view image outside the vehicle are obtained.

[0017] According to the user voice instruction with the added point of interest label, the recommended information is obtained.

[0018] The real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the added point of interest label are input into a pre-stored multi-modal geographical language model to obtain output recommended information.

[0019] Alternatively, before the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the added point of interest label are input into the pre-stored multi-modal geographical language model, the method further comprises:

[0020] The features contained in the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the added point of interest label are mapped into feature vectors of the same dimension.

[0021] Alternatively, the pre-stored multi-modal geographical language model includes a geographical context, which is used to represent structured spatial information, and the structured spatial information is used to represent the association between the geographical location corresponding to the point of interest and the surrounding geographical objects.

[0022] Alternatively, controlling the vehicle according to the recommended information comprises:

[0023] The current driving scene is obtained.

[0024] The navigation route is determined according to the current driving scene and the recommended information.

[0025] The vehicle is controlled to travel along the navigation route.

[0026] In another aspect, the present application also provides a multi-modal car machine voice assistant interaction device, which comprises:

[0027] The voice module is configured to obtain a user voice instruction.

[0028] The judgment module is configured to judge whether the user voice instruction is related to the geographical location.

[0029] The adding module is configured to add a point of interest label to the user voice instruction when it is judged that the user voice instruction is related to the geographical location.

[0030] The recommendation module is configured to obtain recommended information according to the user voice instruction with the added point of interest label.

[0031] The control module is configured to control the vehicle according to the recommendation information.

[0032] Optionally, the adding module is configured to:

[0033] When it is judged that the user voice instruction is related to the geographic location, a geographic enhanced speech recognition model is called to add a point of interest label to the user voice instruction.

[0034] When it is judged that the user voice instruction is not related to the geographic location, a standard speech recognition model is called.

[0035] The device further comprises an image acquisition module configured to:

[0036] Acquire the real-time position of the vehicle and the real-time street view image outside the vehicle.

[0037] The recommendation module is configured to:

[0038] Input the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the point of interest label into a pre-stored multi-modal geographic language model to obtain output recommendation information.

[0039] Optionally, the device further comprises a unification module configured to:

[0040] Before bringing the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the point of interest label into the pre-stored multi-modal geographic language model, map the features contained in the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the point of interest label into feature vectors of the same dimension.

[0041] The pre-stored multi-modal geographic language model includes a geographic context, the geographic context is used to represent structured spatial information, and the structured spatial information is used to represent the association relationship between the geographic location corresponding to the point of interest and the surrounding geographic objects, and the control module is configured to:

[0042] Acquire the current driving scene.

[0043] Determine a navigation route according to the current driving scene and the recommendation information.

[0044] Control the vehicle to travel along the navigation route.

[0045] In another aspect, the present application also provides a multi-modal vehicle machine voice assistant interaction device, comprising a processor, a memory, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the following steps are implemented:

[0046] Acquire a user voice instruction.

[0047] determine whether the user voice instruction is related to a geographic location.

[0048] add a point of interest label to the user voice instruction when it is determined that the user voice instruction is related to a geographic location.

[0049] obtain recommendation information according to the user voice instruction to which the point of interest label is added.

[0050] control the vehicle according to the recommendation information.

[0051] The multi-modal car machine voice assistant interaction method provided in the application can add a point of interest label to a user voice instruction when it is determined that the user voice instruction is related to a geographic location, thereby extracting a point of interest label related to a geographic location in the user voice instruction, and subsequently pushing recommendation information to the user according to the point of interest label and further controlling the vehicle according to the recommendation information, so that the text information in the user voice and other dimensional information such as geographic information and image information are combined, and the car machine can combine multi-dimensional information to interact with the user and control the vehicle, thereby improving the efficiency of car machine interaction and vehicle control. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0053] Figure 1 The flowchart of the multi-modal car machine voice assistant interaction method provided by the embodiment of the application;

[0054] Figure 2 Another flowchart of the multi-modal car machine voice assistant interaction method provided by the embodiment of the application;

[0055] Figure 3 The architecture schematic diagram of the multi-modal car machine voice assistant interaction method provided by the embodiment of the application;

[0056] Figure 4 The flowchart of the POI dynamic loading technology in the multi-modal car machine voice assistant interaction method provided by the embodiment of the application;

[0057] Figure 5 The flowchart of obtaining a key frame in the multi-modal car machine voice assistant interaction method provided by the embodiment of the application;

[0058] Figure 6 The structural diagram of the multi-modal car machine voice assistant interaction device provided by the embodiment of the application;

[0059] Figure 7 A structural diagram of a multi-modal car-machine voice assistant interaction device is provided in the embodiments of the present application. DETAILED DESCRIPTION

[0060] The technical solutions in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0061] The embodiments of the present application provide a multi-modal car-machine voice assistant interaction method, as shown in the method comprises steps S101, S102, S103, S104 and S105, wherein: Figure 1

[0062] In step S101, a user voice instruction is acquired.

[0063] In step S102, it is judged whether the user voice instruction is related to a geographic location.

[0064] In step S103, when it is judged that the user voice instruction is related to a geographic location, a point of interest label is added to the user voice instruction.

[0065] In step S104, recommended information is acquired according to the user voice instruction to which the point of interest label is added.

[0066] In step S105, the vehicle is controlled according to the recommended information.

[0067] In some optional embodiments, when it is judged that the user voice instruction is related to a geographic location, adding a point of interest label to the user voice instruction comprises:

[0068] When it is judged that the user voice instruction is related to a geographic location, a geographic enhanced voice recognition model is called to add a point of interest label to the user voice instruction, and the method further comprises:

[0069] When it is judged that the user voice instruction is not related to a geographic location, a standard voice recognition model is called.

[0070] In some optional embodiments, the method further comprises:

[0071] The real-time position of the vehicle and the real-time street view image outside the vehicle are acquired.

[0072] Acquiring recommended information according to the user voice instruction to which the point of interest label is added comprises:

[0073] ​The vehicle real-time position, the vehicle-outside real-time street view image and the user voice instruction with the added point-of-interest label are input into a pre-stored multi-modal geographic language model to obtain output recommended information.

[0074] In some optional embodiments, before the vehicle real-time position, the vehicle-outside real-time street view image and the user voice instruction with the added point-of-interest label are brought into the pre-stored multi-modal geographic language model, the method further includes:

[0075] Features contained in the vehicle real-time position, the vehicle-outside real-time street view image and the user voice instruction with the added point-of-interest label are mapped into feature vectors of the same dimension.

[0076] In some optional embodiments, the pre-stored multi-modal geographic language model includes a geographic context, and the geographic context is used to represent structured spatial information used to represent the association relationship between the geographic location corresponding to the point-of-interest and the surrounding geographic objects.

[0077] In some optional embodiments, the control of the vehicle according to the recommended information includes:

[0078] A current driving scene is obtained.

[0079] A navigation route is determined according to the current driving scene and the recommended information.

[0080] The vehicle is controlled to travel along the navigation route.

[0081] The multi-modal car-machine voice assistant interaction method provided in the present application can add a point-of-interest label to the user voice instruction when it is determined that the user voice instruction is related to a geographic location, so as to extract the point-of-interest label related to the geographic location in the user voice instruction. Subsequently, recommended information can be pushed to the user according to the point-of-interest label, and the vehicle can be further controlled according to the recommended information, so that the text information in the user voice and other dimensional information such as geographic information and image information are combined, and the car-machine can combine multi-dimensional information to interact with the user and control the vehicle, thereby improving the efficiency of car-machine interaction and vehicle control.

[0082] The embodiments of the present application provide a multi-modal car-machine voice assistant interaction method, as shown in Figure 2 The method includes steps S201, S202, S203, S204, S205, S206 and S207, wherein:

[0083] In step S201, a user voice instruction is obtained.

[0084] It can be understood that the user voice instruction can be collected by a car-mounted microphone array, and the car-mounted microphone array receives the voice (such as the user saying "navigate to West Lake") and is responsible for "listening".

[0085] In step S202, it is judged whether the user voice instruction is related to a geographic location.

[0086] It can be understood that it is first judged whether the user instruction is of a geographic location related type.

[0087] In step S203, when it is judged that the user voice instruction is related to a geographic location, a geographic enhanced voice recognition model is called to add a point of interest label to the user voice instruction.

[0088] It can be understood that when it is judged that the user voice instruction is related to a geographic location (such as "is there a charging station nearby"), a "geographic enhanced Geo-ASR" branch is called.

[0089] In some optional embodiments, the method further comprises:

[0090] When it is judged that the user voice instruction is not related to a geographic location, a standard voice recognition model is called. It can be understood that a regular instruction (such as "play music") → enters a "standard ASR" branch (simple voice recognition).

[0091] In step S204, recommended information is obtained according to the user voice instruction to which the point of interest label is added.

[0092] It can be understood that "geographic enhanced Geo-ASR" is specifically designed to handle geographic related voice and is responsible for labeling POI for user voice instructions ("Qianjiang New City" and "Starbucks" are geographic entities). For example, the user says "Starbucks in Qianjiang New City".

[0093] In step S205, a current driving scene is obtained.

[0094] In step S206, a navigation route is determined according to the current driving scene and the recommended information.

[0095] It can be understood that by recognizing the current driving scene (such as "highway", "urban congestion", and "parking"), the response strategy is adjusted, such as prioritizing "navigation lane change" instructions when on a highway and prioritizing "find nearby parking lot" when in congestion.

[0096] In step S207, the vehicle is controlled to travel along the navigation route.

[0097] It can be understood that the vehicle can also handle multiple tasks at the same time while being controlled to travel along the navigation route, such as:

[0098] Navigation instructions → call high-precision map rendering.

[0099] Vehicle control instructions → call CAN instruction encoding.

[0100] Information query -> invoke voice synthesis TTS.

[0101] Entertainment control -> invoke entertainment system API.

[0102] In some optional embodiments, the method further comprises:

[0103] Obtaining the real-time position of the vehicle and the real-time street view image outside the vehicle.

[0104] It can be understood that high-precision positioning is performed using RTK / INS positioning (more accurate than ordinary GPS), and the real-time position of the vehicle is output, which is responsible for "positioning". The scene outside the vehicle (such as an intersection or a building) is captured by a vehicle-mounted camera, which is responsible for "seeing".

[0105] According to the user voice instruction with the added point of interest label, the recommended information comprises:

[0106] The real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the added point of interest label are input into a pre-stored multi-modal geographic language model to obtain output recommended information.

[0107] It can be understood that the multi-modal geographic language model MGeo can fuse "text, geography, and vision" three types of information for accurate positioning, and the multi-modal geographic language model specifically comprises:

[0108] A text encoder (BERT-style) for understanding text semantics (such as "charging station").

[0109] A geographic encoder (H3+GNN) for processing position data (dividing areas using H3 grid and understanding road relationships using GNN).

[0110] A visual encoder (GAEA model) for analyzing pictures taken by a camera (identifying "intersection traffic light" and "building style").

[0111] In some optional embodiments, before the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the added point of interest label are input into the pre-stored multi-modal geographic language model, the method further comprises:

[0112] Mapping the features contained in the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction with the added point of interest label into feature vectors of the same dimension.

[0113] It can be understood that the visual projector can map the visual features extracted from the real-time street view image outside the vehicle to the space of the text vector in the user voice instruction with the added point of interest label, so that the image and text features can "dialogue". Finally, the visual feature sequence + text feature sequence is obtained, which is fed into the large model.

[0114] In some optional embodiments, the pre-stored multi-modal geographic language model includes a geographic context, and the geographic context is used to represent structured spatial information used to represent the association relationship between the geographic location corresponding to the point of interest and the surrounding geographic objects.

[0115] It can be understood that the geographic context (GC) is defined as the association relationship between the geographic location of the query or the POI and the surrounding geographic objects (such as roads, areas), including spatial relationships (proximity, coverage) and relative positions, which explicitly shows the key role of GC in query-POI matching and solves the limitation of traditional single-modal models that cannot utilize geographic spatial information. The geographic context GC plays a key role in query updating the point of interest (POI) matching, solves the limitation of traditional single-modal models that cannot utilize geographic spatial information, enables the dialogue response involving geographic information location in user interaction, improves the user's satisfaction with the intelligent voice assistant, the intelligent cockpit and the whole vehicle, enhances the user's stickiness, and gradually establishes the brand reputation.

[0116] As shown in FIG. 1, the multi-modal vehicle machine voice assistant interaction method provided by the present application includes the following three-layer architecture: Figure 3

[0117] I. Layer 1 - Multi-modal input layer:

[0118] The role of the multi-modal input layer is to collect three types of data of "sound, position, image" to provide raw materials for subsequent processing.

[0119] The vehicle-mounted microphone array 301 is responsible for receiving voice (such as the user saying "navigate to West Lake") and is responsible for "listening".

[0120] The RTK / INS positioning 302 is responsible for high-precision positioning (better than ordinary GPS), outputs the real-time position of the vehicle, and is responsible for "positioning".

[0121] The vehicle-mounted camera 303 is responsible for shooting the scene outside the vehicle (such as intersections, buildings), and is responsible for "looking".

[0122] After collecting three types of data of "sound, position, image", in step 304, the data is multi-modal aligned, and then input into the consciousness recognition shunt 305 in Layer 2.​

[0123] II. Layer2-Core Processing Engine:

[0124] The role of Core Processing Engine is to understand user intent, and convert "raw data" into "executable instructions", including:

[0125] 1. Intent Recognition Shunt 305:

[0126] First, determine whether the user instruction is related to the geographical location type:

[0127] Geographical related Query (such as "Is there a charging station nearby?") → Enter the "Geo-enhanced Geo-ASR" 306 branch.

[0128] Regular instruction (such as "Play music") → Enter "Standard ASR" 307 branch (simple speech recognition).

[0129] 2. Geo-enhanced Geo-ASR (306):

[0130] Special processing of "geographically related speech", tagging POI for user Query (such as the user saying "Starbucks in Qianjiang New City", "Qianjiang New City" and "Starbucks" are geographical entities), then input the POI-labeled text into the multi-modal geographic language model MGeo (308).

[0131] 3. Multi-modal geographic language model MGeo (308):

[0132] The multi-modal geographic language model MGeo (308) is used to integrate "text, geography, vision" three types of information, for accurate positioning, and the multi-modal geographic language model MGeo (308) includes:

[0133] Text encoder (BERT-style) 309, used to understand the semantic meaning of text (such as "charging station").

[0134] Geographic encoder (H3+GNN) 310, used to process location data (H3 grid is used to divide the area, and GNN is used to understand the road relationship).

[0135] Visual encoder (GAEA model) 311, used to analyze camera pictures (identify "intersection traffic lights" and "building style").

[0136] RTK / INS positioning 302 can also send real-time location to geographic encoder 310, and vehicle-mounted camera 303 can also send key frames to visual encoder 311.

[0137] The outputs of the text encoder 309, the geocoding encoder 310, and the visual encoder 311 are fed into the vehicle-mounted adaptation layer of Layer 3 after cross-modal fusion at step 312:

[0138] III. Layer 3 - Vehicle-mounted Adaptation Layer:

[0139] The role of the vehicle-mounted adaptation layer is to convert the "understood intent" into "vehicle action", including:

[0140] 1. Dynamic Context Manager 313:

[0141] Identify the current driving scenario (such as "high speed", "urban congestion", "parking"), that is, identify the scenario state, so as to adjust the response strategy, such as:

[0142] When driving at high speed, prioritize processing "navigation lane change" instructions; when congested, prioritize responding to "find nearby parking lot".

[0143] 2. Multi-instruction parallel engine 314 (based on DAG scheduling), configured to handle multiple tasks simultaneously, such as:

[0144] Navigation instructions 315→ invoke high-precision map rendering 316.

[0145] Vehicle control instructions 317→ invoke CAN instruction encoding 318.

[0146] Information query 319→ invoke speech synthesis TTS (320).

[0147] Entertainment control 321→ invoke entertainment system API (322).

[0148] The multi-modal vehicle machine voice assistant interaction method provided in the present application focuses on improving the geographical spatial information understanding ability of the existing vehicle machine voice assistant. The following is an expansion of Layer 2.

[0149] POI dynamic loading technology is a key optimization technology in geographical spatial information processing, designed specifically for the real-time needs of vehicle-mounted voice assistants. Its core lies in dynamically loading / updating point of interest (POI) data according to vehicle location and user intent, to balance computational load and response speed. As shown in Figure 4 the following are the implementation steps:

[0150] Step S1: Real-time position sensing and grid division:

[0151] Input: centimeter-level vehicle coordinates (longitude, latitude, heading) provided by the RTK / INS combination system.

[0152] Implementation method: Use the h3 geographical grid system to divide the current location.

[0153] Output: Active grid ID set (for subsequent POI query).

[0154] Step S2: Semantic intent parsing and POI pre-filtering:

[0155] Input: Speech recognition result (geo-labeled text output by Geo-ASR).

[0156] Implementation method: Use NER model to extract POI category keywords in the query, such as "gas station" → amenity = fuel, realize the pointing of first-level classification label in open geographic data standard (such as OpenStreetMap, OSM) to gas station type, if the category is not explicitly specified, predict according to historical data (such as 80% of user's "nearby" query pointing to catering).

[0157] Output: OSM label combination of target POI (for database query).

[0158] Step S3: Hierarchical dynamic loading strategy, load POI according to L0, L1, L2 three levels to current grid, 1 ring grid, 3 ring grid.

[0159] Step S4: Real-time spatial relationship calculation, calculate POI with spatial predicates such as vehicle front 2 kilometers, vehicle right side, etc.

[0160] Step S5: Multi-modal verification and cache update, through visual auxiliary verification:

[0161] Vehicle-mounted camera captures street view, identifies store signs through OCR, and performs fuzzy matching with POI name (such as "Star 8K" → "Starbucks").

[0162] Step S6: Abnormal handling and degradation:

[0163] Network disconnection scenario: Enable local vector tiles (containing basic POI), use differential backup (recently successfully loaded POI set).

[0164] Positioning drift: Correct GPS coordinates through visual positioning (VIO), trigger POI reload threshold: horizontal error > 50m.

[0165] The main features of the multi-modal geographic language model are:

[0166] 1) Formal definition of geographic context (GC), define geographic context (GC) as the association between the geographic location of the query or POI and the surrounding geographic objects (such as roads, areas), including spatial relationship (adjacent, cover) and relative position, clearly the key role of GC in query-POI matching, solve the limitation of traditional single-modal model that cannot use geographic space information.

[0167] 2) Multi-modal fusion architecture:

[0168] Geo-encoder: Encode GC into a new modality, extract features such as ID, shape, map location, spatial relationship of geographic objects, learn context interaction through Transformer.

[0169] Multi-modal interaction module: Fusion of text semantics and geographic features, support text-geography, geography-geography cross-modal association modeling, improve matching accuracy.

[0170] Compatible with GC-free query scenarios, enhance model generalization through optional user location input.

[0171] Specific expansion figure P is a set of POIs, which can contain dozens of candidate POIs, or a large number of POIs in a massive database. Each POI, P contains a text description t p and its geographic location l p . The text description contains its official address and name. The text description is divided into three types, namely common address description, official house number description and casual spoken description.

[0172] Query represents the user's query. According to the Query query POI matching, in location-based services, given a set of POIs P and a user query Query, estimate the POIp that best matches the user's intention.

[0173] According to the size of P, two tasks need to be defined, namely sorting and retrieval. Specifically, for the sorting task, P is a list of candidate POIs, the number is limited, and it contains the most matching POI. As for the retrieval task, P is a massive database containing all POIs, the total number of POIs is huge. Since the cross-encoder is less efficient in handling large-scale P, it only runs in the sorting task. The dual encoder can run in both sorting and retrieval tasks.

[0174] Geographic information systems are built on spatial data that define the geometry of the real world. Let G be a spatial database. Each geographic object o∈G with m vertices contains a series of geographic locations o1,o2,…o m A geographic object is essentially characterized by its absolute position ID and shape o S ∈{LINE,POLYGON} in the map. Specifically, LINE represents a road in the real world, and POLYGON represents a region of interest (ROI).

[0175] We use m to represent the number of vertices in o. Note that given the geographic location of a point of interest (POI) or a query, we can form a list of nearby geographic objects o1,o2,…o ni.e. o1is the nearest geographical object to the POI or query. n is used to denote the number of geographical objects of the geographical location 1pq. We export OSM to PostGIS2 and obtain the geographical context (GC) of the geographical location from it.

[0176] The GC can characterize the relevance between a user Query or POI and its surrounding geographical objects (lines or polygons) and represents a given geographical location of a POI or Query as l pq The features of the GC can be represented using l pq and o1, o2,... o n with the relation type r t ∈ NEAR, COVERED the relative position of l pq to o i .

[0177] When searching for a target point of interest (POI), it is possible to augment the nearby context with spatial data and mention multiple relevant geographical objects in the query.

[0178] The identifiers of OSM are mapped to embedding vectors, the identifier embedding vector of the i-th nearby geographical object o i is denoted as Emb i d . The categorical shape is encoded as a numerical number using one-hot encoding and the shape embedding vector Emd i s is obtained.

[0179] The absolute position of o i in the map Emd i m is used to distinguish it from other geographical objects. The whole map area is divided into a N x N grid to obtain the scale factors S Ing and S lat for its longitude and latitude respectively:

[0180]

[0181] where lng m right and lng m left are the longitude of the right and left side of the map respectively, and lat m top and lat m bottom are the latitude of the top and bottom side of the map respectively.

[0182] Then, o i m left and o im bottom The latitude and longitude of o

[0183]

[0184] o i The discrete location feature of o i m Emb i m left Emb i m bottom Emb i m right Emb i m top} can be represented as Emb

[0185] The relationship type of o i t is also encoded using one-hot encoding, denoted as r i t .

[0186] The approximate matrix of o i is denoted as A

[0187]

[0188] where lng and lat represent the approximate latitude and longitude of the jth grid in o i The relative position r i p is calculated by the normalized distance of each edge of o pq r i p Emb i p left Emb i p bottom Emb i p right Emb i p top} can be encoded as Emb i p Emb i p left Emb​i p bottom , Emb i p right , Emb i p top}.

[0189] o i The geographical features of a geographic object can be summarized as:

[0190] Emb i = Emb i d + Emb i s +∑Emb i m + Emb i t +∑Emb i p

[0191] The intrinsic features of a geographic object are described by three components Emb d , Emb s and Emb m . Emb d is the unique identifier of the geographic object, Emb s is used to distinguish roads from regions of interest (ROI), and Emb m depicts the positional relationship between different geographic objects. The other two components EMB t and EMB p describe the correlation between geographic locations and geographic objects. After encoding the surrounding geographic objects into a sequence \({e_{1},...,e_{m}}\), the geographic encoder employs a multi-layer bidirectional transformer

[33] to learn their interactions. Following the practice of adding CLS at the beginning of the encoder for classification tasks, a GS token is added at the beginning of the encoder, and thus the output of the geographic encoder is represented as h GC , h1, …, h m .

[0192] The multi-modal geographic language model uses the following training task design for the geographic encoder:

[0193] Masked geographic modeling (MGM): Predict masked geographical features (such as OSMID, geometric type), learn the intrinsic representation of geographic objects.

[0194] Geographic contrast learning (GCL): Align the implicit geographic distance with the real geographic distance through KL divergence, and strengthen the semantic representation of spatial position.

[0195] Training loss L of the geocoder g is calculated by the following formula:

[0196] L g = L MGM + L GCL

[0197] The multi-modal geographic language model is mainly characterized in that a multi-modal pre-training strategy is used as follows:

[0198] In combination with a single-modal masked language model (MLM), a multi-modal MLM and a multi-modal geographic language modeling (MGM) task, the hidden spaces of the text and the geographic modal are aligned, and the cross-modal interaction capability is enhanced. The Bi-Encoder and Cross-Encoder architectures are supported, and the retrieval efficiency and matching accuracy are taken into account.

[0199] The street view function of the vehicle-mounted camera is used. When the POI in the user Query directly mentions the direction words such as "vehicle left side" and "front", the corresponding position camera is called to obtain the video stream, and the key frame is extracted in the video stream, and the image geographic information positioning is performed using the key frame.

[0200] The flow of obtaining the image geographic position information according to the direction is as follows:

[0201] S1. Dynamic camera scheduling, related video stream data is called according to the direction word determined by the POI. The direction word mapping logic is as shown in the following table:

[0202] Table 1

[0203] Orientation words Video source Front Front 120-degree wide-angle camera front_120deg_fov Rear Rear 90-degree camera rear_90deg_fov Left Lateral 60-degree narrow-angle camera left_60deg_fov Right Lateral 60-degree narrow-angle camera right_60deg_fov Default Front 120-degree wide-angle camera front_120deg_fov

[0204] S2. Obtain the key frame. A video key frame screening method based on reinforcement learning is used. The core logic is to use deep learning + reinforcement learning, as shown in Figure 5 , the feature extractor 502 is used to pick out the most critical and information-intensive frame from the video frame 501, the ResNet-18 feature is extracted by understanding the single-frame content using ResNet-18, the inter-frame relationship is understood by using the LSTM time sequence analysis 503, and the rules of "selecting which frame is most useful" are learned by using reinforcement learning (policy network 504 + reward calculation 506 based on sparsity and information amount), the action of retaining / dropping the key frame is performed, and finally the key frame queue 505 is screened out, realizing efficient compression or analysis of video information.

[0205] S3. Obtain geographic information according to the key frame image. This step relies on the image geographic information positioning technology, and the open source model GAEA model is used, and the key point is that:

[0206] 1) Input layer (left and right dual input):

[0207] Left side: image input.

[0208] Process the picture with a ViT Encoder (visual Transformer encoder) to convert the image into visual features (extract visual information such as buildings, scenes, etc.).

[0209] Right side: text input.

[0210] Convert the user question (Prompt) into Text Embeddings (text vector) to allow the model to understand the question semantics (such as "What Italian restaurants are nearby?").

[0211] 2) Feature alignment (middle bridge):

[0212] Visual Projector (visual projector): map the visual features extracted by ViT to a "same dimension" space as the text vector, allowing the image and text features to "talk" to each other.

[0213] Finally, get the visual feature sequence + text feature sequence, and feed them into the large model together.

[0214] 3) Core model (Qwen2.5VL):

[0215] Use a multi-modal architecture composed of Qwen2.5 (language large model) + MLP (Multilayer Perceptron, auxiliary multi-modal fusion) + ViT (visual encoding, which has been processed in advance) to understand the "combined image and text semantics".

[0216] The model takes the aligned image and text features and starts reasoning: combining the Venice scene in the image + the text question to find the answer to "Italian restaurant".

[0217] 4) Output layer (generate answer):

[0218] Based on the image and text information, the model enters the aforementioned Layer3 to call TTS to generate a text answer (such as "Farelo Restaurant"), completing the multi-modal question answering.

[0219] The multi-modal car machine voice assistant interaction method provided by the present application can:

[0220] 1) Improve the reasoning and positioning ability of geographic space tasks: by fusing a cross-modal model of a language large model and geographic information location, the user's intent is accurately associated with real-world geographic information, supporting users to ask questions about location in voice interaction, such as "What are the nearby scenic spots?" "What is the recommended Japanese restaurant nearby?" etc.

[0221] 2) Enhance multimodal interaction capabilities: By combining multimodal information such as text, images, and geographic vector data, the voice assistant's ability to understand complex scenarios is improved, resulting in a more natural and intelligent interactive experience.

[0222] 3) Improve the accuracy and security of voice interaction: By optimizing the voice recognition process and combining technologies such as in-vehicle interface image recognition and gesture recognition, we can achieve "see and speak" voice control, improve the accuracy and response speed of the interaction, and reduce driver distraction.

[0223] This application also provides a multimodal in-vehicle voice assistant interaction device, such as... Figure 6 As shown, the device includes:

[0224] The voice module 601 is configured to acquire user voice commands.

[0225] The judgment module 602 is configured to determine whether the user's voice command is related to the geographical location.

[0226] Add module 603, which is configured to add point of interest tags to user voice commands when it is determined that the user's voice command is related to a geographical location.

[0227] Recommendation module 604 is configured to obtain recommendation information based on user voice commands with added point-of-interest tags.

[0228] The control module 605 is configured to control the vehicle based on recommended information.

[0229] In some optional embodiments, the adding module 603 is configured to:

[0230] When it is determined that the user's voice command is related to a geographical location, the geographic-enhanced speech recognition model is invoked to add point of interest tags to the user's voice command.

[0231] When it is determined that the user's voice command is not related to the geographical location, the standard speech recognition model is invoked.

[0232] The device also includes an image acquisition module 606, configured to:

[0233] Acquire the vehicle's real-time location and real-time street view images outside the vehicle.

[0234] Recommended module 604 is configured as follows:

[0235] The system inputs the vehicle's real-time location, real-time street view images outside the vehicle, and user voice commands with added point-of-interest tags into a pre-stored multimodal geolinguistic model to obtain the output recommendation information.

[0236] In some alternative embodiments, the apparatus further includes a unification module 607 configured to:

[0237] Before bringing the vehicle real-time position, the real-time street view image outside the vehicle, and the user voice instruction added with the interest point label into the pre-stored multi-modal geographic language model, the features contained in the vehicle real-time position, the real-time street view image outside the vehicle, and the user voice instruction added with the interest point label are mapped into feature vectors of the same dimension.

[0238] The pre-stored multi-modal geographic language model includes a geographic context, and the geographic context is used to represent structured spatial information, and the structured spatial information is used to represent the association relationship between the geographic location corresponding to the interest point and the surrounding geographic objects, and the control module 605 is configured to:

[0239] Obtain a current driving scene.

[0240] Determine a navigation route according to the current driving scene and the recommendation information.

[0241] Control the vehicle to travel along the navigation route.

[0242] When it is determined that the user voice instruction is related to the geographic location, the interest point label can be added to the user voice instruction, so that the interest point label related to the geographic location in the user voice instruction is extracted, and subsequently, the recommendation information can be pushed to the user according to the interest point label, and the vehicle can be further controlled according to the recommendation information, so that the text information in the user voice and other dimensional information, such as geographic information and image information, are combined, and the car machine can combine multi-dimensional information to interact with the user and control the vehicle, thereby improving the efficiency of car machine interaction and vehicle control.

[0243] As shown in Figure 7 The embodiment of the application further provides a multi-modal car machine voice assistant interaction device 70, which comprises a processor 701, a memory 702, and a computer program stored in the memory 702 and executable on the processor 701, and characterized in that the computer program is executed by the processor 701 to realize:

[0244] Obtain a user voice instruction.

[0245] Determine whether the user voice instruction is related to a geographic location.

[0246] When it is determined that the user voice instruction is related to the geographic location, an interest point label is added to the user voice instruction.

[0247] According to the user voice instruction added with the interest point label, obtain recommendation information.

[0248] Control the vehicle according to the recommendation information.

[0249] In some optional embodiments, when it is judged that the user voice instruction is related to the geographic location, adding the point of interest label to the user voice instruction comprises:

[0250] When it is judged that the user voice instruction is related to the geographic location, calling the geographic enhanced speech recognition model, adding the point of interest label to the user voice instruction, the computer program further implements the following when executed by the processor 701:

[0251] When it is judged that the user voice instruction is not related to the geographic location, calling the standard speech recognition model.

[0252] In some optional embodiments, the computer program further implements the following when executed by the processor 701:

[0253] Obtaining the real-time position of the vehicle and the real-time street view image outside the vehicle.

[0254] According to the user voice instruction to which the point of interest label is added, obtaining the recommended information comprises:

[0255] Inputting the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction to which the point of interest label is added into the pre-stored multi-modal geographic language model to obtain the output recommended information.

[0256] In some optional embodiments, before the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction to which the point of interest label is added are input into the pre-stored multi-modal geographic language model, the computer program further implements the following when executed by the processor 701:

[0257] Mapping the features contained in the real-time position of the vehicle, the real-time street view image outside the vehicle, and the user voice instruction to which the point of interest label is added into feature vectors of the same dimension.

[0258] In some optional embodiments, the pre-stored multi-modal geographic language model comprises a geographic context, and the geographic context is used to represent structured spatial information, and the structured spatial information is used to represent the association relationship between the geographic location corresponding to the point of interest and the surrounding geographic objects.

[0259] In some optional embodiments, controlling the vehicle according to the recommended information comprises:

[0260] Obtaining the current driving scene.

[0261] Determining a navigation route according to the current driving scene and the recommended information.

[0262] Controlling the vehicle to travel along the navigation route.

[0263] The multi-modal vehicle machine voice assistant interaction device provided in the application can add a point of interest label to the user voice instruction when it is determined that the user voice instruction is related to a geographic location, thereby extracting a point of interest label related to the geographic location in the user voice instruction. Subsequently, the device can push recommended information to the user according to the point of interest label, and further control the vehicle according to the recommended information. Thus, the text information in the user voice and other dimensional information, such as geographic information and image information, are combined, and the vehicle machine can interact with the user and control the vehicle in combination with multi-dimensional information, thereby improving the efficiency of vehicle machine interaction and vehicle control.

[0264] The application also provides a computer-readable storage medium, for example, a memory including program code, which can be executed by a processor of a multi-modal vehicle machine voice assistant interaction device to complete the multi-modal vehicle machine voice assistant interaction method in the above embodiments. For example, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0265] Those skilled in the art can understand that all or part of the steps of the above embodiments can be completed by hardware, or by program code related hardware, and the program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0266] In the present application, it should be understood that the terms "first", "second" and the like are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated.

[0267] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. The application is intended to cover any variations, uses or adaptive changes of this application following the general principles thereof and including those expressly disclosed in the specification and those which are not. The specification and examples are to be construed as merely illustrative.

[0268] It should be understood that the application is not limited to the precise construction that has been described and shown in the accompanying drawings, and that various modifications and changes can be effected therein by those skilled in the art without departing from the scope of the application. The scope of the application is to be limited only by the appended claims.

[0269] The above merely describes the technical solutions of the present application for the purpose of enabling those skilled in the art to understand the technical solutions of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A multimodal in-vehicle voice assistant interaction method, characterized in that, The method includes: Obtain user voice commands; Determine whether the user's voice command is related to geographical location; When it is determined that the user's voice command is related to a geographical location, a point of interest tag is added to the user's voice command; Based on the user's voice command with the added point of interest tags, obtain recommendation information; The vehicle is controlled based on the recommended information.

2. The multimodal vehicle-mounted voice assistant interaction method according to claim 1, characterized in that, When it is determined that the user's voice command is related to a geographical location, adding a point of interest tag to the user's voice command includes: When it is determined that the user's voice command is related to a geographical location, a geo-enhanced speech recognition model is invoked to add point-of-interest (POI) tags to the user's voice command. The method further includes: When it is determined that the user's voice command is not related to the geographical location, the standard speech recognition model is invoked.

3. The multimodal vehicle-mounted voice assistant interaction method according to claim 1, characterized in that, The method further includes: Acquire the vehicle's real-time location and real-time street view images outside the vehicle. The step of obtaining recommendation information based on the user's voice command with the added point of interest tags includes: The real-time location of the vehicle, the real-time street view image outside the vehicle, and the user's voice command with the added point of interest tags are input into a pre-stored multimodal geolinguistic model to obtain the output recommendation information.

4. The multimodal vehicle-mounted voice assistant interaction method according to claim 3, characterized in that, Before inputting the real-time vehicle location, the real-time street view image outside the vehicle, and the user's voice command with the added point of interest tags into the pre-stored multimodal geolinguistic model, the method further includes: The features contained in the real-time vehicle location, the real-time street view image outside the vehicle, and the user voice command with the added point of interest tags are mapped into feature vectors of the same dimension.

5. The multimodal vehicle-mounted voice assistant interaction method according to claim 3, characterized in that, The pre-stored multimodal geographic language model includes a geographic context, which is used to represent structured spatial information and characterize the relationship between the geographic location corresponding to a point of interest and its surrounding geographic objects.

6. The multimodal vehicle-mounted voice assistant interaction method according to claim 1, characterized in that, The step of controlling the vehicle based on the recommended information includes: Get the current driving scenario; Determine the navigation route based on the current driving scenario and the recommended information; Control the vehicle to travel along the navigation route.

7. A multimodal in-vehicle voice assistant interaction device, characterized in that, The device includes: The voice module is configured to acquire user voice commands; The judgment module is configured to determine whether the user's voice command is related to a geographical location; The module is configured to add point of interest tags to the user's voice command when it is determined that the user's voice command is related to a geographical location. The recommendation module is configured to obtain recommendation information based on the user's voice command with the added point of interest tags; The control module is configured to control the vehicle based on the recommended information.

8. The multimodal vehicle-mounted voice assistant interaction device according to claim 7, characterized in that, The added module is configured as follows: When it is determined that the user's voice command is related to a geographical location, a geolocation-enhanced speech recognition model is invoked to add point-of-interest (POI) tags to the user's voice command. When it is determined that the user's voice command is not related to the geographical location, the standard speech recognition model is invoked. The device further includes an image acquisition module, configured to: Acquire the vehicle's real-time location and real-time street view images outside the vehicle. The recommendation module is configured as follows: The real-time location of the vehicle, the real-time street view image outside the vehicle, and the user's voice command with the added point of interest tags are input into a pre-stored multimodal geolinguistic model to obtain the output recommendation information.

9. The multimodal vehicle-mounted voice assistant interaction device according to claim 8, characterized in that, The device also includes a unified module configured as follows: Before inputting the vehicle's real-time location, the real-time street view image outside the vehicle, and the user's voice command with the added point of interest (POI) tags into the pre-stored multimodal geolinguistic model, the features contained in the vehicle's real-time location, the real-time street view image outside the vehicle, and the user's voice command with the added POI tags are mapped into feature vectors of the same dimension. The pre-stored multimodal geographic language model includes a geographic context, which represents structured spatial information. This structured spatial information characterizes the relationship between the geographic location of a point of interest and its surrounding geographic objects. The control module is configured to: Get the current driving scenario; Determine the navigation route based on the current driving scenario and the recommended information; Control the vehicle to travel along the navigation route.

10. A multimodal in-vehicle voice assistant interaction device, comprising a processor, a memory, and a computer program stored in the memory and executable on the processor, characterized in that, The computer program is implemented when executed by the processor: Obtain user voice commands; Determine whether the user's voice command is related to geographical location; When it is determined that the user's voice command is related to a geographical location, a point of interest tag is added to the user's voice command; Based on the user's voice command with the added point of interest tags, obtain recommendation information; The vehicle is controlled based on the recommended information.