Navigation guidance method, device, equipment and storage medium

By extracting the navigation text and the image ahead of the vehicle, and using the reference extraction model to generate navigation guidance data, the problem of poor navigation effect of the vehicle navigation software under complex terrain is solved, and the accuracy of navigation guidance and the driver's navigation decision-making ability are improved.

CN120252771APending Publication Date: 2025-07-04DONGFENG MOTOR CO LTD DONGFENG NISSAN PASSENGER VEHICLE CO
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510566620.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing vehicle navigation software has poor navigation performance on complex terrain such as complex intersections and overpasses, making it difficult for drivers to determine the correct navigation road, affecting the accuracy of navigation guidance.

Method used

By performing feature extraction of navigation text and vehicle front images, the reference extraction model is used to generate navigation guidance data, modify or rewrite navigation text to improve the accuracy of navigation guidance, including training the reference extraction model to adapt to complex road conditions.

Benefits of technology

Improves the effectiveness of navigation guidance, allowing drivers to more easily determine lanes, ensuring good navigation guidance can be maintained under complex road conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120252771A_ABST
    Figure CN120252771A_ABST
Patent Text Reader

Abstract

The invention discloses a navigation guiding method and device, equipment and a storage medium, and relates to the technical field of vehicle navigation, and the method comprises the following steps: carrying out text feature extraction on a navigation text to obtain navigation text features, and carrying out visual feature extraction on a vehicle front image to obtain front image features; the navigation text features and the front image features are input into a reference object extraction model for reasoning, navigation guide data are generated, the navigation guide data comprise reasoning process data and a navigation guide text, and the navigation guide text is text data generated after the navigation text is modified or rewritten according to a reference object in the image; and performing vehicle navigation guidance based on the navigation guidance data. Due to the fact that the proper reference object is selected through the model to modify the navigation text provided by the navigation software, a driver can determine the lane where the driver should run more easily according to the reference object, the navigation guiding effect is improved, and the good guiding effect can still be kept even in the face of complex road conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of vehicle navigation, and particularly to a navigation guidance method, device, equipment, and storage medium. Background Art

[0002] With the development of cities, the road environment within cities has become increasingly complex. Currently, the navigation maps of in-vehicle navigation software have poor voice guidance effects for complex terrains. Especially at complex intersections and complex overpasses, even when voice navigation is combined with terrain maps, drivers cannot determine the correct navigation roads based on voice navigation alone, making it difficult to effectively complete navigation guidance and resulting in a poor actual experience for drivers. Summary of the Invention

[0003] The main purpose of this application is to provide a navigation guidance method, device, equipment, and storage medium, aiming to solve the technical problem that current navigation guidance cannot handle complex terrains, resulting in poor navigation effects.

[0004] To achieve the above objective, this application proposes a navigation guidance method, which includes:

[0005] Extract text features from the navigation text to obtain navigation text features, and extract visual features from the image in front of the vehicle to obtain front image features. The navigation text is the text used by the in-vehicle navigation software for voice navigation prompts based on the navigation route and the vehicle's position.

[0006] Input the navigation text features and the front image features into a reference object extraction model for inference to generate navigation guidance data. The navigation guidance data includes inference process data and navigation guidance text. The navigation guidance text is text data generated by modifying or rewriting the navigation text based on the reference objects in the image.

[0007] Perform vehicle navigation guidance based on the navigation guidance data.

[0008] Optionally, before extracting text features from the navigation text to obtain navigation text features and extracting visual features from the image in front of the vehicle to obtain front image features, it further includes:

[0009] Sample the navigation prompt text of the in-vehicle navigation software to obtain a prompt text set, and acquire the front images collected during the vehicle's driving to obtain a front image set.

[0010] Match the data in the prompt text set and the front image set based on the navigation route to generate a vehicle driving sample set.

[0011] Train an initial extraction model based on the vehicle driving sample set to obtain a reference object extraction model.

[0012] Optionally, training the initial extraction model based on the vehicle driving sample set to obtain a reference object extraction model includes:

[0013] Traverse the vehicle driving sample set, and use the traversed vehicle driving sample as the current sample;

[0014] Extract the image in front of the sample and the sample prompt text from the current sample;

[0015] Extract image features from the image in front of the sample to obtain sample image features, and extract text features from the sample prompt text to obtain sample text features;

[0016] Input the sample image features and the sample text features into the initial extraction model for inference to obtain inference guidance data;

[0017] Evaluate the inference guidance data based on the result evaluation index to generate an evaluation reward value;

[0018] Iteratively optimize the initial extraction model according to the evaluation reward value;

[0019] If a preset end condition is met, use the initial extraction model as the reference object extraction model;

[0020] If the preset end condition is not met, return to the step of traversing the vehicle driving sample set and using the traversed vehicle driving sample as the current sample.

[0021] Optionally, the result evaluation index includes road modeling accuracy;

[0022] The evaluating the inference guidance data based on the result evaluation index to generate an evaluation reward value includes:

[0023] Obtain the actual road data corresponding to the current sample, where the actual road data includes the number of lanes, lane categories, lane attributes, and lane positions;

[0024] Extract road modeling data from the inference process data included in the inference guidance data;

[0025] Match the road modeling data with the actual road data to determine modeling errors and omissions;

[0026] Perform result evaluation based on the modeling errors and omissions to generate an evaluation reward value.

[0027] Optionally, the result evaluation index includes environmental inference accuracy;

[0028] Evaluating the inference guidance data based on the result evaluation metrics to generate an evaluation reward value, including:

[0029] Obtain the environmental description data corresponding to the current sample, where the environmental description data includes environmental description categories, specific feature categories when describing each environmental description category, and optional parameters when making specific descriptions;

[0030] Extract the inference environmental description from the inference guidance data;

[0031] Match the inference environmental description with the environmental description data to determine the correct description items and incorrect description items;

[0032] Generate an evaluation reward value based on the correct description items and the incorrect description items.

[0033] Optionally, the result evaluation metrics include guidance instruction accuracy;

[0034] Evaluating the inference guidance data based on the result evaluation metrics to generate an evaluation reward value, including:

[0035] Obtain the target lane corresponding to the current sample, where the sample navigation description includes navigation action description and navigation road description;

[0036] Extract the inference navigation description from the inference guidance data;

[0037] Match the lane corresponding to the inference navigation description with the target lane to generate a lane detection result;

[0038] Generate an evaluation reward value based on the lane detection result.

[0039] Optionally, the matching the lane corresponding to the inference navigation description with the target lane to generate a lane detection result includes:

[0040] Determine at least one reference object description according to the inference navigation description;

[0041] Obtain the lanes corresponding to each reference object description;

[0042] Match the lanes corresponding to each reference object description with the target lane respectively to generate a lane detection result.

[0043] In addition, to achieve the above object, the present application further provides a navigation guidance device, where the navigation guidance device includes:

[0044] An extraction module, configured to extract text features from navigation text to obtain navigation text features, and extract visual features from an image in front of the vehicle to obtain front image features, where the navigation text is the text used by an in-vehicle navigation software for voice navigation prompts according to a navigation route and the vehicle position;

[0045] An inference module, configured to input the navigation text features and the front image features into a reference object extraction model for inference to generate navigation guidance data, where the navigation guidance data includes inference process data and navigation guidance text, and the navigation guidance text is text data generated by modifying or rewriting the navigation text based on reference objects in the image;

[0046] A guidance module, configured to perform vehicle navigation guidance based on the navigation guidance data.

[0047] In addition, to achieve the above object, the present application further provides a navigation guidance device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the navigation guidance method as described above.

[0048] In addition, to achieve the above object, the present application further provides a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the navigation guidance method as described above are implemented.

[0049] In addition, to achieve the above object, the present application further provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the navigation guidance method as described above are implemented.

[0050] One or more technical solutions proposed by the present application have at least the following technical effects:

[0051] Since the model selects appropriate reference objects to modify the navigation text provided by the navigation software, the driver can more easily determine the lane to be traveled based on the reference objects, improving the navigation guidance effect. Even in the face of complex road conditions, a good guidance effect can still be maintained. Description of the Drawings

[0052] The drawings here are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0053] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 It is a schematic flowchart provided for the first embodiment of the navigation guidance method of the present application;

[0055] Figure 2 It is a schematic flowchart provided for the second embodiment of the navigation guidance method of the present application;

[0056] Figure 3 It is a schematic diagram of the model training framework for an embodiment of the present application;

[0057] Figure 4 It is a schematic diagram of the module structure of the navigation guidance device according to the embodiment of the present application;

[0058] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the navigation guidance method according to the embodiment of the present application.

[0059] The implementation, functional features, and advantages of the purpose of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0060] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0061] To better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and the specific implementation manners.

[0062] Based on this, the embodiments of the present application provide a navigation guidance method, referring to Figure 1 , Figure 1 It is a schematic flowchart of the first embodiment of the navigation guidance method of the present application.

[0063] In this embodiment, the navigation guidance method includes steps S10 to S40:

[0064] Step S10: Extract text features from the navigation text to obtain navigation text features, and extract visual features from the image in front of the vehicle to obtain front image features.

[0065] It should be noted that the execution subject of this embodiment can be the vehicle itself or a navigation and guidance device provided in the vehicle. The navigation and guidance device can be a controller provided in the vehicle, such as an ECU controller, or other devices that can achieve the same or similar functions. This embodiment does not limit this. In this embodiment and the following embodiments, the navigation and guidance method of the present application will be described by taking the navigation and guidance device as an example.

[0066] It should be noted that the navigation text can be the text used by the in-vehicle navigation software for voice navigation prompts according to the navigation route and the vehicle position.

[0067] In actual use, when extracting text features from the navigation text and visual features from the image in front of the vehicle, a pre-set feature extraction algorithm or a pre-trained feature extraction model can be used. This embodiment does not limit this.

[0068] For example, for the image I1 in front of the vehicle, a visual encoder can be used to extract features, which can be represented as:

[0069] V = fv(I1) ∈ R*Nv×d

[0070] In the formula, fv is the visual encoder (such as ViT, ResNet), V is the visual embedding feature of the image, and R*Nv×d is a real number matrix with Nv rows and d columns;

[0071] In order to facilitate unifying the dimensions, the final image feature in front can also be obtained through a linear transformation:

[0072] V′ = MLP(V)

[0073] V′ is the visual embedding feature of the final image, that is, the image feature in front. MLP is the multi-layer perceptron (Multilayer Perceptron, MLP) used for linear transformation;

[0074] For the navigation text, the navigation text can be encoded. The input text T is converted into a token sequence through the tokenizer:

[0075] T = Tokenizer(sentence) = [w1, w2,..., wNt]

[0076] After that, a text encoder (such as Transformer, BERT, etc.) is used to extract text features, then there is:

[0077] L = fl(T) ∈ R*Nt×d

[0078] Where \(f_l\) is the text encoder (such as BERT or CLIP-Text), \(L\) is the text token feature, \(N_t\) is the token sequence length, and \(d\) is the dimension;

[0079] Similarly, for the convenience of unifying the dimensions, a linear transformation can also be performed, so there is:

[0080] \(L' = MLP(L)\)

[0081] Where \(L'\) is the navigation text feature, and \(MLP\) is the Multilayer Perceptron (MLP) used for linear transformation.

[0082] Step S20: Input the navigation text feature and the front image feature into the reference object extraction model for inference to generate navigation guidance data.

[0083] It should be noted that the navigation guidance data includes inference process data and navigation guidance text. The navigation guidance text is text data generated by modifying or rewriting the navigation text based on the reference objects in the image.

[0084] In practical applications, the reference object extraction model can be a specialized trained large intelligent model (such as LLaMA, GPT-4, T5).

[0085] In actual use, the reference object extraction model can determine the reference objects available for the navigation process based on the input text features and the front image features, and modify or rewrite the navigation text feature data based on the reference objects to generate navigation guidance text. Then, navigation guidance data is constructed according to the derivation process and the navigation guidance text.

[0086] For example: The navigation text is "Drive along the right road at the intersection ahead", and the navigation guidance text in the generated navigation guidance data can be "At the intersection ahead, you need to drive along the right road, and the right road is the road on which the red vehicle ahead is driving."

[0087] In practical applications, for the convenience of the reference object extraction model to process, the visual information and navigation information can also be aligned, so that the navigation text feature \(L\) and the front image feature \(V\) are projected into the same semantic space, and the feature pair \((V, L)\) after contrast learning is encoded by cross-attention, and the feature pair is deeply fused to be input into the reference object extraction model.

[0088] For example: Using the Cross-Attention mechanism, the front image feature \(V\) is fused into the navigation text feature:

[0089] \(Q = W\) Q \(L'\), \(K = W\) K \(V'\), \(V = W\) V \(V'\)

[0090]

[0091] Among them, Q is generated from navigation text features (query). K and V are generated from the front image features (key and value);

[0092] Softmax normalizes the attention scores to align the text and image features, and finally obtains the fused features:

[0093] H = CrossAttention(L′, V′)

[0094] O = LLM(H)

[0095] The fused feature H is fed into the reference object extraction model LLM to obtain the output O. In the training phase, O is used for evaluation and continues to be trained based on the evaluation; in the inference phase, the output O is the final output result.

[0096] Step S30: Perform vehicle navigation guidance based on the navigation guidance data.

[0097] It should be noted that after obtaining the navigation guidance data, the navigation guidance device can display the navigation guidance data. At the same time, the navigation guidance text in the navigation guidance data is broadcast in a voice announcement manner to achieve vehicle navigation guidance.

[0098] This embodiment provides a navigation guidance method. Since the model is used to select a suitable reference object to modify the navigation text provided by the navigation software, the driver can more easily determine the lane to be traveled according to the reference object, improving the navigation guidance effect. Even in the face of complex road conditions, a good guidance effect can still be maintained.

[0099] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , before step S10, the navigation guidance method further includes steps S01 to S03:

[0100] Step S01: Sample the navigation prompt text of the in-vehicle navigation software to obtain a prompt text set, and obtain the front images collected during the vehicle driving process to obtain a front image set.

[0101] Step S02: Match the data in the prompt text set and the front image set based on the navigation route to generate a vehicle driving sample set.

[0102] It should be noted that a real vehicle can be used to drive according to a navigation route. During this process, the navigation prompt text of the in-vehicle navigation software is collected to obtain a prompt text set. At the same time, the images in front of the vehicle are continuously collected during the vehicle driving process to obtain a front image set. After that, according to the position of the vehicle when the front image is collected and the position of the vehicle when the navigation prompt text is broadcast, the data in the prompt text set and the front image set are matched, so that the navigation prompt text at the same vehicle position corresponds to the front image, and a vehicle driving sample is constructed. Based on this, the obtained vehicle driving samples are aggregated into a vehicle driving sample set.

[0103] Among them, after obtaining the vehicle driving samples, the management personnel of the navigation guidance device can also standardize the vehicle driving samples and set standard navigation guidance texts for them, so as to evaluate the generation effect of the model.

[0104] Step S03: Train the initial extraction model based on the vehicle driving sample set to obtain a reference object extraction model.

[0105] In actual use, the vehicle driving sample set can be used to iteratively train the initial extraction model. When the training reaches the preset end condition, the trained initial extraction model is used as the reference object extraction model.

[0106] Among them, the preset end condition can be set in advance by the management personnel of the navigation guidance device. For example, the preset end condition is set to the training times reaching the preset number of times, the training rounds reaching the preset number of rounds, and / or the model accuracy rate reaching the preset ratio.

[0107] In a specific implementation, in order to reasonably perform model training and ensure the effect of the trained model, step S03 in this embodiment may include:

[0108] Traverse the vehicle driving sample set and use the traversed vehicle driving sample as the current sample;

[0109] Extract the sample front image and the sample prompt text from the current sample;

[0110] Extract image features from the sample front image to obtain sample image features, and extract text features from the sample prompt text to obtain sample text features;

[0111] Input the sample image features and the sample text features into the initial extraction model for inference to obtain inference guidance data;

[0112] Evaluate the inference guidance data based on the result evaluation index to generate an evaluation reward value;

[0113] Iteratively optimize the initial extraction model according to the evaluated reward value;

[0114] If the preset end condition is satisfied, use the initial extraction model as the reference extraction model.

[0115] It should be noted that the specific implementation of image feature extraction for the image in front of the sample and text feature extraction for the sample prompt text is basically the same as that of text feature extraction for the navigation text and visual feature extraction for the image in front of the vehicle. The specific implementation can refer to the above description and will not be elaborated here. The result evaluation indicators may include at least one of road modeling accuracy, environment reasoning accuracy, guidance instruction accuracy, and format standardization.

[0116] In actual use, the management personnel of the navigation guidance device can set at least one result evaluation indicator according to actual needs.

[0117] If only one result evaluation indicator is set, the evaluation can be directly based on this result evaluation indicator, and the obtained evaluation result can be used as the evaluation reward value.

[0118] If multiple result evaluation indicators are set at the same time, the evaluation can be carried out separately based on each result evaluation indicator to generate the reward value corresponding to each result evaluation indicator. Then, the reward values corresponding to each result evaluation indicator are weighted and summed to obtain the final evaluation reward value.

[0119] Among them, the weight coefficients used for weighted summation can be set in advance by the management personnel of the navigation guidance device according to actual needs for the result evaluation indicators, and the weight coefficients of different result evaluation indicators are different.

[0120] In actual use, after obtaining the evaluation reward value, the evaluation reward value can be fed back to the initial extraction model to enable the initial extraction model to clarify the effect of its generated result, and learn based on this effect to adjust its corresponding model strategy or model parameters. Then, it can be detected whether the preset end condition is satisfied;

[0121] If the preset end condition is satisfied, the trained initial extraction model can be used as the reference extraction model; if the preset end condition is not satisfied, it can return to traverse the vehicle driving sample set, use the traversed vehicle driving sample as the current sample, and continue the traversal.

[0122] In the specific implementation, if the result evaluation indicator only includes road modeling accuracy, the step of evaluating the inference guidance data based on the result evaluation indicator to generate the evaluation reward value may include:

[0123] Obtain the actual road data corresponding to the current sample, where the actual road data includes;

[0124] Extract road modeling data from the inference process data included in the inference guidance data;

[0125] Match the road modeling data with the actual road data to determine the modeling omissions and errors;

[0126] Based on the modeling omissions and errors, conduct result evaluation to generate an evaluation reward value.

[0127] It should be noted that the actual road data corresponding to the current sample can extract road information using a dataset or an environment simulation platform to obtain the actual road data. The actual road data can include information such as the number of lanes, lane categories, lane attributes, and lane positions. For example, "There are a total of 3 lanes, where the first lane from left to right is a left-turn lane, the second lane is a straight-through lane, and the third lane is a right-turn lane."

[0128] In actual use, if appropriate navigation guidance text needs to be given, it is necessary to ensure that the model can correctly perform road modeling, that is, correctly determine the number of lanes included in the road, the categories of lanes, the attributes of lanes, and the positions of lanes. In the inference process of the model, this kind of data is generally only used for the inference process. Based on this, in order to correctly conduct evaluation, road modeling data can be extracted from the inference process data included in the inference guidance data. Then, the road modeling data is matched with the actual road data to determine the modeling omissions and errors. Then, based on the modeling omissions and errors, result evaluation is conducted to generate an evaluation reward value.

[0129] For example: The actual road data can contain a total of M keywords such as the number of lanes, the positions of each lane, and the attributes of each lane. If all the keywords are included in the road modeling data extracted from the inference process data, a score of 1 point can be given. If each keyword is missing (i.e., one modeling omission or error), then 1 / M points will be deducted.

[0130] Among them, when deducting points, different scores can also be deducted according to the importance of the keywords. For example, if the number of roads is incorrect, the error is serious and more points can be deducted.

[0131] In specific implementation, if the result evaluation index only includes the accuracy of environment inference, then the step of evaluating the inference guidance data based on the result evaluation index to generate an evaluation reward value can include:

[0132] Obtain the environment description data corresponding to the current sample;

[0133] Extract the inference environment description from the inference guidance data;

[0134] Match the inference environment description with the environment description data to determine the correct description items and incorrect description items;

[0135] Generate an evaluation reward value based on the correct description items and the incorrect description items.

[0136] It should be noted that the environment description data includes environment description categories, specific feature categories when describing each environment description category, and optional parameters when making specific descriptions.

[0137] For ease of understanding, the following is an example for illustration, but it does not limit the present solution. The information that the environment description data can include is shown in the following table:

[0138]

[0139]

[0140] In actual use, the inference guidance data can be comprehensively recognized to determine the environment-related descriptions in the process of model inference and the navigation guidance text obtained by inference, so as to obtain the inference environment description. The environment description data corresponding to the current sample can also be obtained by extracting from a data set or an environment simulation platform.

[0141] In actual application, the inference environment description can be matched with the environment description data to determine the correctness of the description during model inference and the incorrect description items. Finally, based on the correct description items and the incorrect description items, result evaluation is performed to generate an evaluation reward value.

[0142] For example: Score according to the integrity of category description and optional values. In a single category, a complete description gets 1 point. If it contains K features, when one feature is missing, 1 / K points are deducted. Then, the scores of all descriptions are accumulated. When necessary, the accumulated scores can also be normalized.

[0143] In specific implementation, if the result evaluation index only includes the accuracy of the guidance instruction, the step of evaluating the inference guidance data based on the result evaluation index to generate an evaluation reward value may include:

[0144] Obtain the target lane corresponding to the current sample;

[0145] Extract the inference navigation description from the inference guidance data;

[0146] Match the lane corresponding to the inference navigation description with the target lane to generate a lane detection result;

[0147] Generate an evaluation reward value based on the lane detection result.

[0148] It should be noted that the target lane corresponding to the current sample can be the lane that should be traveled according to the actual navigation prompt. The inference navigation description can include a navigation action description and a navigation road description. For example, in the navigation instruction, there are [action] and [road]. Among them, action is the navigation action description, such as follow, turn left, turn right, etc., and road is the road description, such as there is an XX sign, there is an XX intersection, etc.

[0149] In actual use, the inference navigation description can be extracted from the inference guidance data. After that, the lane corresponding to it is determined according to the road description in the inference navigation description. For example, assume that the inference navigation description is "There is / isn't a [color] [stop / pedestrian crossing, etc.] sign on the [left / right] side of the target lane. Pay attention and observe", then the lane to which the sign belongs is used as the lane corresponding to the inference navigation description.

[0150] In actual use, the lane corresponding to the inference navigation description can be matched with the target lane to determine whether they are consistent. After that, an evaluation reward value is generated according to the lane detection result. For example, if the lane detection result shows consistency, the evaluation reward value is 1; if they are inconsistent, the evaluation reward value is 0.

[0151] In a specific implementation, in order to improve the guidance effect, when the model selects reference objects, it may select multiple reference objects. In order to reasonably give a reward value in such a case, the step of matching the lane corresponding to the inference navigation description with the target lane and generating a lane detection result in this embodiment may include:

[0152] Determine at least one reference object description according to the inference navigation description;

[0153] Obtain the lanes corresponding to each reference object description;

[0154] Match the lanes corresponding to each reference object description with the target lane respectively to generate a lane detection result.

[0155] It should be noted that at least one reference object description can be extracted according to the description statement in the inference navigation description. For example, assume that the inference navigation description is "There is an A sign on the right side of the target lane. At the same time, there is a red vehicle driving in the target lane and you can follow the red vehicle", then the reference object descriptions at this time are "There is an A sign on the right side" and "There is a red vehicle".

[0156] In actual use, after obtaining the reference object description, the lane of the reference object corresponding to the reference object description can be obtained, so as to obtain the lane corresponding to the reference object description. After that, the lanes corresponding to each reference object description are matched with the target lane respectively to generate a lane detection result.

[0157] It can be understood that in such a case of multiple reference objects, the method of generating an evaluation reward value based on the lane detection result before may no longer be applicable. At this time, the number of reference object descriptions corresponding to the determined lane being consistent with the target lane can be determined according to the lane detection result, and the evaluation reward value can be determined according to this number.

[0158] For example: for each reference object description where the corresponding lane is consistent with the target lane, the evaluation reward value is increased by 1. At the same time, if there is each reference object description where the corresponding lane is inconsistent with the target lane, the evaluation reward value is decreased by 2 (the minimum value is 0). If there is not a single reference object description consistent with the target lane, the evaluation reward value is set to 0.

[0159] In a specific implementation, to ensure the readability of the generated result, the result evaluation index can include format correctness. If the result evaluation index only includes format correctness, then the inference guidance data is evaluated based on the result evaluation index to generate an evaluation reward value, which can include: detecting whether the format of the inference guidance data conforms to a preset format; if it conforms, the evaluation reward value is 1; if it does not conform, the evaluation reward value is set to 0, or determining the format difference degree, and taking the difference between 1 and the format difference degree as the evaluation reward value.

[0160] For the sake of easy understanding, it is now combined with Figure 3 for illustration, but it does not limit this solution. Figure 3 This is a schematic diagram of the model training framework for this embodiment.

[0161] As Figure 3 shown, the sample prompt text and the sample front image can be extracted from the current sample, the sample prompt text is subjected to language encoding (i.e., text feature extraction), the sample front image is subjected to visual encoding (i.e., image feature extraction), and then, the extracted features are subjected to feature fusion and input into the LLM base (i.e., the initial extraction model). Then, the RL model evaluates the inference guidance data generated by the initial extraction model to generate an evaluation reward value. Then, the evaluation reward value is fed back to the initial extraction model to optimize the objective function, and iterative training is performed in this way until the training is completed (i.e., the preset end condition is satisfied), and the trained initial extraction model is used as the reference object extraction model.

[0162] Among them, at least one result evaluation index can be preset in the RL model. Assuming that the result evaluation index includes 4, namely road modeling accuracy, environmental inference accuracy, guidance instruction accuracy, and format correctness, the reward values corresponding to each result evaluation index are R1, R2, R3, and R4 respectively. Then the final evaluation reward value Rpos = a1*R1 + a2*R2 + a3*R3 + a4*R4, where a1, a2, a3, and a4 are the weight coefficients corresponding to each result evaluation index respectively.

[0163] In practical applications, in order to prevent the model description from hallucinating, that is, describing non-existent things, a penalty term is added to the reward model, introducing a penalty coefficient bi (the penalty coefficient can be negative). When the sub-items in the above reward items do not meet the expectations, a penalty score is added.

[0164] Rreg = b1R1 + b2R2 + b3R3 + b4R4

[0165] Then the final evaluation reward value R at this time is R = Rpos + Rreg.

[0166] This embodiment provides a navigation guidance method. Since the model for selecting reference objects is pre-trained, and at the same time, various preset indicators can be combined for result evaluation during the training process, it is ensured that the model can be reasonably adjusted to ensure the effectiveness of the reference object selection of the finally trained model.

[0167] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the navigation guidance method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0168] This application also provides a navigation guidance device. Please refer to Figure 4 , the navigation guidance device includes:

[0169] An extraction module 10, configured to extract text features from navigation text to obtain navigation text features, and extract visual features from the image in front of the vehicle to obtain front image features. The navigation text is the text used by the in-vehicle navigation software for voice navigation prompts according to the navigation route and the vehicle position;

[0170] An inference module 20, configured to input the navigation text features and the front image features into a reference object extraction model for inference to generate navigation guidance data. The navigation guidance data includes inference process data and navigation guidance text. The navigation guidance text is text data generated after modifying or rewriting the navigation text based on the reference objects in the image;

[0171] A guidance module 30, configured to perform vehicle navigation guidance based on the navigation guidance data.

[0172] The navigation guidance device provided by this application adopts the navigation guidance method in the above embodiment, and can solve the technical problem that the current navigation guidance cannot handle complex terrains, resulting in poor navigation effects. Compared with the prior art, the beneficial effects of the navigation guidance device provided by this application are the same as those of the navigation guidance method provided by the above embodiment, and the other technical features in the navigation guidance device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0173] This application provides a navigation guidance device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the navigation guidance method in the first embodiment above.

[0174] Reference is made below Figure 5 to FIG., which shows a schematic structural diagram of a navigation guidance device suitable for implementing the embodiments of the present application. The navigation guidance device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description: tablet computers), PMPs (Portable Media Player: portable multimedia players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The navigation guidance device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0175] As Figure 5 shown, the navigation guidance device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1002 or the program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the navigation guidance device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the navigation guidance device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a navigation guidance device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be alternatively implemented or had.

[0176] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by a processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.

[0177] The navigation guidance device provided by the present application adopts the navigation guidance method in the above-mentioned embodiment, and can solve the technical problem that the current navigation guidance cannot cope with complex terrains, resulting in poor navigation effects. Compared with the prior art, the beneficial effects of the navigation guidance device provided by the present application are the same as those of the navigation guidance method provided by the above-mentioned embodiment, and other technical features in the navigation guidance device are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.

[0178] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0179] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0180] The present application provides a computer-readable storage medium, which has computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the navigation guidance method in the above-mentioned embodiment.

[0181] The computer-readable storage medium provided by the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0182] The above computer-readable storage medium may be included in a navigation guidance device; or may exist separately without being assembled into the navigation guidance device.

[0183] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the navigation guidance device, the navigation guidance device: extracts text features from the navigation text to obtain navigation text features, and extracts visual features from the image in front of the vehicle to obtain front image features, where the navigation text is the text used by the in-vehicle navigation software for voice navigation prompts according to the navigation route and the vehicle position; inputs the navigation text features and the front image features into a reference object extraction model for inference to generate navigation guidance data, where the navigation guidance data includes inference process data and navigation guidance text, and the navigation guidance text is text data generated by modifying or rewriting the navigation text based on the reference object in the image; and performs vehicle navigation guidance based on the navigation guidance data.

[0184] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0186] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0187] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned navigation guidance method, which can solve the technical problem that the current navigation guidance cannot cope with complex terrains, resulting in poor navigation effects. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the navigation guidance method provided by the above embodiments, and will not be elaborated here.

[0188] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the navigation guidance method as described above.

[0189] The computer program product provided by the present application can solve the technical problem that the current navigation guidance cannot cope with complex terrains, resulting in poor navigation effects. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the navigation guidance method provided in the above embodiments, and will not be elaborated herein.

[0190] The above are only partial embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A navigation guidance method, characterized in that, The navigation guidance method includes: Performing text feature extraction on the navigation text to obtain navigation text features, and performing visual feature extraction on the image in front of the vehicle to obtain front image features, where the navigation text is the text used by the in-vehicle navigation software for voice navigation prompts according to the navigation route and the vehicle position; Inputting the navigation text features and the front image features into a reference object extraction model for inference to generate navigation guidance data, where the navigation guidance data includes inference process data and navigation guidance text, and the navigation guidance text is text data generated by modifying or rewriting the navigation text based on the reference objects in the image; Performing vehicle navigation guidance based on the navigation guidance data.

2. The navigation guidance method according to claim 1, wherein Before performing text feature extraction on the navigation text to obtain navigation text features and performing visual feature extraction on the image in front of the vehicle to obtain front image features, it further includes: Sampling the navigation prompt text of the in-vehicle navigation software to obtain a prompt text set, and acquiring the front images collected during the vehicle driving to obtain a front image set; Matching the data in the prompt text set and the front image set based on the navigation route to generate a vehicle driving sample set; Training an initial extraction model based on the vehicle driving sample set to obtain a reference object extraction model.

3. The navigation guidance method according to claim 2, wherein The training an initial extraction model based on the vehicle driving sample set to obtain a reference object extraction model includes: Traversing the vehicle driving sample set and using the traversed vehicle driving sample as the current sample; Extracting the sample front image and the sample prompt text from the current sample; Performing image feature extraction on the sample front image to obtain sample image features, and performing text feature extraction on the sample prompt text to obtain sample text features; Inputting the sample image features and the sample text features into the initial extraction model for inference to obtain inference guidance data; Evaluating the inference guidance data based on a result evaluation index to generate an evaluation reward value; Iteratively optimizing the initial extraction model according to the evaluation reward value; If a preset end condition is satisfied, using the initial extraction model as the reference object extraction model.

4. The navigation guidance method according to claim 3, wherein The result evaluation index includes road modeling accuracy; The evaluating the inference guidance data based on a result evaluation index to generate an evaluation reward value includes: Obtaining the actual road data corresponding to the current sample, where the actual road data includes the number of lanes, lane categories, lane attributes, and lane positions; Extracting road modeling data from the inference process data included in the inference guidance data; Matching the road modeling data with the actual road data to determine modeling errors and omissions; Performing result evaluation based on the modeling errors and omissions to generate an evaluation reward value.

5. The navigation guidance method according to claim 3, wherein The result evaluation index includes environmental inference accuracy; The evaluating the inference guidance data based on a result evaluation index to generate an evaluation reward value includes: Obtaining the environmental description data corresponding to the current sample, where the environmental description data includes environmental description categories, specific feature categories when describing each environmental description category, and optional parameters when specifically describing; Extract the inference environment description from the inference guidance data; Match the inference environment description with the environment description data to determine the correct description items and incorrect description items; Generate an evaluation reward value based on the correct description items and the incorrect description items.

6. The navigation guidance method according to claim 3, wherein The result evaluation index includes the accuracy of the guidance instruction; Evaluating the inference guidance data based on the result evaluation index to generate an evaluation reward value, including: Obtain the target lane corresponding to the current sample, where the sample navigation description includes a navigation action description and a navigation road description; Extract the inference navigation description from the inference guidance data; Match the lane corresponding to the inference navigation description with the target lane to generate a lane detection result; Generate an evaluation reward value based on the lane detection result.

7. The navigation guidance method according to claim 6, characterized in that, The matching of the lane corresponding to the inference navigation description with the target lane to generate a lane detection result includes: Determine at least one reference object description according to the inference navigation description; Obtain the lanes corresponding to each reference object description; Match the lanes corresponding to each reference object description with the target lane respectively to generate a lane detection result.

8. A navigation guidance device, characterized in that, The navigation guidance device includes: An extraction module, configured to extract text features from navigation text to obtain navigation text features, and extract visual features from the image in front of the vehicle to obtain front image features, where the navigation text is the text used by the in-vehicle navigation software for voice navigation prompts according to the navigation route and the vehicle position; An inference module, configured to input the navigation text features and the front image features into a reference object extraction model for inference to generate navigation guidance data, where the navigation guidance data includes inference process data and a navigation guidance text, and the navigation guidance text is text data generated by modifying or rewriting the navigation text based on the reference objects in the image; A guidance module, configured to perform vehicle navigation guidance based on the navigation guidance data.

9. A navigation guiding device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the navigation guidance method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the navigation guidance method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Navigation method, device, equipment, storage medium and product

    CN121702413A