Method, device and electronic equipment for processing navigation guidance information

By converting the navigation guide image into semantic text and matching it with the text, the problem of inconsistent mode of navigation guide information in map-like applications is solved, and the accurate identification and processing of abnormal navigation information is realized, improving the user's navigation experience.

CN115564968BActive Publication Date: 2025-08-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211356742.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-08-12
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

In the prior art, there may be problems of inconsistent expression meaning between different modes of navigation and guidance information of map-like applications, which leads to the occurrence of yaw situations and affects the normal use of users.

Method used

By obtaining the images and text in the navigation guidance information, converting the image into semantic text, and semantically matching it with the text, we determine whether the navigation guidance information is abnormal navigation guidance information.

Benefits of technology

Effectively identify and handle abnormal navigation and guidance information, improve user navigation experience, and avoid yaw.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115564968B_ABST
    Figure CN115564968B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, device and electronic device for processing navigation guidance information, which relates to the field of data processing technology, and in particular to the field of intelligent transportation, data mining or machine learning technology. The specific implementation scheme is: obtaining navigation guidance information corresponding to the target road area, the navigation guidance information includes a navigation guidance image and a navigation guidance text; determining the semantic text of the navigation guidance image; and determining whether the navigation guidance information is abnormal navigation guidance information based on the semantic text and the navigation guidance text. In this scheme, by converting the navigation guidance image into semantic text, the semantic text and the navigation guidance text are both text-modal information, which facilitates semantic matching of the semantic text and the navigation guidance text, thereby determining whether the navigation guidance information is abnormal navigation guidance information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, in particular to the field of intelligent transportation, data mining or machine learning technology. Specifically, the present disclosure relates to a method, device and electronic device for processing navigation guidance information. Background Art

[0002] With the rapid development of mobile Internet technology, map applications have also been widely used. More and more users are starting to use map applications for navigation.

[0003] Map applications typically display navigation guidance information in multiple modes, such as displaying navigation guidance images and providing voice guidance text. However, existing technologies may present inconsistent meanings in different modes of navigation guidance. In such cases, the navigation guidance information is considered abnormal and can affect user experience.

[0004] If the abnormal navigation guidance information can be effectively determined, the abnormal navigation guidance information can be effectively processed. Therefore, how to effectively determine the abnormal navigation guidance information has become an important technical issue. Summary of the Invention

[0005] In order to solve at least one of the above-mentioned deficiencies, the present disclosure provides a method, device and electronic device for processing navigation guidance information.

[0006] According to a first aspect of the present disclosure, a method for processing navigation guidance information is provided, the method comprising:

[0007] Obtaining navigation guidance information corresponding to the target road area, the navigation guidance information including a navigation guidance image and navigation guidance text;

[0008] Determine the semantic text of the navigation guide image;

[0009] Based on the semantic text and the navigation guidance text, it is determined whether the navigation guidance information is abnormal navigation guidance information.

[0010] According to a second aspect of the present disclosure, a device for processing navigation guidance information is provided, the device comprising:

[0011] A navigation guidance information acquisition module is used to obtain navigation guidance information corresponding to the target road area, the navigation guidance information including a navigation guidance image and a navigation guidance text;

[0012] A semantic text extraction module, used to determine the semantic text of the navigation guide image;

[0013] The abnormal navigation guidance information determination module is used to determine whether the navigation guidance information is abnormal navigation guidance information based on the semantic text and the navigation guidance text.

[0014] According to a third aspect of the present disclosure, an electronic device is provided, including:

[0015] at least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the navigation guidance information processing method.

[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned method for processing navigation guidance information.

[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the above-mentioned method for processing navigation guidance information when executed by a processor.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0022] Figure 1 This is a flowchart of a method for processing navigation guidance information provided by an embodiment of the present disclosure;

[0023] Figure 2 This is a flowchart of a specific implementation of the method for processing navigation guidance information provided by an embodiment of the present disclosure;

[0024] Figure 3 It is a schematic diagram of a page used for manual review of abnormal navigation guidance information;

[0025] Figure 4 This is a structural diagram of a navigation guidance information processing device provided by an embodiment of the present disclosure;

[0026] Figure 5 It is a block diagram of an electronic device used to implement the method for processing navigation guidance information according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] Map applications typically output navigation guidance information in multiple modes when performing navigation, allowing users to better understand the information. However, in existing technologies, navigation guidance information in different modes may have inconsistent meanings, which may lead to deviation (deviation refers to the inconsistency between the actual driving route and the driving route indicated by the navigation guidance information), affecting the user experience. Navigation guidance information in this case can be identified as abnormal navigation guidance information.

[0029] For example, navigation guidance information includes a navigation guidance image displayed on the screen and navigation guidance text displayed via voice broadcast. The two may have inconsistent meanings. For example, the navigation guidance image may indicate a right turn at the intersection ahead, while the navigation guidance text may indicate a right turn at the intersection ahead.

[0030] Navigation information from different modalities may be generated independently by different modules, which can lead to inconsistent meanings among the navigation guidance information in different modalities, i.e., abnormal navigation guidance information. If abnormal navigation guidance information can be effectively identified, it can be effectively handled. Therefore, how to effectively identify abnormal navigation guidance information has become a key technical issue.

[0031] In related technologies, discrepancies in the meaning of navigation guidance information in different modes can lead to yaw. Therefore, it is generally assumed that yaw is caused by discrepancies in the meaning of navigation guidance information in different modes, i.e., abnormal navigation guidance information. However, this analysis method is not accurate because users may actively yaw the vehicle.

[0032] The navigation guidance information processing method, device, and electronic device provided in the embodiments of the present disclosure are intended to solve at least one of the above technical problems in the prior art.

[0033] Figure 1 A flowchart of a method for processing navigation guidance information provided by an embodiment of the present disclosure is shown. Figure 1As shown in , the method may mainly include:

[0034] Step S110: Obtain navigation guidance information corresponding to the target road area, where the navigation guidance information includes a navigation guidance image and navigation guidance text.

[0035] Step S120: Determine the semantic text of the navigation guide image.

[0036] Step S130: Based on the semantic text and the navigation guidance text, determine whether the navigation guidance information is abnormal navigation guidance information.

[0037] The target road area is a road area where the user needs to make a driving decision. The navigation guidance information is used to guide the user to make a driving decision in the target road area.

[0038] In the disclosed embodiments, navigation guidance information includes both image and text modes. The navigation guidance image is navigation guidance information in image mode, which can be an image displayed on a screen of a terminal device such as a mobile phone or vehicle-mounted terminal, and is used to display information such as a map of the target road area and route guidance lines. The navigation guidance text is navigation guidance information in text mode.

[0039] In the disclosed embodiment, the navigation guidance image generally displays information such as a map and a path guide line. By analyzing the navigation guidance image, the semantic text of the navigation guidance image can be determined, and the semantic text can describe the meaning expressed by the navigation guidance image.

[0040] Since the navigation guidance image is converted into semantic text, the semantic text and the navigation guidance text are both text-modal information, which facilitates semantic matching between the semantic text and the navigation guidance text, and determines whether the navigation guidance information is abnormal navigation guidance information based on the semantic matching result.

[0041] Specifically, when the semantic text and the navigation guidance text express the same semantics, it indicates that the meaning expressed by the navigation guidance image is consistent with the meaning expressed by the navigation guidance text, and the navigation guidance information can be determined to be normal. When the semantic text and the navigation guidance text express different semantics, it indicates that the meaning expressed by the navigation guidance image is inconsistent with the meaning expressed by the navigation guidance text, and the navigation guidance information can be determined to be abnormal.

[0042] The disclosed method provides a method for obtaining navigation guidance information corresponding to a target road area, the navigation guidance information comprising a navigation guidance image and navigation guidance text; determining the semantic text of the navigation guidance image; and determining whether the navigation guidance information is abnormal based on the semantic text and the navigation guidance text. In this solution, by converting the navigation guidance image into semantic text, both the semantic text and the navigation guidance text are in textual mode, facilitating semantic matching between the semantic text and the navigation guidance text, thereby determining whether the navigation guidance information is abnormal.

[0043] In an optional manner of the present disclosure, the target road area includes a target intersection area, the navigation guidance image includes a local enlarged image of the target intersection area, and the navigation guidance text is used for voice broadcasting.

[0044] In the disclosed embodiment, when a vehicle is at an intersection, the driver needs to make a driving decision to determine the road to take. Therefore, when the vehicle is about to enter the intersection, the map application generally displays a zoomed-in image to the user and plays a voice navigation guidance text to prompt the user to select the road to take.

[0045] The partially enlarged image will display information about the roads within the intersection area and the target road that the user is guided to.

[0046] By voice broadcasting the navigation guidance text, users can be prompted to make road decisions through voice.

[0047] In the embodiment of the present disclosure, the intersection area may be defined as an area including the area where the intersection is located and an area extending along each road entering or exiting the intersection area by a preset distance.

[0048] As an example, the target road area is a target intersection area. When a vehicle is about to enter the target intersection area, the map application displays a partially enlarged image of the target intersection area to the user and simultaneously voice-plays the navigation guidance text.

[0049] In an optional manner of the present disclosure, determining the semantic text of the navigation guide image includes:

[0050] Extracting image features of the navigation guidance image;

[0051] The image features are input into the pre-trained image-text conversion model to obtain the semantic text of the navigation guidance image.

[0052] In the disclosed embodiment, the image features of the navigation guidance image can be extracted through a pre-trained convolutional neural network.

[0053] In the disclosed embodiment, a pre-trained image-text conversion model can be a temporal neural network model. The image features of the navigation guidance image are input into the image-text conversion model to obtain the semantic text of the navigation guidance image output by the image-text conversion model.

[0054] In an optional embodiment of the present disclosure, the image-text conversion model is obtained by:

[0055] Obtaining sample navigation guidance information corresponding to the sample road area, where the sample navigation guidance information includes a sample navigation guidance image and a sample navigation guidance text;

[0056] Determining semantic text labels based on sample navigation guide text;

[0057] extracting sample image features of the sample navigation guidance image;

[0058] The model is trained based on sample image features and semantic text labels to obtain an image-text conversion model.

[0059] In the disclosed implementation, the sample navigation guidance information can be historical navigation guidance information accumulated online. Semantic text labels can be determined based on the sample navigation guidance text, and sample image features of the sample navigation guidance image can be extracted. The image-to-text conversion model is trained using the sample image features of the sample navigation guidance image as input and the semantic text labels as output.

[0060] In the disclosed embodiment, since the historical navigation guidance information accumulated online is used as training data and the semantic text labels are determined based on the sample navigation guidance text, the process of manual semantic annotation is omitted, avoiding a large amount of manpower consumed in the annotation process.

[0061] If a deep learning model is used to directly map navigation guidance information from different modalities to the same vector dimension, and then determine semantic consistency based on the similarity of the corresponding vectors of navigation guidance information from different modalities, the final quality assessment result will be vector similarity, which has only mathematical significance and is difficult to interpret for its practical business significance. In this solution, by training an image-to-text conversion model and converting the navigation guidance image into semantic text, the image is then semantically matched with the navigation guidance text. Quality assessment is performed based on the semantic matching results, making the final quality assessment result interpretable.

[0062] In an optional manner of the present disclosure, a semantic text label is determined based on the sample navigation guide text in the following manner:

[0063] Based on a preset keyword extraction method, sample semantic keywords of the sample navigation guide text are extracted, and the semantic keywords of the sample navigation guide text are used as semantic text tags.

[0064] In an embodiment of the present disclosure, a preset keyword extraction method may be provided to extract sample semantic keywords from sample navigation guide texts, and use the sample semantic keywords as semantic text tags.

[0065] Specifically, a semantic keyword thesaurus may be constructed based on semantic keywords that may appear in navigation guidance texts corresponding to various road conditions, and semantic keywords may be extracted based on the semantic keyword thesaurus.

[0066] In the embodiment of the present disclosure, the semantic keywords extracted from the sample navigation guidance text may be in the form of a semantic keyword sequence, for example: left turn; three-way fork; middle.

[0067] In the disclosed embodiment, the semantic keywords may include intersection type keywords, turning direction keywords, enhanced prompt keywords, and conjunctions. Among them. Intersection types include turning intersections or forks. Turning direction keywords are used to indicate the turning direction of the vehicle, such as turning left. Enhanced prompt keywords are used to provide enhanced guidance to users when the intersection situation is more complicated to avoid driving into the wrong road. For example, the navigation guidance text is "Turn to the right front at the intersection ahead, not turn right", where "not turn right" is the enhanced prompt keyword. Conjunctions are used to connect the clauses in the navigation guidance text, such as conjunctions "later".

[0068] In an optional manner of the present disclosure, the semantic text of the navigation guidance image includes semantic keywords of the navigation guidance image, and determining whether the navigation guidance information is abnormal navigation guidance information based on the semantic text and the navigation guidance text includes:

[0069] Extracting semantic keywords of navigation guide text based on a preset keyword extraction method;

[0070] Based on the semantic keywords of the navigation guidance image and the semantic keywords of the navigation guidance text, it is determined whether the navigation guidance information is abnormal navigation guidance information.

[0071] In the embodiment of the present disclosure, when the semantic text label is a semantic keyword of the sample navigation guide text, the semantic text predicted based on the image-text conversion model is the semantic keyword.

[0072] As an example, the semantic text may be in the form of a sequence of semantic keywords, such as: left turn; three-way fork; middle.

[0073] In the disclosed embodiments, the navigation guidance text may contain some invalid information, such as titles. Semantic keywords in the navigation guidance text can be extracted based on the above-described predictive keyword extraction method. Semantic keywords in the navigation guidance text are valid information in the navigation guidance text and can represent the semantic meaning of the navigation guidance text.

[0074] Since the semantic keywords of the navigation guidance image can represent the semantics expressed by the navigation guidance image, and the semantic keywords in the navigation guidance text can represent the semantics of the navigation guidance text, the semantic keywords of the navigation guidance image and the semantic keywords of the navigation guidance text can be matched to determine whether the navigation guidance information is abnormal navigation guidance information.

[0075] Specifically, the semantic keywords of the navigation guidance image and the semantic keywords of the navigation guidance text can be directly matched to determine whether the two are exactly the same. If the two are exactly the same, it can be determined that the navigation guidance information is not abnormal navigation guidance information; if the two are not exactly the same, it can be determined that the navigation guidance information is abnormal navigation guidance information.

[0076] In an optional manner of the present disclosure, determining whether the navigation guidance information is abnormal navigation guidance information based on semantic keywords of the navigation guidance image and semantic keywords of the navigation guidance text includes:

[0077] Determining, based on a preset correspondence between semantic keywords and semantic types, a first semantic type corresponding to the semantic keyword of the navigation guidance image and a second semantic type corresponding to the semantic keyword of the navigation guidance text;

[0078] In response to the first semantic type being different from the second semantic type, the navigation guidance information is determined to be abnormal navigation guidance information.

[0079] In the disclosed embodiment, there may be a situation where the semantic keywords of the navigation guidance image are not completely consistent with the semantic keywords of the navigation guidance text, but the meanings actually expressed by the two are consistent. For example, the semantic keywords of the navigation guidance image are "go straight, then, three-way fork, left front", and the semantic keywords of the navigation guidance text are "go straight, then, three-way fork, left front, note that it is not a left turn", and the meanings actually expressed by the two are consistent. In order to avoid misjudgment of abnormal navigation guidance information due to this situation, semantic types can be defined, and semantic keywords that express the same meaning can be established in correspondence with the same semantic type, so that when the semantic keywords of the navigation guidance image are not completely consistent with the semantic keywords of the navigation guidance text, if the semantic keywords of the navigation guidance image and the semantic keywords of the navigation guidance text can be mapped to the same semantic type, then it can be determined that the semantic keywords of the navigation guidance image are consistent with the semantic keywords of the navigation guidance text.

[0080] Specifically, when the first semantic type and the second semantic type are the same, it indicates that the meaning expressed by the navigation guidance image is consistent with the meaning expressed by the navigation guidance text, and in this case, it can be determined that the navigation guidance information is not abnormal. When the first semantic type and the second semantic type are different, it indicates that the meaning expressed by the navigation guidance image is inconsistent with the meaning expressed by the navigation guidance text, and in this case, it can be determined that the navigation guidance information is abnormal.

[0081] As an example, the semantic keyword sequence "go straight, then, three-way fork, left front" can correspond to type A, and "go straight, then, three-way fork, left front, note that it is not a left turn" can also correspond to type A.

[0082] In an optional implementation manner of the present disclosure, after determining that the navigation guidance information is abnormal navigation guidance information in response to the first semantic type being different from the second semantic type, the method further includes:

[0083] An abnormality level of the abnormal navigation guidance information is determined based on the first semantic type and the second voice type.

[0084] In the disclosed embodiment, after abnormal navigation guidance information is determined, the abnormal type of the abnormal navigation guidance information can be determined based on the first semantic type and the second voice type. The abnormal type is associated with a preset abnormality level, so that the abnormality level of the abnormal navigation guidance information can be characterized by the abnormal type. The abnormal type can be used to subsequently obtain abnormal navigation guidance information based on the abnormal type for corresponding processing.

[0085] As an example, the abnormality type may include abnormal turning direction, abnormal intersection type, abnormal enhanced prompt, etc.

[0086] Specifically, when the meanings expressed by the first semantic type and the second voice type are similar, it can be considered that the abnormality level of the abnormal navigation guidance information is low.

[0087] In an optional embodiment of the present disclosure, the present invention further includes:

[0088] In response to receiving the abnormal navigation guidance information acquisition request, target abnormal navigation guidance information is acquired based on the abnormality type specifying information carried in the abnormal navigation guidance information acquisition request.

[0089] In the embodiment of the present disclosure, the user can set the exception type designation information as needed. The exception type designation information is used to specify the target exception type to be obtained, so as to obtain the target exception navigation guidance information corresponding to the target exception type for subsequent use.

[0090] Specifically, the user can configure the abnormal type designation information according to the abnormal degree of the abnormal navigation guidance information to be obtained. For example, if abnormal navigation guidance information with a higher abnormal degree is required, the user can configure the abnormal type designation information corresponding to the abnormal level with a higher abnormal degree.

[0091] In this disclosed embodiment, navigation guidance images can be derived based on a guidance image prediction model, and navigation guidance text can be derived based on a guidance text prediction model. Abnormal navigation guidance information with a high degree of abnormality can be used as training samples for training the guidance image prediction model or the guidance text prediction model. Abnormal navigation guidance information with a high degree of abnormality is more effective, and model training based on this abnormal navigation guidance information can effectively improve the model's prediction performance.

[0092] In an optional implementation manner of the present disclosure, after determining that the navigation guidance information is abnormal navigation guidance information in response to the first semantic type being different from the second semantic type, the method further includes:

[0093] Adjust abnormal navigation guidance information based on preset adjustment strategies.

[0094] In the disclosed embodiment, during navigation, it is possible to detect in real time whether the navigation guidance information is abnormal. After determining that the abnormal navigation guidance information is abnormal, the abnormal navigation guidance information can be adjusted in a timely manner to avoid the occurrence of deviation caused by the abnormal navigation information, thereby improving the user experience.

[0095] As an example, the navigation guide image can be obtained based on the guide image prediction model, and the navigation guide text can be obtained based on the guide text prediction model. Different guide image prediction models and guide text prediction models can be used to regenerate the navigation guide image and navigation guide image text.

[0096] As another example, it may also be determined that there is an erroneous navigation guidance image or navigation guidance image text in the abnormal navigation guidance information, and adjustments may be made to the erroneous navigation guidance image or navigation guidance image text.

[0097] As an example, Figure 2 A flowchart of a specific implementation of the method for processing navigation guidance information provided by an embodiment of the present disclosure is shown in FIG.

[0098] like Figure 2 As shown in , the input picture is the navigation guide image input.

[0099] The encoder convolutional neural network uses a convolutional neural network model as an encoder to perform feature encoding on the navigation guidance image to obtain image features of the navigation guidance image.

[0100] Multimodal intermediate identifiers (visual / semantic), i.e., image features of navigation guidance images extracted by a convolutional neural network model.

[0101] The decoder temporal neural network uses a temporal neural network model as an image-text conversion model to perform semantic analysis on the navigation guidance image to obtain the semantic text of the navigation guidance image.

[0102] The tile at the intersection ahead, pointing to the right, is an example of semantic text in a navigation guidance image.

[0103] In the embodiment of the present disclosure, in order to ensure the accuracy of the abnormal navigation guidance information determined, the abnormal navigation guidance information is also reviewed manually. Figure 3 A schematic diagram of a page for manual review of abnormal navigation guidance information is shown in FIG.

[0104] like Figure 3 As shown in the figure, the right side of the page displays a real-world image of the vehicle's surroundings captured by onboard sensors. The right side of the page also displays navigation guidance text and the semantic text of the navigation guidance image. A human operator determines whether the meaning of the navigation guidance text and the navigation guidance image is consistent based on the real-world image. If they are inconsistent, the user can click the "Validate" button. If they are consistent, the user can click the "Invalidate" button. This allows for manual review of abnormal navigation guidance information.

[0105] Based on Figure 1 The same principle as shown in the method, Figure 4 A schematic diagram of the structure of a navigation guidance information processing device provided by an embodiment of the present disclosure is shown. Figure 4 As shown, the navigation guidance information processing device 40 may include:

[0106] The navigation guidance information acquisition module 410 is used to obtain navigation guidance information corresponding to the target road area, and the navigation guidance information includes a navigation guidance image and a navigation guidance text;

[0107] A semantic text extraction module 420 is used to determine the semantic text of the navigation guide image;

[0108] The abnormal navigation guidance information determining module 430 is configured to determine whether the navigation guidance information is abnormal navigation guidance information based on the semantic text and the navigation guidance text.

[0109] The device provided by the disclosed embodiments obtains navigation guidance information corresponding to a target road area, including a navigation guidance image and navigation guidance text; determines the semantic text of the navigation guidance image; and, based on the semantic text and the navigation guidance text, determines whether the navigation guidance information is abnormal. In this solution, by converting the navigation guidance image into semantic text, both the semantic text and the navigation guidance text are in textual mode, facilitating semantic matching between the semantic text and the navigation guidance text, thereby determining whether the navigation guidance information is abnormal.

[0110] Optionally, the target road area includes a target intersection area, the navigation guidance image includes a local magnified image of the target intersection area, and the navigation guidance text is used for voice broadcasting.

[0111] Optionally, the semantic text extraction module is specifically used to:

[0112] Extracting image features of the navigation guidance image;

[0113] The image features are input into the pre-trained image-text conversion model to obtain the semantic text of the navigation guidance image.

[0114] Optionally, the image-text conversion model is obtained by:

[0115] Obtaining sample navigation guidance information corresponding to the sample road area, where the sample navigation guidance information includes a sample navigation guidance image and a sample navigation guidance text;

[0116] Determining semantic text labels based on sample navigation guide text;

[0117] extracting sample image features of the sample navigation guidance image;

[0118] The model is trained based on sample image features and semantic text labels to obtain an image-text conversion model.

[0119] Optionally, a semantic text label is determined based on the sample navigation guide text in the following manner:

[0120] Based on a preset keyword extraction method, semantic keywords of the sample navigation guide text are extracted, and the semantic keywords of the sample navigation guide text are used as semantic text tags.

[0121] Optionally, the semantic text of the navigation guidance image includes semantic keywords of the navigation guidance image, and the abnormal navigation guidance information determination module is specifically configured to:

[0122] Extracting semantic keywords of navigation guide text based on a preset keyword extraction method;

[0123] Extracting semantic keywords of navigation guide text based on a preset keyword extraction method;

[0124] Based on the semantic keywords of the navigation guidance image and the semantic keywords of the navigation guidance text, it is determined whether the navigation guidance information is abnormal navigation guidance information.

[0125] Optionally, when the abnormal navigation guidance information determining module determines whether the navigation guidance information is abnormal navigation guidance information based on the semantic keywords of the navigation guidance image and the semantic keywords of the navigation guidance text, it is specifically configured to:

[0126] Determining, based on a preset correspondence between semantic keywords and semantic types, a first semantic type corresponding to the semantic keyword of the navigation guidance image and a second semantic type corresponding to the semantic keyword of the navigation guidance text;

[0127] In response to the first semantic type being different from the second semantic type, the navigation guidance information is determined to be abnormal navigation guidance information.

[0128] Optionally, the above device further includes:

[0129] An abnormality type determination module is used to determine the abnormality type of the abnormal navigation guidance information based on the first semantic type and the second voice type after determining the navigation guidance information as abnormal navigation guidance information in response to the first semantic type being different from the second semantic type, wherein the abnormality type is associated with a preset abnormality degree.

[0130] Optionally, the above device further includes:

[0131] The target abnormal navigation guidance information acquisition module is used to obtain the target abnormal navigation guidance information in response to receiving the abnormal navigation guidance information acquisition request and based on the abnormal type specifying information carried in the abnormal navigation guidance information acquisition request.

[0132] Optionally, the above device further includes:

[0133] The abnormal navigation guidance information adjustment module is configured to adjust the abnormal navigation guidance information based on a preset adjustment strategy after determining the navigation guidance information as abnormal navigation guidance information in response to the first semantic type being different from the second semantic type.

[0134] It is understandable that the above modules of the navigation guidance information processing device in the embodiment of the present disclosure have the function of realizing Figure 1The functions of the corresponding steps of the method for processing navigation guidance information in the embodiment shown in . This function can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated with multiple modules. For the functional description of each module of the above navigation guidance information processing device, please refer to Figure 1 The corresponding description of the method for processing navigation guidance information in the embodiment shown in is not repeated here.

[0135] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0136] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0137] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for processing navigation guidance information as provided in the embodiment of the present disclosure.

[0138] Compared to existing technologies, this electronic device obtains navigation guidance information corresponding to a target road area, including a navigation guidance image and navigation guidance text; determines the semantic text of the navigation guidance image; and, based on the semantic text and navigation guidance text, determines whether the navigation guidance information is anomalous. This solution converts the navigation guidance image into semantic text, making both the semantic text and the navigation guidance text in textual mode. This facilitates semantic matching between the semantic text and the navigation guidance text, thereby determining whether the navigation guidance information is anomalous.

[0139] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method for processing navigation guidance information as provided in an embodiment of the present disclosure.

[0140] Compared to existing technologies, this readable storage medium obtains navigation guidance information corresponding to a target road area, including a navigation guidance image and navigation guidance text; determines the semantic text of the navigation guidance image; and, based on the semantic text and navigation guidance text, determines whether the navigation guidance information is anomalous. This solution converts the navigation guidance image into semantic text, making both the semantic text and the navigation guidance text in textual mode. This facilitates semantic matching between the semantic text and the navigation guidance text, thereby determining whether the navigation guidance information is anomalous.

[0141] The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the method for processing navigation guidance information provided in the embodiment of the present disclosure.

[0142] Compared to existing technologies, this computer program product obtains navigation guidance information corresponding to a target road area, including a navigation guidance image and navigation guidance text; determines the semantic text of the navigation guidance image; and, based on the semantic text and navigation guidance text, determines whether the navigation guidance information is anomalous. This solution converts the navigation guidance image into semantic text, making both the semantic text and the navigation guidance text in textual mode. This facilitates semantic matching between the semantic text and the navigation guidance text, thereby determining whether the navigation guidance information is anomalous.

[0143] Figure 5 A schematic block diagram of an example electronic device 50 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0144] like Figure 5 As shown, the electronic device 50 includes a computing unit 510, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 520 or a computer program loaded from a storage unit 580 into a random access memory (RAM) 530. Various programs and data required for the operation of the device 50 can also be stored in the RAM 530. The computing unit 510, the ROM 520, and the RAM 530 are connected to each other via a bus 540. An input / output (I / O) interface 550 is also connected to the bus 540.

[0145] Multiple components in device 50 are connected to I / O interface 550, including: an input unit 560, such as a keyboard, mouse, etc.; an output unit 570, such as various types of displays, speakers, etc.; a storage unit 580, such as a magnetic disk, optical disk, etc.; and a communication unit 590, such as a network card, modem, wireless communication transceiver, etc. The communication unit 590 allows device 50 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0146] The computing unit 510 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 510 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 510 performs the processing method of the navigation guidance information provided in the embodiments of the present disclosure. For example, in some embodiments, the processing method of the navigation guidance information provided in the embodiments of the present disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 580. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 50 via the ROM 520 and / or the communication unit 590. When the computer program is loaded into the RAM 530 and executed by the computing unit 510, one or more steps of the processing method of the navigation guidance information provided in the embodiments of the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 510 can be configured to perform the processing method of the navigation guidance information provided in the embodiments of the present disclosure by any other appropriate means (e.g., by means of firmware).

[0147] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0149] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0150] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0151] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0152] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0153] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0154] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A method for processing navigation guidance information, comprising: Obtaining navigation guidance information corresponding to the target road area, wherein the navigation guidance information includes a navigation guidance image and a navigation guidance text; Determining a semantic text of the navigation guide image, where the semantic text of the navigation guide image includes semantic keywords of the navigation guide image; Extracting semantic keywords of the navigation guide text based on a preset keyword extraction method; Determining, based on a preset correspondence between semantic keywords and semantic types, a first semantic type corresponding to the semantic keyword of the navigation guidance image and a second semantic type corresponding to the semantic keyword of the navigation guidance text; In response to the first semantic type being different from the second semantic type, determining the navigation guidance information as abnormal navigation guidance information; The semantic text of the navigation guide image is generated based on a picture-text conversion model, and the picture-text conversion model is obtained in the following manner: Acquire sample navigation guidance information corresponding to the sample road area, wherein the sample navigation guidance information includes a sample navigation guidance image and a sample navigation guidance text; Extracting semantic keywords of the sample navigation guide text based on the preset keyword extraction method, and using the semantic keywords of the sample navigation guide text as semantic text tags; extracting sample image features of the sample navigation guidance image; Model training is performed based on the sample image features and the semantic text labels to obtain the image-text conversion model.

2. The method according to claim 1, wherein The target road area includes a target intersection area, the navigation guidance image includes a partial enlarged image of the target intersection area, and the navigation guidance text is used for voice broadcasting.

3. The method according to claim 1 or 2, wherein: The determining of the semantic text of the navigation guide image includes: extracting image features of the navigation guidance image; The image features are input into a pre-trained image-text conversion model to obtain the semantic text of the navigation guidance image.

4. The method according to claim 1, wherein After determining the navigation guidance information as abnormal navigation guidance information in response to the first semantic type being different from the second semantic type, the method further includes: An abnormality type of the abnormal navigation guidance information is determined based on the first semantic type and the second semantic type, where the abnormality type is associated with a preset abnormality degree.

5. The method according to claim 4, further comprising: In response to receiving the abnormal navigation guidance information acquisition request, target abnormal navigation guidance information is acquired based on the abnormality type specifying information carried in the abnormal navigation guidance information acquisition request.

6. The method according to any one of claims 1 to 5, wherein After determining the navigation guidance information as abnormal navigation guidance information in response to the first semantic type being different from the second semantic type, the method further includes: The abnormal navigation guidance information is adjusted based on a preset adjustment strategy.

7. A device for processing navigation guidance information, comprising: A navigation guidance information acquisition module is used to obtain navigation guidance information corresponding to the target road area, wherein the navigation guidance information includes a navigation guidance image and a navigation guidance text; a semantic text extraction module, configured to determine the semantic text of the navigation guidance image, wherein the semantic text of the navigation guidance image includes semantic keywords of the navigation guidance image; an abnormal navigation guidance information determination module, configured to extract semantic keywords of the navigation guidance text based on a preset keyword extraction method; determine a first semantic type corresponding to the semantic keywords of the navigation guidance image and a second semantic type corresponding to the semantic keywords of the navigation guidance text based on a preset correspondence between semantic keywords and semantic types; and determine the navigation guidance information as abnormal navigation guidance information in response to the first semantic type being different from the second semantic type; The semantic text of the navigation guide image is generated based on a picture-text conversion model, and the picture-text conversion model is obtained in the following manner: Acquire sample navigation guidance information corresponding to the sample road area, wherein the sample navigation guidance information includes a sample navigation guidance image and a sample navigation guidance text; Extracting semantic keywords of the sample navigation guide text based on the preset keyword extraction method, and using the semantic keywords of the sample navigation guide text as semantic text tags; extracting sample image features of the sample navigation guidance image; Model training is performed based on the sample image features and the semantic text labels to obtain the image-text conversion model.

8. The device according to claim 7, wherein The target road area includes a target intersection area, the navigation guidance image includes a partial enlarged image of the target intersection area, and the navigation guidance text is used for voice broadcasting.

9. The device according to claim 8, wherein The semantic text extraction module is specifically used for: extracting image features of the navigation guidance image; The image features are input into a pre-trained image-text conversion model to obtain the semantic text of the navigation guidance image.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

12. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent navigation apparatus, navigation terminal and its information navigation method

    CN101482420A

  • Determining a navigation destination using a navigation device

    DE102018010101A1