Intersection navigation method based on large language model and related device
Through the intersection navigation method based on the large language model, the street view marking diagram of the intersection is extracted and analyzed, the reference objects of the target road are determined, and accurate intersection navigation guidelines are generated, which solves the problem of inaccurate navigation guidelines in the existing technology and achieves higher navigation accuracy.
Patent Information
- Application Number
- CN202510796322.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-15
AI Technical Summary
When existing map navigation products navigate at intersections, it is difficult to provide more accurate navigation guidance, especially lane-level navigation guidance, which makes it difficult for users to accurately judge the driving direction.
Through the intersection navigation method based on the large language model, the street view marking diagram of the intersection is extracted and the large language model is input. The model is used to understand the marking meaning and intersection guidance in the street view marking diagram to generate instructions, determine the reference object of the target road, and generate accurate intersection navigation guidance.
It significantly improves the accuracy of intersection navigation, making navigation guidelines strongly correlated with intersections, and can guide vehicles to the target road more accurately.
Smart Images

Figure CN120489168A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, specifically to the field of artificial intelligence technology such as image recognition, computer vision, large language models, and intelligent navigation, and more particularly to a large language model-based intersection navigation method, device, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Existing map navigation products, when providing navigation services at intersections, are based on the road network geometry and travel planning of the intersection, and provide navigation instructions such as "go straight, turn left, turn left in front of the right" according to the corresponding driving direction.
[0003] However, due to the sophistication of intersection data modeling and the popularity of lane-level navigation, it is difficult to provide users with more accurate intersection navigation guidance according to simple terms such as "left front, right front, keep left, keep right". Summary of the Invention
[0004] The embodiments of the present disclosure provide a large language model-based intersection navigation method, device, electronic device, computer-readable storage medium, and computer program product.
[0005] In a first aspect, an embodiment of the present disclosure proposes a method for intersection navigation based on a large language model, comprising: in response to a vehicle passing through an intersection according to a travel plan, extracting a street view marker map corresponding to the intersection from the travel plan; wherein the street view marker map includes a first mark of the actual position of the vehicle at the intersection, a second mark of the direction of the target road to be traveled to according to the travel plan, and a third mark of the direction of other roads; the street view marker map, an explanation of the meaning represented by each type of mark contained in the street view marker map, and an intersection guidance generation instruction are collectively input as prompt information into a preset large language model; based on the output information of the large language model, an intersection navigation instruction is determined to guide the vehicle through the intersection toward the target road; wherein, under the intention expressed by the intersection guidance generation instruction, the large language model understands the street view marker map according to the explanation and determines a target reference object for distinguishing the target road from other roads, and generates output information based on the actual position, target reference object and intention.
[0006] In a second aspect, an embodiment of the present disclosure proposes an intersection navigation device based on a large language model, comprising: a street view map extraction unit, configured to extract a street view mark map corresponding to the intersection from the travel plan in response to a vehicle passing through the intersection according to a travel plan; wherein the street view mark map includes a first mark for the actual position of the vehicle at the intersection, a second mark for the direction of the target road to be traveled to according to the travel plan, and a third mark for the direction of other roads; a prompt information input unit, configured to input the street view mark map, an explanation of the meaning represented by each type of mark contained in the street view mark map, and an intersection guidance generation instruction as prompt information into a preset large language model; an intersection navigation instruction generation unit, configured to determine, based on the output information of the large language model, intersection navigation instructions that guide the vehicle to pass through the intersection and travel toward the target road; wherein, under the intention expressed by the intersection guidance generation instruction, the large language model understands the street view mark map according to the explanation and determines a target reference object for distinguishing the target road from other roads, and generates output information based on the actual position, target reference object, and intention.
[0007] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that when executed by the at least one processor, the intersection navigation method based on the large language model as described in the first aspect can be implemented.
[0008] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which are used to enable a computer to implement the intersection navigation method based on a large language model as described in the first aspect when executed.
[0009] In a fifth aspect, an embodiment of the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, can implement the steps of the intersection navigation method based on a large language model as described in the first aspect.
[0010] The large language model-based intersection navigation solution provided by the present disclosure pre-associates a matched street view marker map with the intersections traversed in the trip plan. The street view marker map pre-associates the vehicle's intersection location, the target road that matches the trip plan, and the directions of other non-matching roads in the intersection street view map with corresponding types of markers. The street view marker map, explanations of the meanings represented by each type of marker, and intersection guidance generation instructions are then input into a preset large language model as prompt information, so that the large language model determines the intention expressed by the intersection guidance generation instructions, and interprets the street view marker map based on the explanations based on the intention to find a target reference object that can distinguish the target road from other roads from the dimension of image features. Finally, the large language model can output intersection navigation instructions based on the target reference object to guide the vehicle from the intersection to the target road. Because the intersection navigation instructions are determined based on the target reference object and have a strong association with the intersection, they have strong recognition and can significantly improve the accuracy of intersection navigation instructions.
[0011] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Other features, objects and advantages of the present disclosure will become more apparent from a reading of the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 is an exemplary system architecture in which the present disclosure may be applied; Figure 2 A flowchart of a large language model-based intersection navigation method provided in an embodiment of the present disclosure; Figure 3 A flowchart of a method for pre-generating a street view marker map provided by an embodiment of the present disclosure; Figure 4 A flowchart of a method for generating output information based on input prompt information using a large language model provided by an embodiment of the present disclosure; Figure 5 A flowchart of a method for determining a target reference object using a large language model provided in an embodiment of the present disclosure; Figure 6-1 and Figure 6-2 A schematic diagram of a flow chart of a large language model-based intersection navigation method in an application scenario provided by an embodiment of the present disclosure; Figure 7 A structural block diagram of a large language model-based intersection navigation device provided in an embodiment of the present disclosure; Figure 8A schematic structural diagram of an electronic device suitable for executing a large language model-based intersection navigation method provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0013] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other unless there is a conflict.
[0014] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0015] Figure 1 An exemplary system architecture 100 is shown to which embodiments of the large language model-based intersection navigation method, apparatus, electronic device, and computer-readable storage medium of the present disclosure may be applied.
[0016] like Figure 1 As shown, the system architecture 100 may include a vehicle 101, an onboard terminal 1011 installed thereon, a server 102, and an intersection stop line 1031, other roads 1032 and 1033, a target road 1034 and a target reference object 1035 corresponding to the intersection scene.
[0017] The driver of vehicle 101 can use the on-board terminal 1011 to provide the driver with travel navigation planning based on an electronic map independently or in interaction with the server 102. The above-mentioned target road 1034 is the road that the vehicle 101 should drive to after leaving the intersection stop line 1031 according to the travel navigation plan. The other roads 1032 and 1033 are other roads provided at the intersection in addition to the target road 1034. The target reference object 1035 is a building used to distinguish the target road 1034 from the other roads 1032 and 1033 in the intersection scene.
[0018] Various applications for realizing information communication between the vehicle terminal 1011 and the server 102 may be installed, such as map navigation applications, large language model tool applications, instant messaging applications, etc.
[0019] The vehicle terminal 1011 can independently provide various services to the driver of the vehicle 101 through various built-in applications, and can also provide various services to the driver together with the server 102 in a communication combination. Taking a map navigation application that can provide intersection-based navigation guidance services as an example, the vehicle terminal 1011 can achieve the following effects when running the map navigation application: when it is determined that the vehicle 101 passes through an intersection according to the travel plan, a street view mark map corresponding to the intersection is extracted from the travel plan, and the street view mark map includes a first mark of the actual position of the vehicle 101 at the intersection, a second mark of the direction of the target road 1034 to be traveled to according to the travel plan, and other road 1034. 32 and 1033; then, the street view mark map, the explanation of the meaning represented by the various marks contained in the street view mark map, and the intersection guidance generation instruction are collectively input into the preset large language model as prompt information; finally, based on the output information of the large language model, the intersection navigation guidance for guiding the vehicle 101 to pass through the intersection and drive towards the target road 1034 is determined. The large language model understands the street view mark map according to the explanation under the intention expressed by the intersection guidance generation instruction and determines the target reference object 1035 for distinguishing the target road 1034 from the other roads 1032 and 1033, and generates the output information based on the actual position, the target reference object 1035 and the intention.
[0020] It should be noted that the street view marker map corresponding to each road at each intersection as the target road can be pre-stored in the vehicle terminal 1011 or the server 102, so that the matching street view marker map can be directly selected to participate in the generation of the travel navigation plan when forming the travel navigation plan.
[0021] The intersection navigation method based on the large language model provided in the subsequent embodiments of the present disclosure is generally executed independently by the vehicle terminal 1011 or jointly executed in conjunction with the server 102. Accordingly, the intersection navigation device based on the large language model can generally also be independently set in the vehicle terminal 1011 or split into the vehicle terminal 1011 and the server 102.
[0022] It should be understood that Figure 1 The number of vehicles, vehicle-mounted terminals, servers, stop lines at intersections, target roads, other roads, and target reference objects is merely illustrative. Depending on the needs of implementation, any number of vehicles, vehicle-mounted terminals, servers, stop lines at intersections, target roads, other roads, and target reference objects may be provided.
[0023] Please refer to Figure 2 , Figure 2 A flowchart of a large language model-based intersection navigation method provided in an embodiment of the present disclosure, wherein process 200 includes the following steps: Step 201: In response to a vehicle passing through an intersection according to a travel plan, extracting a street view marker corresponding to the intersection from the travel plan; This step is intended to be performed by the execution body of the intersection navigation method based on the large language model (for example, Figure 1 The vehicle-mounted terminal 1011 shown (acting independently or together with the server 102) can extract the street view marker map corresponding to the current intersection from the travel plan when it finds that the vehicle has passed through any intersection that should be passed according to the travel plan. The street view marker map is pre-formed by adding additional marks on the basis of the street view map of the corresponding intersection. The additional marks added include: a first mark for representing the actual position of the vehicle at the corresponding intersection, a second mark for representing the direction of the target road to be traveled to according to the travel plan, and a third mark for representing the direction of other roads (i.e., roads that are not target roads provided by the intersection as shown in the street view marker map). It should be understood that any road at the intersection may be used as a target road (i.e., to match the different travel needs of different users). Therefore, a corresponding street view marker map can be formed in advance for each road at each intersection when it is used as a target road, so as to facilitate the subsequent participation in the formation of travel planning.
[0024] The first mark, the second mark and the third mark can all be single-point marks, and different single-point marks can be distinguished by different shapes or colors. Since the second mark and the third mark are used to indicate the location of the target road and other roads, they can also be expressed as regional marks formed by polygons and their coverage areas, so as to better cover the road surface areas of the target road and other roads, so as to facilitate the search for target reference objects used to distinguish the target road from other roads under larger and more comprehensive regional marks.
[0025] Step 202: Inputting the street view mark map, explanations of the meanings represented by the various marks contained in the street view mark map, and instructions for generating intersection directions into a preset large language model as prompt information; Based on step 201, this step aims to have the above-mentioned execution entity input the street view mark map, the explanation of the meaning represented by the various marks contained in the street view mark map, and the intersection guidance generation instructions as prompt information into the preset large language model.
[0026] Among them, the explanation is the content that explains the actual meanings of the first mark, second mark and third mark mentioned above. For example, when the first mark is a "circular mark", the second mark is a "square mark" and the third mark is a "triangle mark", the explanation can be specific as follows: the position of the "circular mark" in the street view mark map is the vehicle waiting position at the intersection where the vehicle is located, the direction indicated by the "square mark" is the direction of the target road to be traveled to according to the travel plan, and the direction indicated by the "triangle mark" is the direction of the non-target road in the intersection.
[0027] Among them, the intersection guidance generation instruction is used to inform the large language model that it needs to understand the image content in the intersection street view mark map based on the explanation, and inform it to find reference objects that can distinguish the target road from other roads from the dimension of image features, and generate more accurate guidance content based on the reference objects to enable the driver to drive towards the target road after passing the intersection.
[0028] The prompt information generated by the street view marked image, the explanatory instructions and the intersection guidance generation instructions can be obtained by simply splicing the above three parts, especially the splicing of the explanatory instructions and the intersection guidance generation instructions (in this case, the street view marked image exists as a reference image), or a comprehensive intersection guidance generation instruction can be generated based on the content of the explanatory instructions. For example, it can be specific as: "You are a map navigation expert, please provide unconfusing intersection navigation instructions. The attached figure is a street view picture of an intersection. I am a driver. The vehicle I am driving is at the location of the circular mark in the picture. The goal is to drive towards the target road indicated by the square mark, and it cannot be confused with other roads indicated by the triangular mark. How should accurate driving direction instructions be generated?"
[0029] Furthermore, in order to make the results generated by the large language model better meet the needs of actual situations, some constraints or requirements can be added to it, such as "the generated driving direction guidance should be simple and easy to understand, and cannot be confused with other roads in similar directions."
[0030] Furthermore, to help the large language model clarify the angles from which to search for target reference objects, additional supplementary information can be provided, such as "consider using distinctive buildings, distinctive roads, or distinctive ground signs as target reference objects to enhance the description." The target reference object can include at least one of the following: a building with a first distinguishing feature, a road with a second distinguishing feature, or a ground marking with a third distinguishing feature. The first distinguishing feature includes at least one of height, number of floors, color, signboard size, signboard color, signboard text, and shape. The second distinguishing technical feature includes at least one of the number of lanes, lane type distribution, road width, slope change trend, and whether it connects to a bridge. The third distinguishing feature includes at least one of the color of a ground sign or the text of a ground sign.
[0031] Step 203: Determine intersection navigation guidance for guiding the vehicle to pass through the intersection toward the target road based on the output information of the large language model.
[0032] Building on step 202, this step aims to determine, by the execution entity, intersection navigation directions for guiding the vehicle through the intersection toward the target road based on the output information of the large language model. The large language model interprets the street view labeled image according to the intent expressed in the intersection guidance generation instruction, determines a target reference object that distinguishes the target road from other roads, and generates output information based on the actual location, the target reference object, and the intent.
[0033] For example, when the target reference object is determined by the large language model to be a tall building with a specific color and shape, the intersection navigation guidance determined based on the output information of the large language model can be: "After passing the intersection ahead, please go straight in the direction of the dark green oval-shaped tall building on the left and enter the road ahead."
[0034] For example, when the target reference object is determined by the large language model to be a target road with four lanes compared to other roads, and the two middle lanes of the four lanes are connected to overpasses facing different directions, the intersection navigation guidance determined based on the output information of the large language model can be: "When crossing the intersection ahead, please go straight toward the four-lane road with the two middle lanes connected to the overpass, and drive toward the leftmost lane that does not lead to the bridge."
[0035] For example, when the target reference object is determined by the large language model to be the yellow ground marking text on the target road compared to other roads, the intersection navigation instructions determined based on the output information of the large language model can be: "Go straight in the direction of the yellow "fast lane" written on the ground through the intersection ahead and enter the road ahead."
[0036] It should be noted that when the large language model is able to find the target reference object that distinguishes the target road from other roads from the dimension of image features in the street view mark map based on the prompt information, only the target image features that represent the distinction can be used as the target reference object to form intersection navigation guidance. If it is impossible to find a suitable target image feature as the target reference object from the dimension of image features in some special street view mark maps, the large language model can also find the target reference object that distinguishes the target road from other roads from the dimension of road network features by accessing the authorized road network database, that is, it can also use the target road network features or target road features that represent the distinction (such as at least one of the number of lanes, lane type distribution, road width, slope change trend, and whether it is connected to a bridge mentioned above) as the target reference object to form intersection navigation guidance.
[0037] The large language model-based intersection navigation method provided by the disclosed embodiments pre-associates a matched street view marker map with the intersections traversed in the travel plan. The street view marker map pre-associates the vehicle's intersection location, the target road that matches the travel plan, and the directions of other non-matching roads in the intersection street view map with corresponding types of markers. The street view marker map, explanations of the meanings represented by each type of marker, and intersection guidance generation instructions are then input into a preset large language model as prompt information. The large language model determines the intention expressed by the intersection guidance generation instructions and interprets the street view marker map based on the explanations to find a target reference object that can distinguish the target road from other roads based on the image feature dimension. Finally, the large language model outputs intersection navigation instructions based on the target reference object to guide the vehicle from the intersection to the target road. Because the intersection navigation instructions are determined based on the target reference object and have a strong association with the intersection, they have high recognition and can significantly improve the accuracy of intersection navigation instructions.
[0038] To further understand how to pre-generate Street View labeled images, please refer to Figure 3 , Figure 3 A flowchart of a method for pre-generating a street view labeled map provided in an embodiment of the present disclosure, wherein process 300 includes the following steps: Step 301: Obtain a candidate street view atlas corresponding to the intersection from a street view library; This step aims to enable the execution subject to firstly select a candidate street view atlas corresponding to the selected intersection from the street view atlas library, that is, the candidate street view atlas library contains at least one candidate street view image of the selected intersection.
[0039] Step 302: Filtering out the filtered street view images that contain all roads from the candidate street view image set; Based on step 301, this step aims to allow the aforementioned execution entity to filter out a filtered street view image that includes all roads from the candidate street view image set. Specifically, some candidate street view images that fail to simultaneously reflect all candidate roads corresponding to the intersection are removed. This ensures that the filtered street view image obtained after filtering can simultaneously reflect all drivable roads at the intersection. This allows for comprehensive identification of all other roads while identifying the target road, thereby facilitating the determination of a suitable target reference object.
[0040] Step 303: Determine the stop line at the intersection, the first position of the target road, and the second positions of other roads in the filtered street view image; Based on step 302, this step aims to determine, by the execution entity, the stop line at the intersection, the first position of the target road, and the second positions of other roads in the filtered street view.
[0041] Step 304: Determine a vehicle waiting position based on the position of the stop line in the filtered street view image, and attach a first marker to the vehicle waiting position; Building on step 303, this step involves the execution entity determining a vehicle waiting position based on the position of the stop line in the filtered street view image and attaching a first marker to the vehicle waiting position. The vehicle waiting position is typically located before the stop line (i.e., at a distance from the vehicle's end before the stop line). However, given that some street view images may not display much image information before the stop line, the vehicle waiting position can be simply located on the stop line in such cases.
[0042] Step 305: Add a second mark to the first position of the filtered street view image, and add a third mark to the second position of the filtered street view image to obtain a street view marked image.
[0043] Based on step 303, this step aims to obtain a marked street view image by having the aforementioned execution entity add a second mark to the first position of the filtered street view image and a third mark to the second position of the filtered street view image.
[0044] This embodiment provides a specific solution for pre-generating a marked street view map through steps 301 to 305. Specifically, a matching candidate street view map set is first screened for each intersection. A filtered street view map that shows all roads is then screened out. A vehicle waiting position and the orientation information of the target road and other roads are then determined by identifying stop lines in the filtered street view map. Finally, different markers are added to the vehicle waiting position, the orientation of the target road, and the orientation of other roads to obtain the marked street view map.
[0045] Based on the previous embodiment, when adding the second and third markers, a first road surface area of the target road can be first determined in the first orientation of the filtered street view image, and then the second marker can be added to the first road surface area. Similarly, a second road surface area of other roads can be first determined in the second orientation of the filtered street view image, and then the third marker can be added to the second road surface area. Specifically, the second and third markers should be avoided whenever possible from being placed on objects near the target road and other roads (e.g., signs, roadside fire hydrants, road signs, etc.). Instead, they should be placed on the road surface area of the corresponding road, so as to accurately indicate the relative positional relationship between the target road and other roads relative to the vehicle's position. This also facilitates finding target reference objects near the markers placed on the road surface area, and avoids the situation where the markers cover potential target reference objects.
[0046] Furthermore, in addition to the second and third marks mentioned above being single-point marks, they can also be regional marks. For example, the polygon that includes the first road surface area can be determined as the second mark, and the polygon that includes the second road surface area can be determined as the third mark. This better covers the road surface areas of the target road and other roads, so that it is easier to find the target reference object used to distinguish the target road from other roads under a larger and more comprehensive regional mark. When the second and third marks both use areas framed by polygons, in addition to being able to distinguish the second and third marks by the different polygonal shapes (for example, the second mark is shaped like a long strip and the third mark is shaped like an ellipse), the second mark can also be controlled to have a color feature different from that of the third mark, thereby being able to distinguish the second and third marks in terms of color.
[0047] To further understand how large language models generate output information based on input prompt information, see Figure 4 , Figure 4 This is a flowchart of a method for a large language model to generate output information based on input prompt information provided by an embodiment of the present disclosure. The execution subject of each of the following steps is the large language model. The process 400 includes the following steps: Step 401: The large language model generates instructions based on the intersection guidance in the prompt information to determine the expressed intention; This step aims to determine the intention expressed by the large language model based on the intersection guidance generation instructions in the prompt information. That is, the large language model semantically understands the intersection guidance generation instructions, and then determines the intention of "needing to understand the street view marking map in combination with explanations and find target reference objects used to distinguish the target road from other roads and form intersection navigation instructions based on the target reference objects."
[0048] Step 402: Determine the actual location of the target road, other roads, and vehicles in the street view labeled image based on the interpretation of the intent using the large language model; Based on step 401, this step aims to determine the actual position of the vehicle shown in the street view marked image and the orientation information of the target road and other roads according to the interpretation based on the determined intention by the large language model.
[0049] Step 403: The large language model determines a target reference object in the street view labeled image for distinguishing the target road from other roads based on the intent; Based on step 402, this step aims to use the large language model to determine, based on the determined intent, a target reference object in the street view labeled image used to distinguish the target road from other roads. Once the large language model has determined in step 402 which portion is the target road and which is other roads, the target reference object can be determined by comparing the differences between all image features located near the target road and those located near other roads.
[0050] Step 404: The large language model generates output information based on the actual location, target reference, and intent.
[0051] Based on step 403, this step aims to generate output information based on the actual location, target reference and intention by the large language model.
[0052] This embodiment provides a specific implementation method for a large language model to generate output information based on input prompt information through steps 401 to 404. That is, the large language model first understands the user's task intent based on the prompt information, then fully understands the key location information in the street view marker map according to the explanation under the task intent, and then searches for a target reference object that can reflect the difference when the target road and other roads are clearly identified, and finally generates the output information based on the target reference object.
[0053] To deepen the understanding of the solution for determining the target reference object mentioned in step 403 of the previous embodiment, please refer to Figure 5 , Figure 5 This is a flowchart of a method for determining a target reference object using a large language model provided by an embodiment of the present disclosure. The process 500 includes the following steps: Step 501: The large language model determines a first image feature of a target road within a preset range under the intention; Step 502: The large language model intentionally determines second image features of other roads within a preset range. In steps 501 and 502 , the large language model determines, based on the determined intention, first image features within a preset range of the target road and second image features within preset ranges of other roads.
[0054] Step 503: The large language model determines whether the first image features contain target image features that are different from the second image features. If so, execute step 506; otherwise, execute step 504. This step aims to determine by the large language model whether there is a target image feature in the first image feature that is different from the second image feature under the determined intention, that is, whether the target image feature used as the target reference can be found solely by relying on the dimension of the image feature, and there are two different processing branches based on the judgment result.
[0055] Step 504: The large language model determines a target road feature for distinguishing the target road from other roads based on the road network information; This step is based on the judgment result of step 503 that there is no target image feature that can serve as a target reference object. It is intended to use the large language model to determine the target road features used to distinguish the target road from other roads based on the road network information (such as at least one of the number of lanes, lane type distribution, road width, slope change trend, and whether it is connected to a bridge mentioned above).
[0056] Step 505: The large language model determines the reference object to which the target road feature belongs as the target reference object under the intention; Based on step 504 , this step aims to determine, by the large language model, the reference object to which the target road feature belongs as the target reference object under the determined intention.
[0057] Step 506: The large language model determines the reference object to which the target image feature belongs as the target reference object under the intention.
[0058] This step is based on the case where the judgment result of step 503 is that there is a target image feature that can serve as a target reference object, and is intended to allow the large language model to determine the reference object to which the target image feature belongs as the target reference object under the determined intention.
[0059] This embodiment shows two different processing branches through steps 501 to 506. One branch provides a method for determining the target reference object based only on the target image features determined from the dimensions of the image features. The other branch, when it is impossible to determine the target image features as the target reference object only from the dimensions of the image features, additionally provides a method for determining the target road features as the target reference object from the dimensions of the road network features or road features, so as to obtain the target reference object through as many processing methods as possible, thereby forming accurate intersection navigation guidance.
[0060] It should be understood that the two branches provided in this embodiment can both independently form different embodiments, and this embodiment exists only as a preferred implementation method of connecting the two branches in the same embodiment through a judgment step.
[0061] Based on any of the above embodiments, for the process of generating intersection navigation guidance based on the output information described in step 203, this embodiment further provides the following specific implementation method: First, a preferred guidance type is determined from the driving preferences of the vehicle driver, and the preferred guidance type may include at least one of a text presentation type, an image viewing type, and a voice broadcast type. Then, the output information of the large language model is adjusted according to the preferred guidance type to obtain intersection navigation guidance for guiding the driver to drive the vehicle through the intersection toward the target road.
[0062] For example, when the preference guidance type only includes the image viewing type, the output information can be presented in the form of a corresponding navigation guidance image or animation; when the preference guidance type includes not only the image viewing type but also the voice broadcast type, some characteristics of the voice broadcast can also be combined to generate the content to be broadcast and broadcast it at the appropriate time.
[0063] For example, when the preferred guidance type includes voice broadcast, the aforementioned execution entity may first determine a target broadcast speed and an upper limit on the number of words required for the voice broadcast. The content length of the output information of the large language model may then be adjusted based on the upper limit to obtain the content to be broadcast. Furthermore, the predetermined time required to broadcast the content to be broadcast at the target broadcast speed is determined, resulting in intersection navigation guidance that is configured to be broadcasted a predetermined time before the vehicle is expected to depart from the waiting area at the intersection. In other words, the timing of voice broadcasting should be as far in advance as possible to avoid situations where the driver has already passed the intersection and is unable to select the correct driving direction in time according to the guidance.
[0064] Furthermore, the content to be announced can be announced in voice according to the voice announcement tone selected by the driver in the past to meet the driver's tone preference.
[0065] To deepen understanding, this disclosure also provides a specific intersection navigation guidance solution based on a specific intersection scenario: For this intersection scenario, please refer to Figure 6-1 , where the white vertical line is the stop line at the intersection, and in front of the stop line at the intersection there are a right-turn road, a left-turn road and a straight road in the right front.
[0066] First, when the vehicle driven by the user is about to drive to the intersection according to the travel navigation plan, the vehicle terminal extracts the pre-built-in street view mark map taken in the direction of travel from the travel navigation plan. The street view mark map can be found in Figure 6-2 ,and Figure 6-2 is Figure 6-1 It is obtained by adding multiple star marks on the basis of the above. The star mark marked with A is used to indicate the position where the vehicle may stop at the stop line of the intersection, the star mark marked with B is used to indicate the target straight road that the vehicle should correctly drive to according to the travel navigation plan, and the star mark marked with C is used to indicate the left-turn road and the right-turn road that are different from the straight road.
[0067] Next, the vehicle terminal can call the built-in large language model based on the above content and enter the following prompt: You are a map navigation expert, and your goal is to provide unambiguous intersection navigation instructions. The image above shows a street view of an intersection. I am a driver, and my vehicle is at the five-pointed star location marked A. My goal is to drive to the road marked with the five-pointed star B, and I must not confuse it with the road corresponding to the five-pointed star C. How should I provide driving directions? The instructions should be simple and understandable, and not be confused with other roads in similar directions. Consider using building signs, building colors, building heights, road widths, and other ground marking features to enhance the description.
[0068] Afterwards, the vehicle terminal can receive the answer output by the large language model: At the intersection ahead, go straight toward the right side between the tall blue-green exterior wall (assuming it is actually this color) and the building with the "XX Hotel" sign hanging on the right side, and enter the road ahead. Do not drive into the roads on the left or right.
[0069] As this example demonstrates, the guidance is very intuitive, making it easy for drivers to understand the route and reducing the risk of making mistakes. To improve system performance, relevant text guidance can be pre-generated when generating navigation routes, avoiding the performance issues of temporary text generation. Furthermore, previously generated intersection information can be cached for use by other vehicles traveling the same route.
[0070] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a road intersection navigation device based on a large language model. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0071] like Figure 7As shown, the intersection navigation device 500 based on the large language model of this embodiment may include: a street view image extraction unit 701, a prompt information input unit 702, and an intersection navigation guide generation unit 703. Among them, the street view map extraction unit 701 is configured to extract a street view mark map corresponding to the intersection from the travel plan in response to the vehicle passing through the intersection according to the travel plan; wherein the street view mark map includes a first mark of the actual position of the vehicle at the intersection, a second mark of the direction of the target road to be traveled to according to the travel plan, and a third mark of the direction of other roads; the prompt information input unit 702 is configured to input the street view mark map, the explanation of the meaning represented by each type of mark contained in the street view mark map, and the intersection guidance generation instruction as prompt information into a preset large language model; the intersection navigation guidance generation unit 703 is configured to determine the intersection navigation guidance that guides the vehicle to pass through the intersection and travel toward the target road based on the output information of the large language model; wherein the large language model understands the street view mark map according to the explanation under the intention expressed by the intersection guidance generation instruction and determines the target reference object used to distinguish the target road from other roads, and generates output information based on the actual position, target reference object and intention.
[0072] In this embodiment, in the intersection navigation device 700 based on the large language model, the specific processing of the street view image extraction unit 701, the prompt information input unit 702, and the intersection navigation guidance generation unit 703 and the technical effects thereof can be referred to respectively. Figure 2 The relevant descriptions of steps 201-203 in the corresponding embodiment are not repeated here.
[0073] In some other optional implementations of this embodiment, the intersection navigation device 700 based on the large language model may further include: a street view labeled image generation unit configured to generate a street view labeled image. The street view labeled image generation unit may include: The candidate atlas acquisition subunit is configured to acquire a candidate street view atlas corresponding to the intersection from the street view library; a screening subunit configured to screen out a screened street view image that includes each road from the candidate street view image set; A key position information determination subunit is configured to determine a stop line at an intersection, a first position of a target road, and second positions of other roads in the filtered street view image; a first mark adding subunit configured to determine a vehicle waiting position based on a position of the stop line in the filtered street view image, and to add a first mark to the vehicle waiting position; The second and third mark adding subunits are configured to add a second mark to a first position of the filtered street view image and to add a third mark to a second position of the filtered street view image to obtain a street view mark image.
[0074] In some other optional implementations of this embodiment, the second and third mark adding subunits may include: a second marker adding module configured to determine a first road surface area of the target road at a first position of the filtered street view image and to add a second marker to the first road surface area; The third marker adding module is configured to determine a second road surface area of another road at a second position of the filtered street view image, and to add a third marker to the second road surface area.
[0075] In some other optional implementations of this embodiment, the second mark adding module may be further configured to: determining a polygon included in the first road surface area as a second mark; Correspondingly, the third marking additional module can be further configured as follows: The polygon including the second road surface area is determined as the third marker.
[0076] In some other optional implementations of this embodiment, the second mark has a color characteristic different from that of the third mark.
[0077] In some other optional implementations of this embodiment, the intersection navigation guidance generating unit 703 may include an output information generating subunit configured to generate output information based on the input prompt information using a large language model. The output information generating subunit may include: an intent determination module configured as a large language model to determine the intent expressed based on the generated instructions for the intersection in the prompt information; A location and road determination module is configured to determine the actual location of a target road, other roads, and vehicles in the street view labeled image based on the interpretation of the intent using the large language model; A target reference object determination module is configured to determine a target reference object in the street view labeled image for distinguishing a target road from other roads based on the intention of the large language model; The output information generation module is configured to generate output information based on the actual location, target reference and intention of the large language model.
[0078] In some other optional implementations of this embodiment, the target reference object determination module may be further configured to: The large language model determines the first image feature of the target road within a preset range under the intention; The large language model determines the second image features of other roads within a preset range under the intention; The large language model determines, based on the intention, a target image feature from among the first image features that is distinguishable from the second image feature; The large language model determines the reference object to which the target image feature belongs as the target reference object under the intention.
[0079] In some other optional implementations of this embodiment, the output information generating subunit may further include: a road feature determination module configured to, in response to the large language model failing to determine the target image feature under the intent, determine, based on the road network information, a target road feature for distinguishing the target road from other roads; The reference object determination module based on the road feature is configured to determine the reference object to which the target road feature belongs as the target reference object under the intention of the large language model.
[0080] In some other optional implementations of this embodiment, the target reference object includes at least one of the following: Buildings with a first distinguishing feature, roads with a second distinguishing feature, and ground markings with a third distinguishing feature. The first distinguishing feature includes: at least one of: height, number of floors, color, signboard size, signboard color, signboard text, and shape; the second distinguishing technical feature includes: at least one of: number of lanes, lane type distribution, road width, slope change trend, and whether it connects to a bridge; the third distinguishing feature includes: at least one of: ground sign color and ground sign text.
[0081] In some other optional implementations of this embodiment, the intersection navigation guidance generating unit includes: a preference guidance type determination subunit configured to determine a preference guidance type from the driving preference of the vehicle driver; wherein the preference guidance type includes at least one of a text presentation type, an image viewing type, and a voice broadcast type; The preference adjustment subunit is configured to adjust the output information of the large language model according to the preferred guidance type to obtain intersection navigation guidance for guiding the driver to drive the vehicle through the intersection toward the target road.
[0082] In some other optional implementations of this embodiment, the preference adjustment subunit is further configured to: In response to the preferred guidance type including the voice broadcast type, determining a target broadcast speech speed and an upper limit on the number of words in the voice broadcast; Adjust the content length of the output information of the large language model according to the upper limit of the number of words in the voice broadcast to obtain the content to be broadcast; A preset time required to complete the broadcast content at a target broadcast speed is determined, and intersection navigation guidance is obtained that is set to voice broadcast the broadcast content a preset time before the vehicle is expected to leave the vehicle waiting position at the intersection.
[0083] In some other optional implementations of this embodiment, the preference adjustment subunit may further include: The tone setting module is configured to voice broadcast the content to be broadcast according to the voice broadcast tone selected by the driver in the past.
[0084] This embodiment exists as an apparatus embodiment corresponding to the above-mentioned method embodiment. The large language model-based intersection navigation device provided in this embodiment pre-associates a matched street view marker map with the intersections traversed in the travel plan. The street view marker map pre-associates the vehicle's intersection location, the target road matching the travel plan, and the directions of other non-matching roads in the intersection street view map with corresponding types of markers. The large language model then inputs the street view marker map, explanations of the meanings represented by each type of marker, and intersection guidance generation instructions as prompt information into a preset large language model, so that the large language model determines the intent expressed by the intersection guidance generation instructions and, based on the intent, interprets the street view marker map according to the explanations to find a target reference object that can distinguish the target road from other roads from the dimension of image features. Ultimately, the large language model outputs intersection navigation instructions based on the target reference object to guide the vehicle from the intersection to the target road. Because the intersection navigation instructions are determined based on the target reference object and are strongly associated with the intersection, they have high recognition and can significantly improve the accuracy of intersection navigation instructions.
[0085] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the intersection navigation method based on the large language model described in any of the above embodiments can be implemented when the at least one processor executes the instructions.
[0086] According to an embodiment of the present disclosure, the present disclosure further provides a readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to implement the intersection navigation method based on a large language model described in any of the above embodiments when executed.
[0087] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product, which, when executed by a processor, can implement the intersection navigation method based on a large language model described in any of the above embodiments.
[0088] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0089] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0090] Various components in device 800 are connected to I / O interface 805, including an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0091] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the intersection navigation method based on a large language model. For example, in some embodiments, the intersection navigation method based on a large language model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the intersection navigation method based on a large language model described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the intersection navigation method based on the large language model in any other appropriate manner (eg, by means of firmware).
[0092] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0093] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0094] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0096] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0097] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host. This is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.
[0098] According to the technical solution of the embodiment of the present disclosure, a street view marker map is pre-associated with the intersections traveled in the trip plan. The street view marker map pre-associates the vehicle's intersection location, the target road that matches the trip plan, and the directions of other unmatched roads in the intersection street view map with corresponding types of markers. Then, the street view marker map, the explanation of the meaning represented by each type of marker, and the intersection guidance generation instruction are input into a preset large language model as prompt information, so that the large language model determines the intention expressed through the intersection guidance generation instruction, and understands the street view marker map according to the explanation based on the intention to find a target reference object that can distinguish the target road from other roads from the dimension of image features. Finally, the large language model can output intersection navigation guidance for guiding the vehicle from the intersection to the target road based on the target reference object. Because the intersection navigation guidance is determined based on the target reference object and has a strong association with the intersection, it has a strong recognition and can significantly improve the accuracy of intersection navigation guidance.
[0099] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0100] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A large language model-based intersection navigation method, comprising: In response to a vehicle passing through an intersection according to a travel plan, extracting a marked street view image corresponding to the intersection from the travel plan; wherein the marked street view image includes a first mark indicating the actual position of the vehicle at the intersection, a second mark indicating the direction of a target road determined to be traveled to according to the travel plan, and a third mark indicating the directions of other roads; Inputting the street view mark map, explanations of the meanings represented by various marks contained in the street view mark map, and instructions for generating road directions as prompt information into a preset large language model; Based on the output information of the large language model, intersection navigation guidance is determined to guide the vehicle to pass through the intersection toward the target road; wherein, the large language model understands the street view marker map according to the interpretation under the intention expressed by the intersection guidance generation instruction and determines the target reference object for distinguishing the target road from the other roads, and generates the output information based on the actual position, the target reference object and the intention.
2. The method according to claim 1, wherein The generation process of the street view marked map includes: Obtaining a candidate street view atlas corresponding to the intersection from a street view library; Filtering out a filtered street view image that includes all roads from the candidate street view image set; Determine a stop line at the intersection, a first position of the target road, and a second position of the other roads in the filtered street view image; determining a vehicle waiting position based on the position of the stop line in the filtered street view image, and adding the first mark to the vehicle waiting position; The second mark is added to the first position of the filtered street view image, and the third mark is added to the second position of the filtered street view image to obtain the street view marked image.
3. The method according to claim 2, wherein: Adding the second mark to the first position of the filtered street view image and adding the third mark to the second position of the filtered street view image includes: determining a first road surface area of the target road at a first position of the filtered street view image, and adding the second mark to the first road surface area; A second road surface area of the other road is determined at a second position of the filtered street view image, and the third mark is added to the second road surface area.
4. The method according to claim 3, wherein: Adding the second mark to the first road surface area includes: determining a polygon included in the first road surface area as the second mark; Correspondingly, adding the third mark to the second road surface area includes: The polygon including the second road surface area is determined as the third marker.
5. The method according to claim 3 or 4, wherein: The second marking has a color characteristic different from that of the third marking.
6. The method according to claim 1, wherein The process of generating the output information from the input prompt information by the large language model includes: The large language model determines the expressed intention based on the intersection guidance generation instruction in the prompt information; Determining the actual locations of the target road, other roads, and vehicles in the labeled street view image according to the interpretation based on the intent using the large language model; The large language model determines, based on the intention, a target reference object in the street view labeled image for distinguishing the target road from the other roads; The large language model generates the output information based on the actual location, the target reference object, and the intention.
7. The method according to claim 6, wherein: The large language model determines, based on the intention, a target reference object in the street view labeled image for distinguishing the target road from the other roads, including: The large language model determines a first image feature of the target road within a preset range under the intention; The large language model determines, based on the intention, a second image feature of the other road within a preset range; The large language model determines, based on the intention, a target image feature among the first image features that is different from the second image feature; The large language model determines the reference object to which the target image feature belongs as the target reference object under the intention.
8. The method according to claim 7, further comprising: In response to the large language model failing to determine the target image feature under the intention, the large language model determines, based on road network information, a target road feature for distinguishing the target road from the other roads; The large language model determines the reference object to which the target road feature belongs as the target reference object under the intention.
9. The method according to any one of claims 1, 6-8, wherein The target reference object includes at least one of the following: Buildings with a first distinguishing feature, roads with a second distinguishing feature, and ground markings with a third distinguishing feature, the first distinguishing feature includes: at least one of: height, number of floors, color, signboard size, signboard color, signboard text, and shape; the second distinguishing technical feature includes: at least one of: number of lanes, lane type distribution, road width, slope change trend, and whether it connects to a bridge; the third distinguishing feature includes: at least one of: ground sign color and ground sign text.
10. The method according to claim 1, wherein The determining, based on the output information of the large language model, the intersection navigation guidance for guiding the vehicle to pass through the intersection and travel toward the target road, includes: Determining a preferred guidance type from the driving preference of the driver of the vehicle; wherein the preferred guidance type includes at least one of a text presentation type, an image viewing type, and a voice broadcast type; The output information of the large language model is adjusted according to the preferred guidance type to obtain intersection navigation guidance for guiding the driver to drive the vehicle through the intersection toward the target road.
11. The method according to claim 10, wherein: The step of adjusting the output information of the large language model according to the preferred guidance type to obtain intersection navigation guidance for guiding the driver to drive the vehicle through the intersection toward the target road includes: In response to the preference guidance type including the voice broadcast type, determining a target broadcast speech rate and an upper limit on the number of words in the voice broadcast; Adjusting the content length of the output information of the large language model according to the upper limit of the number of words in the voice broadcast to obtain the content to be broadcast; A preset time required to complete the broadcast content at the target broadcast speech rate is determined, and an intersection navigation guide is obtained that is set to voice broadcast the content to be broadcast within the preset time before the vehicle is expected to leave the vehicle waiting position at the intersection.
12. The method according to claim 11, further comprising: The content to be announced is announced in voice according to the voice announcement tone selected by the driver in the past.
13. A large language model-based intersection navigation device, comprising: A street view image extraction unit is configured to, in response to a vehicle passing through an intersection according to a travel plan, extract a street view marker image corresponding to the intersection from the travel plan; wherein the street view marker image includes a first marker indicating the actual position of the vehicle at the intersection, a second marker indicating the direction of a target road determined to be traveled to according to the travel plan, and a third marker indicating the directions of other roads; a prompt information input unit configured to input the street view mark map, explanations of the meanings represented by various marks contained in the street view mark map, and instructions for generating intersection guidance as prompt information into a preset large language model; An intersection navigation guidance generation unit is configured to determine, based on the output information of the large language model, intersection navigation guidance that guides the vehicle to pass through the intersection toward the target road; wherein, the large language model understands the street view marker map according to the interpretation under the intention expressed by the intersection guidance generation instruction and determines a target reference object for distinguishing the target road from the other roads, and generates the output information based on the actual position, the target reference object and the intention.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the intersection navigation method based on a large language model according to any one of claims 1 to 12.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the intersection navigation method based on a large language model according to any one of claims 1 to 12.
16. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the intersection navigation method based on a large language model according to any one of claims 1 to 12.