Basic model for semantic route planning

By training and fine-tuning a machine learning-based semantic route planning model, the problems of poor performance and high resource consumption in map manipulation applications when generating sightseeing routes and mixed traffic routes are solved, achieving more efficient and accurate route planning.

CN121941901APending Publication Date: 2026-04-28GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2024-09-13
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing map manipulation applications struggle to meet users' semantic needs, particularly exhibiting poor performance and high resource consumption when generating sightseeing routes and mixed transportation routes.

Method used

By training a machine learning-based semantic route planning model, using a large dataset for training and fine-tuning, and adjusting the model parameters to optimize route planning, more accurate and efficient route planning information can be generated.

Benefits of technology

It significantly improves the efficiency of map operation applications, reduces user travel time and resource consumption, and can better understand the semantic meaning of user requests to generate more accurate route planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121941901A_ABST
    Figure CN121941901A_ABST
Patent Text Reader

Abstract

Training data is obtained. The training data includes: (a) route information indicating a route from a starting location to a destination location, where the route includes a plurality of route segments, the plurality of route segments including a first subset of route segments and a second subset of route segments; and (b) route characteristic information describing one or more route characteristics. At least the first subset of route segments and a portion of the route characteristic information associated with the first subset of route segments are processed using a machine-learned semantic route planning model to obtain one or more predicted route segments for a second subset of route segments. One or more parameters of the machine-learned semantic route planning model are adjusted based on an optimization function that evaluates differences between the one or more predicted route segments and the second subset of route segments.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority requirements

[0002] This application is based on and claims priority to U.S. non-provisional application 18 / 468,338, filed on September 15, 2023, which is incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to semantic routing. More specifically, this disclosure relates to underlying machine learning models trained for semantic understanding of map manipulation information. Background Technology

[0004] Base models, such as Large Language Models (LLMs), are models with a large number of parameters that are trained on large datasets to perform multiple tasks. Base models are currently revolutionizing the capabilities of assistive AI techniques in many domains, ranging from everyday conversational assistants to multimodal editing of content such as audio, images, or video. Once trained, base models can perform a wide range of tasks within the context of the training data used to train them. For example, once trained, LLMs can perform a wide variety of language tasks. Summary of the Invention

[0005] Aspects and advantages of embodiments of this disclosure will be set forth in part in the description which follows, or may be learned from the description or by practice of the embodiments.

[0006] One example aspect of this disclosure relates to a computer-implemented method. The method includes obtaining training data by a computing system comprising one or more computing devices, the training data including: (a) route information indicating a route from a starting location to a destination location, wherein the route comprises a plurality of route segments, the plurality of route segments comprising a first subset of route segments and a second subset of route segments; and (b) route characteristic information describing one or more route characteristics. The method includes having the computing system process at least a first subset of route segments and a portion of the route characteristic information associated with the first subset of route segments using a machine learning-based semantic route planning model to obtain one or more predicted route segments for the second subset of route segments. The method includes having the computing system adjust one or more parameters of the machine learning-based semantic route planning model based on an optimization function that evaluates the differences between one or more predicted route segments and the second subset of route segments.

[0007] Another example aspect of this disclosure relates to a computing system comprising: one or more processor devices; and a memory. The memory includes a machine learning semantic route planning model, wherein the machine learning semantic route planning model is trained to process map operation information to generate model output, the model output including suggested route segments and / or information associated with the route segments. The memory includes one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to operate. The operation includes: obtaining one or more inputs from a client computing device for the machine learning semantic route planning model, wherein the one or more inputs include at least one of: request information indicating a requested route segment and / or a request for map operation-related information; or route characteristic information indicating one or more route characteristics. The operation includes processing the one or more inputs to obtain a model output, wherein the model output includes at least one of: (a) route planning information indicating a route including one or more suggested route segments; or (b) semantic map operation information associated with the route. The operation includes providing the model output to the client computing device.

[0008] Another exemplary aspect of this disclosure relates to one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to operate. The operation includes obtaining training data comprising: (a) route information indicating a route from a starting location to a destination location, wherein the route comprises a plurality of route segments, the plurality of route segments comprising a first subset of route segments and a second subset of route segments; and (b) route characteristic information describing one or more route characteristics. The operation includes utilizing a machine learning-based semantic route planning model to process at least the first subset of route segments and a portion of the route characteristic information associated with the first subset of route segments to obtain one or more predicted route segments for the second subset of route segments. The operation includes adjusting one or more parameters of the machine learning-based semantic route planning model based on an optimization function that evaluates the differences between the one or more predicted route segments and the second subset of route segments.

[0009] Other aspects of this disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0010] These and other features, aspects, and advantages of the various embodiments of this disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate exemplary embodiments of the disclosure and, together with the description, serve to explain the relevant principles. Attached Figure Description

[0011] Referring to the accompanying drawings, a detailed discussion of embodiments is set forth in this specification for those skilled in the art, in which:

[0012] Figure 1A A block diagram of an example computational system for training and utilizing a machine learning semantic route planning model according to some implementations of this disclosure is depicted.

[0013] Figure 1B A block diagram depicts an example computing device for training and / or pre-training a machine learning semantic route planning model according to some implementations of this disclosure.

[0014] Figure 1C A block diagram of an example computing device is depicted that utilizes a machine learning-based semantic route planning model to generate map operation information or map operation-related information according to some implementations of this disclosure.

[0015] Figure 2A A data flow diagram for pre-training a machine learning semantic route planning model is depicted according to some implementations of this disclosure.

[0016] Figure 2B The present disclosure describes some implementations of the method for use with Figure 2A Different route information is used to create a data flow diagram for subsequent pre-training iterations of the machine learning semantic route planning model.

[0017] Figure 3 A block diagram depicts an example of a machine learning semantic route planning model 300 based on some implementations of this disclosure.

[0018] Figure 4 A flowchart is depicted illustrating example methods for training and / or fine-tuning a machine learning semantic route planning model, according to some implementations of this disclosure.

[0019] Figure 5 A flowchart is provided illustrating an example method for performing map-related tasks using a machine learning-based semantic route planning model, according to some implementations of this disclosure.

[0020] The repeated reference numerals across multiple figures are intended to identify the same features in various implementations. Detailed Implementation

[0021] Overview

[0022] Generally, this disclosure relates to semantic route planning. More specifically, this disclosure relates to a basic machine learning model trained for semantic understanding of map operation information. In particular, training data can be obtained to train a basic map operation model, such as a machine learning semantic route planning model. The training data may include route information indicating a route, said route comprising multiple route segments (e.g., real routes previously requested and taken by the user). The training data may also include route characteristic information. Route characteristic information may be metadata associated with the route (e.g., preferred route type, mode of transportation, contextual information, etc.). For example, if the route indicated by the route information is a route previously provided to the user, the route characteristic information may include intermediate locations within the route, entities located along the route (e.g., businesses, points of interest (POIs), landmarks, etc.).

[0023] Training data can be used to train or fine-tune a machine learning semantic route planning model. This model can be a foundational model for map-related tasks (e.g., generating routes and / or route segments, semantically analyzing routes, etc.), comprising a large number of parameters and trained on a large corpus of training data (e.g., map manipulation data). To train or fine-tune the model, some route segments closer to the destination can be masked. Training data can be fed into the model to obtain its output. The output can include predicted route segments for the masked segments. An optimization function evaluating the difference between the predicted and masked route segments can be used to tune one or more parameters of the model. In this way, the model can be trained to perform multiple map manipulation tasks.

[0024] This disclosure provides several technical effects and benefits. As an example, the implementation of this disclosure can significantly improve the efficiency of map manipulation applications. In particular, conventional map manipulation applications are generally difficult to navigate and offer users relatively few input options indicating specific routes or routes with specific characteristics. For example, few (if any) map manipulation applications provide users with the ability to request routes that pass through scenic routes along the coast. However, the implementation of this disclosure provides a basic model for generating route planning information based on a semantic understanding of user requests. In this way, the semantic meaning of user requests can be satisfied to provide more accurate and efficient routes for map manipulation applications.

[0025] As another example of technical effect and benefit, conventional map operations applications can sometimes generate optimal routes between starting and ending points, but exhibit relatively poor performance when adjustments to the route are needed while traveling along it, or when generating routes for multimodal traffic (e.g., a mixture of public and private traffic). However, the implementation of this disclosure can predict more efficient route segments of routes already in progress. This, in turn, can significantly reduce the user's travel time and thus significantly reduce the resource consumption required to traverse the route (e.g., energy resources for autonomous vehicles, computing resources for onboard computers, fuel resources for conventional vehicles, etc.).

[0026] Exemplary embodiments of this disclosure will now be discussed in further detail with reference to the accompanying drawings.

[0027] Example devices and systems

[0028] Figure 1A A block diagram of an example computing system 100 for training and utilizing a machine learning semantic route planning model according to some implementations of this disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0029] User computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet), a game console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0030] User computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 114 can store data 116 and instructions 118, which are executed by processor 112 to operate user computing device 102.

[0031] In some implementations, the user computing device 102 may store or include one or more machine learning semantic route planning model models 120. For example, the machine learning semantic route planning model model 120 may be, or may otherwise include, various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., transformer models).

[0032] Specifically, in some implementations, the machine learning semantic route planning model 120 may be, or otherwise include, a part of a model trained to process text content or a part of a model trained to process text content. For example, a model (such as an LLM) is trained to generate text content based on text input. For this purpose, an LLM typically includes a decoder part (and in some instances, and / or a decoder section) that processes the text input to generate an intermediate representation of the text input (e.g., a sequence of lexical units, etc.). This intermediate representation is further processed to ultimately generate text output. However, in some implementations, the decoder part of the LLM may be included in the machine learning semantic route planning model 120, such that text input from the user (e.g., “I want a scenic route to the grocery store that avoids highways”) can be processed along with regular map operation information.

[0033] Similarly, in some implementations, the machine learning semantic route planning model 120 may include encoder or decoder portions from other base models. For example, the machine learning semantic route planning model 120 may include encoder and / or decoder portions from a base computer vision model trained to generate intermediate representations of image data (e.g., video, still images, renderings, textures, etc.). As another example, the machine learning semantic route planning model 120 may include encoder and / or decoder portions from a base audio model trained to generate intermediate representations of audio data (e.g., recordings, speech, etc.). In this way, the machine learning semantic route planning model 120 can utilize multimodal inputs to generate routes more accurately for the user.

[0034] It should be noted that specific base models or parts of base models (e.g., encoder and / or decoder parts, etc.) are discussed as being included in the machine learning semantic route planning model 120. However, in some implementations, such base models or model parts may be stored and instantiated separately from the machine learning semantic route planning model 120. More generally, such base models or model parts may be used in conjunction with the machine learning semantic route planning model 120, but may also be instantiated, executed, stored, etc., separately from the machine learning semantic route planning model 120. Specific parts and layers of the machine learning semantic route planning model 120 will be discussed in further detail in the specification.

[0035] In some implementations, one or more machine learning semantic route planning models 120 may be received from server computing system 130 via network 180, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 may implement multiple parallel instances of a single machine learning semantic route planning model 120 (e.g., parallel semantic route planning across multiple instances of machine learning semantic route planning model 120).

[0036] The machine learning semantic route planning model 120 can perform a wide variety of map operation-related tasks. More specifically, the machine learning semantic route planning model 120 can process map operation information and / or additional map-related information to generate map operation-related model outputs. In some implementations, the machine learning semantic route planning model 120 can process model inputs to generate predicted route segments. These model inputs may include previous or current route segments, information indicating a route planning request from a user, image data depicting the location of the requested route, information describing multiple possible locations, contextual information (e.g., information indicating available transportation options, user preferences, weather conditions, etc.), trip requests, or any other type or manner of input related to map operations.

[0037] Additionally or separately, in some implementations, the machine learning semantic route planning model 120 can process model input to generate semantic map operation information. Generally, semantic map operation information can refer to information describing a semantic understanding of routes, geographic locations, points of interest, etc. For example, model input can be a query requesting the route segment most likely a user will take when searching for a restaurant. Model output can be semantic map operation information including predicted route segments corresponding to the query. As another example, model input can be a query used to identify characteristics of users navigating a specific route segment. Semantic map operation information can indicate common characteristics of users navigating a specific route segment (e.g., type of vehicle used, mode of transportation used, consumer preferences, etc.). As another example, model input can be a request to identify specific characteristics of a route, and semantic map operation information can describe the characteristics of a route or route segment (e.g., sightseeing, dangerous in weather conditions, adjacent to a specific POI, low traffic, etc.). Furthermore, model input can be a request to describe a specific route or route segment, and semantic map operation information can include text content, image content, and / or audio content describing the route or route segment.

[0038] Additionally or alternatively, one or more machine learning semantic route planning models 140 may be included in or otherwise stored and implemented by server computing system 130, which communicates with user computing system 102 according to a client-server relationship. For example, the machine learning semantic route planning model 140 may be implemented by server computing system 130 as part of a web service (e.g., a map manipulation service). Thus, one or more models 120 may be stored and implemented at user computing device 102, and / or one or more models 140 may be stored and implemented at server computing system 130.

[0039] User computing device 102 may also include one or more user input components 122 that receive user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display or touchpad) sensitive to the touch of a user input object (e.g., a finger or stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which the user can provide input.

[0040] Server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138, which are executed by processor 132 to operate server computing system 130.

[0041] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by said one or more server computing devices. In instances where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0042] As described above, server computing system 130 may store or otherwise include one or more machine learning semantic route planning models 140. For example, model 140 may be, or may otherwise include, various machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., transformer models).

[0043] User computing device 102 and / or server computing system 130 can interactively train models 120 and / or 140 via training computing system 150, which is communicatively coupled to network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.

[0044] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., processor core, microprocessor, ASIC, FPGA, controller, microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158, which are executed by the processor 152 to operate the training computing system 150. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by said one or more server computing devices.

[0045] It should be noted that, generally, as described in this article, “training” a model can also refer to “fine-tuning” or additional training / tuning iterations of the model. For example, an initial training session can be performed to train the model and lock its parameters. After utilizing the model for a specific period of time during the inference phase, additional tuning iterations can be performed to “fine-tune” the model for additional tasks or to improve its performance in the current task.

[0046] The training computing system 150 may include a model trainer 160 that uses various training or learning techniques, such as, for example, error backpropagation, to train machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters over multiple training iterations. Alternatively, gradient-free methods may be used to update the parameters.

[0047] In some implementations, error backpropagation can include truncated backpropagation through time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0048] Specifically, model trainer 160 can train machine learning semantic route planning models 120 and / or 140 based on a set of training data 162. Training data 162 can include training data for pre-training and fine-tuning models 120 and / or 140. For example, training data can include multiple sets of pre-training data. Pre-training pairs can include route segments of routes previously traveled by the user of the map operation service or simulated routes traveled by the simulated user. Pre-training pairs can also include route characteristic information. Route characteristic information can be or otherwise include metadata associated with the route traveled by the user. Route characteristic information can indicate the mode of transportation used, total time, traffic flow information, entities located along or near the route (e.g., POIs, businesses, landmarks, residences, etc.), the initial request made by the user, geographic location information (e.g., coordinates of the starting and destination locations), etc.

[0049] In some implementations, pre-training pairs can be created dynamically. Specifically, N-1 training pairs can be generated using a single route comprising N segments, and these N-1 training pairs can be utilized in various ways. For example, given 1, 2, ..., K, a training task could be performed to predict K+1, ..., N (continuous) route segments. Another example is the training task to predict K (mask segments) given 1, 2, ..., K+1, ..., N.

[0050] Model trainer 160 can train machine learning semantic route planning models 120 and / or 140 by masking specific route segments of a pre-trained pair and then processing the pre-trained pair with machine learning semantic route planning models 120 and / or 140 to obtain model output. The model output may include predicted route segments for the masked route segments. Model trainer 160 can then train and evaluate an optimization function that evaluates the difference between the masked route segments and the predicted route segments. Based on the optimization function, the model trainer can adjust the values ​​of the parameters of machine learning semantic route planning models 120 and / or 140.

[0051] Training data 162 may also include fine-tuning data to fine-tune the machine learning semantic route planning models 120 and / or 140. The type or manner of fine-tuning data included in training data 162 may vary depending on the task to which the machine learning semantic route planning models 120 and / or 140 are being fine-tuned. For example, suppose the machine learning semantic route planning models 120 and / or 140 have been pre-trained and are being fine-tuned to generate semantic text descriptions of routes. The set of fine-tuning data may include route planning information indicating sightseeing routes along a coastline, as well as baseline real text descriptions of the routes. Model trainer 160 may utilize the machine learning semantic route planning models 120 and / or 140 to process the route planning information to generate model outputs that include predicted text descriptions of the routes. Model trainer 160 may train the machine learning semantic route planning models 120 and / or 140 based on an optimization function that evaluates the difference between the baseline real text descriptions and the predicted text descriptions.

[0052] In some implementations, training examples can be provided by the user computing device 102 if the user has already given consent. Therefore, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process may be referred to as model personalization.

[0053] Model trainer 160 includes computer logic for providing desired functionality. Model trainer 160 may be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as RAM, a hard disk, or an optical or magnetic medium.

[0054] Network 180 can be any type of communication network, such as a local area network (e.g., intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over network 180 can be carried via any type of wired and / or wireless connection using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).

[0055] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.

[0056] In some implementations, the input to the machine learning model of this disclosure can be image data. The machine learning model can process the image data to generate output. As an example, the machine learning model can process image data to generate image recognition output (e.g., image data identification, latent embedding of image data, encoded representation of image data, hashing of image data, etc.), which can be further used to generate route planning information, map-related information, etc. As another example, the machine learning model can process image data to generate prediction output, which can be further used to generate route planning information, map-related information, etc.

[0057] In some implementations, the input to the machine learning model of this disclosure can be text or natural language data. The machine learning model can process the text or natural language data to generate output. As an example, the machine learning model can process text or natural language data to generate latent text embedding output, which can be further used to generate route planning information, map-related information, etc. As another example, the machine learning model can process text or natural language data to generate translation output, which can be further used to generate route planning information, map-related information, etc. As yet another example, the machine learning model can process text or natural language data to generate semantic intent output, which can be further used to generate route planning information, map-related information, etc.

[0058] In some implementations, the input to the machine learning model of this disclosure can be speech data. The machine learning model can process the speech data to generate output. As an example, the machine learning model can process speech data to generate speech recognition output, which can be further used to generate route planning information, map-related information, etc. As another example, the machine learning model can process speech data to generate speech translation output, which can be further used to generate route planning information, map-related information, etc. As yet another example, the machine learning model can process speech data to generate latent embedding output. As yet another example, the machine learning model can process speech data to generate text representation output (e.g., a text representation of the input speech data), which can be further used to generate route planning information, map-related information, etc.

[0059] In some implementations, the input to the machine learning model of this disclosure can be latent encoded data (e.g., an input latent spatial representation, etc.). The machine learning model can process the latent encoded data to generate an output. As an example, the machine learning model can process the latent encoded data to generate an identification output, which can be further used to generate route planning information, map-related information, etc. As another example, the machine learning model can process the latent encoded data to generate a prediction output, which can be further used to generate route planning information, map-related information, etc.

[0060] In some implementations, the input to the machine learning model of this disclosure can be sensor data. The machine learning model can process the sensor data to generate output. As an example, the machine learning model can process sensor data to generate identification output, which can be further used to generate route planning information, map-related information, etc. As another example, the machine learning model can process sensor data to generate visualization output, which can be further used to generate route planning information, map-related information, etc.

[0061] In some cases, the input includes visual data, and the task is a computer vision task, which can be further used to generate route planning information, map-related information, etc. In other cases, the input includes pixel data from one or more images, and the task is an image processing task. For example, an image processing task could be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the probability that one or more images depict an object belonging to that object class. An image processing task could be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the probability that the region depicts an object of interest. As another example, an image processing task could be image segmentation, where the image processing output defines the probability of each category in a predetermined set of categories for each pixel in one or more images. For example, the set of categories could be foreground and background. As another example, the set of categories could be object classes. As another example, an image processing task could be depth estimation, where the image processing output defines a corresponding depth value for each pixel in one or more images. As another example, an image processing task could be motion estimation, where the network input includes multiple images, and the image processing output defines the motion of the scene depicted at that pixel between the images in the network input for each pixel in one of the input images.

[0062] In some cases, the input includes audio data representing spoken utterance, and the task is a speech recognition task, which can be further used to generate route planning information, map-related information, etc. The output may include text output mapped to spoken utterance. In some cases, the task includes microprocessor performance tasks, such as branch prediction or memory address translation.

[0063] Figure 1A An example computing system that can be used to implement this disclosure is shown. Other computing systems may also be used. For example, in some implementations, user computing device 102 may include a model trainer 160 and a training dataset 162. In such implementations, model 120 can be trained locally on user computing device 102 and both can be used. In some of such implementations, user computing device 102 may implement model trainer 160 to personalize model 120 based on user-specific data.

[0064] Figure 1B A block diagram of an example computing device 10 for training and / or pre-training a machine learning semantic route planning model according to some implementations of this disclosure is depicted. The computing device 10 may be a user computing device or a server computing device.

[0065] The computing device 10 includes multiple applications (e.g., application 1 to application N). Each application contains its own machine learning library and machine learning model. For example, each application may include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.

[0066] like Figure 1B As shown, each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can use an API (e.g., a public API) to communicate with each device component. In some implementations, the API used by each application is application-specific.

[0067] Figure 1C A block diagram of an example computing device 50 is depicted, which utilizes a machine learning-based semantic route planning model to generate map operation information or map operation-related information according to some implementations of this disclosure. The computing device 50 may be a user computing device or a server computing device.

[0068] The computing device 50 includes multiple applications (e.g., application 1 to application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application may use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).

[0069] The central intelligence layer comprises multiple machine learning models. For example, such as... Figure 1C As shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model is managed by a central intelligent layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligent layer can provide a single model for all applications. In some implementations, the central intelligent layer is included within the operating system of the computing device 50 or otherwise implemented by the operating system.

[0070] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository for computing device 50. For example... Figure 1C As shown, the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer may use an API (e.g., a private API) to communicate with each device component.

[0071] Figure 2A A data flow diagram 200A for pre-training a machine learning semantic route planning model is depicted according to some implementations of this disclosure. Specifically, data flow diagram 200 includes training data 202. Training data 202 may be, or otherwise include, data used for pre-training the machine learning semantic route planning model. Training data 202 may include route information 204. Route information 204 may indicate a route from a starting point to a destination. In some implementations, route information 204 may indicate a route previously traveled by a user (such as a user of a map operation application). In this instance, route information 204 is strictly processed to remove any identifying information in order to protect user privacy. Alternatively, in some implementations, route information 204 may indicate a simulated route, or may include media and / or multimedia that generally describe a route. Alternative sources and / or types of route information 204 will vary depending on the implementation. Figure 2B Let's discuss this in more detail.

[0072] Specifically, route information 204 may indicate a route comprising multiple route segments 206A, 206B, 206C, and 206D (generally route segment 206). Each of route segments 206 may include, or otherwise indicate, a starting point, a destination, a mode of transport, the amount of time taken to navigate the route segment, and any other type or manner of information relating to the specific route segment 206 indicated by route information 204. As depicted in the example, route segment 206A may be the first route segment, and therefore, the starting point of route segment 206A may be the same starting point as the route indicated by route information 204. Route segment 206A may have intermediate destinations different from the final destination of the route indicated by route information 204. In some implementations, the intermediate destination of a route segment may be the intermediate starting point of the next route segment immediately following the current route segment. As illustrated in the example, the intermediate destination location of route segment 206A (e.g., 35.78, -78.81) can be the same as the starting intermediate starting location of route segment 206B that immediately follows route segment 206A (e.g., 35.78, -78.81).

[0073] As depicted, in some implementations, route segment 206 may each use latitude / longitude coordinates to indicate a location (e.g., starting location, destination location, etc.). Additionally or alternatively, in some implementations, route segment 206 may indicate a location in an alternative manner. For example, route segment 206 may indicate the location based on the name of an entity associated with the location (e.g., a specific landmark, POI, business, residence, etc.). As another example, route segment 206 may encode the location as an address.

[0074] Additionally or alternatively, as depicted, in some implementations, route segment 206 may indicate a mode of transportation for traversing a particular route segment. For example, route segment 206A indicates the use of a private vehicle (e.g., PRIV_VEH) to traverse the route segment, while route segment 206D indicates the use of public transportation (e.g., PUB_TRANS) to traverse the route segment.

[0075] Additionally or alternatively, as depicted, in some implementations, route segment 206 may indicate the time spent traversing the route segment. For example, route segment 206A indicates that it took 17 minutes and 35 seconds to traverse the route segment. It should be noted that the time spent traversing the route segment can be calculated such that only the time spent traversing the segment directly is evaluated. In other words, the time spent traversing the route segment can exclude the time spent by the user doing other things besides traversing the route segment (e.g., stopping to eat or refuel, shopping, taking detours, etc.).

[0076] The route indicated by route information 204 and route segment 206 can utilize any type or manner of transportation infrastructure, recreational resources, or infrastructure. In some implementations, the route indicated by route information 204 may be a route traversing roads (e.g., streets, highways, bridges, etc.). Alternatively, in some implementations, the route indicated by route information 204 may traverse additional transportation infrastructure inaccessible to conventional ground-based public or private vehicles (such as cars, trucks, buses, etc.). For example, the route or route segment may traverse sidewalks, bike paths, hiking trails, alleyways, buildings (e.g., access to a building via a connector such as an overpass or subway and access to another building via that connector), tunnels, bodies of water, airspace, etc. Alternatively, in some implementations, the route or route segment may be traversed by vehicles other than conventional ground-based vehicles. For example, a route or section of a route can be traversed by trains, subways, trams, personal transportation devices (e.g., bicycles, scooters, unicycles, electronic devices, etc.), ships, ferries, airplanes, helicopters, vertical takeoff and landing (VTOL) vehicles, etc.

[0077] Additionally or alternatively, in some implementations, a route or segment of a route may traverse recreational resources or infrastructure. For example, a route or segment of a route may include ski trails, hiking trails, cycling trails, campsites, etc. Additionally or alternatively, in some implementations, a route or segment of a route may traverse virtual resources or infrastructure. For example, a route or segment of a route may traverse virtual spaces such as video games, simulations, etc. In particular, a route or segment of a route may traverse such spaces in the same manner as corresponding non-virtual transportation resources. For example, suppose a virtualized fictional city in a video game includes transportation infrastructure similar to that of a modern city (e.g., roads, public transportation, bike paths, etc.). A route or segment of a route may traverse the transportation infrastructure of a virtualized fictional city in the same manner as traversing a modern city.

[0078] Additionally or alternatively, in some implementations, a route or segment of a route may utilize, combine with, or otherwise connect with transportation service providers. For example, a segment of a route traversing city roads may be traversed using ride-sharing services, taxis, etc. Another example is a segment of a route traveling from one city to another using commercial flight services. Yet another example is a segment of a route traveling from one city to another using public transportation (e.g., buses, trains, etc.).

[0079] Additionally or alternatively, in some implementations, the route can be a multimodal route, where different segments of the route traverse different types of transport infrastructure and / or utilize different modes of transport. For example, the first segment of the route might traverse a road to reach a train station, and then the second segment might utilize a train to traverse a railway line. Or, for instance, the first segment of the route might utilize a bicycle to traverse a recreational mountain biking trail to reach a parking lot where the user's personal vehicle is parked, and then the second segment might utilize the personal vehicle to return to the user's home.

[0080] Training data 202 may include route characteristic information 208. Route planning characteristic information 208 may be, or otherwise include, metadata or meta-information indicating various segment-level and / or route-level characteristics of the route indicated by route information 204. As illustrated in the example, route characteristic information 208 may include weather characteristics indicating weather conditions when traversing each of route segments 206. Alternatively, route characteristic information 208 may include an initial request from a user causing the generation of a route indicated by route information 204 (e.g., “From home to RDU … nohighway”). As another example, route characteristic information 208 may include traffic flow characteristics indicating traffic flow conditions when traversing each of route segments 206. Therefore, it should be broadly understood that route characteristic information 208 may include any type or manner of information associated with the route indicated by route information 204 or traversing said route on a segment-by-segment or route-by-route basis.

[0081] In some implementations, route characteristic information 208 may include images associated with route segment 206. For example, vehicles traversing route segments 206A and 206B may include camera sensors (e.g., to facilitate autonomous vehicle operation). Route characteristic information 208 may include images captured at the start and / or destination locations of route segments 206A and 206B using the vehicle's camera sensors. Alternatively, route characteristic information 208 may include publicly available street view imagery, traffic flow sensor imagery, satellite imagery, etc., of the start and / or destination locations of route segment 206.

[0082] In some implementations, the route characteristic information 208 may include historical user information specific to users traversing the route indicated by the route information 204. After being processed to remove any identifying information to protect user privacy, the historical user information may indicate various characteristics of the user that can be associated with traversing route segment 206. For example, the historical user information may display the language spoken by the user. As another example, the historical user information may indicate the available modes of transportation for the user (e.g., PRIV_VEH, PUB_TRANS, SELF, etc.) and / or whether the user is capable of walking or using walking-assisted transportation (e.g., bicycle, electric device, mobility device, etc.). Furthermore, the historical user information may include user information embeddings that serve as a potential representation of user information related to map operations and / or route traversal. In this way, the historical user information can capture potentially relevant or potentially irrelevant details in a privacy-preserving manner.

[0083] In some implementations, route characteristic information 208 can be enhanced using additional route characteristic information 209. Specifically, data enhancer 211 can obtain additional route characteristic information 209 and can use it to enhance route characteristic information 208. Additional route characteristic information 209 may include information related to route segment 206, entities located along route segment 206 (e.g., businesses, POIs, landmarks, etc.), user-generated content associated with route segment 206, etc. For example, additional route characteristic information 209 may include user-submitted reviews of businesses along a specific route segment 206. In some implementations, additional route characteristic information 209 may include user reviews specifically for route segment 206. For example, if route segment 206 is recreational (e.g., hiking trails, ski trails, scenic roads, etc.), additional route characteristic information 209 may include reviews of route segment 206 from a user review web source. In some implementations, the additional route characteristic information 209 may include information from map operation services, applications, sources, etc., that are different from the map operation service associated with the same route information 204. For example, the additional route characteristic information 209 may include route metadata from another map operation application or service, a different geographic information system (GIS), a comment platform, etc.

[0084] It should be noted that route characteristic information 208 is shown only as different from route information 204 to more clearly illustrate the various implementations of this disclosure. Conversely, in some implementations, route information 204 may include specific characteristics shown as included in route characteristic information 208 (e.g., segment weather characteristics, segment traffic flow characteristics, etc.), or may be entirely integrated with route characteristic information 208. Additionally or alternatively, in some implementations, route characteristic information 208 may include characteristics or information shown as included in route segment 206 (e.g., traffic mode, time of travel through said segment, origin / destination location, etc.).

[0085] Training data 202 can be used to train and / or fine-tune the machine learning semantic route planning model 210. The machine learning semantic route planning model 210 can be any type or manner of model that is at least partially trained to process route information 204 and route characteristic information 208, as well as data of any type or manner included in the route information 204 and route characteristic information 208. To train and / or fine-tune the machine learning semantic route planning model 210, one or more route segments 206 can be masked before processing using the machine learning semantic route planning model 210. In some implementations, training data 202 can be obtained after masking has been applied to it. Alternatively, in some implementations, training data 202 can be masked after it has been obtained.

[0086] As illustrated in the example, route segment 206C can be masked by blurring or otherwise removing information associated with it. In some implementations, route characteristic information 208 associated with the masked route segment 206C can also be masked. For example, segment-by-segment traffic flow characteristics of route segment 206C can be masked. Alternatively, in some implementations, the route characteristic information 208 associated with the masked route segment 206C can remain unmasked or can be partially masked. In some implementations, a single route segment 206 can be masked. Alternatively, in some implementations, multiple route segments 206 can be masked (e.g., sequential route segments, non-sequential route segments, etc.).

[0087] Training data 202 can be processed using a machine learning semantic route planning model 210 to obtain a predicted route segment 212. The predicted route segment 212 can be a prediction for a masked route segment 206C. A model trainer 214 can evaluate the predicted route segment 212 using an optimization function 216. Specifically, the optimization function 216 can evaluate the difference between the predicted route segment 212 and a baseline true route segment 218. The baseline true route segment 218 can be the route segment 206C before the mask is applied (i.e., the "original" route segment 206C). Based on the optimization function 216, the model trainer 214 can apply adjustments 220 to the parameters (or values ​​of the parameters) of the machine learning semantic route planning model 210. In this way, the machine learning semantic route planning model 210 can be trained or pre-trained to efficiently and accurately predict route segments given a route and information associated with said route.

[0088] Figure 2B The present disclosure describes some implementations of the method for use with Figure 2A The data flow diagram 200B shows the different route information used to perform subsequent pre-training iterations on the machine learning semantic route planning model. In particular, a second set of training data 222 can be obtained for another iteration of fine-tuning, training, or pre-training the machine learning semantic route planning model 210.

[0089] Training data 222 may include route information 224. Unlike the route information 204 of training data 202, the route information 224 of training data 222 may be, or otherwise include, content that semantically indicates a route. As depicted in the example, route information 224 may include a file from a book titled “My Travels Through Vietnam” (e.g., MY_TRAVELS_THROUGH_VIETNAM.epub) and corresponding images extracted from said book (e.g., VIETNAM_IMG_1.png, VIETNAM_IMG_2.png, etc.). It should be noted that although route information 224 does not explicitly delineate specific route segments, such as… Figure 2A The route information 204 contains route segments 206, but the machine learning semantic route planning model 210 can be trained and / or fine-tuned to identify and segment routes semantically indicated by the route information 224.

[0090] Following the illustrated example, suppose the book content included in route information 224 describes a vacation during which the author visited four major cities in Vietnam. The four images included in route information 224 correspond to the four major cities respectively. A machine learning semantic route planning model 210 can process training data 222 to generate four predicted route segments 226. The predicted route segments 226 can correspond to routes between the four major cities visited by the author. In this way, the machine learning semantic route planning model 210 can be trained to recognize the characteristics of routes semantically described in media or multimedia. The machine learning semantic route planning model 210 can also be trained to divide the identified routes into specific segments.

[0091] Model trainer 214 can utilize optimization function 216 to evaluate the predicted route segment 226. Optimization function 216 can evaluate the difference between the predicted route segment 226 and the baseline true route segment 228. The baseline true route segment 228 can be a route segment that accurately corresponds to the route described by the author in route information 224. For example, the baseline true route segment 228 can be created by a user who has read the books included in route information 224. Alternatively, the baseline true route segment 228 can be generated as the output of an LLM that processes at least some of the books included in route information 224, or generated based on the output of said LLM. Based on optimization function 216, model trainer 214 can apply adjustment 230 to the values ​​of the parameters of the machine learning semantic route planning model 210.

[0092] although Figure 2B Books and corresponding images are represented as route information 224, but any type or form of media or multimedia can be used as route information 224 to train and / or fine-tune the machine learning semantic route planning model 210. For example, any type or form of written material can be used as route information 224, such as books, articles, blog posts, social media posts, crawled or aggregated text content, text content generated by a machine learning model, etc. Similarly, any type or form of image or video data can be used as route information 224, such as images, movies, television programs, etc. Furthermore, any type or form of 3D representation, simulated information, etc., can be used to train the machine learning semantic route planning model 210, such as simulated maps or spaces.

[0093] Figure 3 A block diagram depicts an example of a machine learning semantic route planning model 300 based on some implementations of this disclosure. It should be noted that this document relates to... Figure 3Any part of the machine learning semantic route planning model 300 described may or may not be a "component" of the machine learning semantic route planning model 300 or may or may not be integrated with the machine learning semantic route planning model in other ways. In some implementations, the machine learning semantic route planning model 300 may generally refer to a single model comprising multiple model parts (e.g., encoder, decoder, tokenizer, etc.). Alternatively, in some implementations, the machine learning semantic route planning model 300 may refer only to some of the model parts described herein. For example, a computing system may instantiate an embedding layer 302 to generate an intermediate representation 306 and then instantiate a decoder layer 318 to process the intermediate representation 306. As another example, a computing system may provide model input 304 to another computing system that has already instantiated the embedding layer 302 for processing, and in response, may receive the intermediate representation 306 from the other computing system.

[0094] The machine learning semantic route planning model 300 may include one or more embedding layers 302. The embedding layers 302 may process the model input 304 to generate an intermediate representation 306 of the model input 304. In some implementations, the embedding layer 302 may include a pre-trained language encoder 308A (e.g., the decoder part of an LLM). The pre-trained language encoder 308A may be trained to generate an intermediate representation of text content. For example, the model input 304 may include request information 310 and route characteristic information 312. The request information 310 may include text content describing a user request for a specific route, specific route characteristics, information related to specific map operations, etc. The pre-trained language encoder 308A may process the text content to obtain an intermediate representation of the text content. The intermediate representation 306 may include an intermediate representation of the text content, or may be an intermediate representation based on the text content.

[0095] In some implementations, model input 304 may include routes or route segments previously generated for the user. For example, suppose the user has provided feedback with particularly high ratings for two routes. Route information indicating previous routes can be included in the model input to provide "few-shot" hints to the machine learning semantic route planning model 300. In other words, previously highly rated routes can be used as context for the route planning request and can encourage the machine learning semantic route planning model 300 to generate routes similar to those indicated by model input 304.

[0096] Additionally or alternatively, in some implementations, the embedding layer 302 may include a pre-trained audio encoder 308B (e.g., the encoder portion of a base audio model, etc.). The pre-trained audio encoder 308B may be trained to generate an intermediate representation of the audio data. For example, the request information 310 may include audio data comprising recordings of spoken utterances from a user describing a specific route, specific route characteristics, etc. The pre-trained audio encoder 308B may process the audio data to obtain an intermediate representation of the audio data. The intermediate representation 306 may include an intermediate representation of the audio data, or may be based on an intermediate representation of the audio data.

[0097] Additionally or alternatively, in some implementations, the embedding layer 302 may include a pre-trained visual encoder 308C (e.g., the encoder portion of a base visual model, etc.). The pre-trained audio encoder 308C may be trained to generate intermediate representations of image data (e.g., images, video data, etc.). For example, request information 310 may include image data provided by a user depicting the location of a specific destination. The pre-trained visual encoder 308C may process the image data to obtain an intermediate representation of the image data. The intermediate representation 306 may include, or may be based on, an intermediate representation of the image data.

[0098] Additionally or alternatively, in some implementations, the embedding layer 302 may include a mapping encoder 308D. The mapping encoder 308D may be trained, at least in part, using pre-trained data, such as regarding... Figure 2A Training data 202 and Figure 2B The training data 222 is described. Specifically, the map manipulation encoder 308D can be trained to generate intermediate representations of map information and / or route information. For example, request information 310 and / or route feature information 312 may include map information and / or route information. The map manipulation encoder 308D can process the map information and / or route information to obtain an intermediate representation of the map information and / or route information. The intermediate representation 306 may include an intermediate representation of the map information and / or route information, or may be based on an intermediate representation of the map information and / or route information.

[0099] It should be noted that in some implementations, encoders 308A to 308D can be encoder / decoder architectures, decoders, etc. More specifically, encoders 308A to 308D can refer to some type of machine learning model or layer of machine learning model capable of processing model input to generate intermediate representations 306.

[0100] In some implementations, the machine learning semantic route planning model 300 may include a tokenizer 309, or the tokenizer may be accessed in other ways. The tokenizer 309 may be a machine learning model or part of a model capable of generating a stream of lexical units based on the input. In some implementations, the tokenizer 309 may process the model input 304 to obtain an intermediate representation 306, which may be or otherwise include lexical units generated by the tokenizer 309. Alternatively, in some implementations, the tokenizer 309 may process the model input 304 to generate a stream of lexical units, which may be processed by an embedding layer 302 to obtain the intermediate representation 306. Alternatively, in some implementations, the tokenizer 309 may process the intermediate representation 306 to generate a stream of lexical units, which may be processed by an attention layer 314 and / or a layer 316.

[0101] In some implementations, the word segmenter 309 can generate multiple word streams, each specific to a particular type of information. For example, the word segmenter 309 may include an audio word segmenter that processes the model input 304 to generate a stream of audio words (e.g., discretizing the original audio into a series of words at a specific rate). As another example, the word segmenter 309 may include a visual word segmenter that processes the model input 304 to generate a stream of image words (e.g., segmenting an image using 16×16 block slices, etc.).

[0102] In some implementations, the embedding layer 302 may include a single encoder that can process stacked tokens from multiple modalities. For example, suppose the model input 304 includes image data and audio data. The tokenizer 309 can process the model input 304 to obtain an audio token stream and an image token stream. The audio token stream and the image token stream can be stacked and processed by a single encoder to generate an intermediate representation 306.

[0103] In some implementations, the machine learning semantic route planning model 300 may include an attention layer 314. The attention layer 314 may apply weights indicating the relative importance of specific portions of the input. For example, the attention layer 314 may be, or otherwise include, a transformer-based self-attention mechanism. Additionally or alternatively, in some implementations, the machine learning semantic route planning model 300 may include a layer 316. Layer 316 may be a layer of a graph neural network model. Layer 316 may represent routes and segments of routes as a graph of nodes and edges, where nodes represent starting and destination locations, and edges represent route segments navigating between the starting and destination locations. Additionally, in some implementations, layer 316 may include additional relationships between nodes in a route comprising multiple routes (e.g., a route from home to an airport, from the airport to a second airport, from the second airport to a hotel, etc.).

[0104] The machine learning semantic route planning model 300 may include a decoder layer 318. The decoder layer 318 may process an intermediate output (e.g., intermediate representation 306, output from attention layer 314 or layer 316, etc.) to obtain a model output 320. The model output 320 may include route planning information 322 and / or semantic map operation information 324. The route planning information 322 may indicate routes and / or route segments. For example, the model input 304 may indicate a request for an alternative route segment of the planned route. The route planning information 322 may include information indicating suggested route segments in response to the request.

[0105] Semantic map operation information 324 may include information describing routes, route segments, or specific route characteristics. For example, model input 304 may indicate a request for a description of a specific route. Semantic map operation information 324 may include text content describing the specific route in response to the request. As another example, model input 304 may indicate a request to identify the most popular businesses located along the route. Semantic map operation information 324 may include content indicating the most popular businesses along the specific route in response to the request (e.g., images of the businesses, text content describing the businesses, map operations and / or route planning information associated with the businesses, user reviews corresponding to the businesses, etc.).

[0106] In some implementations, decoder layer 318 may include a pre-trained language decoder 326A. For example, the pre-trained language decoder 326A may be a decoder of an LLM model that includes a pre-trained language encoder 308A. The pre-trained language decoder 326A may be trained to generate language output based on some intermediate output (e.g., intermediate representation 306, etc.). For example, suppose request information 310 includes a request to describe a specific route. The pre-trained language decoder 326A may process some intermediate output to generate text content describing the specific route indicated by model input 304. The text content may be included in semantic map operation information 324.

[0107] In some implementations, decoder layer 318 may include a pre-trained audio decoder 326B. For example, the pre-trained audio decoder 326B may be a decoder that includes a base audio model with a pre-trained audio encoder 308B. The pre-trained audio decoder 326B may be trained to generate audio output based on some intermediate output (e.g., intermediate representation 306, etc.). For example, suppose request information 310 includes a request to generate audio describing a specific route. The pre-trained audio decoder 326B may process some intermediate output to generate audio describing the specific route indicated by model input 304. The audio may be included in semantic map operation information 324.

[0108] In some implementations, decoder layer 318 may include a pre-trained visual decoder 326C. For example, the pre-trained visual decoder 326C may be a decoder of a base visual model including a pre-trained visual encoder 308C. The pre-trained visual decoder 326C may be trained to generate image output based on some intermediate output (e.g., intermediate representation 306, etc.). For example, suppose request information 310 includes a request to depict a specific route in a stylized manner (e.g., “draw this route in the same style as a ski map”). The pre-trained visual decoder 326C may process some intermediate output to generate image data depicting the specific route in a stylized manner indicated by model input 304. The image data may be included in semantic map operation information 324.

[0109] In some implementations, the decoder layer 318 may include a mapping decoder 326D. For example, the mapping decoder 326D may be a decoder trained in conjunction with a mapping encoder 308D. The mapping decoder 326D may be trained to generate map manipulation information and / or route planning information based on some intermediate output (e.g., intermediate representation 306, etc.). For example, suppose the request information 310 includes a request for alternative route segments for depicting and generating a specific route segment. The mapping decoder 326D may process some intermediate output to generate map manipulation information and / or route planning information indicating the suggested route segments as alternatives to the route segments. The map manipulation information and / or route planning information may be included in the route planning information 322.

[0110] Additionally or alternatively, in some implementations, decoder layer 318 may generate multimodal model output 320. For example, suppose request information 310 includes: (a) text content describing a request to generate an alternative route segment for a specific route segment in the route, and (b) a request describing the alternative route segment and the differences between the alternative route segment and the specific route segment. Map manipulation decoder 326D may process some intermediate output to generate suggested route segments as alternatives to the specific route segment. Pre-trained language decoder 326A may process some intermediate output to generate a text description of the suggested route segments. The suggested route segments may be included in route information 322, and the text description may be included in semantic map manipulation information 324. In this way, the decoder layer may generate multimodal output including multiple types of data.

[0111] Example Method

[0112] Figure 4 A flowchart 400 depicts an example method for training and / or fine-tuning a machine learning semantic route planning model, according to some implementations of this disclosure. Although Figure 4 The steps are described in a particular order for purposes of illustration and discussion, but the method of this disclosure is not limited to the order or arrangement specifically shown. The various steps of method 400 may be omitted, rearranged, combined and / or adapted in various ways without departing from the scope of this disclosure.

[0113] At 402, the computing system can obtain training data for training (or “pre-training”) a machine learning semantic route planning model. The training data may include route information. Route information may indicate a route from a starting point to a destination. A route may include multiple route segments, comprising a first subset of route segments (e.g., a set of “unmasked” segments) and a second subset of route segments (a set of “masked” route segments). Route information may also include route characteristic information. Route characteristic information may describe one or more route characteristics. Route characteristics may include the type of route (e.g., sightseeing, efficient, fast, safe, etc.), entities located along the route (e.g., businesses, residences, landmarks, POIs, etc.), average traffic flow conditions for the route or a segment of the route, average weather conditions for the route or a segment of the route, information associated with users traversing the route, etc.

[0114] At 404, the computational system can utilize a machine learning semantic route planning model to process at least a first subset of route segments and a portion of route characteristic information associated with the first subset of route segments to obtain one or more predicted route segments for a second subset of route segments. More specifically, the first subset of route segments may be unmasked route segments. The second subset of route segments may be masked or otherwise processed using a machine learning semantic route planning model, such that the model can predict said segments.

[0115] In some implementations, to process a first subset of route segments, the computational system may utilize a first part of a machine learning-based semantic route planning model to process the first subset of route segments and the portion of route characteristic information associated with the first subset of route segments to obtain a potential representation of the first subset of route segments. In some implementations, the first part of the machine learning-based semantic route planning model may be, or otherwise include, an encoder portion of a pre-trained LLM. For example, route information may indicate a second route from a second starting point to a second destination location and indicate a prompt for a description of the second route. The model output may include text content describing the second route.

[0116] At point 406, the computational system can train and / or fine-tune a machine learning-based semantic route planning model based on an optimization function that evaluates the differences between one or more predicted route segments and a second subset of route segments.

[0117] In some implementations, the computational system may utilize a machine learning-based semantic route planning model to process input data and obtain model output. The input data may include route information indicating one or more first route segments of an incomplete route, and prompts indicating a request to generate one or more second route segments with the requested route characteristics for the incomplete route. The model output may include one or more second route segments with the requested route characteristics.

[0118] In some implementations, route information may indicate one or more example routes, a second route comprising multiple second route segments, and a prompt indicating a request to generate one or more alternative route segments for one or more corresponding second route segments of the second route based on the one or more example routes. The model output may include one or more alternative route segments of the second route.

[0119] In some implementations, the computational system may utilize a machine learning-based semantic route planning model to process input data and obtain model output. The input data may include route information describing a second route comprising multiple second route segments. The model output may include classification information that categorizes the route into a first route type among multiple route types (e.g., sightseeing route type, efficiency route type, non-highway route type, toll route type, low traffic flow route type, etc.). In some implementations, the classification information may also categorize the second route segments among multiple second route segments into a first route segment type among multiple route segment types.

[0120] In some implementations, the computational system can utilize a machine learning-based semantic route planning model to process input data and obtain model output. Input data may include the current route segment of an existing route, including intermediate destination locations. Output data may include one or more candidate route segments. The starting position of each of the one or more candidate route segments may correspond to an intermediate destination location of the current route segment.

[0121] In some implementations, the computing system may utilize a machine learning-based semantic route planning model to process input data and obtain model output. The input data may include route request information describing a route from the requested start point to the requested destination, and preferred route characteristic information describing the preferred route characteristics. The model output may include route information indicating a route from the requested start point to the requested destination, including the preferred route characteristics.

[0122] In some implementations, the computational system may utilize a machine learning-based semantic route planning model to process input data and obtain model output. Input data may include one or more images of one or more locations. Output data may include travel information indicating a proposed route including at least one of the one or more locations.

[0123] Figure 5 A flowchart 500 depicts an example method for performing map-related tasks using a machine learning-based semantic route planning model, according to some implementations of this disclosure. Although Figure 5The steps are depicted in a particular order for illustrative and discussion purposes, but the method of this disclosure is not limited to the order or arrangement specifically shown. The various steps of method 500 may be omitted, rearranged, combined and / or adapted in various ways without departing from the scope of this disclosure.

[0124] At point 502, the computing system may obtain one or more inputs from a client computing device for a machine learning semantic route planning model. The machine learning semantic route planning model may be a model trained to process map operation information to generate model output, which includes suggested route segments and / or information associated with those route segments. One or more inputs may include at least one of the following: request information indicating a requested route segment and / or a request for map operation-related information; or route characteristic information indicating one or more route characteristics.

[0125] At point 504, the computing system can process one or more inputs to obtain a model output. The model output may include at least one of the following: (a) route planning information indicating a route including one or more suggested route segments; or (b) semantic map operation information associated with the route. Specifically, the route planning information may be, or otherwise include, data for map operation applications or information describing a particular route or route segment. For example, the route planning information may be data that, when provided to an application programming interface (API) for map operation applications, enables the application to generate routes or modify existing routes based on the model output. Alternatively, the route planning information may be text content describing a route or route segment, or modifications to a route. Or, the route planning information may be an image depicting the route.

[0126] Semantic map operation information generally refers to information describing a particular aspect or characteristic of routes or map operation-related entities (such as businesses, landmarks, residences, POIs, etc.) included in the map. Semantic map operation information can include text content responding to queries. For example, model input may include queries related to a specific route or entity located on the map. Semantic map operation information can also include general descriptions of routes or entities located on the map. For example, semantic map operation information might include a semantic description of a specific route (e.g., "this route is a scenic route along the coast that is unlikely to encounter any heavy traffic or adverse weather conditions"). Similarly, semantic map operation information may include a semantic description of a route segment related to a specific road or route segment (e.g., "the road on which this business is located is mostly frequented by business travelers during normal work hours"). For example, semantic map operation information may be or otherwise include classification outputs that categorize specific routes or segments (e.g., traffic flow: low; collision probability: low; severe weather probability: low, etc.).

[0127] In some implementations, for use cases that utilize machine learning-based semantic route planning models to predict route segments or suggest alternative route segments, the computational system can restrict such predictions or suggestions to known route segments that have been verified to exist. In other words, the model output can select known route segments instead of attempting to generate route segments that may or may not be traversable. In this way, the computational system can ensure that the route segments provided to the user are traversable.

[0128] At point 506, the computing system can provide model output to a client computing device. The client computing device can generally refer to a device that receives route planning or map operation-related services from the computing system, such as a user computing device (e.g., a smartphone, laptop, wearable device, etc.), a second computing system (e.g., a virtualized computing device, a computing or network node, etc.). In some implementations, the computing system can provide model output in response to receiving model input. For example, the computing system can receive route planning or map operation-related requests, including model input, from the client computing device. In response, the computing system can provide model output to the client computing device, or it can provide information describing the model output or otherwise based on the model output.

[0129] Additionally or alternatively, in some implementations, the computing system may iteratively provide model outputs to client computing devices. For example, within multiple iterations, the computing system may receive model inputs, including requests to classify route segment types associated with specific route segments. After classifying a certain number of route segments, the computing system may update the client computing devices by providing information indicating the classified route segments. More generally, if the model output generated by the computing system is applicable to multiple client computing devices (e.g., classifying route segments, classifying business entities, etc.), the computing system may update some (or all) of the client devices based on the model output. For example, the computing system may push such information to client devices by updating a map operation application associated with the computing system.

[0130] Additionally or alternatively, in some implementations, the computational system can utilize the model output to perform an additional action. For example, if the model output categorizes specific route segments or entities (e.g., businesses, residences, landmarks, POIs, etc.) existing within the map operation data, the computational system can store the categorizations in a data store that stores similar categories to enable map operation services.

[0131] In some implementations, obtaining one or more inputs for a machine learning semantic route planning model may include obtaining example route information indicating one or more example routes and a second route including multiple second route segments from a client computing device, as well as prompts indicating a request to generate one or more alternative route segments for one or more corresponding second route segments of the second route based on the one or more example routes. Processing one or more inputs to obtain a model output may include processing the example route information and the prompts to obtain the model output. The model output may include route planning information indicating one or more alternative route segments for one or more corresponding second route segments.

[0132] In some implementations, obtaining one or more inputs for a machine learning semantic route planning model may include obtaining information from a client computing device indicating the current route segment, including intermediate destination locations. Processing one or more inputs to obtain a model output may include processing the information indicating the current route segment to obtain the model output. The model output may include one or more candidate route segments. The starting position of each of the one or more candidate route segments may correspond to an intermediate destination location of the current route segment.

[0133] In some implementations, obtaining one or more inputs for a machine learning semantic route planning model may include obtaining (a) route request information describing a route from a requested start point to a requested destination from a client computing device, and (b) preferred route characteristic information describing preferred route characteristics. Processing one or more inputs to obtain a model output may include processing information indicating the route request information and the preferred route characteristic information to obtain the model output. The model output may include information indicating a route from the requested start point to the requested destination, including the preferred route characteristics.

[0134] In some implementations, obtaining one or more inputs for a machine learning semantic route planning model may include obtaining information indicating one or more requested route segments from a client computing device. The information may include one or more images of one or more locations. Processing the one or more inputs to obtain a model output may include processing the information indicating the current route segment to obtain the model output. The model output may include trip information indicating a proposed route including at least one of the one or more locations.

[0135] Additional Publication

[0136] This paper discusses technologies related to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a wide variety of possible configurations, combinations, and partitions of tasks and functionality between and within components. For example, the processes discussed in this paper can be implemented using a single device or component, or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0137] While various specific example embodiments of this subject matter have been described in detail, each example is provided by way of explanation and not limitation. Modifications, variations, and equivalents of such embodiments will be readily apparent to those skilled in the art upon acquiring an understanding of the foregoing. Therefore, this disclosure does not exclude the inclusion of such modifications, variations, and / or additions to the subject matter as will be apparent to those skilled in the art. For example, features shown or described as part of an embodiment may be used with another embodiment to produce further embodiments. Therefore, this disclosure is intended to cover such modifications, variations, and equivalents.

Claims

1. A computer-implemented method, comprising: Training data is obtained by a computing system comprising one or more computing devices, the training data including: (a) Route information indicating a route from a starting location to a destination location, wherein the route comprises multiple route segments, the multiple route segments comprising a first subset of route segments and a second subset of route segments; and (b) Route characteristic information describing one or more route characteristics; The computing system uses a machine learning-based semantic route planning model to process at least a first subset of route segments and a portion of the route characteristic information associated with the first subset of route segments to obtain one or more predicted route segments for a second subset of route segments; and The computational system adjusts one or more parameters of the machine learning semantic route planning model based on an optimization function that evaluates the differences between the one or more predicted route segments and the second subset of route segments.

2. The computer-implemented method of claim 1, wherein processing the first subset of at least the route segments and the portion of the route characteristic information associated with the first subset of the route segments using the machine learning semantic route planning model comprises: The computing system uses a first part of the machine learning semantic route planning model to process the first subset of at least the route segments and the part of the route characteristic information associated with the first subset of the route segments to obtain a potential representation of the first subset of the route segments.

3. The computer implementation method of claim 2, wherein the first part of the machine learning semantic route planning model includes an encoder and / or decoder part of a pre-trained large language model (LLM).

4. The computer-implemented method of claim 3, wherein the method further comprises: The computing system uses the machine learning-based semantic route planning model to process the input data to obtain the model output, wherein the input data includes: Route information indicating a second route from a second starting point to a second destination; and The prompt indicates a request to describe the second route; and The model output includes text content describing the second route.

5. The computer-implemented method of claim 3, wherein the method further comprises: The computing system uses the machine learning semantic route planning model to process the input data to obtain the model output, wherein the input data includes text content describing the requested route, and wherein the model output includes route information indicating the requested route.

6. The computer-implemented method as described in claim 5, wherein, Before processing the input data using the machine learning-based semantic route planning model to obtain the model output, the method includes: The computing system uses the machine learning-based semantic route planning model to process second training data, including the training route, to obtain a textual description of the training route; and The computing system adjusts one or more parameters of the machine learning semantic route planning model based on a loss function that evaluates the text description of the training route and the corresponding benchmark real text description of the training route.

7. The computer-implemented method of claim 2, wherein the portion of processing the first subset of at least the route segments and the route characteristic information associated with the first subset of the route segments using the first portion of the machine learning semantic route planning model further comprises: The computational system processes the latent representation using the graph-based portion of the machine learning semantic route planning model to obtain a graph output, wherein the graph output includes: Multiple nodes, wherein the multiple nodes represent the starting position, the destination position, and an intermediate position between the starting position and the destination position; and Multiple edges, where each edge represents a route segment.

8. The computer-implemented method of claim 1, wherein the method further comprises: The computing system uses the machine learning-based semantic route planning model to process the input data to obtain the model output, wherein the input data includes: Route information indicating one or more first route segments of an incomplete route; and A prompt indicating a request to generate one or more second route segments with the characteristics of the requested route for the incomplete route; and The model output includes one or more second route segments having the characteristics of the requested route.

9. The computer-implemented method of claim 1, wherein the method further comprises: The computing system uses the machine learning-based semantic route planning model to process the input data to obtain the model output, wherein the input data includes: Route information indicating one or more example routes and a second route including multiple second route segments; and A prompt indicating a request to generate one or more alternative route segments for one or more corresponding second route segments of the second route based on the one or more example routes; and The model output includes one or more alternative route segments of the second route.

10. The computer-implemented method of claim 1, wherein the method further comprises: The computing system uses the machine learning semantic route planning model to process the input data to obtain the model output, wherein the input data includes route information describing a second route comprising multiple second route segments, and wherein the model output includes classification information classifying the route into a first route type among multiple route types.

11. The computer-implemented method of claim 10, wherein the classification information further classifies the second route segments of the plurality of second route segments into a first route segment type of the plurality of route segment types.

12. The computer-implemented method of claim 1, wherein the method further comprises: The computing system processes the input data using the machine learning-based semantic route planning model to obtain the model output, wherein the input data includes the current route segment of the existing route, including intermediate destination locations; and The model output includes one or more candidate route segments, and the starting position of each of the one or more candidate route segments corresponds to the intermediate destination position of the current route segment.

13. The computer-implemented method of claim 12, wherein the method further comprises: The computing system uses the machine learning-based semantic route planning model to process the input data to obtain the model output, wherein the input data includes: Route request information describing the route from the requested starting point to the requested destination; and Preferred route characteristic information describing the preferred route characteristics; and The model output includes route information indicating a route from the requested start location to the requested destination location, including the preferred route characteristics.

14. The computer-implemented method of claim 1, wherein the method further comprises: The computing system uses the machine learning-based semantic route planning model to process the input data to obtain the model output, wherein the input data includes one or more images of one or more locations; and The model output includes travel information indicating a proposed route that includes at least one of the one or more locations.

15. A computing system, comprising: One or more processor devices; The memory includes: A machine learning semantic route planning model, wherein the machine learning semantic route planning model is trained to process map operation information to generate model output, the model output including suggested route segments and / or information associated with route segments; One or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, including: One or more inputs are obtained from a client computing device for the machine learning semantic route planning model, wherein the one or more inputs include at least one of the following: A request for information indicating the requested route segment and / or a request for information related to map operations; or Route characteristic information indicating one or more route characteristics; Process the one or more inputs to obtain a model output, wherein the model output includes at least one of the following: (a) Indicates route planning information including one or more suggested route segments; or (b) Semantic map operation information associated with the route; and The model output is provided to the client computing device.

16. The computing system of claim 15, wherein obtaining the one or more inputs for the machine learning semantic route planning model comprises: The client computing device obtains example route information indicating one or more example routes and a second route including multiple second route segments, as well as a prompt indicating a request to generate one or more alternative route segments for one or more corresponding second route segments of the second route based on the one or more example routes; and Processing the one or more inputs to obtain the model output includes: The example route information and the prompts are processed to obtain the model output, wherein the model output includes route planning information indicating one or more alternative route segments for the one or more corresponding second route segments.

17. The computing system of claim 15, wherein obtaining the one or more inputs for the machine learning semantic route planning model comprises: Information indicating the current route segment, including intermediate destination locations, is obtained from the client computing device; and Processing the one or more inputs to obtain the model output includes: The information indicating the current route segment is processed to obtain the model output, wherein the model output includes route planning information indicating one or more candidate route segments, and wherein the starting position of each of the one or more candidate route segments corresponds to the intermediate destination position of the current route segment.

18. The computing system of claim 15, wherein obtaining the one or more inputs for the machine learning semantic route planning model comprises: Obtained from the client computing device: Route request information describing the route from the requested starting location to the requested destination location; as well as Preferred route characteristic information describing the preferred route characteristics; and Processing the one or more inputs to obtain the model output includes: The information indicating the route request information and the preferred route characteristic information is processed to obtain the model output, wherein the model output includes route planning information indicating a route from the requested start location to the requested destination location, including the preferred route characteristics.

19. The computing system of claim 15, wherein obtaining the one or more inputs for the machine learning semantic route planning model comprises: The client computing device obtains the information indicating the requested route segment, wherein the information includes one or more images of one or more locations; and Processing the one or more inputs to obtain the model output includes: The one or more images of the one or more locations are processed to obtain the model output, wherein the model output includes route planning information indicating a proposed route including at least one of the one or more locations.

20. One or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations, including: Obtain training data, the training data including: (a) Route information indicating a route from a starting location to a destination location, wherein the route comprises multiple route segments, the multiple route segments comprising a first subset of route segments and a second subset of route segments; and (b) Route characteristic information describing one or more route characteristics; Utilizing a machine learning-based semantic route planning model to process at least a first subset of route segments and a portion of the route characteristic information associated with the first subset of route segments to obtain one or more predicted route segments for a second subset of route segments; and One or more parameters of the machine learning semantic route planning model are adjusted based on an optimization function that evaluates the differences between the one or more predicted route segments and the second subset of route segments.