Training method of model, method, device, equipment and medium for generating driving text

By introducing spatial text prompt samples and time token modules in the training set of driving text generation model, the problem of insufficient space and time perception in the prior art is solved, and the accuracy of driving text generation and driving decision accuracy is improved.

CN117763115BActive Publication Date: 2025-06-27SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311796419.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27
Estimated Expiration
2043-12-25

AI Technical Summary

Technical Problem

The prior art is difficult to fully explore the perception of space and time in driving scenarios, resulting in low accuracy of generated driving texts, which in turn affects the accuracy of driving decisions.

Method used

By determining the training set for setting the driving text generation model, including the first training set of the pre-training stage and the second training set of the fine-tuning stage, the spatial text prompt sample is used for pre-training, and the model is fine-tuned through the time token module to enhance the model's spatial and time perception capabilities.

Benefits of technology

It improves the accuracy of generation of driving text, enhances the spatial and temporal perception of the model in driving scenarios, and thus improves the accuracy of driving decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117763115B_ABST
    Figure CN117763115B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, device and medium for training a model and generating driving text. The method includes: determining a training set corresponding to a set driving text generation model; wherein the training set includes a first training set and a second training set; pre-training the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; fine-tuning the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein the set driving text generation model includes a time token module. In the embodiments of the present disclosure, by pre-training the set driving text generation model with a first training set containing spatial text prompt samples and fine-tuning the pre-trained set driving text generation model through a time token module, the perception abilities of space and time can be enhanced, and the generation accuracy of driving text can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a method, apparatus, device, and medium for training a model and generating driving text. Background Art

[0002] In the field of intelligent agents (such as autonomous vehicles, robots, and drones), embodied understanding enables these agents to interpret instructions and analyze scenarios based on their experiences. However, this is a crucial but challenging task that has not been solved yet. In recent years, significant progress has been made in general vision by leveraging the extensive knowledge and causal reasoning capabilities of vision-language models (VLMs). However, in driving scenarios, traditional VLMs are limited to generating narrative phrases, i.e., scene descriptions, and their perception of space and time has not been fully explored. That is, there are still limitations in the perception of space and time in the prior art, resulting in low accuracy of the generated driving text and thus potential deficiencies in driving decisions. Summary of the Invention

[0003] Embodiments of the present invention provide a method, apparatus, device, and medium for training a model and generating driving text, which can improve the perception ability of space and time, thereby improving the accuracy of generating driving text.

[0004] In a first aspect, an embodiment of the present invention provides a method for training a model, including: determining a training set corresponding to a set driving text generation model; wherein, the training set includes a first training set and a second training set; pre-training the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample in a driving scenario, a spatial text prompt sample, and a corresponding first driving text; wherein, the first target data sample includes a first target image and / or a first target video; fine-tuning the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample in a driving scenario, a time text prompt sample, and a corresponding second driving text; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a time token module; the time token module is configured to: for each group of second sample pairs, determine a target image feature from the second target data sample based on the time text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the time text prompt sample, and the corresponding second driving text.

[0005] In a second aspect, an embodiment of the present invention further provides a method for generating a driving text, including: obtaining a target text prompt and target data; wherein, the target data includes a video and / or a picture; inputting the target text prompt and the target data into a target set driving text generation model to output a target driving text; wherein, the target set driving text generation model is obtained by the model training method according to any one of the above.

[0006] In a third aspect, an embodiment of the present invention further provides a model training device, including: a training set determination module, configured to determine a training set corresponding to a set driving text generation model; wherein, the training set includes a first training set and a second training set;

[0007] A pre-training module for pre-training the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample, a spatial text prompt sample, and a corresponding first driving text in a driving scenario; wherein, the first target data sample includes a first target image and / or a first target video; A fine-tuning module for fine-tuning the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample, a time text prompt sample, and a corresponding second driving text in a driving scenario; wherein the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a time token module; the time token module is configured to: for each group of second sample pairs, determine a target image feature from the second target data sample based on the time text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the time text prompt sample, and the corresponding second driving text.

[0008] In a fourth aspect, an embodiment of the present invention further provides a driving text generation device, including: an acquisition module for acquiring a target text prompt and target data; wherein, the target data includes a video and / or a picture; a target driving text output module for inputting the target text prompt and the target data into a target set driving text generation model to output a target driving text; wherein, the target set driving text generation model is obtained by the training method of any one of the models.

[0009] In a fifth aspect, an embodiment of the present invention further provides an electronic device, the electronic device includes:

[0010] At least one processor; and

[0011] A memory communicatively connected to the at least one processor; wherein,

[0012] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training of the model or the driving text generation method according to any one of the embodiments of the present invention.

[0013] In a sixth aspect, an embodiment of the present invention further provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions for causing a processor to execute the training of the model or the driving text generation method according to any one of the embodiments of the present invention when executed.

[0014] The technical solution provided by the embodiments of the present invention determines a training set corresponding to a set driving text generation model; wherein, the training set includes a first training set and a second training set; pre-train the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample in a driving scenario, a spatial text prompt sample, and a corresponding first driving text; wherein, the first target data sample includes a first target image and / or a first target video; fine-tune the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample in a driving scenario, a time text prompt sample, and a corresponding second driving text; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a time token module; the time token module is used for: for each group of second sample pairs, determine a target image feature from the second target data sample based on the time text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the time text prompt sample, and the corresponding second driving text. In the embodiments of the present disclosure, by pre-training the set driving text generation model with the first training set containing spatial text prompt samples and fine-tuning the pre-trained set driving text generation model through the time token module, the spatial and temporal perception capabilities of the target set driving text generation model can be enhanced, thereby improving the generation accuracy of driving texts. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a flowchart of a method for training a model provided by an embodiment of the present invention;

[0016] Figure 2 It is a training schematic diagram of the pre-trained set driving text generation model in the fine-tuning stage provided by an embodiment of the present invention;

[0017] Figure 3 It is a flowchart of a method for generating a driving text provided by an embodiment of the present invention;

[0018] Figure 4 It is a schematic structural diagram of a model training device provided by an embodiment of the present disclosure;

[0019] Figure 5 It is a schematic structural diagram of a driving text generation device provided by an embodiment of the present disclosure;

[0020] Figure 6 It is a schematic structural diagram of an electronic device implementing the embodiments of the present invention. Detailed implementation manners

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0022] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0023] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0024] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.

[0025] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0026] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0027] Figure 1 It is a flowchart of a method for training a model provided by an embodiment of the present invention. This embodiment is applicable to the situation of generating driving text from text prompts and images or from text prompts and videos in a given driving scenario. This method can be executed by a model training device, and specifically includes the following steps:

[0028] S110. Determine a training set corresponding to a set driving text generation model.

[0029] Among them, the training set includes a first training set and a second training set. The first training set is used for the pre-training stage, and the second training set is used for the fine-tuning stage. The first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample in a driving scenario, a spatial text prompt sample, and a corresponding first driving text; the first driving text can be understood as a label, equivalent to the correct answer, and the first target data sample and the spatial text prompt sample are the contents that the set driving text generation model needs to learn in the pre-training stage. The difference between the first training set and the second training set is that the first training set does not include a time text prompt sample. The first training set can be sourced from an open-world data corpus (diverse data in multiple fields).

[0030] Among them, the first sample pair is not limited to including a spatial text prompt sample, and can also include a general text prompt sample, etc. The general text prompt sample can be understood as a description of the driving scenario, or some common problems in real life, such as "Turn left or right?", "Is the traffic light ahead red or green?" The spatial text prompt sample can be understood as a question about spatial description, which can include spatial distance, spatial position, etc. For example, "Is the distance between the vehicle ahead and our vehicle 30 meters?" "Where is the current vehicle located?" etc. The first driving text can be "The distance between the vehicle ahead and our vehicle is 50 meters", "The position of the current vehicle is (30, 50)", etc.

[0031] Among them, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample in a driving scenario, a time text prompt sample, and a corresponding second driving text. The second driving text can be understood as a label, equivalent to the correct answer, and the second target data sample and the time text prompt sample are the contents that the set driving text generation model needs to learn in the fine-tuning stage. The second sample pair is not limited to including a time text prompt sample, and can also include a general text prompt sample, a spatial text prompt sample, etc. Specifically, the second training set can vary according to different downstream tasks, that is, the second training set can be sorted according to specific downstream tasks. The time text prompt sample can be understood as a question about time description, such as "What was the road sign that appeared 5 seconds ago?", "What happened 3 seconds ago?" etc. The second driving text can be "The road sign that appeared 5 seconds ago is a speed limit of 60 sign", "An animal crossed the road 3 seconds ago", etc.

[0032] Among them, the second target data sample includes a second target image and / or a second target video; the first target data sample includes a first target image and / or a first target video.

[0033] Optionally, determine the training set corresponding to the set driving text generation model, including: obtaining an original data sample set in a driving scenario; the original data sample set includes an original image set and / or an original video set; cleaning each original data sample in the original data sample set; forming a target data sample set by all the cleaned original data samples that meet the set quality conditions, inputting the target data sample set into a set vision-language model, and outputting a driving text set corresponding to the target data sample set; if the driving texts in the driving text set do not meet the semantic conditions of the corresponding original data samples, input the corresponding original data samples back into the set vision-language model, and re-output the corresponding driving texts until the driving text set meets the semantic conditions of the corresponding original data samples, to obtain a target driving text set; determining a target text prompt sample set corresponding to the target data sample set through a first set language model; wherein, the target data sample set includes a target image set and / or a target video set, and the target text prompt sample set is a spatial text prompt sample set or a temporal text prompt sample set.

[0034] Wherein, the training set is composed of a target data sample set, a target text prompt sample set, and the corresponding target driving text set. There is a one-to-one correspondence between the target data sample set, the target text prompt sample set, and the target driving text set. The training set can also be understood as including at least one set of sample pairs, and each set of sample pairs includes a target data sample, a target text prompt sample, and a target driving text.

[0035] In this embodiment, the training sets in the pre-training stage and the fine-tuning stage can be determined in the following manner: Exemplarily, taking the training set in the pre-training stage as an example, in order to achieve spatial positioning and retain the ability to describe driving scenarios, an original data sample set can be collected from various fields. The data sources include datasets such as nuScenes, Waymo, YouTube, and Ego4D, covering cities, rural areas, and various weather conditions, etc. The original data sample set includes original images and / or original videos. The original images and original videos can both include road, traffic, etc. conditions from all over the world. After collecting the original data sample set, each original data sample in the original data sample set is cleaned, and the cleaned original data sample set is inspected to make the cleaned original images and / or original videos meet the set quality conditions, where the set quality conditions can be that the clarity, damage degree, etc. all meet the corresponding requirements. And all the cleaned original data samples that meet the set quality conditions are grouped into a target data sample set, and the target data sample set is input into a set vision-language model to output the driving text set corresponding to the target data sample set; in this embodiment, no limitation is imposed on the specific set vision-language model. For example, it can be the parameter-efficient vision instruction model of LLaMA-AdapterV2. Semantic inspection is performed on the driving text set to determine whether the driving texts in the driving text set conform to the semantics of the corresponding original data samples. If the driving texts in the driving text set do not meet the semantic conditions of the corresponding original data samples, that is, the driving texts do not conform to the semantics of the corresponding original data samples, the corresponding original data samples are re-input into the set vision-language model to re-output the corresponding driving texts until each driving text in the driving text set meets the semantic conditions of the corresponding original data sample, and a target driving text set is obtained. Among them, each driving text meeting the semantic conditions of the corresponding original data sample can be understood as that for each piece of data, the driving text conforms to the semantics of the corresponding original data sample. After obtaining the target data sample set and the corresponding driving text set, a plurality of preset questions (such as dozens of questions, hundreds of questions) are input into the first set language model, and each question outputs a plurality of different statements (that is, a plurality of and unique text prompt templates). For example, one question outputs more than 1000 statements, and the plurality of statements output are semantically consistent with the input question. Then, one statement can be randomly selected from the plurality of statements of the plurality of questions as a target text prompt, so that a target text prompt sample set corresponding to the target data sample set can be obtained. Finally, the training set is composed of the target data sample set, the target text prompt sample set, and the corresponding target driving text set. No limitation is imposed on the specific first set language model in this embodiment. For example, it can be any series of models of the ChatGPT language model. Among them, the target data sample set includes a target image set and / or a target video set, and the target text prompt sample set is a spatial text prompt sample set.

[0036] In this embodiment, by collecting data from multiple fields, including open-world data such as nuScenes, Waymo, YouTube, and Ego4D, a large and diverse first training set is formed to improve the robustness and generalization ability of the model.

[0037] If the training set in the fine-tuning stage is constructed in the above manner, the original data sample set needs to be reorganized, which can be organized according to specific downstream tasks. Correspondingly, the obtained target text prompt sample set is a time text prompt sample set.

[0038] In this embodiment, quality inspection is performed on the cleaned original data sample set and semantic inspection is performed on the driving text set. Through two rounds of inspections, it is possible to ensure the generation of images (or videos) and labels with excellent quality, injecting a large number of high-quality training samples and descriptive labels into the training set, thereby improving the model's understanding ability of driving scenarios and thus improving the overall performance of the model.

[0039] In this embodiment, a large number of unique text prompt templates are generated by the first set language model, and then an automatic annotation process of human-computer collaboration is carried out, which can enhance the model's spatial understanding. Descriptive labels are introduced, covering the overall scene, traffic elements, and driving decisions.

[0040] S120. Pre-train the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model.

[0041] Among them, the set driving text generation model further includes a second set language model. This embodiment does not limit the specific second set language model. For example, it can be the Flan-T5 language model (Instruction finetuning-T5, Flan-T5). In this embodiment, the set driving text generation model can be pre-trained based on each group of first samples to obtain a pre-trained set driving text generation model.

[0042] In this embodiment, through the first training set, that is, a large-scale open-world data corpus, especially the training set collected from datasets such as nuScenes and Waymo in the field of autonomous driving, the target set driving text generation model is endowed with strong spatial positioning ability. Compared with the prior art, the pre-training in the present invention enables the model to understand the driving scenario more comprehensively and accurately, improving the ability to grasp spatial information. The first training set in the present invention is a large and diverse training set. Compared with the use of a single dataset in the prior art, the diverse dataset helps to improve the robustness and generalization ability of the model, enabling it to better adapt to various complex driving scenarios.

[0043] Optionally, pre-train the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model, including: performing image encoding on the first target data sample to obtain a first image encoding feature; performing text encoding on the spatial text prompt sample to obtain a first text encoding feature; pre-training a second set language model based on the first image encoding feature and the first text encoding feature to obtain a pre-trained set driving text generation model.

[0044] Among them, the first target data sample includes a first target image and / or a first target video. In this embodiment, image encoding can be performed on the first target image and / or the first target video, text encoding can be performed on the spatial text prompt sample to obtain a first text encoding feature, and the first image encoding feature and the first text encoding feature are input into a second set language model for training until the corresponding iteration stop condition is met, obtaining a pre-trained set driving text generation model. The iteration stop condition can be that the loss value is less than the loss threshold for a continuously set number of times, or the number of iterations is greater than the set iteration number threshold. Among them, the second set language model includes an encoder and a decoder.

[0045] In this embodiment, pre-training the set driving text generation model with a training set containing spatial text prompts can enhance the spatial perception ability of the set driving text generation model, enabling the set driving text generation model to have functions such as spatial positioning, etc., and can solve the deficiency of spatial perception in the prior art.

[0046] Optionally, performing text encoding on the spatial text prompt sample to obtain a first text encoding feature includes: for non-numeric text in the spatial text prompt sample, performing non-numeric text encoding on the non-numeric text to obtain a non-numeric text encoding feature; for numeric text in the spatial text prompt sample, rounding the numeric text; performing numeric text encoding on the rounded numeric text to obtain a numeric text encoding feature.

[0047] Among them, the first text encoding feature includes a non-numeric text encoding feature and a numeric text encoding feature; in this embodiment, for non-numeric text in the training set, the non-numeric text can be directly encoded as non-numeric text, that is, the non-numeric text is converted into a corresponding token, while for numeric text in the training set, a set method can be used for processing. Specifically, the RT2-like tokenizer method can be used to divide the three-dimensional space into a network with a 1-meter resolution, and the position of the target point is quantified as the index of the network to solve the problem that the set driving text generation model is not sensitive to numbers.

[0048] Specifically, first round the digital text; then convert the rounded digital text into the corresponding token. Exemplarily, the digital text is "1.5", after rounding, it becomes the number "2", and the number "2" is converted into the corresponding token.

[0049] S130. Fine-tune the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model.

[0050] In this embodiment, each group of second sample pairs can be input into the time token module. For each group of second sample pairs, through the time token module, the target image features can be determined from the second target data samples based on the time text prompt samples, so that the target image features, time text prompt samples, and corresponding second driving texts based on each group of second sample pairs are used to fine-tune the pre-trained set driving text generation model.

[0051] Among them, fine-tuning can be understood as continuing to train the pre-trained set driving text generation model based on the second training set (such as the training set of the downstream task). During the fine-tuning process, the pre-trained parameters can be frozen, and the pre-trained set driving text generation model is continuously trained based on the second training set.

[0052] Optionally, fine-tuning the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model includes: performing image encoding on the second target data samples to obtain second image encoding features; performing text encoding on the time text prompt samples to obtain second text encoding features; inputting the second image encoding features and the second text encoding features into the time token module to output target image features; inputting the target image features and the second text encoding features into the pre-trained second set language model to output predicted driving texts; determining a loss value based on the predicted driving texts and the corresponding second driving texts; and performing iterative training on the pre-trained set driving text generation model based on the loss value until the corresponding iterative training stop condition is met to obtain a target set driving text generation model.

[0053] Specifically, in the fine-tuning stage, the second target data sample is image-encoded (e.g., encoded using a Transformer model. For a video, each video frame can be converted into a fixed-length token feature) to obtain the second image encoding feature; the time text prompt sample is text-encoded (e.g., encoded using a Bert model) to obtain the second text encoding feature; wherein, the second text encoding feature includes a text encoding feature and a timestamp encoding feature; text encoding includes non-numeric text encoding and numeric text encoding, which are the same as those in the above embodiments and will not be elaborated here. For timestamp encoding, if the time text prompt sample carries a timestamp, and the timestamp can be from 0 to 20 seconds, then the times between 0 and 20 seconds can be respectively timestamp-encoded to obtain corresponding time tokens (e.g., each second corresponds to a time token). During the training process, the pre-trained set driving text generation model will give more attention to the corresponding timestamp tokens based on the time in the time text prompt sample.

[0054] After encoding, the second image encoding feature and the second text encoding feature are input into the time token module to output the target image feature; for the time token module, if the second target data sample is an image, the time token module directly outputs the feature of the image. If the second target data sample is a video, the corresponding target image is determined from the second target data sample based on the second text encoding feature (including the feature corresponding to the timestamp), and the target image feature is output. Finally, the target image feature and the second text encoding feature (excluding the feature corresponding to the timestamp) are input into the pre-trained set driving text generation model to output the predicted driving text; the loss value is determined based on the predicted driving text and the corresponding second driving text; the pre-trained set driving text generation model is iteratively trained based on the loss value until the corresponding iterative training stop condition is met, and the target set driving text generation model is obtained. The iterative training stop condition can be that the loss value is less than the loss threshold for a continuously set number of times or the number of iterations is greater than the set iteration number threshold. Among them, the pre-trained set driving text generation model can include a pre-trained second set language model.

[0055] Optionally, input the second image encoding feature and the second text encoding feature into a temporal token module to output a target image feature, including: mapping the timestamp encoding feature and the text encoding feature into a hidden space through the first mapping layer to obtain corresponding timestamp mapping features and text mapping features; processing the preset random matrix and the text mapping feature through a first preset cross-attention mechanism to obtain an attention feature; encoding the second image encoding feature again through a preset encoding model to obtain a new image encoding feature; mapping the new image encoding feature into the hidden space through the second mapping layer to obtain a corresponding image mapping feature; concatenating the image mapping feature and the timestamp mapping feature to obtain a concatenated feature; processing the concatenated feature, the attention feature, and the second image encoding feature through a second preset cross-attention mechanism to obtain a target image feature.

[0056] Among them, the second preset language model includes a first mapping layer and a second mapping layer. Since the second text encoding feature includes a text encoding feature and a timestamp encoding feature, the timestamp encoding feature and the text encoding feature can be respectively mapped into a hidden space through the first mapping layer to obtain corresponding timestamp mapping features and text mapping features; using the preset random matrix as a query Query, and inputting the text mapping feature as a key Key and a value Value into the first preset cross-attention mechanism for processing to obtain an attention feature, where the preset random matrix can be a learnable preset random matrix obtained through random initialization. Encoding the second image encoding feature again through a preset encoding model to obtain a new image encoding feature. In this embodiment, no limitation is imposed on the specific preset encoding model. For example, it can be a lightweight query transformer Q-former model. Mapping the new image encoding feature into the hidden space through the second mapping layer to obtain a corresponding image mapping feature; concatenating the image mapping feature and the timestamp mapping feature to obtain a concatenated feature; inputting the concatenated feature as a key Key, the attention feature as a query Query, and the second image encoding feature as a value Value into the second preset cross-attention mechanism for processing to obtain a target image feature. Among them, both the first preset cross-attention mechanism and the second preset cross-attention mechanism can be the same cross-attention mechanism.

[0057] In this embodiment, by processing the second image encoding feature and the second text encoding feature during the training process through the time token module and adaptively selecting the target image feature, long-term sequence data can be effectively processed, enabling the target set driving text generation model to have the ability to accurately process time cues, solving the deficiency of time perception in the prior art, and providing a key innovation for time understanding in the driving scenario. Compared with the prior art, the time token module enables the model to more effectively understand and predict time-related information in the driving scenario, improving the utilization efficiency of sequence information.

[0058] Optionally, inputting the target image feature and the second text encoding feature into the pre-trained second set language model to output the predicted driving text includes: freezing the parameters of the second set language model trained in the pre-training stage; inputting the target image feature and the text encoding feature into the set language model module to output the predicted driving text.

[0059] Among them, the second set language model further includes a set language model module. In this embodiment, there is no limitation on the set language model in the set language model module. For example, it can be the LoRA language model (Low-Rank Adaptation of Large Language Models, LoRA). Exemplarily, freeze the parameters of the second set language model trained in the pre-training stage; use the target image feature and the text encoding feature to train the LoRA language model in the set language model module.

[0060] Exemplarily, Figure 2 is a training schematic diagram of the pre-trained set driving text generation model provided by the embodiment of the present invention in the fine-tuning stage. As Figure 2 shown, perform image encoding on the second target data sample to obtain the second image encoding feature, that is, Figure 2 the image token Token in, perform text encoding on the time text prompt sample to obtain the second text encoding feature, and the second text encoding feature includes the text encoding feature and the timestamp encoding feature, that is, Figure 2The timestamp token and the text token in can be mapped to the hidden space by the first mapping layer respectively to obtain the corresponding timestamp mapping feature and text mapping feature; taking the set random matrix as the query Query, and taking the text mapping feature as the key Key and value Value and inputting them into the first set cross-attention mechanism for processing to obtain the attention feature; encoding the image token again through the set encoding model to obtain a new image encoding feature. Mapping the new image encoding feature to the hidden space through the second mapping layer to obtain the corresponding image mapping feature; splicing the image mapping feature and the timestamp mapping feature to obtain a spliced feature; taking the spliced feature as the key Key, taking the attention feature as the query Query, and taking the image token as the value Value and inputting them into the second set cross-attention mechanism for processing to obtain the target image feature. The target image feature is also the target image token. Inputting the target image feature and the second text encoding feature into the pre-trained set driving text generation model to output the predicted driving text, so that the pre-trained set driving text generation model can be iteratively trained based on the predicted driving text and the corresponding second driving text until the corresponding iterative training stop condition is met, and the target set driving text generation model is obtained.

[0061] Taking the Q-former model as an example for the set encoding model, taking the video as an example for the second target data sample, taking the Flan-T5 language model as the second set language model, the specific implementation formula of the time token module is as follows:

[0062]

[0063] Among them, represents the video token (also called the image token) before encoding in the Q-former model at time t, represents the video frame at time t, d represents the dimension of the image, which can also be called the visual dimension, and SA() represents the SlotAttention slot attention mechanism module. d' represents the dimension of text encoding, and q l represents the learnable encoding used to convert video content into text information; represents the video token after passing through the Q-former model at time t (i.e., the new image encoding feature).

[0064] q mid =MHCA[q i ,T5 Enc (T p ),T5 Enc (Tp )];

[0065]

[0066] Among them, q i represents the learnable query Query (i.e., a set random matrix), and T5 Enc (T p ) can be understood as a text mapping feature; T p represents the text prompt (not encoded by the text), and T5 Enc () represents the encoder of the Flan-T5 language model; MHCA() represents the multi-head cross-attention module (i.e., the cross-attention mechanism Cross-attention), and q mid represents an intermediate token variable used to integrate the text prompt, and can also be understood as an attention feature.

[0067] T t represents the timestamp (not encoded by the text); q v represents the entire video token before encoding by the Q-former model, represents the entire video token after encoding by the Q-former model; concat() represents concatenation. E vis represents the target image feature.

[0068] In this embodiment, a time token module is introduced to dynamically select the video tokens most relevant to the given prompt through the learnable query Query (set random matrix). These selected tokens (i.e., target image features) are merged into the language model as visual inputs. The learnable query Query is introduced to understand the input prompt through the cross-attention mechanism, while the timestamp and visual embedding (i.e., concatenated features) serve as keys in the cross-attention module. This ensures the alignment of their embeddings in the text feature space.

[0069] The technical solution provided by the embodiments of the present invention determines a training set corresponding to a set driving text generation model; wherein, the training set includes a first training set and a second training set; pre-train the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample in a driving scenario, a spatial text prompt sample, and a corresponding first driving text; wherein, the first target data sample includes a first target image and / or a first target video; fine-tune the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample in a driving scenario, a time text prompt sample, and a corresponding second driving text; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a time token module; the time token module is used for: for each group of second sample pairs, determine a target image feature from the second target data sample based on the time text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the time text prompt sample, and the corresponding second driving text. In the embodiments of the present disclosure, by pre-training the set driving text generation model with the first training set containing spatial text prompt samples and fine-tuning the pre-trained set driving text generation model with the time token module, the spatial and temporal perception capabilities of the target set driving text generation model can be enhanced, thereby improving the generation accuracy of driving texts.

[0070] The training process of the set driving text generation model in the present invention mainly includes: pre-training based on the first training set and fine-tuning on diverse tasks. Different tasks may have different second training sets. At the same time, the first training set includes multiple groups of first sample pairs, and the first sample pairs contain spatial text prompt samples, which can endow the target set driving text generation model with the ability to perform spatial positioning in a driving scenario while retaining its description ability. In fine-tuning, the pre-trained set driving text generation model takes as input a video (or image) and a text prompt (including a timestamp). After encoding into tokens, the time token module collects appropriate tokens according to the text prompt and sends the appropriate tokens (target image features) to the pre-trained set driving text generation model to output a predicted driving text.

[0071] It should be noted that since the first training set can include not only spatial text prompt samples but also general text prompt samples, and the second training set can include not only temporal text prompt samples but also spatial text prompt samples and general text prompt samples, etc., the target-setting driving text generation model proposed by the present invention can achieve the description of the environment, the positioning of objects, the memory of specific events, and the prediction of future scenarios, enabling more comprehensive and accurate driving decisions to be made in complex driving scenarios. Therefore, the target-setting driving text generation model provided by the present invention expands the capabilities of traditional vision-language models, not only having descriptive tasks but also providing more comprehensive and in-depth language understanding within a large spatio-temporal range.

[0072] In an embodiment of the present invention, by constructing a training set containing spatial text prompt samples and pre-training the setting driving text generation model with the first training set containing spatial text prompt samples, a robust positioning ability can be obtained in a large-scale space, enabling the target-setting driving text generation model to more accurately understand the surrounding environment.

[0073] In an embodiment of the present invention, by introducing a temporal token module, long-term sequential data can be effectively processed. By encoding each frame of the image into sparse tokens and establishing a token library, combined with learnable queries, the target-setting driving text generation model can efficiently retrieve the most relevant specific moments and specific content clues related to a given instruction from long-term memory, thereby achieving effective long-term information retrieval.

[0074] Figure 3 It is a flowchart of a method for generating driving text provided by an embodiment of the present invention. The specific steps are as follows:

[0075] S310. Obtain a target text prompt and target data.

[0076] Among them, the target text prompt can be a text prompt in actual applications or in a test set. In this embodiment, no specific restrictions are imposed on the text prompt, which can include any description of the driving scenario, such as descriptions regarding space and time. The target data includes videos and / or pictures in the driving scenario.

[0077] S320. Input the target text prompt and the target data into the target-setting driving text generation model, and output a target driving text.

[0078] Among them, the target-setting driving text generation model is obtained through the model training method in the above embodiment.

[0079] In this embodiment, by processing the target text prompt through the target-setting driving text generation model, the accuracy of the output target driving text is higher, which can improve the user experience.

[0080] Exemplarily, the present invention has been experimented on large-scale autonomous driving datasets such as nuScenes, Ego4D, Youtube, and Waymo datasets. Among them, nuScenes, Ego4D, and Waymo belong to public datasets, and the Youtube autonomous driving data is the data obtained according to the method for determining the training set provided by the present invention.

[0081] Table 1 Embodied scene understanding tasks for autonomous driving

[0082]

[0083]

[0084] By using long videos in the Ego4D dataset to supplement the evaluation of long-term memory, which is missing in the autonomous driving dataset. Scene description and egocentric description tasks have been applied to ordinary vision-language models. S: Spatial span; R: Spatial resolution; T: Total duration; F: Number of frames; #: Number of question-answer QA pairs (sample pairs).

[0085] Table 2 Statistics of the pre-training data of the present invention and comparison with other datasets

[0086]

[0087]

[0088] It can be seen that the pre-training data of the present invention exceeds general vision (LLaVA, VideoChat, and Vid-ChatGPT) and autonomous driving (nuScences-QA, DriveGPT4, and LLM-driver) in terms of both quantity and diversity. #: Number of sample pairs.

[0089] Table 3 Experimental results in six tasks of the nuScenes dataset

[0090]

[0091]

[0092] It can be seen that in the six tasks of the nuScenes dataset, the goal-setting driving text generation model exceeds the previously optimal method in most metrics, verifying the universality of this model. Pr@1 refers to the precision within 1m, Pr@2 refers to the precision within 2m, C represents the CIDEr evaluation metric, R represents the ROUGE-L evaluation metric, and B represents the BLEU evaluation metric.

[0093] As shown in Table 1, Table 2, and Table 3, the target-setting driving text generation model achieved SOTA (state-of-the-art, the best performance) on most key metrics with a relatively small number of parameters across multiple datasets and tasks.

[0094] Table 4 Experimental Results of the Temporal Token Module

[0095]

[0096]

[0097] As shown in Table 4, the model was extended to the Ego4D dataset, and the generality of the temporal token module was verified on four tasks.

[0098] Table 5 Comparison Results of the Number of Parameters

[0099] Method Parameter Quantity (Bytes) BLIP2-opt 2.7B BLIP2-flant5 2.7B LLaMA-Ada 7B LLaVA 7B Otter 7B VideoChat 7B Vid-ChatGPT 7B Target-Setting Driving Text Generation Model 2.7B

[0100] Table 5 shows the comparison of the number of parameters of the adopted large language models (LLMs). As shown in Table 3, Table 4, and Table 5, the target-setting driving text generation model achieved better performance on most key metrics with a relatively small number of parameters compared to other existing methods in ten tasks.

[0101] In this embodiment, to verify the performance of the target-setting driving text generation model in all aspects, an evaluation suite consisting of ten different tasks was constructed. These tasks cover the evaluation of single and comprehensive capabilities in aspects such as description, localization, memory, and prediction, aiming to comprehensively evaluate the performance of the target-setting driving text generation model in driving scene understanding. Experiments have proved that through experiments on the performance of the target-setting driving text generation model on the newly constructed evaluation suite, it is confirmed that the target-setting driving text generation model has excellent performance in cross-domain driving scenarios compared to other methods (such as LLaMA-AdapterV2, LLaVA, Otter, VideoChat, etc.). The visualization of the experimental results of the ten tasks shows that the target-setting driving text generation model has made significant improvements in all aspects compared to the previous BLIP2-flant5 method.

[0102] The present invention constructs an evaluation suite containing ten different tasks to comprehensively evaluate the performance of the target-setting driving text generation model in aspects such as description, localization, memory, and prediction. Compared with the relatively narrow evaluation metrics in the prior art, this new task evaluation suite provides a more comprehensive and accurate standard for the comprehensive evaluation of driving scene understanding.

[0103] Figure 4Schematic structural diagram of a training device for a model provided by an embodiment of the present disclosure. The device includes: a training set determination module 410, a pre-training module 420, and a fine-tuning module 430;

[0104] The training set determination module 410 is configured to determine a training set corresponding to a set driving text generation model; wherein, the training set includes a first training set and a second training set;

[0105] The pre-training module 420 is configured to pre-train the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample, a spatial text prompt sample, and a corresponding first driving text in a driving scenario; wherein, the first target data sample includes a first target image and / or a first target video;

[0106] The fine-tuning module 430 is configured to fine-tune the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample, a time text prompt sample, and a corresponding second driving text in a driving scenario; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a time token module; the time token module is configured to: for each group of second sample pairs, determine a target image feature from the second target data sample based on the time text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the time text prompt sample, and the corresponding second driving text.

[0107] The technical solution provided by the embodiment of the present invention determines a training set corresponding to a set driving text generation model through a training set determination module; wherein, the training set includes a first training set and a second training set; a pre-training module pre-trains the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample in a driving scenario, a spatial text prompt sample, and a corresponding first driving text; wherein, the first target data sample includes a first target image and / or a first target video; a fine-tuning module fine-tunes the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample in a driving scenario, a time text prompt sample, and a corresponding second driving text; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a time token module; the time token module is configured to: for each group of second sample pairs, determine a target image feature from the second target data sample based on the time text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the time text prompt sample, and the corresponding second driving text. In the embodiment of the present disclosure, by pre-training the set driving text generation model through the first training set containing the spatial text prompt sample and fine-tuning the pre-trained set driving text generation model through the time token module, the spatial and temporal perception capabilities of the target set driving text generation model can be enhanced, thereby improving the generation accuracy of the driving text.

[0108] Among them, the training set is composed of a target data sample set, a target text prompt sample set, and a corresponding target driving text set. Optionally, the training set determination module is specifically configured to: obtain an original data sample set in a driving scenario; the original data sample set includes an original image set and / or an original video set; clean each original data sample in the original data sample set; form all the cleaned original data samples that meet the set quality conditions into a target data sample set, input the target data sample set into a set vision-language model, and output a driving text set corresponding to the target data sample set; if the driving texts in the driving text set do not meet the semantic conditions of the corresponding original data samples, re-input the corresponding original data samples into the set vision-language model, re-output the corresponding driving texts until all the driving texts in the driving text set meet the semantic conditions of the corresponding original data samples, and obtain a target driving text set; determine a target text prompt sample set corresponding to the target data sample set through a first set language model; among them, the target data sample set includes a target image set and / or a target video set, and the target text prompt sample set is a spatial text prompt sample set or a temporal text prompt sample set.

[0109] Among them, the set driving text generation model further includes a second set language model; optionally, the pre-training module is specifically configured to: perform image encoding on the first target data sample to obtain a first image encoding feature; perform text encoding on the spatial text prompt sample to obtain a first text encoding feature; pre-train the second set language model based on the first image encoding feature and the first text encoding feature to obtain a pre-trained set driving text generation model.

[0110] Among them, the first text encoding feature includes a non-numeric text encoding feature and a numeric text encoding feature; optionally, the pre-training module is further configured to: for the non-numeric text in the spatial text prompt sample, perform non-numeric text encoding on the non-numeric text to obtain a non-numeric text encoding feature; for the numeric text in the spatial text prompt sample, round the numeric text; perform numeric text encoding on the rounded numeric text to obtain a numeric text encoding feature.

[0111] Optionally, the fine-tuning module is specifically configured to: perform image encoding on the second target data sample to obtain a second image encoding feature; perform text encoding on the time text prompt sample to obtain a second text encoding feature; input the second image encoding feature and the second text encoding feature into a time token module to output a target image feature; input the target image feature and the second text encoding feature into the pre-trained second preset language model to output a predicted driving text; determine a loss value based on the predicted driving text and the corresponding second driving text; perform iterative training on the pre-trained preset driving text generation model based on the loss value until a corresponding iterative training stop condition is met, and obtain a target preset driving text generation model.

[0112] Among them, the second text encoding feature includes a text encoding feature and a timestamp encoding feature; among them, the pre-trained preset driving text generation model includes the pre-trained second preset language model, and the second preset language model includes a first mapping layer and a second mapping layer; optionally, the fine-tuning module is further configured to: respectively map the timestamp encoding feature and the text encoding feature into a hidden space through the first mapping layer to obtain corresponding timestamp mapping features and text mapping features; process a preset random matrix and the text mapping feature through a first preset cross-attention mechanism to obtain an attention feature; re-encode the second image encoding feature through a preset encoding model to obtain a new image encoding feature; map the new image encoding feature into the hidden space through the second mapping layer to obtain a corresponding image mapping feature; splice the image mapping feature and the timestamp mapping feature to obtain a spliced feature; process the spliced feature, the attention feature, and the second image encoding feature through a second preset cross-attention mechanism to obtain a target image feature.

[0113] Among them, the second preset language model further includes a preset language model module; optionally, the fine-tuning module is further configured to: freeze the parameters of the second preset language model trained in the pre-training stage; input the target image feature and the text encoding feature into the preset language model module to output a predicted driving text.

[0114] Figure 5 This is a schematic structural diagram of a driving text generation device provided by an embodiment of the present disclosure. The device includes: an acquisition module 510 and a target driving text output module 520;

[0115] The acquisition module 510 is configured to acquire a target text prompt and target data; among them, the target data includes a video and / or a picture;

[0116] A target driving text output module 520 is configured to input the target text prompt and target data into a target-set driving text generation model, and output a target driving text.

[0117] The above device can execute the methods provided in all the foregoing embodiments of the present invention, and has corresponding functional modules and beneficial effects for executing the above methods. For technical details not described in detail in this embodiment, reference may be made to the methods provided in all the foregoing embodiments of the present invention.

[0118] Figure 6 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0119] As Figure 6 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0120] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0121] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the training of the method model and the generation of driving text.

[0122] In some embodiments, the training of the method model and the generation of driving text can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training of the method model and the generation of driving text described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the training of the method model and the generation of driving text in any other suitable manner (e.g., by means of firmware).

[0123] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0124] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0125] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0126] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0127] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0128] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0129] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0130] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A training method for a model, characterized in that, Including: Determine a training set corresponding to a set driving text generation model; wherein, the training set includes a first training set and a second training set; Pre-train the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample, a spatial text prompt sample, and a corresponding first driving text in a driving scenario; wherein, the first target data sample includes a first target image and / or a first target video; Fine-tune the pre-trained set driving text generation model based on the second training set to obtain a target set driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample, a temporal text prompt sample, and a corresponding second driving text in a driving scenario; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the set driving text generation model includes a temporal token module; the temporal token module is configured to: for each group of second sample pairs, determine a target image feature from the second target data sample based on the temporal text prompt sample, so as to fine-tune the pre-trained set driving text generation model based on the target image feature, the temporal text prompt sample, and the corresponding second driving text.

2. The method according to claim 1, characterized in that, Wherein, The training set is composed of a target data sample set, a target text prompt sample set, and a corresponding target driving text set. Determining a training set corresponding to a set driving text generation model includes: Obtain an original data sample set in a driving scenario; the original data sample set includes an original image set and / or an original video set; Clean each original data sample in the original data sample set; Form all the cleaned original data samples that meet the set quality conditions into a target data sample set, input the target data sample set into a set vision-language model, and output a driving text set corresponding to the target data sample set; If the driving texts in the driving text set do not meet the semantic conditions of the corresponding original data samples, re-input the corresponding original data samples into the set vision-language model to re-output the corresponding driving texts until the driving text set meets the semantic conditions of the corresponding original data samples, and obtain a target driving text set; Determine a target text prompt sample set corresponding to the target data sample set through a first set language model; wherein, the target data sample set includes a target image set and / or a target video set, and the target text prompt sample set is a spatial text prompt sample set or a temporal text prompt sample set.

3. The method according to claim 1, characterized in that, Wherein, The set driving text generation model further includes a second set language model; pre-training the set driving text generation model based on the first training set to obtain a pre-trained set driving text generation model includes: Perform image encoding on the first target data sample to obtain a first image encoding feature; Perform text encoding on the spatial text prompt sample to obtain a first text encoding feature; Pre-train the second preset language model based on the first image encoding feature and the first text encoding feature to obtain a pre-trained preset driving text generation model.

4. The method according to claim 3, wherein Among them, the first text encoding feature includes a non-numeric text encoding feature and a numeric text encoding feature; Performing text encoding on the spatial text prompt sample to obtain a first text encoding feature, including: For the non-numeric text in the spatial text prompt sample, performing non-numeric text encoding on the non-numeric text to obtain a non-numeric text encoding feature; For the numeric text in the spatial text prompt sample, rounding the numeric text; Performing numeric text encoding on the rounded numeric text to obtain a numeric text encoding feature.

5. The method according to claim 4, characterized in that Among them, the pre-trained preset driving text generation model includes the pre-trained second preset language model; fine-tuning the pre-trained preset driving text generation model based on the second training set to obtain a target preset driving text generation model, including: Performing image encoding on the second target data sample to obtain a second image encoding feature; Performing text encoding on the time text prompt sample to obtain a second text encoding feature; Inputting the second image encoding feature and the second text encoding feature into a time token module to output a target image feature; Inputting the target image feature and the second text encoding feature into the pre-trained second preset language model to output a predicted driving text; Determining a loss value based on the predicted driving text and the corresponding second driving text; Performing iterative training on the pre-trained preset driving text generation model based on the loss value until a corresponding iterative training stop condition is met to obtain a target preset driving text generation model.

6. The method according to claim 5, characterized in that, Among them, the second text encoding feature includes a text encoding feature and a timestamp encoding feature; the second preset language model includes a first mapping layer and a second mapping layer; inputting the second image encoding feature and the second text encoding feature into a time token module to output a target image feature, including: Mapping the timestamp encoding feature and the text encoding feature to a hidden space through the first mapping layer respectively to obtain corresponding timestamp mapping features and text mapping features; Processing a preset random matrix and the text mapping feature through a first preset cross-attention mechanism to obtain an attention feature; Encoding the second image encoding feature again through a preset encoding model to obtain a new image encoding feature; Mapping the new image encoding feature to a hidden space through the second mapping layer to obtain a corresponding image mapping feature; Concatenating the image mapping feature and the timestamp mapping feature to obtain a concatenated feature; Processing the concatenated feature, the attention feature, and the second image encoding feature through a second preset cross-attention mechanism to obtain a target image feature.

7. The method according to claim 6, wherein Among them, The second preset language model further includes a preset language model module; inputting the target image features and the second text encoding features into the pre-trained second preset language model, and outputting a predicted driving text, including: Freezing the parameters of the second preset language model obtained during the pre-training phase; Inputting the target image features and the text encoding features into the preset language model module, and outputting a predicted driving text.

8. A method for generating driving text, characterized in that, Including: Obtaining a target text prompt and target data; wherein, the target data includes video and / or pictures; Inputting the target text prompt and target data into a target preset driving text generation model, and outputting a target driving text; wherein, the target preset driving text generation model is obtained by the training method of the model according to any one of claims 1-7.

9. A training device for a model, characterized in that, Including: A training set determination module, configured to determine a training set corresponding to the preset driving text generation model; wherein, the training set includes a first training set and a second training set; A pre-training module, configured to pre-train the preset driving text generation model based on the first training set to obtain a pre-trained preset driving text generation model; wherein, the first training set includes at least one group of first sample pairs, and each group of first sample pairs includes a first target data sample in a driving scenario, a spatial text prompt sample, and a corresponding first driving text; wherein, the first target data sample includes a first target image and / or a first target video; A fine-tuning module, configured to fine-tune the pre-trained preset driving text generation model based on the second training set to obtain a target preset driving text generation model; wherein, the second training set includes at least one group of second sample pairs, and each group of second sample pairs includes a second target data sample in a driving scenario, a temporal text prompt sample, and a corresponding second driving text; wherein, the second target data sample includes a second target image and / or a second target video; wherein, the preset driving text generation model includes a temporal token module; the temporal token module is configured to: for each group of second sample pairs, determine target image features from the second target data sample based on the temporal text prompt sample, so as to fine-tune the pre-trained preset driving text generation model based on the target image features, the temporal text prompt sample, and the corresponding second driving text.

10. A driving text generation device, characterized in that, Including: An acquisition module, configured to acquire a target text prompt and target data; wherein, the target data includes video and / or pictures; A target driving text output module, configured to input the target text prompt and target data into a target preset driving text generation model, and output a target driving text; wherein, the target preset driving text generation model is obtained by the training method of the model according to any one of claims 1-7.

11. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the training method of the model according to any one of claims 1-7 or the generation method of the driving text according to claim 8.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing, when executed by a processor, the training method of the model according to any one of claims 1-7 or the generation method of the driving text according to claim 8.

Citation Information

Patent Citations

  • User identity recognition method and device based on artificial intelligence, terminal and medium

    CN111988294A

  • Speech generation method and device based on pre-training language model, equipment and medium

    CN116364055A