Image processing method and vehicle
By quantifying the correlation between visual tokens and preset visual tokens in an autonomous driving system, redundant information is eliminated and key information is retained, thus solving the problem of low accuracy in visual token pruning operations and improving the computational efficiency and decision-making accuracy of the autonomous driving system.
Patent Information
- Application Number
- CN202511083369.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-21
AI Technical Summary
Existing image processing methods have low accuracy in pruning visual tokens in autonomous driving systems, resulting in wasted computing resources and inference delays, which affect the real-time performance and computational efficiency of autonomous driving systems.
By acquiring and encoding the target image, the degree of association between the visual token and the preset visual token is determined. Redundant background information is removed, and visual tokens that are highly related to the preset business type are retained. A deep learning model and self-attention mechanism are used to quantify the degree of association, thereby achieving precise pruning.
It reduces the waste of computing resources, improves the computing efficiency and decision-making accuracy of autonomous driving systems, ensures that pruning operations are applicable to autonomous driving scenarios, and enhances the processing capabilities of end-to-end autonomous driving models.
Smart Images

Figure CN120997796A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing and automatic driving, in particular to an image processing method and a vehicle. BACKGROUND
[0002] The rapid development of the field of automatic driving puts forward higher requirements for visual processing technology. End-to-end automatic driving systems, such as systems based on visual-language-action models, show great potential in driving decisions. However, traditional visual-language-action models need to process a large number of visual tokens, which will bring significant computational overhead and inference delay, becoming a technical obstacle restricting the real-time performance and computational efficiency of automatic driving systems.
[0003] Currently, in the image processing process, the pruning method for visual tokens mainly focuses on general tasks such as visual question answering and image description. These methods do not fully consider the uniqueness of the automatic driving scene and have low adaptability to the automatic driving scene. When applied to the automatic driving scene, pruning errors may occur, such as a small pruning degree, which will cause a large computational resource overhead, and a large pruning degree, which will cause the loss of key information and affect the execution of subsequent automatic driving tasks, i.e., the accuracy of the related technology for pruning the visual tokens of the image is low.
[0004] At present, there is no effective solution to the above problems. SUMMARY
[0005] The embodiments of the present application provide an image processing method and a vehicle to at least solve the technical problem of low accuracy of related technologies for pruning the visual tokens of the image.
[0006] According to an aspect of an embodiment of the present application, an image processing method is provided, comprising: obtaining a target image and a preset business type, wherein the target image reflects environmental information of an environment around a mobile device; encoding the target image to obtain a plurality of visual tokens, wherein different visual tokens are used to represent features of different image blocks in the target image; determining an association degree between the plurality of visual tokens and a preset visual token, wherein the preset visual token reflects an element attribute corresponding to the preset business type; performing a pruning operation on the plurality of visual tokens based on the association degree to obtain at least one target visual token, wherein the association degree between the target visual token and the preset visual token is greater than a preset degree.
[0007] According to another aspect of the embodiments of the present application, there is also provided an image processing apparatus, comprising: an obtaining module, configured to obtain a target image and a preset service type, wherein the target image reflects environmental information of an environment surrounding a mobile device; an encoding module, configured to encode the target image to obtain a plurality of visual tokens, wherein different visual tokens are used to represent features of different image blocks in the target image; a determining module, configured to determine degrees of association between the plurality of visual tokens and a preset visual token, wherein the preset visual token reflects an element attribute corresponding to the preset service type; and a rejecting module, configured to perform a rejecting operation on the plurality of visual tokens based on the degrees of association to obtain at least one target visual token, wherein the degree of association between the target visual token and the preset visual token is greater than a preset degree.
[0008] According to another aspect of the embodiments of the present application, there is also provided a vehicle, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program, when running, performs the method described above.
[0009] According to another aspect of the embodiments of the present application, there is also provided an electronic device, comprising: a memory storing an executable program; and a processor configured to run the program, wherein the program, when running, performs the method in the various embodiments of the present application.
[0010] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, comprising a stored executable program, wherein the executable program, when running, controls a device where the computer-readable storage medium is located to perform the method in the various embodiments of the present application.
[0011] According to another aspect of the embodiments of the present application, there is also provided a computer program product, comprising a computer program, wherein the computer program, when executed by a processor, implements the method in the various embodiments of the present application.
[0012] According to another aspect of the embodiments of the present application, there is also provided a computer program product, comprising a non-volatile computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method in the various embodiments of the present application.
[0013] According to another aspect of the embodiments of the present application, there is also provided a computer program, wherein the computer program, when executed by a processor, implements the method in the various embodiments of the present application.
[0014] In the embodiment of the present application, first, the target image and the preset service type can be acquired, the target image reflecting the environmental information of the surrounding environment of the mobile device; then, the target image can be encoded to obtain a plurality of visual tokens, different visual tokens being used to represent the features of different image blocks in the target image; then, the association degrees between the plurality of visual tokens and the preset visual token can be determined, the preset visual token reflecting the element attribute corresponding to the preset service type; finally, the plurality of visual tokens can be pruned based on the association degrees to obtain at least one target visual token, the association degree between the target visual token and the preset visual token being greater than a preset degree. It is easy to note that, after the target image of the surrounding environment of the mobile device is encoded to obtain a plurality of visual tokens, the association degrees between the plurality of visual tokens and the preset visual token are determined, and the plurality of visual tokens are pruned based on the association degrees, the redundant background information is removed, and the visual tokens highly related to the preset service type are retained, so that the target visual token obtained can be used as the input of subsequent decision-making, such as path planning and obstacle avoidance, so that the number of target visual tokens obtained by the pruning operation is greatly reduced relative to the number of the plurality of visual tokens, and the visual information contributing to driving decision-making is retained, so that the pruning operation of the visual token is more suitable for the automatic driving scene of the mobile device, and the technical effects of reducing the waste of computing resources, improving the computing efficiency of the automatic driving system, and enabling the end-to-end automatic driving model to more efficiently process visual input and make more accurate driving decisions are achieved, thereby solving the technical problem of low accuracy of the pruning operation of the visual token of the image in the related art. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the present application and are incorporated in and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0016] Figure 1 is a flowchart of an image processing method according to an embodiment of the present application;
[0017] Figure 2 is a schematic diagram of training data according to an embodiment of the present application;
[0018] Figure 3 is a schematic diagram of a pruner training process according to an embodiment of the present application;
[0019] Figure 4 is a schematic diagram of generating a target driving path of a mobile device according to an embodiment of the present application;
[0020] Figure 5 is a schematic diagram of an image processing device according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to make the person skilled in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.
[0022] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0023] According to the embodiments of the present application, an embodiment of an image processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.
[0024] Embodiments of the present application provide an image processing method. The image processing method can be used to provide image processing functions for a preset application scenario. The preset application scenario can include application scenarios of various mobile devices. The mobile device can be a movable vehicle such as a vehicle, a motorcycle, an aircraft, an airplane, a ship, or a driving experience device such as an intelligent cockpit for driving experience, a driving experience game device, and the like, but is not limited thereto. For the vehicle field, the preset application scenario can be the following scenarios: an autonomous driving scenario, an artificial intelligence (AI) driving scenario for a family car, a user driving a vehicle scenario, an assisted navigation driving scenario, an intelligent navigation guided pilot (NGP) scenario in an urban area or a high-speed area. In addition, the preset application scenario can include, but is not limited to, an image processing scenario for assisting navigation of a truck in the logistics transportation field, an image processing scenario for assisting navigation of an agricultural vehicle in the agricultural machinery field, an image processing scenario for controlling a drone, and an image processing scenario for controlling an intelligent robot such as a cleaning robot, a service robot, and a delivery robot.
[0025] When the preset application scenario is a scenario other than the vehicle field, those skilled in the art should understand that the vehicle in the image processing method can be replaced by other objects (such as agricultural machinery, drones, robots, and the like), and the image processing function applied to the vehicle can be replaced by an image processing function related to other objects. On this basis, the specific embodiments of the image processing method are exemplarily described in the embodiments of the present application taking the vehicle field as an example.
[0026] Figure 1 is a flowchart of an image processing method according to an embodiment of the present application, as shown in Figure 1 , the method comprises the following steps:
[0027] In step S102, a target image and a preset service type are obtained.
[0028] The target image reflects environmental information of an environment around the mobile device.
[0029] The mobile device described above can be a device with autonomous or non-autonomous movement capability, such as a vehicle, a truck, a motorcycle, a drone, an unmanned ship, a robot, etc., but is not limited thereto and can be determined according to the actual application scenario. In the embodiments of the present application, the vehicle scenario is taken as an example for detailed description, and the vehicle here can be various types of vehicles, which are not specifically limited in the present application and can be determined according to the actual application scenario. For example, for the use dimension, the vehicle described above can be a passenger car, a commercial vehicle, an agricultural machinery vehicle, an engineering vehicle, etc.; for the power system dimension, the vehicle described above can be a fuel vehicle, an electric vehicle, a hybrid vehicle, etc.; for the structure dimension, the vehicle described above can be a sedan, a truck, a passenger car, etc.; for the driving mode dimension, the vehicle described above can be an autonomous driving vehicle, a manual driving vehicle, an assisted driving vehicle, etc.
[0030] The target image described above can refer to an image of the surrounding environment captured by the visual sensor of the mobile device. The target image can contain various elements of the mobile device in the driving scenario, such as roads, vehicles, pedestrians, traffic signs, lane lines, etc., and the target image can serve as a data basis for the mobile device to perceive, understand, plan and control.
[0031] The preset business type described above can refer to a specific application scenario of the target image for the mobile device, for example, the target image can be used to generate a target visual token to plan a path for the mobile device, avoid obstacles, etc., and the preset business type can also be determined according to actual needs, which is not limited here.
[0032] As an optional implementation, the mobile device can be equipped with visual perception sensors such as cameras, etc., and a multi-angle, multi-focal camera array can also be installed on the mobile device to cover the front, both sides, rear, etc. field of view, so as to comprehensively perceive the complex road conditions. Specifically, real-time environment images can be captured by cameras on the mobile device, i.e. target images are obtained. These cameras can include but are not limited to front-view cameras, surround-view cameras, night-vision cameras, etc. to ensure that clear and relatively complete visual information can be obtained under different light conditions and angles. At the same time, the captured target images can be pre-processed, which can include but is not limited to denoising, color correction, distortion correction, etc. to improve the accuracy and efficiency of subsequent feature extraction. Image pre-processing can also include image format conversion to ensure that images output by various cameras can be in a unified format for subsequent processing. For the obtained target images, region of interest detection can also be performed to reduce the scope of subsequent encoding and pruning, so as to facilitate processing of areas that have a greater impact on driving decisions.
[0033] In the above process, the target image of the environment around the mobile device can truly reflect the environment around the mobile device, providing a data basis for efficient automatic driving. Rapid and accurate acquisition of the target image helps to speed up the response time of the entire end-to-end driving system and enhance the real-time nature of automatic driving decisions.
[0034] In step S104, the target image is encoded to obtain a plurality of visual tokens.
[0035] Different visual tokens are used to represent the features of different image blocks in the target image.
[0036] The visual token mentioned above can refer to a series of feature vectors obtained after the target image is processed by a visual encoder. Each visual token can represent the visual features of a specific patch in the target image. The visual token can serve as a high-dimensional representation of visual information, which can include color, texture, shape, and other information, and can be used for subsequent model inference and decision-making.
[0037] As an optional implementation, the acquired target image can be encoded to obtain a plurality of visual tokens. The encoding process can include dividing the target image into a plurality of image blocks, encoding the features of each image block through a deep learning model, and finally outputting a plurality of visual tokens (tokens). Each visual token can represent the abstract features of an image block. Specifically, the target image can be divided into a plurality of image blocks, and each image block can be converted into a fixed-dimensional image block vector through an embedding layer for subsequent model processing. Then, a deep learning model can be used to encode the features of the image block vector. For example, a self-attention mechanism can be used to capture the correlation between image block vectors and enhance the feature representation capability. The obtained visual token can contain global context information of the image block. Considering that image blocks of different sizes carry different information richness, a multi-scale encoder can also be used to encode image blocks of different resolutions, and then a specific fusion strategy can be used to aggregate multi-level features to generate visual tokens with stronger comprehensive performance.
[0038] In the above process, the target image is encoded to obtain a plurality of visual tokens, which realizes the conversion of image information into machine-readable feature representation. Through deep abstraction of the feature encoding model, the expression ability and decision value of the image information are enhanced, and the obtained visual token effectively condenses the feature information of the image block, which is conducive to accelerating the calculation and transmission, and also reduces the storage demand. The encoded visual token can have a unified format and dimension, which simplifies the interface design of the subsequent module and also provides convenience for subsequent pruning operations.
[0039] In step S106, the association degree between each visual token and a preset visual token is determined.
[0040] The preset visual token reflects an element attribute corresponding to the preset business type.
[0041] The preset visual token can be an element attribute token corresponding to the preset business type. In a case where the preset business type is to generate a target visual token by using a target image to perform path planning, obstacle avoidance, and the like on a mobile device, the preset visual token can be an object in the target image that has a correlation with a mobile device driving scene, such as a foreground object. In an automatic driving scene of a vehicle, the preset visual token can include other vehicles, pedestrians, lane lines, traffic signal lights, traffic signs, and the like, and the preset visual token is relatively important for driving decisions, such as avoiding pedestrians, following a preceding vehicle, and identifying traffic signals. Compared with other background objects (such as buildings and the sky), the preset visual token has more guiding significance, and the preset visual token needs to be highlighted and identified in visual information processing.
[0042] The preset visual token in the present application can be a reference visual token obtained by training and used to represent a foreground object feature. In the pruner of the present application, the preset visual token can guide the pruner to identify which visual tokens have a high correlation with the foreground object.
[0043] The correlation degree can be a numerical result obtained by quantifying the correlation degree between the plurality of visual tokens and the foreground object, such as a quantification result, and can specifically be a saliency score of each visual token. The higher the saliency score, the stronger the correlation between the visual token and the foreground object, and the greater the contribution to driving decisions. The correlation degree can be used as a basis for the pruning strategy of the present application and can be used to guide the pruner to screen and retain visual tokens.
[0044] As an optional implementation, in order to be more suitable for the automatic driving scene of the mobile device, the correlation between each visual token and the foreground object related to the driving scene of the mobile device can be quantified, it can be determined which visual token belongs to the information important to the driving decision and which visual token belongs to the background or redundant information. Specifically, a preset visual token can be obtained by pre-training, and the preset visual token can represent the known characteristics of the foreground object, such as the characteristics of vehicles, pedestrians, traffic lights and the like. The preset visual token can be a statistical feature obtained by learning a large amount of labeled data, or a feature obtained by preprocessing through a specific visual model. Then, the feature interaction and quantitative comparison between each visual token and the preset visual token can be performed, which can be completed by calculating the similarity or correlation between each visual token and the preset visual token, for example, using dot product, cosine similarity, self-attention mechanism and the like. And a specific function or model, such as a feedforward neural network, can be used to calculate the saliency score of each visual token as a quantitative result. In the quantification process, a scene adaptability adjustment mechanism can also be introduced, and the weights of the preset visual tokens can be dynamically adjusted according to different driving environments, such as highways, urban roads and the like, to realize the flexibility and scene specificity of the quantification process.
[0045] In the above process, by quantifying the correlation between each visual token and the foreground object, the visual information contributing to the driving decision can be quantified, which is more suitable for the automatic driving scene of the mobile device, and can provide accurate data basis for subsequent pruning and decision making. The quantification result can guide the pruning operation, eliminate redundant background information, retain more critical foreground information, reduce the waste of computing resources, and improve the computing efficiency of the automatic driving system.
[0046] Step S108, based on the correlation degree, the plurality of visual tokens are pruned to obtain at least one target visual token.
[0047] Among them, the correlation degree between the target visual token and the preset visual token is greater than the preset degree.
[0048] The above pruning operation can refer to the process of discarding visual tokens with low correlation with the foreground object based on the quantification result. The pruning operation on part of the plurality of visual tokens can reduce the consumption of computing resources and improve the efficiency of model inference, while trying to retain more important information for driving decisions.
[0049] The at least one target visual token described above can refer to the visual token retained after the pruning operation on part of the plurality of visual tokens.
[0050] The preset degree can refer to a predetermined standard degree or threshold for measuring the degree of association between the target visual token and the preset visual token in the process of eliminating the visual token. The preset degree can be preset according to the specific needs of the autonomous driving scene, which is not limited here.
[0051] As an optional implementation, the pruning operation can be performed on part of the plurality of visual tokens based on the quantization result, that is, the pruning operation can be performed on the visual tokens containing non-key information, and the target visual tokens that are more important for driving decision can be reserved. Specifically, it can be determined which visual tokens can be reserved and which visual tokens can be eliminated based on the quantization result. For example, the quantization result of each visual token can be compared with a preset threshold, and the visual token with a quantization result lower than the preset threshold can be eliminated, that is, the visual token with a quantization result higher than or equal to the preset threshold can be reserved as the target visual token, or the visual tokens can be arranged in descending or ascending order according to the obtained quantization result of each visual token. Such a sorting strategy can ensure that the foreground visual tokens related to driving decisions are located at one end of the list, facilitating selective reservation. Then, based on the sorting result, the visual tokens with higher quantization results can be reserved according to a preset proportion or quantity, for example, 25% of the visual tokens with higher quantization results can be reserved, and the target visual tokens obtained after the pruning operation can include foreground objects such as vehicles, pedestrians, and traffic signals, obtaining more important information for driving decisions. The pruning threshold of the quantization result can also be adjusted to dynamically determine the degree of pruning, which can be flexibly adapted to different driving conditions, such as different scenes on highways and urban streets, to achieve dynamic and adaptive resource allocation. Eliminating redundant visual tokens can significantly reduce the computational load and memory requirements of the model, which helps to improve the computational efficiency.
[0052] In the above process, by retaining the visual tokens highly related to the foreground objects, a simplified token sequence can be formed, and the target visual tokens obtained can be used as inputs for subsequent decisions such as path planning and obstacle avoidance. The number of visual tokens can be significantly reduced compared to the plurality of visual tokens before the elimination operation, but the key information can be retained, thereby improving the computational efficiency and improving the decision performance. Through such a simplified target visual token, the end-to-end autonomous driving model can more efficiently process visual input and make more accurate driving decisions.
[0053] In the embodiment of the present application, first, the target image and the preset service type can be acquired, the target image reflecting the environmental information of the environment around the mobile device; then, the target image can be encoded to obtain a plurality of visual tokens, different visual tokens being used to represent the features of different image blocks in the target image; then, the association degrees between the plurality of visual tokens and the preset visual tokens can be determined, the preset visual tokens reflecting the element attributes corresponding to the preset service type; finally, the plurality of visual tokens can be pruned based on the association degrees to obtain at least one target visual token, the association degree between the target visual token and the preset visual token being greater than a preset degree. It is easy to note that, after the target image of the environment around the mobile device is encoded to obtain a plurality of visual tokens, the association degrees between the plurality of visual tokens and the preset visual tokens are determined, and the plurality of visual tokens are pruned based on the association degrees, the redundant background information is removed, and the visual tokens highly related to the preset service type are retained, so that the target visual token obtained can be used as the input of subsequent decision-making, such as path planning and obstacle avoidance, so that the number of target visual tokens obtained by the pruning operation is greatly reduced relative to the number of the plurality of visual tokens, and the visual information contributing to driving decision-making is retained, so that the pruning operation of the visual token is more suitable for the automatic driving scene of the mobile device, and the technical effects of reducing the waste of computing resources, improving the computing efficiency of the automatic driving system, and enabling the end-to-end automatic driving model to process visual input more efficiently and make more accurate driving decisions are achieved, thereby solving the technical problem of low accuracy of pruning the visual tokens of the image in the related art.
[0054] Optionally, determining the association degrees between the plurality of visual tokens and the preset visual tokens comprises: mapping, by a target pruner, the plurality of visual tokens from an image feature space to a target feature space to obtain a plurality of target tokens, wherein the image feature space is used to represent the visual information of the target image, and the target feature space is a feature space corresponding to the preset visual tokens; and mapping, by the target pruner, the plurality of target tokens from the target feature space to a numerical space to obtain the association degrees.
[0055] The image feature space described above can be a space in which a feature vector generated after the target image is processed by a visual encoder is located. In the image feature space, each visual token can represent the visual information of a different part of the target image, and can include feature descriptions in multiple dimensions such as color, texture, shape, etc.
[0056] The foreground object feature space described above can be a space used to represent the attributes of the foreground objects strongly related to driving decision-making. By comparing the visual tokens with the preset visual tokens, the tokens in the image feature space can be mapped to the foreground object feature space, thereby facilitating the extraction of features related to the foreground objects.
[0057] The numerical space mentioned above can refer to a space for quantitatively describing the saliency of the foreground token. The numerical space can reflect the importance of a certain visual token to the autonomous driving decision. By mapping the target token to the numerical space, a specific saliency score can be assigned to each visual token, facilitating subsequent sorting and pruning operations.
[0058] As an optional implementation, a plurality of visual tokens are mapped from an image feature space to a foreground object feature space by using a preset visual token to obtain a plurality of target tokens. Each visual token can be interacted with the preset visual token in terms of features, and each visual token can be spliced with the preset visual token and then input into a query neural network module for implementation. The query neural network module can enhance the features related to the foreground object in each visual token by using deep learning techniques such as self-attention mechanism, and map the plurality of visual tokens from the image feature space to the foreground object feature space. Then, the plurality of target tokens can be mapped from the foreground object feature space to the numerical space to obtain the quantization results of the plurality of visual tokens. For example, a feedforward neural network can be used to further process the target token in the foreground object feature space, compress the target token into a numerical form, and obtain the foreground saliency score of each visual token, which can be used as the quantization result.
[0059] In the above process, by mapping each visual token to the foreground object feature space, more attention can be paid to the foreground object information closely related to the driving decision, which helps to improve the accuracy of the driving decision. Mapping the target token to the numerical space converts the abstract feature representation into a quantitative and intuitive numerical form, which facilitates the subsequent sorting and pruning process.
[0060] Optionally, the plurality of visual tokens are mapped from the image feature space to the target feature space by using a target pruner to obtain the plurality of target tokens, including: splicing the plurality of visual tokens with a preset visual token to obtain a spliced token; inputting the spliced token into a query neural network module in the target pruner, and using the query neural network module to interact the plurality of visual tokens with the preset visual token in terms of features to obtain the plurality of target tokens.
[0061] The spliced token mentioned above can be a combined token formed by connecting the plurality of visual tokens with the preset visual token. The spliced token obtained by splicing can be used as an input token for subsequent processing units. Splicing the plurality of visual tokens with the preset visual token to obtain the spliced token can enable each visual token to carry additional foreground object feature information, thereby more effectively enhancing the feature expression related to the foreground object in the subsequent feature interaction.
[0062] The query neural network module can be a neural network structure used to extract specific information from input data. The query neural network module can be used to perform complex interactions between input features to obtain higher-level semantic information. For example, the query neural network module can be based on a self-attention mechanism, which can establish associations between different parts of the input data, thereby improving the richness and accuracy of feature representation.
[0063] As an optional implementation, the plurality of visual tokens and the preset visual token can be spliced to obtain a spliced token. The splicing process can be to directly connect the feature vectors of the visual tokens and the preset visual token, and input them in a combined form into the query neural network module. The query neural network module can process the spliced token through an attention mechanism or the like to capture the interaction relationship between the plurality of visual tokens and the preset visual token in the spliced token. Finally, the query neural network module can output a target token as a high-level feature representation after feature interaction.
[0064] In the above process, the query neural network module can obtain additional contextual information or structural guidance through the preset visual token, thereby helping to improve the effect of feature interaction. By using the query neural network module through a self-attention mechanism, long-distance dependency relationships between the plurality of visual tokens and the preset visual token in the spliced token can be captured, thereby making the feature representation of the output target token more comprehensive and accurate.
[0065] Optionally, the plurality of target tokens are mapped from the target feature space to the numerical space by using the target pruning device to obtain the association degree, including: enhancing the association degree between the plurality of target tokens and the preset visual token to obtain a plurality of fusion tokens; inputting the plurality of fusion tokens into a feedforward neural network module in the target pruning device, and mapping the plurality of fusion tokens from the target feature space to the numerical space by using the feedforward neural network module to obtain the association degree.
[0066] The fusion token can be a token obtained by performing feature enhancement on the feature in the target token for representing the foreground object, and the fusion token can more prominently represent the foreground object information associated with the driving decision in the target image, such as vehicles, pedestrians, traffic signals, and the like.
[0067] The feedforward neural network module can be a neural network structure in which input data is directly transmitted from an input layer to an output layer through a series of hidden layers. The feedforward neural network module can be responsible for mapping the enhanced fusion token from the foreground object feature space to the numerical space for significance quantization, and the output quantization result is used for subsequent visual token pruning.
[0068] As an optional implementation, the features in the plurality of target tokens for representing the foreground object are enhanced by using the preset visual token, to obtain a plurality of fusion tokens, and the features associated with the foreground object are strengthened by the attention mechanism, so that each fusion token more concentratedly represents the attributes of the foreground object. Then, the plurality of fusion tokens obtained can be input into the feedforward neural network module, and the feedforward neural network module can convert the fusion token from the foreground object feature space to the numerical space through the weighted summation of multiple layers of neurons, output the quantitative result in the numerical form, and represent the foreground saliency of each visual token.
[0069] In the above process, the features in the plurality of target tokens for representing the foreground object are enhanced to obtain a plurality of fusion tokens, and the features associated with the foreground object are strengthened by the attention mechanism, so that each fusion token more concentratedly represents the attributes of the foreground object. Through the mapping of the feedforward neural network module, the foreground features of the fusion token are converted into the saliency quantitative result in the numerical space, the visual tokens can be sorted and pruned based on the quantitative result in the numerical form, and the visual tokens with more foreground information can be ensured to be retained, thereby improving the utilization of computing resources.
[0070] Optionally, the plurality of visual tokens are removed based on the correlation degree to obtain at least one target visual token, including one of the following: removing the visual token with a quantitative result less than a preset threshold from the plurality of visual tokens to obtain at least one target visual token; sorting the plurality of visual tokens in descending order of the quantitative result to obtain a visual token sequence, and removing a preset number of visual tokens at the rear of the sequence to obtain at least one target visual token, wherein the proportion of the preset number of visual tokens in the visual token sequence satisfies a preset proportion, and the preset proportion is determined based on the scene type of the driving scene.
[0071] The preset threshold mentioned above can be a predetermined standard threshold for distinguishing the foreground saliency of the visual token in the process of removing the visual token. The preset threshold can be preset according to the specific needs of the autonomous driving scene, which is not limited here.
[0072] As an optional implementation, the quantization results of each visual token can be compared with a preset threshold, and the visual token with a quantization result lower than the preset threshold can be removed, that is, the visual token with a quantization result higher than or equal to the preset threshold can be retained as a target visual token. The selection of the preset threshold can affect the strength of pruning and the amount of retained visual information. The preset threshold can be set in combination with different factors, which can include but are not limited to the complexity of the current scene, available computing resources, etc., to adapt to different driving scenes. For example, on a wide road, since the scene is relatively simple and there are fewer dynamic obstacles, the preset threshold can be set higher, so as to remove more background tokens and reduce the computational burden; while on a complex and busy urban street, since there are many pedestrians, vehicles and obstacles, the preset threshold can be set lower to ensure that more key foreground information can be retained.
[0073] In the above process, the setting of the preset threshold can help to retain more critical foreground information and remove the remaining redundant background information, which is helpful for the autonomous driving decision system to make more accurate decisions within limited time and resources. The preset threshold can be flexibly adjusted according to specific scenes and conditions, so that information filtering and pruning can be adaptively performed in different driving scenes, maintaining stable and efficient decision performance.
[0074] The above visual token sequence can refer to a set of visual tokens arranged from high to low according to the quantization results. The more critical foreground information in the visual token sequence, such as pedestrians, vehicles, traffic lights, etc., can be located at the front end of the visual token sequence, and the more redundant background information can be located at the rear end of the visual token sequence, providing a clear data basis for subsequent pruning operations.
[0075] The above preset number can refer to the number of visual tokens to be removed in the pruning operation. The preset number can be determined in advance according to the performance of the autonomous driving decision system, computing resources and specific requirements of the driving scene, which is not limited here.
[0076] The above preset ratio can refer to the proportion of the preset number in the visual token sequence. The preset ratio can be dynamically adjusted according to the type of the driving scene, such as a highway, an urban road, etc. For example, in a complex scene, the preset ratio can be set lower to retain more critical foreground information; while in a simple scene, the preset ratio can be set higher to reduce redundant computation.
[0077] As another optional implementation, the visual tokens can be ranked from high to low based on the quantification results to form a visual token sequence. Then, a preset proportion can be determined as an indicator of the pruning strength according to the type of the current driving scene. Next, a preset number of visual tokens can be calculated, i.e., the preset number can be determined by multiplying the length of the visual token sequence by the preset proportion. Finally, the last preset number of visual tokens in the visual token sequence can be removed, and the remaining visual tokens in the visual token sequence, i.e., the front-end visual tokens in the visual token sequence, can be used as the target visual tokens.
[0078] In the above process, the dynamic setting of the preset proportion enables the pruning operation to be individually adjusted in different driving scenes, which can avoid the loss of key information due to excessive pruning in complex scenes and reduce unnecessary calculations in simple scenes, thereby achieving scene adaptability of the pruning strategy.
[0079] Optionally, the target image is encoded to obtain a plurality of visual tokens, including: performing a blocking operation on the target image to obtain a plurality of image blocks; inputting the plurality of image blocks into a visual encoder, and respectively encoding the plurality of image blocks by using the visual encoder to obtain a plurality of visual tokens, wherein the plurality of visual tokens have a corresponding relationship with the plurality of image blocks.
[0080] The visual encoder described above can be a deep learning model used for processing image or video data. The visual encoder can convert an input image into a set of compact and abstract feature representations, i.e., visual tokens. The visual encoder can be determined according to actual needs, which is not limited here.
[0081] As an optional implementation, the target image can be subjected to a blocking operation to divide the target image into a plurality of smaller image blocks. The blocking operation can adapt to the input requirements of the visual encoder, for example, each image block can be a fixed-size region, etc., which is not limited here. Then, these image blocks can be respectively input into the visual encoder for encoding processing. The visual encoder can extract key visual features from each image block by applying a series of convolutional layers, self-attention mechanisms or other deep learning techniques, and convert the image block into a visual token. The obtained visual token contains abstract information of the image block. The visual token retains the visual features of the corresponding image block and also captures the relevance between image blocks through the global attention mechanism of the encoder, which helps to understand the scene structure in the image, such as the relationship between the image foreground and the image background.
[0082] In the above process, by performing the blocking and encoding operations on the target image, the local features of each image block are extracted, and the global attention mechanism can also integrate the correlation information across the image blocks, which helps the autonomous driving decision system to understand more complex scenarios in the driving environment. The blocking operation enables flexible processing of target images of different sizes and resolutions, providing support for consistency and robustness in different vehicles and different camera configurations. The resulting visual tokens effectively condense the feature information of the image blocks, which is conducive to accelerating the calculation and transmission, while also reducing the storage requirements. The encoded visual tokens can have a unified format and dimension, simplifying the interface design of subsequent modules and providing convenience for subsequent pruning operations.
[0083] Optionally, the training step of the target pruner includes the following steps: obtaining training data, wherein the training data includes a training image and a mask image, and the mask image is used to represent whether a plurality of pixel points in the training image belong to an element corresponding to a preset business type; encoding the training image to obtain a plurality of training visual tokens, wherein different training visual tokens are used to represent the features of different training image blocks in the training image; determining the correlation degree between the plurality of training visual tokens and an initial visual token using an initial pruner, wherein the initial visual token reflects the element attribute corresponding to the preset business type; performing image reconstruction based on the training quantization result and the plurality of training visual tokens to obtain a first reconstructed image and a second reconstructed image, wherein the first reconstructed image is used to represent the image of the region where the element corresponding to the preset business type is located, and the second reconstructed image is used to represent the image of the region other than the region where the element corresponding to the preset business type is located; adjusting the parameters of the initial pruner and the initial visual token based on the training image, the mask image, the first reconstructed image, and the second reconstructed image to obtain the target pruner and a preset visual token.
[0084] The training data mentioned above can refer to a data set used to train the pruner. The training data can include two parts: training images and corresponding mask images. The training data can be used to teach the pruner how to distinguish and quantify the foreground object information in the image.
[0085] The training image mentioned above can refer to an input image used to train the pruner. The training image can come from various driving scenarios and can cover elements closely related to driving decisions such as roads, pedestrians, vehicles, and traffic signs. This can enable the pruner to learn to recognize and process these key information in the real world.
[0086] The mask image mentioned above can refer to an image format used to represent which pixels in the training image belong to the training object, i.e., the foreground element that needs to be focused on. The mask image can be a black and white image, where white pixels can represent the foreground object position, and black pixels can represent the background or non-object area. The mask image provides a visual way for the pruner to learn which visual tokens are associated with driving decisions, thereby guiding the training of the pruner.
[0087] As an optional implementation, training data can be obtained, which can be composed of training images and mask images, and can provide the initial pruner with rich scenario information, so that the initial pruner learns to distinguish which visual tokens are more important to driving decisions and which visual tokens can be safely regarded as redundant information. The training data can be obtained in the following way: a large number of training images can be obtained through various ways, which can cover various driving scenarios that the mobile device may encounter, including but not limited to urban streets, highways, rural roads, driving environments under adverse weather conditions, etc. Secondly, each training image can be carefully labeled to generate a mask image, which clearly indicates which pixels belong to the training object and which pixels belong to the background. This process can be completed using professional tools or manual labeling to ensure the accuracy and fineness of the mask and provide high-quality supervision signals for the training of the pruner. The obtained training data can also undergo a series of preprocessing steps such as image scaling, normalization, format conversion, etc. to ensure compliance with the input requirements of the pruner.
[0088] In the above process, the acquisition of training data provides the pruner with rich visual information, and also clearly indicates the position and range of the foreground object through the mask image, helping the pruner to learn to accurately identify visual elements highly related to driving decisions during the training process, thereby more effectively filtering and retaining key visual tokens in actual application. The training data can contain training images of multiple scenarios and multiple weather conditions, which can ensure that the pruner can maintain high performance under various conditions.
[0089] The training visual token mentioned above can refer to the feature representation generated by the visual encoder after encoding the training image. Each training visual token can correspond to a specific region of the training image, i.e., a training image block. The training visual token can contain visual features of the region, such as color, texture, shape, etc.
[0090] As an optional implementation, the training image can be encoded by a visual encoder to generate a series of training visual tokens, which can belong to the feature abstraction of different regions in the training image, and the training visual tokens convert the visual information into a form that can be processed by the model, providing a basis for subsequent pruning and decision-making. The visual encoder can divide the training image into multiple image blocks, and the visual information of each image block can be encoded into a vector by a neural network, i.e., a training visual token.
[0091] In the above process, the training image is converted into a training visual token in vector form by the visual encoder, realizing feature abstraction and data dimension reduction, providing more compact and efficient visual information representation for the pruner, facilitating subsequent calculation and analysis.
[0092] The initial visual token mentioned above can refer to a pre-set feature used to initialize the cognition of the pruner on the training object, i.e., a pre-set visual token that has not been trained or has not been trained completely.
[0093] The initial pruner mentioned above can refer to an untrained pruner, which can be used to preliminarily quantify the association between the visual token and the foreground object. The initial pruner analyzes the initial visual token and the training visual token to try to understand which tokens are highly related to the training object, such as pedestrians and vehicles, providing a starting point for subsequent training.
[0094] The training quantization result mentioned above can refer to a numerical evaluation result of the degree of association between the training visual token and the training object. Through the processing of the initial pruner, each training visual token can have a training quantization result, reflecting the degree of contribution to the driving decision, which can be used as data basis for further training and improvement of the pruner.
[0095] As an optional implementation, the degree of association between the training visual token and the training object can be quantified based on the initial visual token and the initial pruner, which can evaluate the contribution of each visual token to the identification and understanding of the key foreground object, providing a basis for subsequent selection and pruning. By splicing the training visual token with the initial visual token and inputting it into the query neural network module, the interaction and fusion of features can be realized. By utilizing the computing power of the neural network, the weight of the visual token is dynamically adjusted, making the visual token more close to the foreground feature associated with the training object. The visual token after feature interaction can be further processed, mapped to a numerical space by the feedforward neural network module, and thus the training quantization result of each training visual token is obtained. This quantization process converts the abstract visual features into comparable numerical values, providing a quantitative basis for pruning.
[0096] In the above process, by quantifying the degree of association, the pruner can more accurately identify which visual tokens are more important to driving decisions, thereby prioritizing the preservation of this information in subsequent processes, which can significantly enhance the pruner's ability to identify key foreground objects. The quantification process allows the pruner to automatically select to retain or discard based on the saliency of the visual tokens, avoiding the processing of a large amount of insignificant background information, which can reduce the consumption of computing resources, thereby improving the efficiency and real-time performance of the autonomous driving system.
[0097] The first reconstructed image mentioned above can refer to an image reconstructed based on visual tokens with high association with the training object, i.e., high saliency tokens, which can be a reconstructed foreground image. The first reconstructed image can be an image of the region where the foreground object is located, reconstructed by inputting visual tokens with high association with the training object into the image reconstruction module, which helps the pruner learn how to accurately retain foreground information.
[0098] The second reconstructed image mentioned above can refer to an image reconstructed based on visual tokens with low association with the training object, i.e., low saliency tokens, which can be a reconstructed background image. The second reconstructed image focuses on reproducing visual information that is considered non-critical or belongs to the background, which can help the pruner understand which information can be safely discarded without affecting driving decisions.
[0099] As an optional implementation, image reconstruction through an adversarial foreground-background reconstruction strategy can enhance the pruner's ability to distinguish between foreground information and background information. Specifically, according to the training quantification results, multiple training visual tokens can be divided into two groups: one group is high saliency tokens with high association with the training object, and the other group is low saliency tokens with low association. High saliency tokens can correspond to foreground elements in the image, while low saliency tokens can represent more background information. Then, high saliency tokens and low saliency tokens can be used for image reconstruction. For high saliency tokens, the goal of reconstruction can be to accurately restore the details of the foreground object to form a first reconstructed image, which is achieved through a reconstruction decoder, which can strengthen the pruner's ability to identify and retain foreground objects. For low saliency tokens, the goal of reconstruction can be to generate a second reconstructed image, which mainly represents background elements and allows for a certain degree of reconstruction error, aiming to weaken the pruner's attention to background information.
[0100] In the above process, by accurately reconstructing high saliency tokens and improving the parameters of the pruner, the perception ability of the pruner for key foreground information can be significantly enhanced, and accurate driving decisions can be made even in complex driving scenarios. Allowing a larger error in the background token during image reconstruction can achieve the strategy of suppressing background information, which helps the pruner to learn more valuable discriminant criteria for foreground tokens and reduces the processing of redundant background tokens, thereby further reducing the consumption of computing resources in practical applications.
[0101] The target pruner mentioned above can refer to a pruner that can effectively identify and retain key visual tokens while eliminating redundant information after training. The target pruner performs better in quantifying the foreground saliency of visual tokens and can accurately filter out information that is more important for driving decisions. The target pruner can work together with the preset visual tokens to form an efficient visual token pruning scheme for autonomous driving scenarios.
[0102] As an optional implementation, the parameters of the initial pruner and the initial visual tokens can be fine-tuned based on the feedback of the training images, the mask images, the first reconstructed images, and the second reconstructed images to achieve better pruning effect and performance. This adjustment process can ensure that the pruner can effectively distinguish between foreground and background, thereby achieving efficient visual information processing in practical applications. To guide the improvement of the pruner and the visual tokens, a comprehensive loss function including foreground reconstruction error and background reconstruction error can be designed, where a smaller foreground reconstruction error indicates that the pruner has successfully retained key foreground information related to driving decisions, and a larger background reconstruction error indicates that the pruner has effectively identified and eliminated redundant background information. By adjusting the weights of the loss function, the retention of foreground information and the pruning of background information can be balanced. Then, based on the calculated loss function value, the initial pruner parameters and the initial visual tokens can be adjusted using the backpropagation algorithm, and the value of the loss function can be gradually reduced through gradient descent and other improvement algorithms, so that the pruner can retain key foreground information while eliminating background redundancy and improve parameters. This adjustment process can be iterative until the performance indicators of the pruner meet the predetermined standards, for example, the retained foreground information can enable the autonomous driving system to maintain or approach the performance level of the original model. Through multiple iterations, the pruner and the visual tokens can continuously learn and improve, achieving accurate foreground token pruning and obtaining the required target pruner and preset visual tokens.
[0103] In the above process, the improvement adjustment process enables the pruner to more accurately identify foreground information in different scenarios, improving the generalization ability of the pruner. The trained pruner can better handle issues such as occlusion and lighting changes in images, effectively identifying and retaining key foreground tokens even in harsh weather or complex lighting conditions, ensuring the stability and reliability of the decision-making process.
[0104] Optionally, the training data is obtained, including: obtaining an original training set, wherein the original training set contains training images; performing image segmentation on training objects in the training images to obtain position information of the training objects in the training images; and generating a mask image based on the position information of the training objects.
[0105] The position information can refer to the specific spatial position of a training object, such as a vehicle, a pedestrian, or a traffic signal, in a training image, which can be presented in the form of pixel coordinates.
[0106] As an optional implementation, an image segmentation algorithm can be used to accurately segment foreground objects in the training images. A specific algorithm such as semantic segmentation or instance segmentation can be used to identify and distinguish different types of foreground objects to obtain pixel-level position information of each foreground object. Based on the position information obtained by segmentation, a mask image can be generated. The mask image can be a binary image with the same size as the training image, in which the pixel positions of the foreground objects are marked as 1 and the pixel positions of the background positions are marked as 0. This implementation generates a corresponding mask for each training object to accurately distinguish the foreground region and the background region of the image.
[0107] In the above process, the acquisition of position information and the generation of mask images can significantly enhance the strength of the supervision signal, making the training process more efficient and the pruner learning more accurate. The pruner can effectively distinguish between foreground and background and improve its performance. The mask image provides pixel-level accurate labeling, ensuring that the pruner can accurately identify and learn which visual tokens are foreground objects related to driving decisions and which are background information that can be safely pruned.
[0108] Optionally, the image is reconstructed based on the training quantization result and the plurality of training visual tokens to obtain a first reconstructed image and a second reconstructed image, including: determining a first training visual token set and a second training visual token set based on the training quantization result and the plurality of training visual tokens, wherein the training quantization result of a first training visual token in the first training visual token set is greater than the training quantization result of a second training visual token in the second training visual token set; and reconstructing the image based on the first training visual token set to obtain the first reconstructed image; and reconstructing the image based on the second training visual token set to obtain the second reconstructed image.
[0109] The first training visual token set can refer to a set of visual tokens that are highly related to key foreground objects in an autonomous driving scene, such as vehicles, pedestrians, and traffic signals, after quantitative analysis during the training process, i.e., a token set with a high foreground saliency score.
[0110] The second training visual token set can be a set of visual tokens that are less associated with the foreground object and more represent background information after quantitative analysis. The foreground saliency scores of the visual tokens in the second training visual token set are lower, and the direct contribution to driving decisions is smaller. The second training visual token set can be used as a pruner to identify and remove objects in practical applications.
[0111] As an optional implementation, the training visual tokens can be sorted from high saliency to low saliency based on the obtained training quantitative results. The sorting basis can be the numerical value of the training quantitative results. A high numerical value can represent a high degree of association with the foreground object, and a low numerical value can represent a low degree of association with the foreground object. The visual tokens can be divided into a first training visual token set and a second training visual token set according to the sorting results. The division can be based on a preset threshold or a preset proportion to ensure that the foreground information more relevant to driving decisions is retained. The first training visual token set and the second training visual token set can be used to reconstruct images using a reconstruction decoder. The first training visual token set can be used to generate a first reconstructed image to obtain a high-precision foreground region image, aiming to enhance the ability of the pruner to capture key foreground information. The second training visual token set can be used to generate a second reconstructed image to obtain a low-precision background region image, allowing a larger reconstruction error to weaken the attention of the pruner to background information.
[0112] In the above process, high-precision reconstruction of the first training visual token set can significantly enhance the perception and understanding of key foreground information by the pruner, and low-precision reconstruction of the second training visual token set can effectively suppress the processing of background information, reduce the waste of computing resources, and improve the overall efficiency of autonomous driving decisions.
[0113] Optionally, based on the training quantitative results, the first training visual token set and the second training visual token set are determined, including: grouping the plurality of training visual tokens based on the training quantitative results to obtain at least one first training visual token and at least one second training visual token, wherein the training quantitative result of the first training visual token is greater than the training quantitative result of the second training visual token; padding the at least one first training visual token to obtain the first training visual token set, wherein the number of training visual tokens in the first training visual token set is the same as the number of the plurality of training visual tokens; and padding the at least one second training visual token to obtain the second training visual token set, wherein the number of training visual tokens in the second training visual token set is the same as the number of the plurality of training visual tokens.
[0114] As an optional implementation, the plurality of training visual tokens can be divided into high saliency and low saliency groups based on the training quantization result. The grouping can be based on a preset threshold or a saliency score ranking, ensuring that the first training visual token set contains foreground information related to driving decisions. In order to maintain consistency with the number of output visual tokens, padding operations can be performed on the grouped tokens. That is, the first training visual tokens can be padded to obtain the first training visual token set, so that the number of training visual tokens in the first training visual token set is the same as the number of the plurality of training visual tokens, retaining key foreground information. At the same time, the second training visual tokens can be padded to obtain the second training visual token set, also adjusting the size of the set, most of which are background or low importance information, and the number of training visual tokens in the second training visual token set is the same as the number of the plurality of training visual tokens.
[0115] In the above process, by constructing the first training visual token set, it can be ensured that the foreground information more important to driving decisions is retained, thereby providing accurate and reliable visual input in subsequent decision making, improving the decision accuracy and reliability of the pruner. By constructing the second training visual token set and allowing errors in the background information in the reconstruction, the computing resources can be more efficiently concentrated on processing the key foreground information, significantly reducing the computational overhead caused by processing the background information, and improving the overall efficiency and real-time performance of the pruner.
[0116] Optionally, based on the training image, the mask image, the first reconstructed image and the second reconstructed image, the parameters of the initial pruner and the initial visual token are adjusted to obtain the target pruner and the preset visual token, including: based on the training image and the mask image, determining an object image of a training object and a scene image in the training image except the object image; based on the error between the first reconstructed image and the object image and the error between the second reconstructed image and the scene image, determining a total loss function value; based on the total loss function value, adjusting the parameters of the initial pruner and the initial visual token to obtain the target pruner and the preset visual token.
[0117] The object image mentioned above can refer to the image part of the training object such as pedestrians, vehicles, traffic lights, etc. after being filtered by the mask image. The object image set embodies the visual information more important to driving decisions in the autonomous driving scene.
[0118] The scene image mentioned above can refer to the part of the mask image that is not marked as a training object, such as buildings, sky, etc.
[0119] The total loss function value can be a loss function value for evaluating the performance of the pruner. The total loss function value can be composed of two parts, including the error between the first reconstructed image and the object image, and the error between the second reconstructed image and the scene image. Through the comprehensive consideration of the two error parts, the total loss function value can more comprehensively reflect the recognition and reconstruction ability of the pruner for foreground information and background information, and provide clear guidance for the adjustment of the pruner parameters.
[0120] As an optional implementation, the training image can be segmented into an object image and a scene image according to the mask image. The object image can be an accurate extraction of the foreground object in the training image, and the scene image can contain background information other than the foreground object. Then, the first training visual token set and the second training visual token set can be used to respectively reconstruct the object image and the scene image, to obtain the first reconstructed image and the second reconstructed image. By comparing the errors between the first reconstructed image and the object image, and the errors between the second reconstructed image and the scene image, the reconstruction ability of the pruner for foreground and background information can be quantified. Then, based on the above reconstruction errors, a total loss function can be constructed. The total loss function can perform weighted calculation on the two groups of errors. By adjusting the weights, the recognition and reconstruction ability of the pruner for foreground and background information can be balanced, to ensure that the pruner can accurately retain key visual information and effectively eliminate redundant background information. Finally, according to the feedback of the total loss function value, the parameters of the initial pruner and the initial visual token can be adjusted using the back propagation algorithm. Through the minimization of the total loss function value, the target pruner obtained through iterative improvement can retain key visual information while reducing the processing amount of background information, to improve the running efficiency and accuracy of the autonomous driving system.
[0121] In the above process, by calculating the error between the first reconstructed image and the object image, the pruner can be prompted to more accurately retain visual information closely related to driving decisions, to ensure that the autonomous driving system can make accurate judgments when processing complex scenes. The error control between the second reconstructed image and the scene image helps the pruner to learn how to identify and eliminate background information, to reduce unnecessary consumption of computing resources. The construction of the total loss function and the adjustment of the parameters can achieve a better balance between foreground information and background information processing. Through iterative improvement, the target pruner can retain more critical foreground information processing while significantly reducing the processing amount of redundant visual tokens, to effectively promote the implementation of efficient and reliable autonomous driving solutions.
[0122] Optionally, the total loss function value is determined based on an error between the first reconstructed image and the object image and an error between the second reconstructed image and the scene image, including: determining a first loss function value based on the error between the first reconstructed image and the object image; determining a second loss function value based on the error between the second reconstructed image and the scene image; and performing weighted processing on the first loss function value and the second loss function value to obtain the total loss function value.
[0123] The first loss function value can be a loss function value derived from a difference between the first reconstructed image and the object image. The first loss function value can quantify the performance of the pruner in retaining and reconstructing key foreground information. The first loss function value can be determined according to actual needs, and is not limited herein. For example, the first loss function value can be a mean squared error (MSE) or a structural similarity index measure (SSIM).
[0124] The second loss function value can be a loss function value of an error between the second reconstructed image and the scene image. The second loss function value can reflect the ability of the pruner in eliminating background information and reducing computational burden. By allowing a certain range of errors, the pruner can more effectively identify and retain key foreground information while eliminating redundant background information. The second loss function value can be determined according to actual needs, and is not limited herein. For example, the second loss function value can be a mean squared error or a structural similarity index measure.
[0125] As an optional implementation, after the pruner groups the visual tokens and completes image reconstruction, errors between the first reconstructed image and the object image and between the second reconstructed image and the scene image can be calculated respectively. Based on the errors, the first loss function value and the second loss function value can be constructed. The first loss function value can minimize the restoration error of the foreground information to ensure accurate retention of key visual information. The second loss function value controls the reconstruction error of the background information to promote the reasonable elimination of the background information by the pruner. Since the foreground and background information have different importance in the automatic driving task, the first loss function value and the second loss function value are weighted to obtain the total loss function value. The weight can be set or adjusted based on experience to find an optimal value to balance the accuracy of foreground information retention and the efficiency of background information elimination, so that the pruner can accurately identify key objects in actual driving scenarios and effectively reduce computational burden.
[0126] In the above process, the weighted total loss function value can guide the improvement of the pruner parameters, ensuring that the pruner accurately retains key foreground information while eliminating redundant background information. During the pruner training process, through the comprehensive consideration of the first loss function value and the second loss function value, the pruning strategy can be continuously adjusted and improved to enhance the performance of the pruner while reducing the consumption of computing resources, meeting the real-time and resource constraints in the mobile device environment. By calculating the error in different scenarios and adjusting the loss function, the pruner can learn a general strategy for distinguishing foreground and background in various driving environments, enhancing the generalization ability of the pruner when facing complex environments.
[0127] Optionally, based on the at least one target visual token, the mobile device is subjected to a service task corresponding to a preset service type, and an execution result of the service task is obtained.
[0128] Optionally, the preset service type is a path planning type, the service task is a task of generating a target driving path of the mobile device, and the execution result is a result of controlling the mobile device to drive based on the target driving path.
[0129] The target driving path mentioned above can refer to a safe and feasible route planned by an automatic driving decision system for the mobile device after analyzing and deciding based on the target visual token and other information such as radar point cloud data. The target driving path can serve as the direction and trajectory that the mobile device needs to follow in the future to achieve smooth, efficient and safe driving from the current location to the destination.
[0130] As an optional implementation, the target visual token can be parsed to extract feature information of specific objects or environmental elements represented by each target visual token, such as the position of pedestrians, the motion direction of vehicles, the state of traffic signals, etc. The feature information in the target visual token can be combined with dynamic data such as the position, speed, direction and target destination of the mobile device to form a complete and real-time environmental perception picture. Subsequently, a path planning algorithm can be used to analyze the above environmental perception picture to find an optimal path from the current location to the destination, which can take into account the avoidance of dynamic obstacles, the compliance with traffic rules and the improvement of driving efficiency to ensure the safety and efficiency of the target driving path. As the mobile device travels and the environment changes, new target visual tokens can be used to update environmental perception and re-plan the path to ensure the real-time and accuracy of the driving path.
[0131] In the above process, path planning based on the target visual token containing more key foreground information can avoid decision bias caused by processing a large amount of redundant background information, ensuring the accuracy of driving decisions. Path planning based on the simplified target visual token can significantly reduce the computational burden and improve the speed of driving decisions. By focusing on the foreground elements related to driving decisions, path planning can quickly respond to potential driving risks such as sudden pedestrians or obstacles, adjust driving strategies in real time, and enhance driving safety.
[0132] As an optional implementation, the target driving path can be converted into a series of fine control instructions, which can include steering angle, acceleration or deceleration instructions, braking instructions, etc., to guide the mobile device how to drive along the planned target path. Subsequently, these control instructions can be sent to the underlying control system of the mobile device, including steering system, power system and braking system, etc., to perform specific physical operations. During the driving of the mobile device, changes in the environment around the mobile device can also be continuously monitored, including real-time updates of the target visual token, as well as data provided by other sensors such as lidar and millimeter wave radar, to adjust the control strategy in a timely manner to deal with unexpected situations such as unexpected obstacles in front. The control instructions can also be dynamically adjusted according to the deviation of the actual driving state of the mobile device from the target driving path, to ensure that the mobile device can safely and smoothly drive on the planned path.
[0133] In the above process, the generation of precise control instructions based on the target driving path, combined with the efficient execution of the underlying control system of the mobile device, improves the following accuracy of the mobile device on the planned path, enabling the mobile device to drive more accurately on the desired target driving path. Since the pruning of the visual token has been performed in the path planning stage, the computational burden in the control process is effectively reduced, making the generation and execution of control instructions more rapid, reducing the delay, and improving the overall response speed of autonomous driving decisions.
[0134] The technical scheme proposed in the application is described below in combination with an optional embodiment. The application proposes a foreground token pruning for focusing on an autonomous driving scene to realize an efficient end-to-end driving scheme, which can realize a token pruning scheme suitable for the autonomous driving scene. Specifically, a large-scale dataset specially annotated with key foreground elements of autonomous driving, such as vehicles, pedestrians, lane lines, etc., is constructed to lay a foundation for accurately identifying foreground tokens strongly related to driving decisions; at the same time, a pruner with dynamic quantification of token saliency is designed, which, in combination with an adversarial foreground-background reconstruction strategy, can greatly improve the perception accuracy of the pruner for key foreground information and effectively overcome the alignment error caused by the spatial offset of visual tokens and original images. On this basis, by screening and retaining a small number of foreground tokens with the most value, such as retaining only 25%, which is not limited here, while greatly reducing the consumption of computing resources, such as reducing the processing amount of visual tokens by 75%, the core performance of the end-to-end autonomous driving model, such as trajectory planning accuracy and collision avoidance ability, can be ensured without being significantly affected, which can meet the dual constraints of limited computing resources and real-time requirements of the vehicle-mounted platform, provide key technical support for the transition of the end-to-end autonomous driving system from laboratory research to actual vehicle-mounted deployment, and promote the landing application of efficient and reliable autonomous driving solutions.
[0135] The efficient visual token pruning technical scheme for the end-to-end autonomous driving scene proposed in the application realizes accurate screening of key foreground tokens through an object-centered design concept, in combination with a special dataset, an innovative pruner and a training strategy, while greatly reducing the computing overhead and retaining the core performance of the pruner. The specific technical scheme is as follows: For the foreground elements in the autonomous driving scene, such as vehicles, pedestrians and lane lines, which are important for driving decisions, and the background elements, such as the sky and buildings, which have high redundancy, the goal is to retain a small number of key foreground tokens and eliminate a large number of background redundant tokens to meet the computing resource and real-time requirements of the vehicle-mounted platform.
[0136] To provide accurate foreground-background supervision signals, a large-scale dataset containing image-mask pairs can be constructed based on the original dataset (Sparse-nuScenes) by using an image segmentation tool to segment images from multiple camera perspectives, and annotating the foreground regions strongly related to autonomous driving, including pedestrians, vehicles, lanes, traffic signals and traffic obstacles. The dataset clearly distinguishes between foreground and background, providing fine-grained spatial annotations for the pruner and solving the technical problem of lack of scene-based supervision in general datasets.
[0137] Figure 2 is a schematic view of a training data according to an embodiment of the application, such as Figure 2As shown, the training data can include target images collected at multiple angles around the mobile device and corresponding foreground images of the target images, and the multiple angles are illustrated as 6 angles around the mobile device, which can be determined according to actual needs, and are not limited herein; and the foreground objects can include pedestrians, lanes, vehicles, traffic signals and traffic obstacles.
[0138] Figure 3 is a schematic diagram of a pruning device training process according to an embodiment of the present application, as shown in Figure 3 In the pruning device training process, the visual encoder encoding process can obtain multiple training visual tokens, and the initial visual token and the multiple training visual tokens can be processed by the query neural network module in the initial pruning device to obtain multiple target tokens. The initial visual token can be used to enhance the multiple target tokens to obtain multiple fusion tokens; and the multiple fusion tokens can be processed by the feedforward neural network module to obtain training quantization results. Based on the multiple training visual tokens, the training quantization results and the padding visual token, selection and padding can be performed to obtain a first training visual token set and a second training visual token set, the first training visual token set can be processed by the reconstruction decoder to obtain a first reconstructed image, and the second training visual token set can be processed by the reconstruction decoder to obtain a second reconstructed image.
[0139] Figure 4 is a schematic diagram of generating a target driving path of a mobile device according to an embodiment of the present application, as shown in Figure 4 In the process of generating the target driving path of the mobile device, the target image can be processed by the visual encoder to obtain multiple visual tokens; the multiple visual tokens can be processed by the target pruning device to obtain quantization results; and the multiple visual tokens can be subjected to a rejection operation based on the quantization results to obtain at least one target visual token. The text prompt can be processed to obtain a text token. The at least one target visual token and the text token can be processed by the large language model to obtain the target driving path.
[0140] The text prompt described above can be as follows: "You are an autonomous driving agent. You can access the front-view camera image of a car Your task is to predict the future path points of the vehicle in the next 3 time steps", and the text prompt can be determined according to actual needs, and is not limited herein.
[0141] The pruning device (ReconPruner) in the present application can be a plug-and-play visual token pruning device, which can be embedded between the visual encoder and the subsequent processing module of the visual-language-action model as an independent module to realize dynamic filtering of visual tokens. The pruning device can include a single-layer large language model layer, i.e., a query neural network module, a single-layer feedforward neural network, i.e., a feedforward neural network module, and a trainable query token, i.e., a preset visual token. The token (token) output by the visual encoder, i.e., multiple visual tokens, is spliced with the query token (token), i.e., the preset visual token, and input into the large language model layer, then the output query token (token) is multiplied with the visual token (token) to obtain a fusion token (token), and finally the fusion token (token) is input into the feedforward neural network to obtain the foreground saliency score of each visual token (token), i.e., the quantization result of the multiple visual tokens. Then, based on a preset proportion, such as 25%, tokens with higher saliency scores can be retained, and redundant background tokens can be removed, and the simplified token sequence can be output for use by subsequent decision modules.
[0142] To improve the perception accuracy of the pruning device (ReconPruner) for foreground tokens, an adversarial training mechanism is designed in the training method of the adversarial foreground-background reconstruction strategy in the present application. Specifically, foreground reconstruction reinforcement can be used to train the pruning device to perform high-precision image reconstruction on the retained foreground tokens, such as foreground saliency scores greater than 0, which are not limited here, by minimizing the reconstruction loss, i.e., the mean square error and the structural similarity index measure loss, to strengthen the model's ability to capture foreground features; background reconstruction suppression can be used to train the pruning device to perform low-precision reconstruction on the removed background tokens, such as foreground saliency scores less than 0, which are not limited here, by allowing larger errors, by maximizing the reconstruction loss, i.e., the same as the foreground reconstruction loss, to weaken the model's attention to background features; an adversarial loss function can be used to weight the foreground and background reconstruction losses, so that the pruning device learns more valuable discriminative ability for foreground tokens, while alleviating the spatial offset problem between visual tokens and original images, and realizing the quantization and accurate ordering of token saliency.
[0143] The data preprocessing in this application utilizes image-mask pairs from the original dataset (Sparse-nuScenes) to generate pixel-level annotations for the foreground region, serving as a supervisory signal for pruning training. Pruner training involves jointly training the ReconPruner with a vision-language-action model, improving pruner parameters through an adversarial foreground-background reconstruction strategy to achieve accurate identification and retention of highly saliency foreground tokens. In the inference phase, during autonomous driving inference, the ReconPruner receives tokens output from the visual encoder in real-time, calculates the saliency score of each token, retains preceding tokens according to a preset ratio, and inputs the simplified sequence into subsequent language understanding and action decision modules to complete the end-to-end driving task. This solution provides supervision through a scenario-based dataset, achieves accurate selection through an innovative pruner, and strengthens foreground perception through adversarial strategies. Ultimately, based on the open-loop planning benchmark of the original dataset (Sparse-nuScenes), it can maintain 98.1% of the original model's performance (i.e., trajectory planning accuracy) while retaining only 25% of the visual tokens, effectively meeting the onboard deployment requirements of end-to-end autonomous driving systems.
[0144] Based on the core concept of this application, focusing on foreground token pruning in autonomous driving scenarios to achieve efficient end-to-end driving, from the overall technical framework to core features, the following alternative or modified solutions exist. These solutions all revolve around the inventive logic of accurately filtering key foreground information, reducing computational overhead while preserving performance, but differ in their implementation paths: An adaptive pruning framework based on dynamic thresholds can avoid relying on a fixed percentage of token pruning, such as a fixed 25% retention (this is not limited here), and can dynamically adjust the pruning threshold according to the scene complexity. For example, in simple scenarios such as highways, fewer tokens can be retained, such as 15% (this is not limited here), while in complex scenarios such as urban intersections, more tokens can be retained, such as 30% (this is not limited here). By evaluating the density of foreground elements in the scene in real time, such as the number of vehicles and pedestrian density, the pruning intensity is automatically adapted. This transforms static proportional pruning into dynamic scene-aware pruning, better meeting the actual needs of fluctuating scene complexity in autonomous driving, and further improving the flexibility of computational efficiency.
[0145] This multimodal fusion foreground token selection framework, in addition to visual modalities, can also integrate LiDAR point cloud data to assist in foreground token recognition. By leveraging the 3D spatial information of the point cloud, such as distance and velocity, it enhances the identification of dynamic foregrounds, such as moving vehicles and pedestrians, overcoming the limitations of pure vision under occlusion and lighting changes. Furthermore, it combines visual tokens for pruning. This expands foreground recognition from pure vision to multimodal cross-validation, improving the robustness of foreground token selection in complex environments and making it applicable to harsh weather or highly occluded scenarios.
[0146] The object-centered scene adaptability design in the application is specifically for the characteristics of the autonomous driving scene, takes foreground objects such as vehicles, pedestrians, lane lines and other key elements for driving decision as the pruning core, explicitly removes background redundant information such as sky and building, realizes the deep adaptation of the pruning strategy and the autonomous driving task. The construction and application of the original dataset (Sparse-nuScenes), through the image segmentation tool, the original dataset (Sparse-nuScenes) is segmented, the original dataset (Sparse-nuScenes) containing image-mask pairs is constructed, and the foreground and background areas in the autonomous driving scene are explicitly labeled, which provides accurate supervision signal for the pruner. The pruner (ReconPruner) and the adversarial foreground-background reconstruction strategy, the plug-and-play pruner (ReconPruner) is designed, and the adversarial foreground-background reconstruction strategy is innovatively introduced, the model is strengthened in the reconstruction foreground, and the attention is weakened in the reconstruction background, the foreground saliency of each visual token is quantified, and the spatial offset problem of the visual token and the original image is solved, and accurate pruning is realized.
[0147] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.
[0148] According to another aspect of the embodiments of the application, an image processing device is also provided, which can execute the image processing method of the above-mentioned embodiments, and the specific implementation method and preferred application scenario are the same as those of the above-mentioned embodiments, which will not be repeated here.
[0149] Figure 5 is a schematic diagram of an image processing device according to an embodiment of the application, as shown in Figure 5 The device includes the following: an acquisition module 502, an encoding module 504, a determination module 506 and a removal module 508.
[0150] The obtaining module is configured to obtain a target image and a preset service type, wherein the target image reflects environmental information of an environment around the mobile device; the encoding module is configured to encode the target image to obtain a plurality of visual tokens, wherein different visual tokens are used to represent features of different image blocks in the target image; the determining module is configured to determine degrees of association between the plurality of visual tokens and a preset visual token, wherein the preset visual token reflects an element attribute corresponding to the preset service type; and the eliminating module is configured to perform an eliminating operation on the plurality of visual tokens based on the degrees of association to obtain at least one target visual token, wherein the degree of association between the target visual token and the preset visual token is greater than a preset degree.
[0151] The quantifying module is further configured to map the plurality of visual tokens from an image feature space to a foreground object feature space by using the preset visual token to obtain a plurality of target tokens, wherein the image feature space is used to represent visual information of the target image, and the foreground object feature space is used to represent object information of the foreground object; and map the plurality of target tokens from the foreground object feature space to a numerical space to obtain quantization results of the plurality of visual tokens.
[0152] The quantifying module is further configured to splice the plurality of visual tokens and the preset visual token to obtain a spliced token; input the spliced token to the query neural network module; and perform feature interaction on the plurality of visual tokens and the preset visual token by using the query neural network module to obtain the plurality of target tokens.
[0153] The quantifying module is further configured to enhance features used to represent the foreground object in the plurality of target tokens by using the preset visual token to obtain a plurality of fusion tokens; input the plurality of fusion tokens to the feedforward neural network module; and map the plurality of fusion tokens from the foreground object feature space to the numerical space by using the feedforward neural network module to obtain the quantization results of the plurality of visual tokens.
[0154] The eliminating module is further configured to perform one of the following: eliminate, from the plurality of visual tokens, a visual token whose quantization result is less than a preset threshold to obtain at least one target visual token; sort the plurality of visual tokens in descending order of the quantization results to obtain a visual token sequence, and eliminate a preset number of visual tokens at the back of the visual token sequence to obtain at least one target visual token, wherein a proportion of the preset number of visual tokens in the visual token sequence satisfies a preset proportion, and the preset proportion is determined based on a scene type of a driving scene.
[0155] The eliminating module is configured to perform a blocking operation on the target image to obtain a plurality of image blocks; input the plurality of image blocks to the visual encoder; and encode the plurality of image blocks by using the visual encoder to obtain the plurality of visual tokens, wherein the plurality of visual tokens have a corresponding relationship with the plurality of image blocks.
[0156] According to another aspect of the embodiments of the present application, a vehicle is also provided, including a memory storing an executable program, and a processor configured to execute the program, wherein the program performs the method when executed.
[0157] According to another aspect of the embodiments of the present application, an electronic device is also provided, wherein the electronic device is deployed on a vehicle, or the electronic device is a server.
[0158] Embodiments of the present application also provide a computer readable storage medium including a stored executable program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to perform the method in various embodiments of the present application when the executable program is executed.
[0159] The computer storage medium described above can refer to a medium in a computer memory for storing certain discrete physical quantities, and the computer storage medium mainly includes semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc. The stored program included in the computer readable storage medium can be a set of instructions that can be recognized and executed by a computer, and the program runs on an electronic computer to meet the informationization tool of a certain demand of people.
[0160] Embodiments of the present application also provide a computer program product including a computer program, and the computer program is executed by a processor to implement the method in various embodiments of the present application.
[0161] The computer program product described above can refer to a software program that has been written, tested and released, and can run on a computer or other device. The computer program product can include application programs, operating systems, tool software, etc., and is used to implement specific functions or solve specific problems.
[0162] Embodiments of the present application also provide a computer program product including a non-volatile computer readable storage medium for storing a computer program, and the computer program is executed by a processor to implement the method in various embodiments of the present application.
[0163] The non-volatile computer readable storage medium described above can refer to a medium for storing data, and the non-volatile computer readable storage medium can keep the data from being lost when power is off. The non-volatile computer readable storage medium can be used to store long-term saved data such as operating systems, application programs and user files. The non-volatile storage medium can include hard disk drives, solid state drives, optical discs and flash memory storage devices, etc.
[0164] Embodiments of the present application also provide a computer program, and the computer program is executed by a processor to implement the method in various embodiments of the present application.
[0165] The computer program described above can refer to a set of instructions for telling a computer to perform a specific task or operation. The computer program can be written by a programmer using a specific programming language, and can include algorithms, data structures, logic and control flow, etc. The computer program can be used for various purposes, including application software, operating systems, etc.
[0166] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0167] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, and can be electrical or other forms.
[0168] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.
[0169] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0170] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0171] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. An image processing method, characterized by, The method comprises the following steps: acquiring a target image and a preset business type, wherein the target image reflects environmental information of an environment around a mobile device; encoding the target image to obtain a plurality of visual tokens, wherein different visual tokens are used to represent features of different image blocks in the target image; determining a degree of association between the plurality of visual tokens and a preset visual token, wherein the preset visual token reflects an element attribute corresponding to the preset business type; performing a culling operation on the plurality of visual tokens based on the degree of association to obtain at least one target visual token, wherein the degree of association between the target visual token and the preset visual token is greater than a preset degree.
2. The method of claim 1, wherein, The method further comprises the following steps: mapping the plurality of visual tokens from an image feature space to a target feature space using a target pruner to obtain a plurality of target tokens, wherein the image feature space is used to represent visual information of the target image, and the target feature space is a feature space corresponding to the preset visual token; mapping the plurality of target tokens from the target feature space to a numerical space using the target pruner to obtain the degree of association.
3. The method of claim 2, wherein, The method further comprises the following steps: splicing the plurality of visual tokens and the preset visual token to obtain spliced tokens; inputting the spliced tokens into a query neural network module in the target pruner, and performing feature interaction on the plurality of visual tokens and the preset visual token using the query neural network module to obtain the plurality of target tokens.
4. The method of claim 2, wherein, The method further comprises the following steps: enhancing the degree of association between the plurality of target tokens and the preset visual token to obtain a plurality of fusion tokens; inputting the plurality of fusion tokens into a feedforward neural network module in the target pruner, and mapping the plurality of fusion tokens from the target feature space to the numerical space using the feedforward neural network module to obtain a quantization result of the plurality of visual tokens.
5. The method according to any one of claims 2 to 4, characterized in that, The training steps of the target pruner comprise the following steps: acquiring training data, wherein the training data comprises a training image and a mask image, and the mask image is used to represent whether a plurality of pixel points in the training image belong to an element corresponding to the preset business type; encoding the training image to obtain a plurality of training visual tokens, wherein different training visual tokens are used to represent features of different training image blocks in the training image; determining a degree of association between the plurality of training visual tokens and an initial visual token using an initial pruner, wherein the initial visual token reflects an element attribute corresponding to the preset business type; reconstruct the image based on the training quantization result and the plurality of training visual tokens to obtain a first reconstructed image and a second reconstructed image, wherein the first reconstructed image is used to represent an image of a region where the element corresponding to the preset business type is located, and the second reconstructed image is used to represent an image of a region other than the region where the element corresponding to the preset business type is located; adjust the parameters of the initial pruner and the initial visual token based on the training image, the mask image, the first reconstructed image and the second reconstructed image to obtain the target pruner and the preset visual token.
6. The method of claim 5, wherein, The reconstructing the image based on the training quantization result and the plurality of training visual tokens to obtain a first reconstructed image and a second reconstructed image comprises: determining a first training visual token set and a second training visual token set based on the training quantization result and the plurality of training visual tokens, wherein the training quantization result of a first training visual token in the first training visual token set is greater than the training quantization result of a second training visual token in the second training visual token set; reconstructing the image based on the first training visual token set to obtain the first reconstructed image; reconstructing the image based on the second training visual token set to obtain the second reconstructed image.
7. The method of claim 6, wherein, The determining the first training visual token set and the second training visual token set based on the training quantization result comprises: grouping the plurality of training visual tokens based on the training quantization result to obtain at least one first training visual token and at least one second training visual token, wherein the training quantization result of the first training visual token is greater than the training quantization result of the second training visual token; filling the at least one first training visual token to obtain the first training visual token set, wherein the number of training visual tokens in the first training visual token set is the same as the number of the plurality of training visual tokens; filling the at least one second training visual token to obtain the second training visual token set, wherein the number of training visual tokens in the second training visual token set is the same as the number of the plurality of training visual tokens.
8. The method of claim 5, wherein, The adjusting the parameters of the initial pruner and the initial visual token based on the training image, the mask image, the first reconstructed image and the second reconstructed image to obtain the target pruner and the preset visual token comprises: determining an object image of the training object and a scene image other than the object image in the training image based on the training image and the mask image; determining a total loss function value based on an error between the first reconstructed image and the object image and an error between the second reconstructed image and the scene image; adjusting the parameters of the initial pruner and the initial visual token based on the total loss function value to obtain the target pruner and the preset visual token.
9. The method of claim 8, wherein, The total loss function value is determined based on errors between the first reconstructed image and the object image and errors between the second reconstructed image and the scene image, including: A first loss function value is determined based on errors between the first reconstructed image and the object image; A second loss function value is determined based on errors between the second reconstructed image and the scene image; The first loss function value and the second loss function value are weighted to obtain the total loss function value.
10. The method of claim 1, wherein, The culling operation is performed on the plurality of visual tokens based on the quantization result to obtain at least one target visual token, including one of: From the plurality of visual tokens, the visual token whose quantization result is less than a preset threshold is removed to obtain the at least one target visual token; The plurality of visual tokens are sorted in descending order of the quantization result to obtain a visual token sequence, and a preset number of visual tokens at the back of the visual token sequence are removed to obtain the at least one target visual token, wherein a proportion of the preset number of visual tokens in the visual token sequence meets a preset proportion, and the preset proportion is determined based on a scene type of the driving scene.
11. The method of claim 1, wherein, The method further includes: Based on the at least one target visual token, a business task corresponding to a preset business type is performed on the mobile device to obtain an execution result of the business task.
12. The method of claim 11, wherein, The preset business type is a path planning type, the business task is a task of generating a target driving path of the mobile device, and the execution result is a result of controlling the mobile device to drive based on the target driving path.
13. A vehicle characterized by comprising: It includes: A memory storing an executable program; A processor configured to run the program, wherein the program runs to perform the method of any one of claims 1 to 12.
Citation Information
Cited By
Vehicle control method and device, vehicle and computer readable storage medium
CN121912996A