Vehicle control method and device, vehicle, storage medium and program product

By loading the first module of the visual language model, local parameters are fine-tuned according to the scene classification results, the performance limitations of a single visual language model in different scenarios are solved, and the reliability of intelligent driving and the accuracy of inference results are improved.

CN120096598APending Publication Date: 2025-06-06XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510504168.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The performance of a single visual language model in different scenarios has limitations and it is difficult to ensure the accuracy of its inference results.

Method used

By loading the first module of the visual language model, the module is determined based on the scene classification results and the file is loaded through parameter adjustments, and then the target image is processed using the visual sub-model and language sub-model to generate more accurate inference results.

Benefits of technology

It improves the reliability of intelligent driving of vehicles, makes the visual language model better adapt to the scene of the target image, and generates more accurate inference results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120096598A_ABST
    Figure CN120096598A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent driving, in particular to a vehicle control method and device, a vehicle, a storage medium and a program product. The method comprises the steps of determining a target image; processing the target image by using a visual language model to obtain a reasoning result; and controlling intelligent driving of the vehicle according to the reasoning result. Wherein a first module of the visual language model is loaded, the first module is determined according to the scene classification result, and the scene classification result is obtained by the visual language model according to the target image. Thus, through loading of the first module, the visual language model can better adapt to the scene of the target image, a more accurate reasoning result is generated, and therefore the reliability of intelligent driving of the vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of intelligent driving, and in particular to a vehicle control method, device, vehicle, storage medium and program product. Background Art

[0002] With the rapid development of science and technology, intelligent driving technology has gradually become a research hotspot in the automotive industry. In an intelligent driving system, vehicles need to accurately perceive the surrounding environment to achieve safe and efficient intelligent driving.

[0003] In recent years, the rise of Visual Language Models (VLMs) has brought new opportunities to the field of intelligent driving. However, the performance of a single VLM in different scenarios has limitations, and it is difficult to ensure the accuracy of its reasoning results. Summary of the invention

[0004] To overcome the problems existing in the related art, the present disclosure provides a vehicle control method, device, vehicle, storage medium and program product.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a vehicle control method, comprising: Determine the target image; Processing the target image using a visual language model to obtain an inference result; Controlling intelligent driving of the vehicle according to the inference result; Wherein, a first module of the visual language model is loaded, wherein the first module is determined according to a scene classification result obtained by the visual language model according to the target image.

[0006] In some possible implementations, the visual language model includes a visual sub-model and a language sub-model, and the scene classification result is obtained by the visual sub-model based on the target image; the first module is a module in the language sub-model, and the first module is used to implement local parameter fine-tuning associated with the scene classification result.

[0007] In some possible implementations, the first module is a module determined according to the scene classification result using a routing mechanism and loaded through a parameter adjustment file.

[0008] In some possible implementations, the parameter adjustment file is generated in the following manner: For each scene in the pre-built scene database, locally fine-tuning training is performed on the visual language model that is not loaded with the first module to obtain local fine-tuning parameters corresponding to each scene; The parameter adjustment file including the local fine-tuning parameters corresponding to each of the scenes is generated.

[0009] In some possible implementations, the processing the target image using a visual language model to obtain an inference result includes: Using the visual sub-model, obtaining an image feature vector according to the target image; The inference result is obtained according to the image feature vector using the language sub-model loaded with the first module.

[0010] In some possible implementations, obtaining the inference result according to the image feature vector using a language sub-model loaded with the first module includes: The inference result is obtained according to the image feature vector and the text used to indicate the target task, using the language sub-model loaded with the first module.

[0011] In some possible implementations, the method further includes: Before loading the first module, determining and caching keys and values ​​of the visual language model based on the target image; When reasoning is performed using the visual language model loaded with the first module, attention calculation is performed based on the cached keys and values, wherein the keys and values ​​of different target image caches are different.

[0012] In some possible implementations, the method further includes at least one of the following: Displaying the inference result and / or the inference process corresponding to the inference result using a display device in the vehicle; The inference result and / or the inference process corresponding to the inference result are voice broadcasted using a speaker in the vehicle.

[0013] According to a second aspect of an embodiment of the present disclosure, there is provided a visual vehicle control device, comprising: A determination module, used for determining a target image; An inference module, used to process the target image using a visual language model to obtain an inference result; A control module, used for controlling the intelligent driving of the vehicle according to the inference result; Wherein, a first module of the visual language model is loaded, wherein the first module is determined according to a scene classification result obtained by the visual language model according to the target image.

[0014] According to a third aspect of an embodiment of the present disclosure, a vehicle is provided, comprising: processor; a memory for storing processor-executable instructions; Among them, the processor is configured to execute the executable instructions in the memory to implement the steps of the vehicle control method provided in the first aspect of the present disclosure.

[0015] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the vehicle control method provided in the first aspect of the present disclosure are implemented.

[0016] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the vehicle control method provided in the first aspect of the present disclosure.

[0017] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects: Determine the target image; process the target image using the visual language model to obtain an inference result; and control the intelligent driving of the vehicle according to the inference result. Among them, the first module of the visual language model is loaded, and the first module is determined according to the scene classification result, and the scene classification result is obtained by the visual language model according to the target image. In this way, by loading the first module, the visual language model can be better adapted to the scene of the target image, and a more accurate inference result can be generated, thereby improving the reliability of the intelligent driving of the vehicle.

[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0020] Figure 1 The present invention is a flow chart of a vehicle control method according to an exemplary embodiment.

[0021] Figure 2 is a schematic diagram of a vehicle control method according to an exemplary embodiment.

[0022] Figure 3 is a schematic diagram of a vehicle control method according to an exemplary embodiment.

[0023] Figure 4 The present invention is a flow chart of a vehicle control method according to an exemplary embodiment.

[0024] Figure 5The figure is a schematic diagram showing a visual language model reasoning process according to an exemplary embodiment.

[0025] Figure 6 is a block diagram of a vehicle control device according to an exemplary embodiment.

[0026] Figure 7 is a block diagram of a vehicle according to an exemplary embodiment.

[0027] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment.

[0028] Fig. 9 is a block diagram of a chip according to an exemplary embodiment. DETAILED DESCRIPTION

[0029] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0030] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the device is located and with the authorization given by the owner of the corresponding device.

[0031] Figure 1 FIG. 1 is a flow chart of a vehicle method according to an exemplary embodiment. Figure 1 As shown, the method may include steps S101 to S103.

[0032] In step S101 , a target image is determined.

[0033] For example, an image acquisition device (such as a camera) pre-installed on the vehicle may be used to acquire an environmental image, and the acquired environmental image may be determined as a target image.

[0034] In step S102, the target image is processed using the visual language model to obtain an inference result.

[0035] For example, the target image can be input into the visual language model, and the output of the visual language model is the inference result. In step S102, the first module of the visual language model used is loaded, and the first module is determined according to the scene classification result, and the scene classification result is obtained by the visual language model according to the input target image.

[0036] For example, the above-mentioned scene classification results may include information for characterizing the scene. For vehicles, the aforementioned scenes may be scenes such as urban road driving scenes, highway driving scenes, and country road driving scenes. Multiple modules may be pre-set, and the corresponding first module may be determined from the multiple modules according to the scene classification results. Different scenes may correspond to different first modules. For example, in an urban road driving scene, the first module may be a module for enhancing pedestrian recognition and traffic sign understanding capabilities; while in a highway driving scene, the first module may be a module for enhancing vehicle detection capabilities. Loading different first modules can enable the visual language model to generate more accurate reasoning results for different scenes, so that intelligent driving can better adapt to various different scenes and improve user experience.

[0037] In step S103, the intelligent driving of the vehicle is controlled according to the inference result.

[0038] For example, the inference result may include data for controlling the vehicle state, such as the recommended speed, the recommended distance from the vehicle in front, acceleration and merging, pedestrian / vehicle avoidance, etc. The inference result may provide a more reliable decision-making basis for the intelligent driving vehicle and improve the reliability of intelligent driving.

[0039] In the above technical solution, a target image is determined; the target image is processed using a visual language model to obtain an inference result; and the intelligent driving of the vehicle is controlled according to the inference result. Among them, the first module of the visual language model is loaded, and the first module is determined according to the scene classification result, and the scene classification result is obtained by the visual language model according to the target image. In this way, by loading the first module, the visual language model can be better adapted to the scene of the target image, and a more accurate inference result can be generated, thereby improving the reliability of the intelligent driving of the vehicle.

[0040] In an optional embodiment, the visual language model includes a visual sub-model (ViT, Vision Transformer) and a language sub-model (LLM, Large Language Model), and the scene classification result is obtained by the visual sub-model based on the target image; the first module is a module in the language sub-model, and the first module is used to implement local parameter fine-tuning associated with the scene classification result.

[0041] The visual sub-model can be used to analyze the input target image, and the language sub-model can be used to generate inference results based on the output content of the visual sub-model. The first module can be used to implement local fine-tuning of model parameters. For example, the first module can be used to implement local adjustment of model parameters by introducing a low-rank matrix into the model without comprehensively updating the entire visual language model.

[0042] Loading the first module of the language sub-model in the visual language model can achieve fine-tuning of the language sub-model without retraining the visual language model. Compared with comprehensive adjustment of the parameters of the entire visual language model, it can effectively reduce the consumption of computing resources. In addition, the computational complexity of local fine-tuning is relatively low. Therefore, the visual language model can complete the loading of the first module more quickly, so as to adapt to the scene of the new target image in a timely manner, and improve the response speed and update frequency of the model in practical applications. In addition, the module used to achieve local fine-tuning occupies less memory space. Therefore, using the visual language model provided by the present disclosure, by loading different first modules, without affecting the original capabilities of the visual language model, it can not only quickly adapt to different scenes, but also effectively reduce the memory usage and avoid resource waste.

[0043] In the above technical solution, the local fine-tuning of the language sub-model is guided by the scene classification results output by the visual sub-model, which can strengthen the association between visual and language modalities. According to the scene classification results, the first module of the language sub-model is loaded in a targeted manner. On the one hand, it can avoid global updates to the entire language sub-model, thereby effectively reducing the time required for updates and the consumption of computing resources; on the other hand, it can enable the visual language model to generate more accurate and relevant language descriptions or responses in different scenarios. This scene-driven fine-tuning mechanism can effectively improve the performance of the visual language model in different scenarios and enhance the generalization and practicality of the model.

[0044] In an optional embodiment, the first module is a module determined based on the scene classification result using a routing mechanism and loaded through a parameter adjustment file.

[0045] Among them, the routing mechanism can dynamically select a suitable first module according to the input data (i.e., the scene classification result). For example, the first module can be determined by the routing mechanism implemented by dynamic slice. In this way, the routing mechanism can be used to dynamically select the first module, effectively reducing the task switching time associated with the scene of the target image, thereby improving the overall reasoning efficiency and enabling the visual language model to efficiently adapt to the current scene.

[0046] In one embodiment, the parameter adjustment file may be generated in the following manner: For each scene in the pre-built scene database, locally fine-tuning training is performed on the visual language model that is not loaded with the first module to obtain local fine-tuning parameters corresponding to each scene; Generate a parameter adjustment file including local fine-tuning parameters corresponding to each scene.

[0047] For example, the scenario database may be a scenario database for intelligent vehicle driving, and each scenario in the scenario database may correspond to a scenario classification result. Figure 2 As shown, the scene database of vehicle intelligent driving and the initial visual language model (i.e., the visual language model when the first module is not loaded) can be used for local fine-tuning training to generate a parameter adjustment file. After that, the parameter adjustment file and the initial visual language model can be used to obtain the visual language model loaded with the first module for subsequent reasoning.

[0048] The visual language model corresponds to multiple modules, such as Figure 3 As shown, module 1 corresponds to intelligent driving scene 1, module 2 corresponds to intelligent driving scene 2, and module 3 corresponds to intelligent driving scene 3. The scene classification result can be determined according to the target image, that is, the specific intelligent driving scene corresponding to the target image is determined. Taking the scene of the determined target image as intelligent driving scene 2 as an example, the first module at this time is module 2. Module 2 can be called out from the visual language model through the routing mechanism and the local fine-tuning parameters corresponding to the intelligent driving scene 2 can be called from the parameter adjustment file to load module 2, merge the called local fine-tuning parameters and the parameters of the initial visual language model, and obtain a visual language model adapted to the scene of the current target image.

[0049] Figure 4 is a flowchart of a visual language model reasoning process according to an exemplary embodiment. Figure 4 , we can more clearly understand the reasoning process of the visual language model provided by the present disclosure. Figure 4 As shown, the reasoning process may include steps S301 to S302.

[0050] In step S301, an image feature vector is obtained according to a target image using a visual sub-model.

[0051] In step S302, an inference result is obtained according to the image feature vector using the language sub-model loaded with the first module.

[0052] In one embodiment, obtaining the inference result according to the image feature vector by using the language sub-model loaded with the first module may include: According to the image feature vector and the text used to indicate the target task, the language sub-model loaded with the first module is used to obtain the inference result.

[0053] For example, the input text may be a Prompt instruction for describing the target task to guide the model to generate specific output content. The text may be pre-set by relevant personnel based on the target task.

[0054] In an optional embodiment, the visual language model can load the first module in the process of acquiring the target image for reasoning. Figure 5 As shown, step S303 may be included between step S301 and step S302, and step S301 may include step S3011.

[0055] In step S3011, the visual sub-model is used to obtain an image feature vector and a scene classification result based on the target image.

[0056] In step S303, according to the scene classification result, the first module is determined by using the routing mechanism, and the first module is loaded in the language sub-model through the parameter adjustment file.

[0057] In step S302, an inference result is obtained according to the image feature vector using the language sub-model loaded with the first module.

[0058] In this way, the visual language model loaded with the first module can better adapt to the scene of the input target image, improve the visual language model's ability to analyze the current scene, and then accurately generate reasoning results that are highly consistent with the target task, thereby improving the accuracy of the reasoning results.

[0059] In addition, the specific implementation of the above-mentioned step S3011 and step S303 has been described in detail above, and the repeated contents will not be repeated here.

[0060] In an optional embodiment, the vehicle control method provided by the present disclosure may further include: Before loading the first module, determine and cache the key (K, Key) and value (V, Value) of the visual language model based on the target image; When performing reasoning using the visual language model loaded with the first module, attention calculation is performed based on the cached keys and values.

[0061] For example, the visual language model without the first module loaded can be used to decompose the target image into a series of tokens, and the KV value corresponding to each token can be calculated. The KV values ​​corresponding to these tokens can be cached in the KV Cache of the Transformer layer. In this way, when reasoning with the visual language model loaded with the first module, the cache can be read directly, and the Transformer layer can be controlled to use the cached keys and values ​​for attention calculation. On the one hand, the calculation time can be reduced, thereby improving the reasoning efficiency; on the other hand, the current visual language model can be prevented from being affected by other previous scenes, ensuring that the currently used visual language model can adapt to the scene of the current target image.

[0062] Different target images have different cache keys and values, which can flexibly adapt to dynamic inputs and ensure that the model can maintain efficiency and accuracy when processing different target images.

[0063] In an optional embodiment, the vehicle control method provided by the present disclosure may further include at least one of the following: Displaying the inference result and / or the inference process corresponding to the inference result using a display device in the vehicle; The inference result and / or the inference process corresponding to the inference result are voice broadcasted using a loudspeaker in the vehicle.

[0064] For example, the display device in the vehicle may be an on-board display screen. By using the display device and / or the speaker to synchronize the user's reasoning results and / or the reasoning process corresponding to the reasoning results, prompts to the user can be achieved. In this way, the user can choose to control the vehicle based on the acquired reasoning results, thereby improving the driving safety of the vehicle and improving the user experience of the vehicle; and providing the user with the reasoning process can help the user clarify the reasoning logic and improve the user's understanding and trust in the reasoning results.

[0065] In one embodiment, displaying the inference result and / or the inference process corresponding to the inference result using a display device in a vehicle may include: in response to receiving an instruction that the user intends to obtain the inference result and / or the inference process corresponding to the inference result, displaying the content that the user intends to obtain using a display device in the vehicle.

[0066] In one embodiment, using a speaker in a vehicle to voice broadcast an inference result and / or an inference process corresponding to the inference result may include: in response to receiving an instruction that a user intends to obtain an inference result and / or an inference process corresponding to the inference result, using a speaker in the vehicle to voice broadcast the content that the user intends to obtain.

[0067] For example, when receiving a question from a user asking "What is the recommended speed?", it can be determined that the user intends to obtain the recommended speed in the inference result. At this time, the display device in the vehicle can be used to display the recommended speed in the inference result, or the speaker can be used to voice broadcast the recommended speed in the inference result. In this way, the user's intention can be fully respected, the user can obtain the content he needs, and the question-and-answer communication between the user and the vehicle can be realized, thereby improving the user experience of the vehicle.

[0068] Based on the same inventive concept, the present disclosure also provides a vehicle control device. Figure 6 is a block diagram of a vehicle control device 500 according to an exemplary embodiment. Figure 6 , the vehicle control device 500 may include: A determination module 501 is used to determine a target image; The reasoning module 502 is used to process the target image using a visual language model to obtain a reasoning result; A control module 503, used to control the intelligent driving of the vehicle according to the inference result; Wherein, a first module of the visual language model is loaded, wherein the first module is determined according to a scene classification result obtained by the visual language model according to the target image.

[0069] In the above technical solution, a target image is determined; the target image is processed using a visual language model to obtain an inference result; and the intelligent driving of the vehicle is controlled according to the inference result. Among them, the first module of the visual language model is loaded, and the first module is determined according to the scene classification result, and the scene classification result is obtained by the visual language model according to the target image. In this way, by loading the first module, the visual language model can be better adapted to the scene of the target image, and a more accurate inference result can be generated, thereby improving the reliability of the intelligent driving of the vehicle.

[0070] In some possible implementations, the visual language model includes a visual sub-model and a language sub-model, and the scene classification result is obtained by the visual sub-model based on the target image; the first module is a module in the language sub-model, and the first module is used to implement local parameter fine-tuning associated with the scene classification result.

[0071] In some possible implementations, the first module is a module determined according to the scene classification result using a routing mechanism and loaded through a parameter adjustment file.

[0072] In some possible implementations, the parameter adjustment file is generated in the following manner: For each scene in the pre-built scene database, locally fine-tuning training is performed on the visual language model that is not loaded with the first module to obtain local fine-tuning parameters corresponding to each scene; The parameter adjustment file including the local fine-tuning parameters corresponding to each of the scenes is generated.

[0073] In some possible implementations, the reasoning module 502 includes: A first reasoning submodule, configured to obtain an image feature vector according to the target image using the visual submodel; The second reasoning submodule is used to obtain the reasoning result according to the image feature vector using the language submodel loaded with the first module.

[0074] In some possible implementations, the vehicle control device 500 includes: A cache module, used for determining and caching keys and values ​​of a Transformer layer in the visual language model according to the target image before loading the first module; A calling module is used to perform attention calculation based on the cached keys and values ​​when reasoning is performed using the visual language model loaded with the first module, wherein the keys and values ​​of different target image caches are different.

[0075] In some possible implementations, the inference result includes data for controlling the vehicle state.

[0076] In some possible implementations, the vehicle control device 500 further includes at least one of the following: A first prompt module, configured to display the inference result and / or the inference process corresponding to the inference result by using a display device in the vehicle; The second prompt module is used to use a speaker in the vehicle to voice broadcast the reasoning result and / or the reasoning process corresponding to the reasoning result.

[0077] In some possible embodiments, the first prompt submodule is used to, in response to receiving an instruction that the user intends to obtain the inference result and / or the inference process corresponding to the inference result, display the content that the user intends to obtain using a display device in the vehicle; the second prompt submodule is used to, in response to receiving an instruction that the user intends to obtain the inference result and / or the inference process corresponding to the inference result, voice broadcast the content that the user intends to obtain using a speaker in the vehicle.

[0078] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0079] Figure 7 6 is a block diagram of a vehicle 600 according to an exemplary embodiment. For example, the vehicle 600 may be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 600 may be an intelligent driving vehicle or a semi-intelligent driving vehicle.

[0080] Reference Figure 7 , the vehicle 600 may include various subsystems, for example, an infotainment system 610, a perception system 620, a decision control system 630, a drive system 640, and a computing platform 650. The vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 600 may be interconnected by wire or wireless means.

[0081] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, and a navigation system, etc.

[0082] The perception system 620 may include several sensors for sensing information about the environment around the vehicle 600. For example, the perception system 620 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter wave radar, an ultrasonic radar, and a camera.

[0083] The decision control system 630 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.

[0084] The drive system 640 may include components that provide powered motion for the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting energy provided by the energy source into mechanical energy.

[0085] Some or all functions of the vehicle 600 are controlled by a computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652, and the processor 651 may execute instructions 653 stored in the memory 652.

[0086] The processor 651 may be any conventional processor, such as a commercially available CPU. The processor may also include a graphics processor (Graphic Process Unit, GPU), a field programmable gate array (Field Programmable Gate Array, FPGA), a system on chip (System on Chip, SOC), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC) or a combination thereof.

[0087] The memory 652 may be implemented by any type of volatile or nonvolatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0088] In addition to the instructions 653 , the memory 652 may also store data, such as road maps, route information, and data such as the location, direction, and speed of the vehicle. The data stored in the memory 652 may be used by the computing platform 650 .

[0089] In the embodiment of the present disclosure, the processor 651 can execute instruction 653 to complete all or part of the steps of the above-mentioned vehicle control method.

[0090] Figure 8 8 is a block diagram of an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a tablet device, a personal digital assistant, and the like.

[0091] Reference Figure 8 , the electronic device 800 may include one or more of the following components: a first processing component 802 , a first memory 804 , a first power component 806 , a multimedia component 808 , an audio component 810 , a first input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0092] The first processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The first processing component 802 may include one or more first processors 820 to execute instructions to complete all or part of the steps of the above-mentioned vehicle control method. In addition, the first processing component 802 may include one or more modules to facilitate the interaction between the first processing component 802 and other components. For example, the first processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the first processing component 802.

[0093] The first memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The first memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0094] The first power source component 806 provides power to various components of the electronic device 800. The first power source component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0095] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0096] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the first memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0097] The first input / output interface 812 provides an interface between the first processing component 802 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0098] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of the components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0099] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0100] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned vehicle control method.

[0101] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 804 including instructions, and the instructions can be executed by a first processor 820 of the electronic device 800 to complete the above-mentioned vehicle control method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0102] The present disclosure also provides a computer program product, which includes a computer program that can be executed by a programmable device, and has a code portion for executing the above-mentioned vehicle control method when the computer program is executed by the programmable device.

[0103] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and the program instructions, when executed by a processor, implement the steps of the vehicle control method provided by the present disclosure.

[0104] Some embodiments of the present disclosure also provide a chip system, such as Fig. 9 As shown, the chip system includes at least one second processor 1301 and at least one interface circuit 1302. The second processor 1301 and the interface circuit 1302 can be interconnected through lines. For example, the interface circuit 1302 can be used to receive signals from other devices (such as the memory of the electronic device). For another example, the interface circuit 1302 can be used to send signals to other devices (such as the second processor 1301). Exemplarily, the interface circuit 1302 can read the instructions stored in the memory and send the instructions to the second processor 1301. When the instructions are executed by the second processor 1301, the vehicle control device can execute the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, and some embodiments of the present disclosure do not specifically limit this.

[0105] In some embodiments of the present disclosure, the interface circuit 1302 can obtain data, program instructions and / or information from the internal storage area of ​​the chip system; it can also obtain data, program instructions and / or information from outside the chip system.

[0106] Optionally, the chip system may further include a second memory for storing necessary computer programs and data.

[0107] Those skilled in the art may also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application may be implemented by electronic hardware, computer software, or a combination of the two. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the functions described for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present application.

[0108] It should be understood that, unless otherwise specifically noted, the features of some embodiments of the various present disclosures described herein may be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of any two or more; similarly, "at least one of . . . " includes any one of the related listed items and any combination of any two or more.

[0109] Although terms such as "first", "second" and "third" can be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, component, region, layer or section from another component, component, region, layer or section. Therefore, without departing from the teachings of each example, the first component, component, region, layer or section mentioned in the examples described herein may also be referred to as the second component, component, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Thus, the features defined as "first" and "second" may expressly or implicitly include at least one of the features. In the description herein, the meaning of "multiple" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0110] In addition, the word "exemplary" is used herein to indicate serving as an example, instance, or diagram. Any aspect or design described as "exemplary" in this article is not necessarily understood to be advantageous compared to other aspects or designs. On the contrary, the use of the word exemplary is intended to present concepts in a specific way. As used herein, the term "or" is intended to represent an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to represent any one of the natural inclusive arrangements. That is, if X applies A; X applies B; or X applies both A and B, "X applies A or B" is satisfied under any of the aforementioned examples. In addition, unless otherwise specified or clearly pointed to a singular form from the context, the articles "one" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0111] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art after reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components (e.g., elements, resources, etc.) described above, unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific functions of the described components, even if the structure is not equivalent to the disclosed structure. In addition, although specific features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and beneficial to any given or specific application. In addition, with respect to "including", "having", "having", "having", or variations thereof used in a specific embodiment or claim, such terms are intended to be inclusive in a manner similar to the term "comprising".

[0112] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

[0113] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A vehicle control method, characterized in that: The method comprises: Determine the target image; Processing the target image using a visual language model to obtain an inference result; Controlling intelligent driving of the vehicle according to the inference result; Wherein, a first module of the visual language model is loaded, wherein the first module is determined according to a scene classification result obtained by the visual language model according to the target image.

2. The method according to claim 1, characterized in that The visual language model includes a visual sub-model and a language sub-model, and the scene classification result is obtained by the visual sub-model according to the target image; the first module is a module in the language sub-model, and the first module is used to implement local parameter fine-tuning associated with the scene classification result.

3. The method according to claim 1, characterized in that The first module is a module that is determined based on the scene classification result using a routing mechanism and loaded through a parameter adjustment file.

4. The method according to claim 3, characterized in that The parameter adjustment file is generated in the following way: For each scene in the pre-built scene database, locally fine-tuning training is performed on the visual language model that is not loaded with the first module to obtain local fine-tuning parameters corresponding to each scene; The parameter adjustment file including the local fine-tuning parameters corresponding to each of the scenes is generated.

5. The method according to claim 2, characterized in that: The process of processing the target image using the visual language model to obtain an inference result includes: Using the visual sub-model, obtaining an image feature vector according to the target image; The inference result is obtained according to the image feature vector using the language sub-model loaded with the first module.

6. The method according to claim 5, characterized in that The obtaining the inference result according to the image feature vector by using the language sub-model loaded with the first module includes: The inference result is obtained according to the image feature vector and the text used to indicate the target task, using the language sub-model loaded with the first module.

7. The method according to claim 1, characterized in that The method further comprises: Before loading the first module, determining and caching keys and values ​​of the visual language model based on the target image; When reasoning is performed using the visual language model loaded with the first module, attention calculation is performed based on the cached keys and values, wherein the keys and values ​​of different target image caches are different.

8. The method according to any one of claims 1 to 7, characterized in that The method further comprises at least one of the following: Displaying the inference result and / or the inference process corresponding to the inference result using a display device in the vehicle; The inference result and / or the inference process corresponding to the inference result are voice broadcasted using a speaker in the vehicle.

9. A vehicle control device, characterized in that: The device comprises: A determination module, used for determining a target image; An inference module, used to process the target image using a visual language model to obtain an inference result; A control module, used for controlling the intelligent driving of the vehicle according to the inference result; Wherein, a first module of the visual language model is loaded, wherein the first module is determined according to a scene classification result obtained by the visual language model according to the target image.

10. A vehicle, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions in the memory to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.