Vehicle light management method and system based on multiple modes

By adopting multimodal large models and multi-wheel cross-attention mechanisms in vehicle lighting management, combining knowledge graphs and correlation matrix, the problem of mode correlation in the existing technology is solved, and more efficient data fusion and more stable vehicle lighting management are achieved.

CN120096439APending Publication Date: 2025-06-06ZHEJIANG DISHI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510214409.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing multimodal data processing technology fails to fully consider the correlation between modes in vehicle lighting management, resulting in poor data fusion effect and difficult to efficiently process on equipment with limited resources.

Method used

A large model based on multimodality is adopted to materialize the modal data through a knowledge graph, and data fusion is used using multiple rounds of cross attention mechanisms and preset shared matrices to introduce correlation matrix to capture the correlation between modalities.

Benefits of technology

It improves the entity alignment accuracy of multimodal data, enhances the effect of data fusion, reduces the demand for computing resources, and makes vehicle lighting management more stable and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120096439A_ABST
    Figure CN120096439A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal-based vehicle light management method, which is characterized in that firstly, multi-modal data of a target vehicle are acquired, the multi-modal data at least comprise one of visual perception information, vehicle navigation information and radar detection information, the first modal data is one of the multi-modal data, and the second modal data is one of the multi-modal data; the second modal data is modal data except the first modal data in the multi-modal data; the first modal data or the second modal data are vehicle navigation information; the first modal data and the second modal data are input into a pre-trained multi-modal large model, the multi-modal large model uses a cross attention mechanism to identify the first modal data and / or the second modal data, and the current driving environment of the vehicle or the current real-time position of the vehicle is confirmed; the controller module controls the irradiation range of the LED lamp set according to the current driving environment of the vehicle or the current real-time position of the vehicle. And finally, vehicle lamps can be accurately managed through vehicle navigation information or environment information, so that the vehicle is safer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of vehicle light control, and in particular to a vehicle light management method and system based on multi-modality. Background Art

[0002] At present, with the popularity of vehicle lighting management systems, multimodal data processing technology is increasingly being used in vehicle lighting management. Multimodal data processing technology refers to the simultaneous processing of data from multiple sources, such as visual data, auditory data, text data, etc. These data can provide rich information about user behavior and environmental conditions, thereby improving the automation level and user experience of vehicle lighting management systems. At the same time, the lighting control system is a very important part of the vehicle lighting management system. It can automatically adjust the brightness and color of the lights according to user needs and environmental changes, thereby creating a comfortable lighting environment.

[0003] Existing multimodal data processing technologies usually use knowledge graphs to entity modal data, and then use multi-round cross-attention mechanisms and cross-modal correlation modeling to fuse data of different modalities. Knowledge graphs are a data representation method based on graph structures that can structure entities and their relationships, making it easier for machine learning models to understand and process them. Multi-round cross-attention mechanisms can make data of different modalities complement each other, thereby improving the effect of data fusion. Cross-modal correlation modeling can capture the correlation between different modalities, thereby more comprehensively understanding and utilizing multimodal data.

[0004] There are still some problems with existing multimodal data processing technologies in practical applications. First, existing technologies usually only consider the complementarity between modalities, but do not fully consider the correlation between modalities. Although the multi-round cross-attention mechanism can capture some complementarity between modalities, it cannot fully capture the correlation between modalities. Second, existing technologies usually require a lot of computing resources when processing multimodal data, which is a big challenge for some devices with limited resources.

[0005] Vehicle navigation information can obtain more comprehensive information. With the popularization of 5G technology, data transmission is faster and more convenient. Therefore, vehicle navigation information can add more road information, such as the size of the curve, the inclination of the curve, etc. This information is more important for controlling the headlights, but the vehicle navigation information will be delayed when the signal is poor. Therefore, relying solely on vehicle navigation information to control vehicle lights is very unstable and has poor reliability.

[0006] The invention patent with Chinese patent application number "2024111749525" and patent name "A method, device, storage medium and equipment for automobile lighting management" proposes to judge the vehicle's environment through multi-modal perception data, which can effectively respond to changes in road conditions and improve driving safety, and use large models to perform in-depth semantic analysis of perception data, so as to generate more intelligent lighting control strategies. However, due to the large amount of data processed, the reaction time from obtaining data information to controlling automobile lights is long, and it is impossible to quickly detect oncoming vehicles and quickly turn off the corresponding lighting module. Summary of the invention

[0007] The present disclosure provides a vehicle lighting management method and system based on multi-modality to at least solve one technical problem existing in the prior art.

[0008] According to a first aspect of the present disclosure, a vehicle light management method based on multimodality is provided, characterized in that multimodal data of a target vehicle is first acquired, the multimodal data comprising at least one of visual perception information, vehicle navigation information and radar detection information, the first modal data is one of the multimodal data, the second modal data is one of the multimodal data except the first modal data; the first modal data or the second modal data is the vehicle navigation information; Inputting the first modal data and the second modal data into a pre-trained multimodal large model, the multimodal large model uses a cross-attention mechanism to identify the first modal data and / or the second modal data, and confirm the current driving environment of the vehicle or the current real-time position of the vehicle; The controller module controls the illumination range of the LED light group according to the current driving environment of the vehicle or the current real-time position of the vehicle.

[0009] In one possible implementation, the multimodal large model uses a knowledge graph to entityize the modal data, specifically represented as each submodality M∈{r,a,v,n,g} of the entity ei, where r represents the relationship, a represents the attribute, v represents the vision, n represents the value, and g represents the graph structure; the L2 norm is used for normalization. The purpose of L2 norm normalization is to scale the feature vector to unit length. The formula is as follows:

[0010] in, is the eigenvector of the modality m of entity ei, and d is the dimension of the eigenvector; The multimodal large model uses a multi-round cross-attention mechanism. In the first round, the first modal data is used as the query set, and the second modal data is used as the key value set in turn. In the second round, the second modal data is used as the query set, and the first modal data is used as the key value set in turn. The preset shared matrices Wq, Wk, and Wv are used to parameterize them to obtain Q, K, and V. Then, the Sigmoid activation function is used to normalize them to obtain the cross-attention weights. Finally, the attention distribution is weighted and summed to obtain the modal output. The formula is as follows:

[0011]

[0012]

[0013] Where, dk represents the dimension of Q matrix, KT represents the device of K matrix, σ represents Sigmoid activation function, m, p∈M represent different modes, Represents entity e i The m-modal embedding, W q The matrix is ​​used to map the features of the input modality to a query vector (Q), which represents the target to be found, W k The matrix is ​​used to map the features of the second modality data into a key vector (K), which is used to compare with the query vector to determine the correlation between the features of different modalities, W v The matrix is ​​used to map the features of the second modality data into a value vector (V), which contains the selected feature information and is weighted and summed according to the matching degree between the query and the key. The above formulas are used to map the features of modality m through W. q Mapped to query vector Q, the features of mode p are passed through W k and W v Mapped into a key vector K and a value vector V, where m and p represent different modalities, for example, m can be a visual modality and p can be a textual modality; After multiple rounds of cross-attention mechanism processing, each modality m has taken into account the complementarity with other modalities, and these complementarily enhanced embedding vectors are concatenated to form a complementary embedding matrix, which is expressed as follows:

[0014] in, represents the concatenation operation, M is the set of all modes, Represents entity e i m-modal embedding of; After attention weighting, the embeddings of different modalities have a certain degree of complementarity. Each modality considers the contribution of other modal features and adds them to its own embedding. Cross-modal correlation modeling is used to capture the connection and shared information between different modalities. When modeling modal complementarity, the correlation matrix S∈RNm×Nm is introduced, where Nm represents the number of modalities, to adjust the modal embedding so that each entity considers the complementary relationship between modalities while also comprehensively considering the association relationship between modalities. After introducing cross-modal correlation modeling, the above formula is further updated to obtain the fusion matrix as follows:

[0015]

[0016] where · represents the dot product of vectors, Represents entity e i The m-modal embedding of is the correlation score between mode m and mode p, It is entity e i Embedding vector on modality p.

[0017] In one possible implementation, the multimodal data includes visual perception information, radar detection information, and cloud perception information.

[0018] In one embodiment, the data for managing vehicle lights also includes Ambient light sensor information: brightness of the vehicle's driving environment; Vehicle speed information: the speed of the vehicle; Vehicle steering information: the vehicle's steering angle and direction; Vehicle status information: including the vehicle's driving status, parking and starting status, parking and not starting status, and starting status; User input information: information input by the driver through the vehicle control panel or the control buttons on the steering wheel to manually control the lights; Vehicle navigation information: The navigation system provides route information and traffic information; Vehicle communication information: communication information between vehicles and other vehicles or infrastructure; Vehicle diagnostic information: The vehicle's diagnostic system monitors the working status of the lights, including switch status and brightness status; Vehicle external condition information: including weather conditions and road conditions.

[0019] In one embodiment, the high beam is turned off or a portion of the high beam is shielded when the following conditions are detected: When the ambient light intensity exceeds a certain intensity, such as 6000lx, or the vehicle speed is less than a certain speed, such as 40km / h, or the navigation system shows that it is in a city road and the visual perception information shows that the street lights are on, or the visual perception information detects an oncoming vehicle, or the visual perception information detects that there is a face within the high beam illumination range.

[0020] In one possible implementation manner, when the vehicle steering information shows that the outer wheel steering angle exceeds 10°, the outer high beam headlights are turned off or the corresponding lamp beads of the high beam headlights with an illumination range on the outer side are shielded.

[0021] In one possible implementation, the vehicle external condition information indicates that the high beam is turned off when the vehicle's climbing angle exceeds 6° or the descending angle exceeds 6°.

[0022] In one possible implementation, the hazard lights are turned on at a high frequency when the vehicle brakes suddenly, in heavy rain, or when the vehicle behind does not maintain a safe distance.

[0023] In one possible implementation, the visual perception module includes a head motion camera installed in the vehicle. After the head motion camera obtains image data, it inputs a trained head motion recognition algorithm to determine the driver's movements. When the algorithm determines that the driver has the following movements, an alarm light is immediately emitted: continuously opening and closing eyes, slowly shaking the head within a certain range, and lowering the head and then quickly raising the head.

[0024] In one possible implementation, the following conditions turn on the dangerous color, such as red ambient light in the car: When visual perception information shows that the vehicle is driving on a solid line; when visual perception information or radar detection information shows that the vehicle is not maintaining a safe distance from the vehicle being followed; when the vehicle's external condition information shows heavy rain or fog; when the vehicle's navigation information shows that a traffic accident has occurred on the road ahead; when the vehicle's navigation information shows that a traffic accident is likely to occur on the road ahead, when the vehicle's diagnostic information shows an abnormality, and when the vehicle's communication information is not processed.

[0025] According to a second aspect of the present disclosure, a multi-modal vehicle lighting system is provided, comprising a data acquisition module, a controller module and an LED matrix light group, wherein the data acquisition module comprises: Visual perception module, which collects visual perception information, including visible light videos and / or pictures; A radar detection module collects radar detection information, including one of laser radar, ultrasonic radar, and microwave radar; The cloud sensing module collects cloud sensing information, collects and sends information through the mobile network, and obtains control strategies in the cloud; Ambient light sensor module: collects the brightness of the vehicle's driving environment; Vehicle speed module: obtain the vehicle's speed; Vehicle steering module: obtains the steering angle and direction of the vehicle; Vehicle status module: confirm the status of the vehicle; User input module: including vehicle control panel or control buttons; Vehicle navigation module: obtain route information and traffic information; Vehicle communication module: communication information between the vehicle and other vehicles or infrastructure; Vehicle diagnostic module: obtain the working status of the vehicle lights; Vehicle external condition module: obtains rainfall information and road angle information, including rain sensor and vehicle body tilt sensor; The controller module analyzes and processes the information from the data acquisition module and controls the illumination range of the LED matrix light group according to the control strategy.

[0026] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory in communication with the at least one processor; wherein, The memory stores instructions that can be executed by at least one processor. The instructions are executed by the at least one processor so that the at least one processor can perform the method of the present disclosure.

[0027] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method of the present disclosure.

[0028] Compared with the prior art, the vehicle lighting management method based on multi-modality disclosed in the present invention has the following beneficial effects: Compared with the existing technology, the beneficial effects of this technical solution are as follows: Compared with the management of vehicle lights by vehicle navigation information alone, the multimodal vehicle light management method is more stable and reliable. The vehicle navigation information can be partially downloaded to the system, and other modal data, such as visual perception information and radar detection information, can be obtained when the vehicle is driving. Although the vehicle navigation information can obtain the position of the vehicle through the positioning device, the positioning will be delayed when the cloud layer is thick or other obstructions are encountered. In this way, it is impossible to perform real-time light management on the vehicle. Visual perception information and radar detection information cannot obtain the vehicle's position information. However, if the vehicle navigation information collects panoramic photos of the corresponding position when collecting data, the information obtained by comparing with the visual perception information and radar detection information can confirm the position of the vehicle and the environmental status of the road surface on which the vehicle is driving. For example, if the current vehicle is in a right turn state, the headlight illumination angle can be adjusted to the right to obtain a better line of sight. There are certain differences between visual perception information and radar detection information and panoramic photos. Therefore, choosing a cross-attention mechanism can better discover the common points between them to more accurately identify the vehicle's environment. Finally, the headlights can be accurately managed through vehicle navigation information or environmental information. For example, when the vehicle turns right, the headlights are turned to the right in advance, and the headlights are turned on before the vehicle enters the tunnel, making the vehicle safer in complex driving environments.

[0029] In the process of multimodal data processing, the present invention not only considers the complementarity between modalities, but also fully considers the correlation between modalities. By introducing the correlation matrix, the present invention can more comprehensively understand and utilize multimodal data, thereby improving the accuracy of entity alignment. This is in sharp contrast to the prior art method that only considers the complementarity of modalities. The fusion method of the present invention is more comprehensive and accurate.

[0030] The present invention uses a preset shared matrix to share parameters between different modes, which can not only improve the efficiency of the model, but also reduce the number of model parameters. This is a great advantage for some devices with limited resources. The prior art usually requires a large amount of computing resources when processing multimodal data, and the present invention is more efficient.

[0031] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, in which: In the drawings, the same or corresponding reference numerals represent the same or corresponding parts.

[0033] Figure 1 A schematic diagram of an implementation flow of a vehicle lighting management method based on multi-modality according to an embodiment of the present disclosure is shown; Figure 2 A schematic structural diagram of a multi-modal vehicle lighting system according to an embodiment of the present disclosure is shown; Figure 3 A schematic diagram of an LED matrix light group based on a multi-modal vehicle lighting system according to an embodiment of the present disclosure is shown; Figure 4 A schematic diagram of a process flow of determining the driver's state and a corresponding processing method in a vehicle lighting management method based on a multi-modal embodiment of the present disclosure is shown; Figure 5 A schematic diagram of navigation map information of a vehicle light management method based on multimodality according to an embodiment of the present disclosure is shown; Figure 6 A schematic diagram of a photo taken at position A1 by a visual perception module of a vehicle light management method based on multimodality according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0034] In order to make the purpose, features, and advantages of the present disclosure more obvious and easy to understand, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.

[0035] refer to Figure 1, a flowchart of a vehicle lighting management method based on multimodality, obtains multimodal data of a target vehicle; and materializes the modal data using a knowledge graph; uses L2 norm for normalization; uses multiple rounds of cross-attention mechanism, each round first uses the first modal data as a query set, such as the first round uses visual perception information, the second round uses radar detection information, and other modalities are used as key-value sets in turn, that is, other information except visual perception information in the first round, and information except radar detection information in the second round, and a preset shared matrix is ​​used to parameterize it; then the Sigmoid activation function is used for normalization to obtain the cross-attention weight, and the attention weight can also be obtained through the softmax function, and finally the attention distribution is weighted and summed to obtain the modal output; after multiple rounds of cross-attention mechanism processing, each modality has considered the complementarity with other modalities, and these embedding vectors enhanced by complementarity are spliced ​​together to form a complementary embedding matrix, and the vehicle lighting is managed according to the complementary embedding matrix data. For example, when there is sufficient light during the day, all lights are turned off without manual lighting, that is, no other factors are considered; when there is insufficient light, the vehicle speed is less than 20 kilometers per hour, only the low beam can be turned on, and the high beam can be turned off. If necessary, the high beam can be turned on manually or controlled by voice commands; when the vehicle speed is higher than 100 kilometers per hour, the high beam will be automatically turned on when there is insufficient light; based on the feedback of navigation information, the vehicle can turn on the lights before entering the tunnel. Lighting management is not only about high and low beam lights, but also about ambient lights, turn signals, etc.

[0036] Knowledge Graph (KG) aims to represent relevant knowledge in the real world in a structured way, where concepts are represented as nodes and the relationships between concepts are represented as edges. These structured knowledge can also be represented as triples[1].<s,p,o> The representation form means <entity, relationship, entity> in the relationship triple and <entity, attribute, attribute value> in the attribute triple.

[0037] The modal data is materialized using the knowledge graph, specifically represented as each sub-modality M∈{r,a,v,n,g} of the entity ei, where r represents the relationship, a represents the attribute, v represents the vision, n represents the value, and g represents the graph structure; using the L2 norm for normalization means dividing a vector by its L2 norm so that the length of the vector (L2 norm) is equal to 1. This operation is also called "unitization" or "normalization". The purpose of L2 norm normalization is to scale the feature vector to unit length, and the formula is as follows:

[0038] in, is the feature vector of the modality m of entity ei, and d is the dimension of the feature vector. Normalizing the vector can avoid the scale difference between features of different modalities, making the feature vectors of different modalities comparable when calculating similarity. In multimodal learning, normalization helps the model better capture the complementarity and correlation between different modalities. Through this formula, we can normalize the feature vectors of different modalities to unit length, preparing for subsequent feature fusion and entity alignment.

[0039] The task of the attention mechanism is to obtain local attention information. The introduction of the attention mechanism allows the task to know which parts of the input data are more worthy of attention. It is usually expressed as Q, K, V, namely Query, Key and Value. All three are derived from the input features themselves and are vectors generated based on the input features. Q can be regarded as a vector of a single input feature. When a set of Q is directly input into the network for training, the network is a network without the introduction of the attention mechanism. However, if the attention mechanism is introduced, it is necessary to multiply this set of Q by a set of weights W (K, V) respectively, so that the local input features can be paid attention to. W (K, V) is used to calculate the similarity between K and V. The common one is the dot-product attention mechanism that uses dot product calculation, that is, Q is the vector representing the input feature, K and V are the feature vectors for calculating the weights of the attention mechanism, and all three are obtained from the input features. It means that the similarity of the current Key and Value is calculated, and this similarity value is calculated through the Softmax layer to obtain a set of weights. The Value value under the attention mechanism is obtained by summing this set of weights and the corresponding Value and product.

[0040] Using multiple rounds of cross attention mechanism, each round firstly takes the first modality data as the query set, and the second modality data as the key value set in turn, that is, in the first round, the first modality data is taken as the query set, and the second modality data is taken as the key value set in turn, in the second round, the second modality is taken as the query set, and the first modality data is taken as the key value set in turn, and so on, to achieve multiple rounds; and use the preset shared matrix W q ,W k ,W v It is parameterized to get Q, K, V, and then normalized using the Sigmoid activation function to get the cross attention weight. Finally, the attention distribution is weighted and summed to get the modal output. The formula is as follows:

[0041]

[0042]

[0043] Among them, d krepresents the dimension of the Q matrix, K T represents the device of K matrix, σ represents Sigmoid activation function, m, p∈M represent different modes, Represents entity e i The Sigmoid function is a common S-shaped function in biology, also known as the S-shaped growth curve. [1] In information science, due to its monotonic and inverse monotonic properties, the Sigmoid function is often used as an activation function of a neural network to map variables between 0 and 1. q The matrix W is used to map the features of the input modality to a query vector (Q), which represents the target to be found. k The matrix is ​​used to map the features of the second modality data into a key vector (K), which is used to compare with the query vector to determine the association between the features of different modalities. v The matrix is ​​used to map the features of the second modality data into a value vector (V), which contains the selected feature information and is weighted and summed according to the matching degree between the query and the key. The above formulas are used to map the features of modality m through W q Mapped to query vector Q, the features of mode p are passed through W k and W v Mapped into a key vector K and a value vector V. Here m and p represent different modalities. For example, m can be a visual modality, and p can be a text modality. The text modality can come from information on a navigation map, other interactive information on the network, or information based on a big data platform. In this way, the model can learn complementary information between different modalities. The benefit of using a shared matrix is ​​that it allows the model to share parameters between different modalities, which can improve the efficiency and generalization ability of the model while reducing the number of parameters of the model. In addition, the shared matrix can also help the model learn common features across modalities, which is very useful for understanding and fusing multimodal data.

[0044] After multiple rounds of cross-attention processing, each modality m has taken into account the complementarity with other modalities. These complementarity-enhanced embedding vectors are concatenated to form a complementary embedding matrix, which is expressed as follows:

[0045] in, represents the concatenation operation, M is the set of all modes, Represents entity e im-modal embedding. After attention weighting, the embeddings of different modalities have a certain degree of complementarity. Each modality considers the contribution of other modal features and adds them to its own embedding. Cross-modal correlation modeling is used to capture the connection and shared information between different modalities. When modeling modal complementarity, the correlation matrix S∈RNm ×Nm is introduced, where Nm represents the number of modalities, to adjust the modal embedding so that each entity considers the complementary relationship between modalities while also comprehensively considering the association relationship between modalities. After introducing cross-modal correlation modeling, the above formula is further updated as follows:

[0046]

[0047] where · represents the dot product of vectors, Represents entity e i The m-modal embedding of is the correlation score between mode m and mode p, It is entity e i Embedding vector on modality p. In this way, the model not only considers the complementarity between different modalities, for example, the visual modality may provide the physical features of the entity, while the textual modality provides the semantic description of the entity. The correlation matrix helps the model understand how these modalities complement each other. It also considers the correlation between modalities, so as to understand and fuse multimodal data more comprehensively. For example, in video content, visual information (video frames) and audio information (dialogue, background music) are closely related. For example, when a vehicle enters a tunnel, there will be a voice prompt "The vehicle enters the tunnel, please turn on the lights", and the video information will also appear in a short period of black screen. When passing through a dangerous area, there will be an alarm sound, and the video information will show complex road conditions such as large bends or ramps. In foggy weather, corresponding warnings will also be heard, and the video information will become unclear. At this time, it may be necessary to turn on the position lights to warn the following vehicles. The correlation matrix can help the model identify and utilize these relationships. This fusion method helps to improve the accuracy of entity alignment because it integrates rich information from different modalities.

[0048] For adaptive headlights, when the vehicle is traveling on a curve, it will turn to the position where the road disappears so that the illumination range is farther, and road surface information at a farther position can be observed. If multimodal data is not used, the accuracy of judging the curvature of the road based solely on image data is poor. Since the illumination range of vehicle headlights is limited at night, it is difficult to use visible light data alone to judge the extension position of the road. Therefore, radar data can be used for complementarity. Radar data and visible light image data are quite different, so they need to be fused. In addition, the information of the navigation map can be used to judge the road surface, but the navigation data may be delayed, so the weight cannot be too large, or the judgment of whether there is a delay in the navigation data can be further introduced. The weight of the navigation data can be directly determined according to the delay of the navigation data.

[0049] Whether there is a delay in navigation data depends on whether there is a delay in positioning. Positioning is closely related to signal reception. For example, in an open location, the positioning signal is good and there is basically no delay. However, if it is blocked or in rainy weather, the positioning information will be deviated. Therefore, the navigation data delay will generally last for a certain period of time, so there is no need for real-time detection. For example, when the vehicle's steering direction changes, the vehicle's inclination angle changes, which can be compared with the navigation data. Of course, the navigation map data also needs to include lane corner radius and inclination related information for comparison.

[0050] Multimodal large models can be deep learning algorithms, for example: NLP is a technology that studies language problems in human-computer interaction. GCN, the full name of which is Graph Convolutional Networks, is a deep learning model that specializes in processing graph structured data. It processes graph data by defining graph convolution operations, which allow the model to learn embedded representations of nodes while considering both the feature information of the nodes and the structural information of the graph. GCN can be used for tasks such as node classification, graph classification, and edge prediction, and obtains embedded representations of the graph. This network structure is particularly useful in processing non-Euclidean structured data such as social networks and chemical molecular structures, because it does not require a fixed input format and can process irregularly structured data.

[0051] refer to Figure 5 and Figure 6 In one embodiment, the multimodal data of the vehicle including visual perception information and panoramic navigation map information is obtained, the vehicle position is determined by using the Beidou or GPS positioning system, and the panoramic navigation map information of the current position of the vehicle is obtained, that is, a panoramic camera is used to first take panoramic photos of different positions on the map, and in order to reduce the amount of data, a group of panoramic photos are taken at intervals of 5 to 10 meters, that is, Figure 5In the figure, panoramic photos are taken at positions A1, A2 and A3 respectively. The visual perception information obtained by the vehicle can also be taken with a similar wide-angle lens, so that the similarity of the photos is higher. Since the vehicle has a certain speed, the objects on both sides are not clear, especially at night, so only the picture information in front is obtained, such as Figure 6 As shown, the deep learning algorithm is used to compare the visual perception information, i.e., the corresponding position photo and the image information of the panoramic navigation map information to find the closest panoramic image shooting position, so as to determine the position of the vehicle and judge whether there is a delay in the vehicle positioning information. If there is no delay, the navigation map data can be used directly to manage the headlights. If there is a delay, the vehicle position is determined according to the panoramic photo shooting position closest to the visual perception information. Of course, the determination of the vehicle position must be continuous. If there is a jumpy conclusion, it means that the detection fails and needs to be repeatedly confirmed or manually confirmed before managing the headlights according to the corresponding positioning information. Controlling the headlights according to the panoramic navigation map information mainly adjusts the left and right up and down angles of the headlights, so as to obtain better lighting effects on curved roads and undulating mountain roads. The panoramic navigation map information requires manual measurement and collection of information such as the road slope and the bending radius, so that as long as the position information is obtained, the headlights can be controlled according to the position and the vehicle's driving direction, and better lighting effects can be obtained when driving at night.

[0052] In this embodiment, if the visual perception information is used as the first modal data and the panoramic navigation map information is used as the second modal data, as the vehicle moves, the visual perception information, i.e., the photographs taken, changes in real time, and the panoramic navigation map information also provides different panoramic photos according to different positioning, and can provide multiple panoramic photos with similar positioning for comparison. After the calculation of the cross-attention mechanism, it is confirmed which of the photographs taken is closest to the panoramic photo, and the real position of the vehicle and the current driving environment can be confirmed, such as Figure 6 As shown, the vehicle is in a right turn state, and the angle of the headlights needs to be adjusted to the z1 position to obtain a farther field of vision, detect abnormal conditions of the road surface in advance, and make driving safer.

[0053] If visual perception information is used as the second modal data and the panoramic navigation map information is used as the first modal data, a cross-attention mechanism calculation can be performed on the panoramic photos obtained through positioning and the photos taken with the visual perception information obtained a certain time before and after. By confirming the time when the closest photo was taken, the navigation delay time can be known. After that, the navigation positioning can be adjusted to continue the comparison. After the navigation information has no delay, the headlights can be directly controlled according to the navigation information. For example, if the navigation shows a right turn ahead, the headlights will be deflected to the right by a certain angle according to the navigation information. If the navigation can record the slope information of the road, the headlights can be further adjusted according to the slope information.

[0054] In one embodiment, the vehicle navigation information also collects control information of the headlights, that is, when a vehicle is manually driven on a target road, not only a panoramic photo of the road is collected, but also the headlight control information is manually collected, that is, the headlights of the vehicle are most recently adjusted at this position. In this way, if there is no delay when using navigation, the headlights can be directly controlled according to the information collected by the navigation. If there is a delay, the positioning information can be adjusted using the technical solution of the present application to avoid the delay of the navigation information. The technical solution of the present application can accurately know the delay time to adjust the navigation information.

[0055] The management of car lights is not only about the management of high beams and low beams, but also includes the integrated control of interior ambient lights, interior lighting, rear brake lights, turn signals, position lights, etc. Therefore, multi-modal data analysis is required to obtain more intelligent control. Of course, different lights require different data for targeted control, which is more efficient. First of all, when the light intensity is high, the high beam, low beam and interior lighting do not need to be turned on, so the information data for controlling the high beam, low beam and interior lighting is not collected. Therefore, some information has a higher priority for car light management and does not need to be fused. However, for the judgment of the same feature, multi-modal fusion is required, such as judging whether there are vehicles and pedestrians on the road.

[0056] Visual perception information mainly obtains visible light pictures or video information through camera equipment, and uses various deep learning algorithms to analyze and identify each direction.

[0057] Radar detection information is obtained directly through radar equipment and can be converted into image information or other modal information.

[0058] The cloud perceives information, mainly by uploading some of the information obtained to the cloud and obtaining certain feedback information from the cloud. The cloud can also select some useful information for the vehicle from the big data and send it to the vehicle, such as the area prone to landslides, the real-time weather information of the area, and some information that cannot be recognized locally in the area; use big data for analysis and obtain corresponding results, and also obtain some solutions from the cloud. In some special cases, the vehicle can be controlled directly in the cloud. In this way, the intelligent solution for headlight management and the update of headlight management strategy can be obtained from the cloud.

[0059] In addition to visual perception information, radar detection information, and cloud perception information, the information in the vehicle that controls the lights also includes: Ambient light sensor information: The ambient light sensor can detect the brightness of the surrounding environment, thereby controlling the automatic turning on and off of the headlights. For example, turn on the headlights when entering a tunnel, and of course turn off the high beam in the tunnel depending on the situation.

[0060] Vehicle speed information: The vehicle speed information can be used to control the brightness or mode of the lights. For example, when driving at high speed, that is, exceeding 90KM / h and the light intensity is less than 6000lx, the lights will automatically switch to high beam.

[0061] Vehicle steering information: The vehicle's steering angle can be used to control the follow-up steering function of the headlights to improve the lighting effect when turning. If the outer headlights have a small rotation range and cannot rotate at the same angle as the wheels, the meaningless part of the high beam illumination range can be directly controlled by closing or shielding. Vehicle steering information can also confirm whether navigation information exists, that is, whether the positioning is accurate.

[0062] Vehicle status information: including the vehicle's driving status, parking status, etc. This information can be used to control the lighting mode of the headlights in different states. The parking status includes the parking start state, the parking not start state, and the starting state. In principle, the starting state forces the headlights with high power consumption to be turned off.

[0063] User input information: information input by the driver through the vehicle control panel or the control buttons on the steering wheel to manually control the lights. Some vehicles can also realize voice control, which can also be considered as user input information. This control has a higher priority and can change the automatic control mode to the custom control mode or manual control mode. The manual mode will not interfere with the operation of the lights except in dangerous situations, and the custom mode realizes partial intelligent control according to the defined functions.

[0064] Vehicle navigation information: The navigation system can provide route information, and combined with road condition information, the headlight system can adjust the lighting strategy in advance. For example, in areas with a lot of vehicles, on roads with close following distances, and when waiting for traffic lights, the high beams can be turned off. The high beams can also be turned on in advance when the vehicle is about to enter a tunnel. The headlights can be turned off in the tunnel according to the light intensity.

[0065] Vehicle communication information: Communication information between a vehicle and other vehicles or infrastructure, such as V2X (Vehicle-to-Everything) technology, can be used for coordinated control of headlights. For example, parallel vehicles can choose one vehicle to turn on the high beam depending on the situation, while the rear vehicle can turn off the high beam if there are other vehicles in front or behind.

[0066] Vehicle diagnostic information: The vehicle's diagnostic system can monitor the status of the headlights, including switch status and brightness status, bulb life, faults, etc., to ensure the normal operation of the headlight system.

[0067] Vehicle location information: The vehicle's GPS location information can be used to control the lighting mode of the headlights in specific geographical locations, such as tunnels or specific road sections. Since vehicle navigation information usually requires a positioning device to confirm the location, the vehicle's location information can be obtained by obtaining generalized vehicle navigation information, and the vehicle's location information can no longer be obtained separately. When the signal is poor, the positioning may be inaccurate, which can be used as a reference for managing the headlights.

[0068] Vehicle external condition information: such as weather conditions (rain, fog, snow, etc.), which can be used to adjust the lighting strategy of the headlights to improve driving safety. Road conditions include the road's inclination, road flatness, and other road information, such as whether the road has isolation strip lighting hardware conditions, which can be intelligently identified using visual devices.

[0069] In one embodiment, the high beam is turned off or a portion of the high beam is shielded when the following conditions are detected: When the ambient light intensity exceeds 6000lx, this value can also be adjusted according to personal vision. If the vision is good, it can be adjusted lower, and if the vision is poor, it can be adjusted higher. After setting, it has little effect on the control strategy. Or when the vehicle speed is less than 40km / h, or the navigation system shows that it is in a city road and the visual perception information shows that the street lights are on, or when the visual perception information detects an oncoming vehicle, or when the visual perception information detects that there is a face in the high beam illumination range. When the speed is close to 40km / h but the vehicle is accelerating, the high beam needs to be turned on.

[0070] In addition to high beams, vehicles also have low beams to provide lighting, so the high beams can be automatically turned off in many road conditions. Of course, the working status of the low beams must be tested first. If the low beams cannot be turned on due to damage or other reasons, the high beams cannot be automatically turned off under certain conditions. For example, when the vehicle speed is higher than 90KM / h and the ambient light intensity is lower than 4000lx, all automated intelligent settings should consider safety as the first factor.

[0071] In one embodiment, when the vehicle steering information shows that the outer wheel steering angle exceeds 10°, the outer high beam lamps are turned off or the corresponding lamp beads of the high beam lamps with an illumination range on the outer side are shielded.

[0072] In one embodiment, the vehicle external condition information indicates that the high beam is turned off when the vehicle's climbing angle exceeds 6° or the descending angle exceeds 6°.

[0073] In one embodiment, the fast-frequency double flash lights are turned on when the vehicle brakes suddenly, in heavy rain, or when the rear vehicle does not maintain a safe distance. The fast flash frequency exceeds 1 time per second, which is different from ordinary double flash lights. This can warn the rear vehicle that there is a certain danger. Of course, ordinary braking only requires turning on the brake lights. When the braking acceleration is greater than 6 m / s², the double flash lights can be used for warning.

[0074] In one embodiment, the red ambient light in the vehicle is turned on in the following states: When the visual perception information shows that the vehicle is driving on the solid line; when the visual perception information or radar detection information shows that the vehicle is not keeping a safe distance from the vehicle; when the vehicle external condition information shows that it is raining or foggy; when the vehicle navigation information shows that a traffic accident has occurred on the road ahead; when the vehicle navigation information shows that a traffic accident is likely to occur on the road ahead, when the vehicle diagnostic information shows an abnormality, and when the vehicle communication information is not processed. The red ambient light in the car can remind the driver to pay attention to danger and drive carefully. The color of the light can be set to purple, orange, yellow, or other colors that the vehicle can provide according to personal preference.

[0075] refer to Figure 2 , a vehicle lighting system based on multi-modality, including a data acquisition module, a controller module and an LED matrix light group, the data acquisition module includes: Visual perception module, which collects visual perception information, including visible light videos and / or pictures, mainly one or more cameras; A radar detection module collects radar detection information, including one of laser radar, ultrasonic radar, and microwave radar; The cloud sensing module collects cloud sensing information, collects and sends information through the mobile network, and obtains control strategies in the cloud; Vehicle navigation module: obtain route information and road condition information, and use GPS or Beidou positioning sensors; Vehicle communication module: communication information between the vehicle and other vehicles or infrastructure; Ambient light sensor module: collects the brightness of the vehicle's driving environment; Vehicle speed module: obtain the vehicle's speed; Vehicle steering module: obtains the steering angle and direction of the vehicle, that is, sets the steering sensor; Vehicle status module: confirm the status of the vehicle; User input module: including vehicle control panel or control buttons; Vehicle diagnostic module: obtain the working status of the vehicle lights; Vehicle external condition module: obtains rainfall information and road angle information, including rain sensor and vehicle body tilt sensor; The controller module analyzes and processes the information from the data acquisition module and controls the illumination range of the LED matrix light group according to the control strategy.

[0076] refer to Figure 3The LED matrix light group is composed of multiple independent lamp beads, so some of the lamp beads can be turned on at will to control the lighting range. Of course, for the management mode of car light control with a higher frequency, the lighting range of the car lights can be controlled by blocking, or by changing the lighting angle of the LED lights, or by shielding part of the lights to control the lighting range.

[0077] refer to Figure 4 According to this process, the driver's status can be judged and the corresponding processing method can be carried out. The visual perception module includes a head motion camera installed in the car. After the head motion camera obtains image data, it inputs the trained head motion recognition algorithm to determine the driver's action. When the algorithm determines that the driver has the following actions, the alarm light will be immediately emitted: continuously opening and closing eyes, slowly shaking the head within a certain range, and lowering the head and then quickly raising the head.

[0078] The camera determines the driver's driving status through the driver's actions, and judges whether the driver is driving fatigued, sleepy, or has closed eyes. If the driver closes his eyes for more than three seconds, the camera immediately adjusts the interior lights to the brightest state and flashes continuously. If the vehicle has an automatic driving function, it can simultaneously slow down and pull over. If the driver is in a state of fatigue driving, the interior ambient light and voice prompt the driver to pull over to rest. The ambient light can flash twice in a cycle and then stay on for a period of time before going out, and then enter the next cycle. This cycle continues until the driver manually turns it off or the camera's intelligent judgment determines that the driver is refreshed and can drive normally. Generally speaking, you can use the ambient light for prompts first, and then use the sound warning if there is no effect, and then use emergency braking if there is no effect. If there is an assisted driving function, you can pull over.

[0079] The following situations can be used to determine that the driver is in poor condition, such as continuously opening and closing eyes, slowly shaking the head within a certain range, lowering the head and then quickly raising it. Opening and closing eyes can be judged by ordinary face recognition, which can be trained using multiple pictures. Action recognition requires the use of models related to action detection technology after training. In action detection, posture estimation and motion tracking are core technologies. Posture estimation refers to identifying the joint position and motion state of the human body from pictures or videos. Motion tracking refers to detecting and tracking targets by tracking the target's motion trajectory. The models that can be used are: Posture estimation model based on deep learning: Deep learning models can automatically learn parameters from a large amount of training data to achieve high-precision posture estimation. Commonly used deep learning models include convolutional neural networks (CNN) and recurrent neural networks (RNN), etc., and can also be predicted using OpenCV software and algorithm models such as YOLOV8 and slowfast. These algorithms can be trained to identify actions such as continuously opening and closing eyes, slowly shaking the head within a certain range, and lowering the head and then quickly raising it, and then judge the driver's driving status.

[0080] By using one of the above models, the driver's actions can be identified to determine whether the driver is driving fatigued, and the driver can be reminded in time in case of fatigue driving to avoid major accidents. In addition to lights, the reminder can also be a sharp sound. This can greatly improve safety.

[0081] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0082] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0083] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0084] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0085] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0086] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0087] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0088] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0089] The above is only a specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present disclosure, which should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.

Claims

1. A vehicle lighting management method based on multi-modality, characterized in that: First, multimodal data of a target vehicle is obtained, wherein the multimodal data includes at least one of visual perception information, vehicle navigation information, and radar detection information, the first modal data is one of the multimodal data, and the second modal data is a modal data in the multimodal data other than the first modal data; the first modal data or the second modal data is the vehicle navigation information; Inputting the first modal data and the second modal data into a pre-trained multimodal large model, wherein the multimodal large model uses a cross-attention mechanism to recognize the first modal data and / or the second modal data, and confirm the current driving environment of the vehicle or the current real-time position of the vehicle; The controller module controls the illumination range of the LED light group according to the current driving environment of the vehicle or the current real-time position of the vehicle.

2. The multi-modal vehicle lighting management method according to claim 1, characterized in that: The multimodal big model uses the knowledge graph to entityize the modal data, which is specifically represented as entity e i Each sub-modality M∈{r,a,v,n,g}, where r represents the relationship, a represents the attribute, v represents the vision, n represents the value, and g represents the graph structure; the L2 norm is used for normalization, and the formula is as follows: in, It is entity e i is the eigenvector of mode m, where d is the dimension of the eigenvector.

3. The multi-modal vehicle lighting management method according to claim 2, characterized in that: After normalization using the L2 norm, multiple rounds of cross attention mechanism are used. In the first round, the first modality data is first used as the query set, and the second modality data is sequentially used as the key value set. In the second round, the second modality data is first used as the query set, and the first modality data is sequentially used as the key value set; and the preset shared matrix W is used. q ,W k ,W v It is parameterized to get Q, K, V, and then normalized using the Sigmoid activation function to get the cross attention weight. Finally, the attention distribution is weighted and summed to get the modal output. The formula is as follows: Among them, d k represents the dimension of the Q matrix, K T represents the device of K matrix, σ represents Sigmoid activation function, m, p∈M represent different modes, Represents entity e i The m-modal embedding, W q The matrix is ​​used to map the features of the first modal data into a query vector (Q), which represents the target to be found. The Wk matrix is ​​used to map the features of the second modal data into a key vector (K), which is used to compare with the query vector to determine the correlation between the features of different modalities. v The matrix is ​​used to map the features of the second modality data to a value vector (V), which contains the selected feature information and is weighted and summed according to the matching degree between the query and the key. The above formulas map the features of modality m to the query vector Q through Wq and the features of modality p to the query vector Q through Wq. k and W v Mapped into a key vector K and a value vector V, where m and p represent different modalities, for example, m can be a visual modality and p can be a textual modality; After multiple rounds of cross-attention mechanism processing, each modality m has taken into account the complementarity with other modalities, and these complementarily enhanced embedding vectors are concatenated to form a complementary embedding matrix, which is expressed as follows: in, represents the concatenation operation, M is the set of all modes, Represents entity e i m-modal embedding of; After attention weighting, the embeddings of different modalities have a certain degree of complementarity. Each modality considers the contribution of other modal features and adds them to its own embedding. Cross-modal correlation modeling is used to capture the connection and shared information between different modalities. When modeling modal complementarity, the correlation matrix S∈RNm×Nm is introduced, where Nm represents the number of modalities, to adjust the modal embedding so that each entity considers the complementary relationship between modalities while also comprehensively considering the association relationship between modalities. After introducing cross-modal correlation modeling, the above formula is further updated as follows: where · represents the dot product of vectors, Represents entity e i The m-modal embedding of is the correlation score between mode m and mode p, It is entity e i Embedding vector on modality p.

4. The multi-modal vehicle lighting management method according to any one of claims 1 to 3, characterized in that: The data for managing vehicle lights also includes: Ambient light sensor information: brightness of the vehicle's driving environment; Vehicle speed information: the speed of the vehicle; Vehicle steering information: the vehicle's steering angle and direction; Vehicle status information: including the vehicle's driving status, parking and starting status, parking and not starting status, and starting status; User input information: information input by the driver through the vehicle control panel or the control buttons on the steering wheel to manually control the lights; Vehicle navigation information: The navigation system provides route information and traffic information; Vehicle communication information: communication information between vehicles and other vehicles or infrastructure; Vehicle diagnostic information: The vehicle's diagnostic system monitors the working status of the lights, including switch status and brightness status; Vehicle external condition information: including weather conditions and road conditions.

5. The multi-modal vehicle lighting management method according to claim 4, characterized in that: Turn off or partially block the high beam when the following conditions are detected: When the ambient light intensity exceeds a certain intensity, or the vehicle speed is less than a certain value, or the navigation system shows that it is in a city road and the visual perception information shows that the street lights are on, or the visual perception information detects an oncoming vehicle, or the visual perception information detects that there is a face in the high beam illumination range.

6. The multi-modal vehicle lighting management method according to claim 5, characterized in that: The vehicle steering information shows that when the outer wheel steering angle exceeds a certain angle, the outer high beam headlights are turned off or the corresponding lamp beads of the high beam headlights with an illumination range outside are shielded.

7. The multi-modal vehicle lighting management method according to claim 6, characterized in that: The vehicle external condition information shows that the high beam is turned off when the vehicle's climbing angle exceeds a certain angle or the descending angle exceeds a certain angle.

8. The multi-modal vehicle lighting management method according to claim 7, characterized in that: The visual perception module includes a head motion camera installed in the car. After the head motion camera obtains image data, it inputs the trained head motion recognition algorithm to determine the driver's movements. When the algorithm determines that the driver has the following movements, the alarm light will be immediately emitted: continuously opening and closing eyes, slowly shaking the head within a certain range, and lowering the head and then quickly raising it.

9. The multi-modal vehicle lighting management method according to any one of claims 5 to 8, characterized in that: The following conditions turn on the interior hazard color ambient light: When the visual perception information shows that the vehicle is driving on a solid line; when the visual perception information or radar detection information shows that the vehicle is not keeping a safe distance from the vehicle being followed; when the vehicle external condition information shows heavy rain or fog; when the vehicle navigation information shows that a traffic accident has occurred on the road ahead; when the vehicle navigation information shows that a traffic accident is likely to occur on the road ahead; when the vehicle diagnostic information shows an abnormality; when the vehicle communication information is not processed.

10. A vehicle lighting system based on multi-modality, characterized in that: It includes a data acquisition module, a controller module and an LED light group, and the data acquisition module includes: Visual perception module, which collects visual perception information, including visible light videos and / or pictures; A radar detection module collects radar detection information, including one of laser radar, ultrasonic radar, and microwave radar; The cloud sensing module collects cloud sensing information, collects and sends information through the mobile network, and obtains control strategies in the cloud; Ambient light sensor module: collects the brightness of the vehicle's driving environment; Vehicle speed module: obtain the vehicle's speed; Vehicle steering module: obtains the steering angle and direction of the vehicle; Vehicle status module: confirm the status of the vehicle; User input module: including vehicle control panel or control buttons; Vehicle navigation module: obtain route information and traffic information; Vehicle communication module: communication information between the vehicle and other vehicles or infrastructure; Vehicle diagnostic module: obtain the working status of the vehicle lights; Vehicle external condition module: obtains rainfall information and road angle information, including rain sensor and vehicle body tilt sensor; The controller module analyzes and processes the information of the data acquisition module and controls the illumination range of the LED light group according to the control strategy.