Vehicle control method and system, and computer equipment

By acquiring multimodal data and mapping it into light language semantic labels, the problems of environmental interference and signal conflict in vehicle light signal recognition are solved, collaborative decision-making and interaction among multiple vehicles are achieved, and the accuracy and safety of vehicle control are improved.

CN120388357BActive Publication Date: 2025-09-09CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510887972.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-09
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing technologies are easily affected by environmental interference when recognizing headlight signals, have difficulty processing headlight signals in complex scenarios, lack the ability to interact with multiple vehicles in two directions, and are unable to resolve signal conflicts.

Method used

By acquiring multimodal data, including headlight image information, vehicle location information, and environmental information, the light language features are determined and mapped into light language semantic labels using the headlight language coding table. Vehicle control is then performed to execute or terminate driving intentions, achieving multi-vehicle two-way collaborative interaction and signal conflict resolution.

Benefits of technology

It effectively avoids misjudgment caused by environmental interference on a single sensor, realizes collaborative decision-making and signal conflict processing among multiple vehicles, and reduces the probability of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388357B_ABST
    Figure CN120388357B_ABST
Patent Text Reader

Abstract

The present application provides a vehicle control method and system, and computer equipment, including: obtaining multimodal data of a first target vehicle, and determining light language features based on the multimodal data, and then using a pre-obtained vehicle light language coding table to map the light language features into light language semantic labels, and controlling a second target vehicle according to the light language semantic labels, so that the first target vehicle, based on the control result of the second target vehicle, executes or terminates the driving intention corresponding to the light language semantic label of the first target vehicle. By obtaining multimodal data for light language signal recognition, the present application can solve the problem that a single sensor is susceptible to environmental interference and may make misjudgments. At the same time, it predicts the vehicle's driving behavior or driving intention based on the light language features, so as to form a response strategy in advance, effectively avoid sudden or high-risk situations, and reduce the probability of traffic accidents. In addition, it can combine traffic participant information to achieve two-way collaborative interaction among multiple vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle technology, and in particular to a vehicle control method and system, and computer equipment. Background Art

[0002] When recognizing headlight signals, related technologies primarily rely on single sensors (e.g., cameras, radars, etc.), which are susceptible to environmental interference, such as billboard lights, rainy and foggy weather, leading to misjudgments. Furthermore, these technologies struggle to interpret headlight signals in real time in complex scenarios, such as continuous flashing and multi-vehicle interaction. Furthermore, in some scenarios, when signal lights are present or headlights are obscured by objects like trees or preceding vehicles, these technologies struggle to enhance driver perception through auxiliary means. Furthermore, these technologies only support one-way headlight recognition, such as the vehicle's one-way recognition of preceding vehicle signals, and lack the ability to interact with multiple vehicles in two directions (e.g., preceding vehicle feedback and multi-vehicle platooning negotiation), making collaborative decision-making difficult in complex scenarios. Furthermore, when multiple vehicles send simultaneous headlight signals, these technologies are unable to resolve signal conflicts, such as when a request to overtake conflicts with a yield instruction. Summary of the Invention

[0003] In view of the above-mentioned shortcomings of the prior art, the purpose of this application is to provide a vehicle control method and system, and a computer device to solve the technical problems existing in the related art.

[0004] To achieve the above objectives and other related objectives, the present application provides a vehicle control method, comprising the following steps:

[0005] Acquire multimodal data of a first target vehicle, the multimodal data including vehicle light image information, vehicle position information, and environmental information, the environmental information including road information and traffic participant information;

[0006] Determining light language features of the first target vehicle based on the multimodal data, the light language features including light flashing frequency, light color, light area, and light language intent;

[0007] The light language features are mapped into light language semantic labels using a pre-obtained vehicle light language coding table, and the second target vehicle is controlled according to the light language semantic labels, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic labels based on the control results of the second target vehicle; wherein the first target vehicle and the second target vehicle are located within a preset range.

[0008] In one embodiment of the present application, the process of controlling the second target vehicle according to the light language semantic label so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the control result of the second target vehicle includes:

[0009] Controlling the lights of the second target vehicle according to the light language semantic label, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the light control result of the second target vehicle;

[0010] or,

[0011] The lights and speed of the second target vehicle are controlled according to the light language semantic label, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the light and speed control results of the second target vehicle.

[0012] In one embodiment of the present application, the process of controlling the second target vehicle according to the light language semantic label so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the control result of the second target vehicle includes:

[0013] The vehicle light language code table is used as an optical coding interaction protocol between a transmitting vehicle and a receiving vehicle; wherein the transmitting vehicle and the receiving vehicle are determined from the traffic participant information, the first target vehicle, and the second target vehicle;

[0014] The receiving vehicle identifies the optical signal parameters and generates a feedback code or a warning code based on the identification result of the optical signal parameters; wherein the optical signal parameters are obtained by the sending vehicle mapping the optical code representing the driving intention of the sending vehicle according to the vehicle light language coding rules;

[0015] The feedback code or the warning code is parsed by the sending vehicle, and the driving intention is executed according to the parsing result of the feedback code, or the driving intention is terminated according to the parsing result of the warning code.

[0016] In one embodiment of the present application, the process of identifying optical signal parameters by the receiving vehicle and generating a feedback code or a warning code based on the identification result of the optical signal parameters includes:

[0017] Capturing vehicle light signals from the optical signal parameters by the receiving vehicle, including extracting vehicle light flashing frequency and vehicle light color; and

[0018] Performing point cloud data verification on the optical signal parameters by the receiving vehicle, including locating the sending vehicle by using the point cloud data and verifying the motion trajectory of the sending vehicle;

[0019] Performing multi-spectral fusion from the optical signal parameters by the receiving vehicle, including collecting visible light and near-infrared band images, and filtering ambient light based on band differences;

[0020] The feedback code or the warning code is generated according to the vehicle light signal capture result, the point cloud data verification result and the multi-spectral fusion result.

[0021] In one embodiment of the present application, the method further includes: performing time window competition according to the priorities of the transmitting vehicle and the receiving vehicle, and transmitting a coding instruction according to the time window competition result; wherein the coding instruction includes: the optical signal parameter, the feedback code or the warning code;

[0022] And / or, within a preset time window, the vehicle with the highest priority among the traffic participant information, the first target vehicle and the second target vehicle is used as the sending vehicle, and the remaining vehicles are used as receiving vehicles, and only the vehicle with the highest priority is allowed to send the optical signal parameters.

[0023] In one embodiment of the present application, the process of obtaining the vehicle light language code table includes:

[0024] Define the meaning of light language based on the type of light, the position of the vehicle and the number of times the light flashes; and

[0025] Determining a vehicle light language encoding rule based on a pre-determined or real-time determined vehicle light language field, a vehicle light language field encoding length, and a vehicle light language field description, and obtaining a complete vehicle light language encoding based on the vehicle light language encoding rule; wherein the vehicle light language field includes a signal type, a flashing mode, a flashing number, a light source direction, a color identifier, a priority, and a preset reserved bit;

[0026] The light language meaning definition, the vehicle light language field and the vehicle light language complete code are associated to form a vehicle light language code table.

[0027] In one embodiment of the present application, the process of controlling the second target vehicle according to the light language semantic label includes:

[0028] Displaying the light language semantic label through a pre-configured target display in the second target vehicle; wherein the target display is located in the field of view of the driver of the second target vehicle;

[0029] Based on the light language semantic label and the occlusion information of the first target vehicle, an audio warning is issued to the second target vehicle; and when the distance between the first target vehicle and the second target vehicle is less than a preset distance, the second target vehicle is braked and the brake lights of the second target vehicle are controlled to flash; wherein the occlusion information is identified by an enhanced display device.

[0030] In one embodiment of the present application, the process of determining the light language feature of the first target vehicle based on the multimodal data includes:

[0031] Calculating a headlight brightness change frequency based on pixel changes in the headlight image information to obtain a flashing frequency of the first target vehicle; and

[0032] Segmenting the headlight image information into red and yellow regions according to a preset color space, and obtaining the headlight color of the first target vehicle based on the red and yellow region segmentation results; wherein the headlight color corresponding to the red region segmentation result is red, and the headlight color corresponding to the yellow region segmentation result is yellow; and,

[0033] Predicting the light language intention of the first target vehicle based on the vehicle light image information within a preset time period and the vehicle motion trajectory of the first target vehicle within a preset continuous time period; and

[0034] When the headlights in the headlight image information are blocked, predicting the headlight positions of the first target vehicle using the point cloud data of the first target vehicle; and

[0035] The three-dimensional data frame of the point cloud data is projected into a two-dimensional image to locate the headlight area of ​​the first target vehicle.

[0036] The present application also provides a vehicle control system, the system comprising:

[0037] a data acquisition module, configured to acquire multimodal data of a first target vehicle, the multimodal data including headlight image information, vehicle position information, and environmental information, the environmental information including road information and traffic participant information;

[0038] a light language feature module, configured to determine light language features of the first target vehicle based on the multimodal data, the light language features including light flashing frequency, light color, light area, and light language intent;

[0039] A vehicle control module is used to map the light language features into light language semantic labels using a pre-obtained light language coding table, and to control a second target vehicle according to the light language semantic labels, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the control result of the second target vehicle; wherein the first target vehicle and the second target vehicle are located within a preset range.

[0040] The present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any one of the vehicle control methods described above.

[0041] The present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any one of the above-mentioned vehicle control methods when executed by a processor.

[0042] As described above, the present application provides a vehicle control method and system, and a computer device, which have the following beneficial effects: the present application obtains multimodal data of a first target vehicle, and then determines the light language characteristics of the first target vehicle based on the multimodal data, and then uses a pre-obtained light language coding table to map the light language characteristics into light language semantic labels, and controls the second target vehicle according to the light language semantic labels, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label of the first target vehicle based on the control result of the second target vehicle; wherein the multimodal data includes headlight image information, vehicle position information and environmental information, the environmental information includes road information and traffic participant information, and the light language characteristics include headlight flashing frequency, headlight color, headlight area and light language intention; the first target vehicle and the second target vehicle are located within a preset range. It can be seen from this that when controlling a vehicle, the present application can solve the problem that a single sensor is easily affected by environmental interference and misjudgments by acquiring multimodal data for light language signal recognition. At the same time, based on the light language characteristics of the first target vehicle, the driving behavior or driving intention of the first target vehicle can be predicted, so that a response strategy can be formed in advance for the second target vehicle within a preset range of the first target vehicle, thereby effectively avoiding sudden or high-risk situations and reducing the probability of traffic accidents. In addition, the present application can also realize two-way collaborative interaction of multiple vehicles in combination with traffic participant information, and when multiple vehicles send light language at the same time, the signal conflict problem can be resolved in a priority manner, so as to realize collaborative decision-making in complex scenarios and reduce vehicle driving misjudgments. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A schematic flow chart of a vehicle control method provided in one embodiment of the present application;

[0044] Figure 2A schematic diagram of the process flow of vehicle cooperative interactive control provided in an embodiment of the present application;

[0045] Figure 3 A schematic flow chart of a vehicle control method provided in another embodiment of the present application;

[0046] Figure 4 A schematic diagram of the hardware structure of a vehicle control system provided in one embodiment of the present application;

[0047] Figure 5 The figure is a schematic diagram of the hardware structure of a computer device suitable for implementing one or more embodiments of the present application. DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present application by specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It is understood that, in the absence of conflict, the features in the following embodiments and embodiments can be combined with each other. In addition, it is understood that the illustrations provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner, and only the components related to the present application are shown in the drawings rather than being drawn according to the number, shape and size of the components during actual implementation. The type, quantity and proportion of each component during actual implementation can be a kind of arbitrary change, and its component layout type may also be more complicated.

[0049] Figure 1 A flow chart of a vehicle control method is shown. Specifically, in an exemplary embodiment, as Figure 1 As shown, this embodiment provides a vehicle control method, which includes the following steps:

[0050] S110, acquiring multimodal data of a first target vehicle, where the multimodal data includes vehicle headlight image information, vehicle position information, and environmental information, where the environmental information includes road information and traffic participant information;

[0051] S120, determining light language features of the first target vehicle based on the multimodal data, where the light language features include light flashing frequency, light color, light area, and light language intent;

[0052] S130, using a pre-obtained vehicle light language coding table to map the light language features into light language semantic labels, and controlling the second target vehicle according to the light language semantic labels, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic labels based on the control result of the second target vehicle; wherein the first target vehicle and the second target vehicle are located within a preset range.

[0053] In some embodiments, the first target vehicle and the second target vehicle may be vehicles including assisted driving functions or intelligent driving functions. The first target vehicle and the second target vehicle may be fuel vehicles, new energy vehicles, or hybrid vehicles. As some examples, the first target vehicle may be referred to as the other vehicle, the preceding vehicle, or the surrounding vehicle, and the second target vehicle may be referred to as the own vehicle, the own vehicle, or the following vehicle.

[0054] In some embodiments, the preset range can be selected or set based on the actual scenario and is not specifically limited herein. For example, the preset range can be selected or set based on the capture range of the image capture device in the first target vehicle and / or the second target vehicle, or based on the perception range of the LiDAR (Light Laser Detection and Ranging) in the first target vehicle and / or the second target vehicle, or based on the capture range of the image capture device and / or the perception range of the LiDAR of other vehicles in the traffic participant information.

[0055] In some exemplary embodiments, before acquiring multimodal data from the first target vehicle, the method may further include optically encoding the first target vehicle's driving intent using a light language coding table, and controlling the first target vehicle's lights according to the optical encoding result of the first target vehicle's driving intent. Thus, before acquiring multimodal data from the first target vehicle, the lights are controlled according to the first target vehicle's driving intent, so that the subsequently acquired light image information of the first target vehicle is light image information associated with the driving intent.

[0056] In some exemplary embodiments, the process of acquiring multimodal data of a first target vehicle may include: capturing the lights of the first target vehicle using a predetermined or real-time image capture device to obtain image information of the lights of the first target vehicle. The image capture device includes, but is not limited to, an onboard camera. For example, the onboard camera may capture infrared or RGB images of the headlights, taillights, turn signals, and other lights of the first target vehicle. Furthermore, point cloud data of the first target vehicle may be collected using a predetermined or real-time laser radar (LiDAR) device, and vehicle position information and the headlight area of ​​the first target vehicle may be determined based on the point cloud data. For example, the three-dimensional position of the first target vehicle may be determined using the point cloud data of the first target vehicle. Furthermore, the headlight area of ​​the first target vehicle may be determined based on the point cloud density of the headlight area. Determining the headlight area using point cloud data is also applicable to lights in obstructed scenarios. Furthermore, the image capture device may capture road information, traffic participant information, and the point cloud data of the road information, traffic participant information, and the like may be obtained using the LiDAR device. Thus, environmental information such as road information and traffic participant information may be obtained using the image capture device and the LiDAR device of the first target vehicle. Environmental information may also include vehicle control information such as the steering wheel angle, throttle position, and brake position of the first target vehicle. Traffic participant information may include other vehicles, people, and obstacles within a preset range. When acquiring multimodal data on the first target vehicle, augmented reality (AR) devices may be used to assist in identification. The specific process of assisting identification using AR devices can be found in related technologies and will not be further elaborated here.

[0057] In some exemplary embodiments, after acquiring multimodal data of the first target vehicle, the process may further include preprocessing the multimodal data, including filtering the headlight image information to remove noise in rain, fog, or at night; aligning the point cloud data with the image coordinate system of the image capture device to generate fused data consisting of the headlight image information and the point cloud data; and performing signal synchronization on the multimodal data to align the timestamps of the multimodal data. Specifically, during image denoising, Gaussian filtering may be used to remove noise in the headlight image information in rain, fog, or low light conditions at night. During signal synchronization, the multimodal data may be aligned using timestamps to ensure a signal error of less than or equal to 10ms. As one example, the light signal characteristics of the first target vehicle may be determined based on the pre-processed multimodal data. As another example, the light signal characteristics of the first target vehicle may be determined based on the pre-processed multimodal data. As yet another example, the light signal characteristics of the first target vehicle may be determined based on both the pre-processed and pre-processed multimodal data.

[0058] In some exemplary embodiments, the process of determining the light language characteristics of the first target vehicle based on multimodal data includes: calculating the frequency of headlight brightness changes based on pixel changes in headlight image information to obtain the flashing frequency of the first target vehicle; and, segmenting the headlight image information into red areas and yellow areas according to a preset color space, and obtaining the headlight color of the first target vehicle based on the red area segmentation results and the yellow area segmentation results; wherein the headlight color corresponding to the red area segmentation result is red, and the headlight color corresponding to the yellow area segmentation result is yellow; and, based on the headlight image information within a preset time period and the vehicle motion trajectory of the first target vehicle in a preset continuous time period, predicting the light language intention of the first target vehicle; and, when the headlights in the headlight image information are blocked, using the point cloud data of the first target vehicle to predict the position of the headlights of the first target vehicle; and, projecting the three-dimensional data frame of the point cloud data into a two-dimensional image to locate the headlight area of ​​the first target vehicle. Specifically, as an example, the optical flow method can be used to determine pixel variations in the headlight image information. The frequency of headlight brightness variations can then be calculated based on the pixel variations in the headlight image information to obtain the flashing frequency of the first target vehicle. The flashing frequency can be expressed in Hz. As an example, the headlight image information can be segmented into red and yellow regions using the HSV color space. The red region segmentation result is used to represent a braking vehicle, with the corresponding headlight color being red; the yellow region segmentation result is used to represent a turn signal, with the corresponding headlight color being yellow. As an example, the preset time period can be selected or set based on the actual scenario and is not limited to a specific value. For example, the preset time period can be determined based on the headlight flashing frequency. As an example, when the headlights in the headlight image information are obscured, such as by trees or other objects in a headlight image captured by an on-board camera, the headlight location can be inferred using the density distribution of the point cloud data. Furthermore, the 3D data frame generated from the point cloud data can be projected onto a 2D image to accurately locate the headlight area. As an example, when determining the light signal characteristics, the light signal signals of consecutive frames can also be modeled through a long short-term memory network (LSTM) to identify dynamic patterns, such as the "double flash warning" intermittent flashing at 2Hz.

[0059] In some exemplary embodiments, the process of mapping light language features into light language semantic labels may include: defining the meaning of light language based on the type of light, the position relationship of the vehicle, and the number of times the light flashes; and determining the light language coding rules based on the light language field, the light language field coding length, and the light language field description determined in advance or in real time, and obtaining the complete light language coding based on the light language coding rules; wherein the light language field includes signal type, flashing mode, number of flashes, light source direction, color identification, priority, and preset reserved bits; associating the light language meaning definition, the light language field, and the light language complete coding to form a light language coding table; using the light language coding table to match the light language features, and mapping the light language features to light language semantic labels.

[0060] Specifically, as an example, when the meaning of light language is defined according to the type of light, the position relationship of the vehicle, and the number of times the light flashes, a light language interpretation table can be obtained, as shown in Table 1 below.

[0061] Table 1 Interpretation of car light language

[0062]

[0063] When the vehicle light language coding rules are determined according to the vehicle light language field, the vehicle light language field coding length and the vehicle light language field description determined in advance or in real time, a vehicle light language coding rule table can be obtained, as shown in Table 2 below.

[0064] Table 2 Car light language coding rules

[0065]

[0066] Based on the vehicle light language coding rules, the complete vehicle light language coding is obtained. At the same time, the light language meaning definition, vehicle light language field and vehicle light language complete coding are associated to form a vehicle light language coding table, as shown in Table 3 below.

[0067] Table 3 Car light language coding table

[0068]

[0069] In some exemplary embodiments, the vehicle control method may further include: controlling the lights of a second target vehicle based on the light language semantic label; after the second target vehicle completes the light control, obtaining multimodal data of the second target vehicle based on the light control results of the second target vehicle through the first target vehicle; determining the light language characteristics of the second target vehicle based on the multimodal data of the second target vehicle; and mapping the light language characteristics of the second target vehicle to light language semantic labels using a light language coding table; then, the first target vehicle controls itself based on the light language semantic labels of the second target vehicle, executing or suspending the driving intention corresponding to the light language semantic labels of the first target vehicle; and, after the first target vehicle executes or suspends the driving intention corresponding to the light language semantic labels of the first target vehicle, the first target vehicle controls the lights again, so that the second target vehicle can confirm whether the first target vehicle executes or suspends its driving intention after reacquiring the multimodal data of the first target vehicle. As an example, in a two-vehicle overtaking scenario, the leading vehicle is designated as the first target vehicle and the trailing vehicle is designated as the second target vehicle. When a trailing vehicle detects that the preceding vehicle is traveling slower than its own speed and that overtaking is permitted on the current road, it triggers an overtaking intention and controls its lights accordingly, flashing its right turn signal at a 3Hz rate. Meanwhile, the vehicle uses LiDAR to detect the distance to the preceding vehicle and the position of vehicles to the side and rear. The preceding vehicle then determines the following vehicle's light language signature based on the multimodal data from the preceding vehicle's right turn signal flashing at 3Hz. The following vehicle then uses a light language encoding table to map these signatures to semantic labels. Based on the semantic labels, the following vehicle controls itself, such as turning on its left turn signal and hazard lights, and simultaneously slowing down the preceding vehicle by 10%. Based on the preceding vehicle's activation of its left turn signal and hazard lights, as well as its simultaneous 10% speed reduction, the following vehicle can determine that the preceding vehicle has approved its overtaking intention. The following vehicle can then increase its speed while maintaining a 3Hz flashing pattern on its right turn signal and using its lidar to monitor its distance to the preceding vehicle and the position of vehicles to the side and rear. This process continues until the following vehicle has passed the preceding vehicle and is at a distance greater than or equal to the safety distance set by the preceding vehicle. At this point, the following vehicle controls its taillights to display a signal indicating the overtaking is complete (e.g., "taillights display √"), notifying the preceding vehicle of the completion of the overtaking attempt. The preceding vehicle also stops decelerating and resumes its pre-deceleration speed. This demonstrates that after the first target vehicle initiates its driving intention through its lights, the second target vehicle recognizes the first target vehicle's lights and then controls its own lights. This allows the first target vehicle to determine whether the second target vehicle agrees with its driving intention based on the second target vehicle's light control results, enabling coordinated control of the first and second vehicles through their lights.

[0070] In some exemplary embodiments, the vehicle control method may also include: controlling the lights of the second target vehicle according to the light language semantic label, and after the second target vehicle completes the light and speed control, obtaining the multimodal data of the second target vehicle based on the light control result and speed control result of the second target vehicle through the first target vehicle, and determining the light language characteristics of the second target vehicle based on the multimodal data of the second target vehicle, and mapping the light language characteristics of the second target vehicle to the light language semantic label using the light language coding table, and then the first target vehicle controls itself according to the light language semantic label of the second target vehicle, and executes or terminates the driving intention corresponding to the light language semantic label of the first target vehicle; and after the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label of the first target vehicle, the first target vehicle controls the lights again, so that after the second target vehicle obtains the multimodal data of the first target vehicle again, it can confirm whether the first target vehicle executes or terminates its driving intention. It can be seen from this that after the first target vehicle initiates the driving intention through the headlights, the second target vehicle recognizes the headlights of the first target vehicle, and then controls the headlights and speed of the second target vehicle itself, so that the first target vehicle can know whether the second target vehicle agrees with its execution of the corresponding driving intention based on the headlight control results and speed control results of the second target vehicle, so that the first target vehicle and the second target vehicle can achieve coordinated control through the headlights.

[0071] In some exemplary embodiments, controlling the second target vehicle based on the light signal semantic tag may include: displaying the light signal semantic tag via a preconfigured target display in the second target vehicle; wherein the target display is located in the field of view of the driver of the second target vehicle; issuing an audio warning to the second target vehicle based on the light signal semantic tag and the obstruction information of the first target vehicle; and, when the distance between the first target vehicle and the second target vehicle is less than a preset distance, braking the second target vehicle and flashing its brake lights; wherein the obstruction information is identified by an enhanced display device. As an example, if the first target vehicle and the second target vehicle are vehicles that include assisted driving or intelligent driving functions, when the assisted driving or intelligent driving functions are activated, controlling the second target vehicle based on the light signal semantic tag may include highlighting the light signal semantic tag of the first target vehicle in the driver's field of view via a preconfigured HUD (Head Up Display) in the second target vehicle, such as marking the "emergency braking" vehicle with a red box in the HUD. Simultaneously, issuing an audio warning to the second target vehicle based on the light signal semantic tag and the obstruction information of the first target vehicle. When an audio warning is issued to the second target vehicle, a graded alert can be triggered based on the semantic urgency. For example, a "beep" indicates a normal warning, while a continuous beep indicates a collision risk warning. The obstruction information is identified by an enhanced display device. The enhanced display device here operates on the same principle as the enhanced display device in some of the aforementioned embodiments, so the process of identifying obstruction information by the enhanced display device will not be further described here. When the second target vehicle detects that the distance between it and the first target vehicle is less than a preset distance, AEB (Automatic Emergency Braking) is triggered to brake the second target vehicle. The preset distance can be selected or set based on the actual scenario and is not limited to a specific value. For example, the preset distance can be a safe distance for the second target vehicle. Simultaneously, the brake lights of the second target vehicle are controlled to flash, transmitting a signal to the vehicle behind. This means that the second target vehicle's brake lights flash synchronously at a preset frequency when braking. Furthermore, the second target vehicle can display the first target vehicle's signal interpretation results on its onboard screen and provide the user with "Confirm" or "Ignore" options for interactive feedback to prevent accidental activation. From this, we can see that controlling the second target vehicle can precisely locate the target vehicle's light signal area by combining point cloud data with image fusion technology. Then, combined with traffic participant status information, a warning or intelligent driving action can be triggered. The light signal analysis results are shown in Table 4 below.

[0072] Table 4 Light language analysis results

[0073]

[0074] In some exemplary embodiments, when mapping light signal features to semantic labels, semantic classification can also be performed. This involves fusing headlight image information and point cloud data using a neural network architecture, such as the Transformer architecture. Specifically, the headlight image information is input into a Convolutional Neural Network (CNN), which outputs image features; and the point cloud data is input into a PointNet, which outputs point cloud features. The image and point cloud features are then used as inputs to the Transformer architecture, which outputs a probability distribution for the light signal labels. For example, the probability distribution for the light signal label "emergency braking" is 0.92, and the probability distribution for the light signal label "left turn" is 0.05. Finally, the classification confidence threshold is adjusted based on environmental conditions (such as visibility), for example, increasing the probability distribution for the rain threshold from 0.8 to 0.9. The core of the Transformer architecture is an encoder-decoder structure that uses a self-attention mechanism to achieve parallel processing of sequential data and global dependency modeling. PointNet is a network that directly processes point cloud data and can perform classification and segmentation.

[0075] In some exemplary embodiments, the vehicle control method may further include: selecting two vehicles from traffic participant information, a first target vehicle and a second target vehicle, and selecting one of the vehicles as a sending vehicle and the other as a receiving vehicle; using a pre-obtained vehicle light language coding table as an optical coding interaction protocol between the sending vehicle and the receiving vehicle; the sending vehicle maps the optical code representing the driving intention of the sending vehicle according to the vehicle light language coding rules to obtain optical signal parameters; the receiving vehicle identifies the optical signal parameters and generates a feedback code or a warning code based on the identification result of the optical signal parameters; the sending vehicle parses the feedback code or the warning code and executes the driving intention based on the parsing result of the feedback code, or terminates the driving intention based on the parsing result of the warning code. Among them, the traffic participant information may include other vehicles, personnel and obstacles that are within a preset range with the first target vehicle and the second target vehicle. The process of identifying optical signal parameters by the receiving vehicle and generating a feedback code or a warning code based on the identification results of the optical signal parameters may include: capturing headlight signals from the optical signal parameters by the receiving vehicle, including extracting the headlight flashing frequency and headlight color; performing point cloud data verification from the optical signal parameters by the receiving vehicle, including locating the sending vehicle through the point cloud data and verifying the movement trajectory of the sending vehicle; performing multispectral fusion from the optical signal parameters by the receiving vehicle, including collecting visible light and near-infrared band images and filtering ambient light interference through band differences; and generating the feedback code or the warning code based on the headlight signal capture results, the point cloud data verification results, and the multispectral fusion results. In addition, in some examples, when transmitting coding instructions such as optical signal parameters, feedback codes, or warning codes, the priority of the sending vehicle and the receiving vehicle can also be obtained, and time window competition can be performed according to the priority, and the coding instructions can be transmitted according to the time window competition results; wherein the coding instructions include but are not limited to optical signal parameters, feedback codes, or warning codes. In other examples, within a preset time window, the vehicle with the highest priority among the traffic participant information, the first target vehicle, and the second target vehicle is used as the sending vehicle, and the remaining vehicles are used as the receiving vehicles, and within the preset time window, only the vehicle with the highest priority is allowed to send optical signal parameters.

[0076] In some examples, the transmitting and receiving vehicles can utilize a coded command multi-vehicle interaction protocol for collaborative interactive control. For example, basic command encoding can be performed based on predefined light language encoding rules, such as those in a light language encoding table, as part of the coded command multi-vehicle interaction protocol. Alternatively, the basic command encoding can be expanded to support dot-matrix LEDs (Light-Emitting Diodes) displaying dynamic graphics (e.g., arrows, numbers, etc.), and then a convolutional neural network can be used to recognize the semantics of the graphics as part of the coded command multi-vehicle interaction protocol.

[0077] As an example, for two adjacent vehicles, the rear vehicle can be used as the sending vehicle and the front vehicle as the receiving vehicle. The process of vehicle cooperative interactive control between the sending vehicle and the receiving vehicle can be as follows: Figure 2As shown. Specifically, the sending vehicle uses the image capture device, laser radar, etc. in the vehicle to perceive the environment, and then forms the driving intention of the sending vehicle (such as an overtaking request, etc.) based on the environmental perception results. Among them, the specific process of the sending vehicle forming the driving intention based on the environmental perception results can be found in other related technologies and will not be described in detail here. For example, when the sending vehicle and the receiving vehicle are both in the second lane, and the environmental perception results of the sending vehicle include that there are no other vehicles, pedestrians, or other objects 200 meters before and after the parallel position of the sending vehicle in the first lane, the sending vehicle can form an overtaking request driving intention based on the environmental perception results. The sending vehicle then uses the light language coding rules in the coding rule library to map the corresponding driving intention into optical signal parameters, and synchronously controls the vehicle lights according to the optical signal parameters corresponding to the driving intention, such as turning on the left turn signal. Among them, the light language coding rules in the coding rule library are obtained based on the vehicle light language coding table and will not be described in detail here. When the transmitting and receiving vehicles engage in collaborative interactive control, conflicting coding instructions (such as optical signal parameters, feedback codes, or warning codes) may occur between the two vehicles. This necessitates time window preemption. During time window preemption, vehicles can compete for the current time window based on their priority. Furthermore, within a preset time window, only the vehicle with the highest priority is allowed to send coding instructions. Vehicle priority includes, but is not limited to, vehicle type, lane position, and the audio and visual warnings of specialized vehicles. For example, an ambulance has a higher priority than other ordinary vehicles, allowing it to interrupt the current time window. If no coding instruction signal conflict occurs, the corresponding vehicles can take turns sending corresponding coding instructions in a 100ms / segment time window. As an example, the time window preemption mechanism can be shown in Table 5 below. In Table 5, vehicle C has a higher priority than both vehicle A and vehicle B, and vehicle A has a higher priority than vehicle B. When vehicle C interrupts the current time window at 150ms, vehicle B, which previously occupied the 100ms-200ms window, stops sending coding commands. Since each vehicle takes turns sending coding commands within a 100ms window, vehicle C now takes 150ms-250ms to send coding commands. In this case, vehicle C could be a special vehicle, such as an ambulance, while vehicles A and B could be ordinary vehicles.

[0078] Table 5 Time window preemption mechanism

[0079]

[0080] The receiving vehicle then performs signal capture, including: taking an image of the headlights of the sending vehicle through an on-board camera, and then extracting headlight features such as the flashing frequency and headlight color of the sending vehicle based on the photographed headlight image. The receiving vehicle then performs LiDAR verification, including: locating the signal source vehicle through LiDAR point cloud data, and verifying the motion trajectory of the signal source vehicle; for example, the motion trajectory of the signal source vehicle can be verified by verifying the acceleration and steering angle of the signal source vehicle. By performing point cloud data verification, when the optical signal of the receiving vehicle is partially blocked, the motion trajectory of the signal source vehicle can be analyzed through point cloud data verification to infer its driving intention, thereby identifying whether the signal source vehicle and the sending vehicle are the same vehicle. The receiving vehicle then performs multispectral fusion, including: collecting image features of visible light and near-infrared band images, and then filtering ambient light interference (such as billboard reflections) through band differences, including: , where Represents the image features after multispectral fusion, Represents the image features of visible light band (400-700nm) images, Represents the image features of near-infrared band (800-1000nm) images, α represents the dynamic weight, ; and dynamic weight It can automatically adjust according to the ambient light intensity, for example, in a night environment, It can be 0.3, in which case the image features of near-infrared band images can be used first.

[0081] The receiving vehicle then performs multimodal perception and recognition using the multispectral fusion image features and verified LiDAR trajectory data. It then performs semantic parsing, including inputting the multimodal perception and recognition results into a CNN classification model, which then outputs corresponding light signal semantic labels (e.g., overtaking requests) and other instruction labels. Feedback is then generated, including making decisions based on the light signal semantic labels, generating feedback codes or warning codes, and transmitting these codes to the sending vehicle. The semantic parsing process can be similar to the semantic classification process described in some of the aforementioned embodiments and will not be further elaborated here.

[0082] The sending vehicle performs another environmental perception exercise to determine whether to execute or abort the driving intention based on the perceived results. If the perceived results meet the driving intention's execution conditions, the CNN classification model performs semantic analysis on the feedback or warning code, outputting corresponding light language semantic labels (e.g., "overtaking allowed," "overtaking denied," and other instruction labels). Simultaneously, the sending vehicle executes the corresponding coordinated action based on the semantic analysis results. For example, if overtaking is permitted, the sending vehicle accelerates to overtake; if overtaking is denied, the sending vehicle maintains a safe distance from the receiving vehicle.

[0083] As can be seen, by combining traffic participant information to achieve multi-vehicle bidirectional collaborative interaction, collaborative decision-making can be achieved in complex scenarios, reducing vehicle misjudgments. Furthermore, during multi-vehicle collaborative interaction, if multiple vehicles simultaneously transmit light signals, signal conflicts between vehicles can be resolved based on vehicle priority. Therefore, in scenarios without V2X (Vehicle to Everything) or cellular networks, multi-vehicle bidirectional collaborative interaction can be achieved by combining traffic participant information and light signal recognition.

[0084] In another example embodiment, Figure 3 As shown, a vehicle control method is provided, comprising the following steps:

[0085] After the target vehicle activates the assisted driving function or intelligent driving function, the target vehicle's onboard camera and lidar acquire, in real time, the target vehicle's light signals, location information, and environmental information. Environmental information includes road information, traffic participant information, and vehicle control information. Vehicle control information includes steering wheel angle, throttle position, and brake position. Augmented reality (AR) equipment is also used to assist in identifying obstructed signal lights. The target vehicle's light signals, location information, and environmental information can be obtained from headlight images captured by the onboard camera and / or headlight point cloud data generated by the lidar. As an example, the on / off status and on / off time of the target vehicle's headlights, taillights, and turn signals can be determined from the headlight images. Alternatively, the target vehicle's three-dimensional position can be determined from the lidar point cloud data. The light area or position can then be determined from the target vehicle's three-dimensional position based on the point cloud density of the light area. For lights in obstructed scenarios, AR equipment can be used to assist in identifying obstructed lights, for example, by generating a prompt image using an image completion algorithm. The light signal includes, but is not limited to, the type of vehicle lights, the on / off status and duration of the lights, the area of ​​the lights, or the location of the lights. Furthermore, the process of obtaining road information and traffic participant information from environmental information using vehicle light images captured by the vehicle camera and / or light point cloud data generated by the LiDAR can be found in some of the aforementioned embodiments and will not be further elaborated here.

[0086] The camera image from the vehicle's camera, the lidar point cloud from the lidar, and the prompt image from the AR device are fed into a convolutional neural network. Feature extraction is performed using attention mechanisms such as spatial attention, channel attention, and spatiotemporal attention to obtain light language features such as flashing frequency, light color, and the light language intent used to represent the trajectory. Semantic matching of the light language features is then performed based on the light language features, mapping them to semantic labels for light language such as emergency braking, left lane change, yielding, and high beam switching. Furthermore, feature fusion extraction is performed on the camera image from the vehicle's camera and the lidar point cloud from the lidar, and a vectorized map is generated based on the feature fusion extraction results.

[0087] The semantic labels of vehicle light language and vectorized maps are combined to form a universal regulatory control framework for the target vehicle; this universal regulatory control framework can be used for vehicle prediction, vehicle planning, and vehicle control of the target vehicle.

[0088] In summary, the present application provides a vehicle control method, which obtains multimodal data of a first target vehicle, then determines the light language characteristics of the first target vehicle based on the multimodal data, and then uses a pre-obtained light language coding table to map the light language characteristics into light language semantic labels, and controls the second target vehicle according to the light language semantic labels, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label of the first target vehicle based on the control result of the second target vehicle; wherein the multimodal data includes headlight image information, vehicle position information and environmental information, the environmental information includes road information and traffic participant information, and the light language characteristics include headlight flashing frequency, headlight color, headlight area and light language intention; the first target vehicle and the second target vehicle are located within a preset range. It can be seen that when performing vehicle control, this method can solve the problem that a single sensor is easily affected by environmental interference and makes misjudgments by acquiring multimodal data for light language signal recognition. At the same time, based on the light language characteristics of the first target vehicle, the driving behavior or driving intention of the first target vehicle can be predicted, so that a response strategy can be formed in advance for the second target vehicle within a preset range of the first target vehicle, thereby effectively avoiding sudden or high-risk situations and reducing the probability of traffic accidents. Moreover, this method can also realize two-way collaborative interaction of multiple vehicles by combining traffic participant information. When multiple vehicles send light language at the same time, the signal conflict problem can be resolved according to priority, and collaborative decision-making can be achieved in complex scenarios to reduce vehicle driving misjudgments. Therefore, the vehicle control method recorded in this method can not only realize vehicle collaborative interactive control in traffic scenarios with V2X or cellular networks, but also realize vehicle collaborative interactive control by controlling the lights between vehicles even in traffic scenarios without V2X or cellular networks.

[0089] In another exemplary embodiment of the present application, Figure 4As shown, this embodiment also provides a vehicle control system, including:

[0090] A data acquisition module 410 is configured to acquire multimodal data of a first target vehicle, the multimodal data including vehicle headlight image information, vehicle position information, and environmental information, the environmental information including road information and traffic participant information;

[0091] Light signal feature module 420, configured to determine light signal features of the first target vehicle based on the multimodal data, where the light signal features include light flashing frequency, light color, light area, and light signal intent;

[0092] The vehicle control module 430 is used to map the light language feature into a light language semantic label and control the second target vehicle according to the light language semantic label; wherein the first target vehicle and the second target vehicle are located within a preset range.

[0093] It can be understood that the vehicle control system provided in the above embodiment and the vehicle control method provided in the above embodiment belong to the same concept, wherein the specific manner in which the vehicle control method performs operations has been described in detail in the above embodiment and will not be repeated here. In actual applications, the vehicle control system provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the vehicle control system into different functional modules, and then implement all or part of the functions of the corresponding functional modules through the vehicle control method described in the above embodiment. For example, all or part of the functions of the data acquisition module 410 can be implemented or executed through the relevant step process of step S110, all or part of the functions of the light language feature module 420 can be implemented or executed through the relevant step process of step S120, and all or part of the functions of the vehicle control module 430 can be implemented or executed through the relevant step process of step S130. The specific implementation or execution process can be referred to the above embodiment and will not be described in detail here.

[0094] Therefore, the present application provides a vehicle control system, which obtains multimodal data of a first target vehicle through a data acquisition module, and then determines the light language features of the first target vehicle based on the multimodal data through a light language feature module, and then uses a pre-obtained light language coding table of the vehicle light through a vehicle control module to map the light language features into light language semantic labels, and controls the second target vehicle according to the light language semantic labels, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label of the first target vehicle based on the control result of the second target vehicle; wherein, the multimodal data includes headlight image information, vehicle position information and environmental information, the environmental information includes road information and traffic participant information, and the light language features include headlight flashing frequency, headlight color, headlight area and light language intention; the first target vehicle and the second target vehicle are located within a preset range. As can be seen, when controlling vehicles, this system acquires multimodal data to identify light signals, resolving the problem of single sensors being susceptible to environmental interference and resulting in misjudgments. Furthermore, based on the light signal characteristics of the first target vehicle, the system can predict the first target vehicle's driving behavior or intention, allowing a response strategy to be formed in advance for a second target vehicle within a preset range of the first target vehicle, effectively avoiding sudden or high-risk situations and reducing the probability of traffic accidents. Furthermore, this system can also combine information from traffic participants to achieve two-way collaborative interaction among multiple vehicles. When multiple vehicles simultaneously transmit light signals, signal conflicts can be resolved based on priority, enabling collaborative decision-making in complex scenarios and reducing vehicle driving misjudgments. Therefore, the vehicle control method described in this system can achieve collaborative interactive control of vehicles not only in traffic scenarios with V2X or cellular networks, but also in traffic scenarios without V2X or cellular networks, by controlling vehicle lights between vehicles.

[0095] In another exemplary embodiment of the present application, a computer device is further provided. The computer device may include a memory, a processor, and a computer program stored in the memory. The processor executes the computer program so that the computer device performs Figure 1 The steps of the vehicle control method. Figure 5 FIG1 shows a schematic diagram of the structure of a computer device 1000. Figure 5 As shown, the computer device 1000 includes: a processor 1010 , a memory 1020 , a power supply 1030 , a display unit 1040 , and an input unit 1060 .

[0096] The processor 1010 is the control center of the computer device 1000. It connects various components using various interfaces and lines, and performs various functions of the computer device 1000 by running or executing computer programs / instructions stored in the memory 1020, thereby monitoring the computer device 1000 as a whole. In the embodiment of the present application, when the processor 1010 calls the computer program stored in the memory 1020, it executes the following Figure 1 The steps of the vehicle control method are as follows. Optionally, processor 1010 may include one or more processing units; preferably, processor 1010 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and applications, and the modem processor primarily processes wireless communications. In some embodiments, the processor and memory may be implemented on a single chip; in some embodiments, they may also be implemented on separate chips.

[0097] The memory 1020 may primarily include a program storage area and a data storage area. The program storage area may store an operating system, various applications, and the like; the data storage area may store instruction data and the like generated based on the use of the computer device 1000. Furthermore, the memory 1020 may include a high-speed random access memory and a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device.

[0098] The computer device 1000 also includes a power supply 1030 (such as a battery) for supplying power to various components. The power supply can be logically connected to the processor 1010 through a power management system, thereby managing functions such as charging, discharging, and power consumption through the power management system.

[0099] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the computer device 1000. In the embodiment of the present application, it is mainly used to display the display interface of each application in the computer device 1000 and objects such as text and images displayed on the display interface. The display unit 1040 may include a display panel 1050. The display panel 1050 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0100] The input unit 1060 can be used to receive user input, such as numbers or characters. The input unit 1060 may include a touch panel 1070 and other input devices 1080. The touch panel 1070, also known as a touch screen, can receive user touch operations on or near it (e.g., operations performed by a user using a finger, stylus, or any other suitable object or accessory on or near the touch panel 1070).

[0101] Specifically, the touch panel 1070 can detect user touch operations and the signals generated by the touch operations, convert these signals into touch point coordinates, and transmit them to the processor 1010. Furthermore, the touch panel 1070 can receive and execute commands from the processor 1010. Furthermore, the touch panel 1070 can be implemented using various types, such as resistive, capacitive, infrared, and surface acoustic wave. Other input devices 1080 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick.

[0102] Of course, the touch panel 1070 can cover the display panel 1050. When the touch panel 1070 detects a touch operation on or near it, it transmits it to the processor 1010 to determine the type of touch event. Then the processor 1010 provides corresponding visual output on the display panel 1050 according to the type of touch event. Figure 5 In the embodiment, the touch panel 1070 and the display panel 1050 are two independent components to realize the input and output functions of the computer device 1000, but in some embodiments, the touch panel 1070 and the display panel 1050 can be integrated to realize the input and output functions of the computer device 1000.

[0103] The computer device 1000 may further include one or more sensors, such as a pressure sensor, a gravity acceleration sensor, a proximity light sensor, etc. Of course, according to the needs of specific applications, the computer device 1000 may also include other components such as a camera.

[0104] In another exemplary embodiment of the present application, a computer-readable storage medium is further provided, wherein the storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the above-mentioned device can execute the above-mentioned method in the present application. Figure 1 The steps of the vehicle control method.

[0105] It will be understood by those skilled in the art that Figure 5This is merely an example of a computer device and does not constitute a limitation on the device. The device may include more or fewer components than shown, or a combination of certain components, or different components. For ease of description, the above sections are divided into modules (or units) based on their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware components. For example, as some examples, the aforementioned computer device can be a vehicle, a vehicle computer, etc.

[0106] It will be understood by those skilled in the art that the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be applied to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0107] It is understood that although the terms "first," "second," etc. may be used to describe target vehicles in the embodiments of the present application, these terms are merely used to distinguish target vehicles from one another. For example, a first target vehicle may also be referred to as a second target vehicle, and similarly, a second target vehicle may also be referred to as a first target vehicle without departing from the scope of the embodiments of the present application.

[0108] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.

Claims

1. A vehicle control method, characterized in that: The method comprises the following steps: Acquire multimodal data of a first target vehicle, the multimodal data including vehicle light image information, vehicle position information, and environmental information, the environmental information including road information and traffic participant information; Determining light language features of the first target vehicle based on the multimodal data, the light language features including light flashing frequency, light color, light area, and light language intent; The light language features are mapped into light language semantic labels using a pre-obtained light language coding table, and a second target vehicle is controlled according to the light language semantic label so that the first target vehicle executes or terminates a driving intention corresponding to the light language semantic label based on a control result of the second target vehicle. The method includes: identifying optical signal parameters by a receiving vehicle, generating a feedback code or a warning code based on the identification result of the optical signal parameters, and executing the driving intention according to the feedback code, or terminating the driving intention according to the warning code. The identification result of the optical signal parameters includes a light signal capture result obtained by capturing the light signal from the optical signal parameters, a point cloud data verification result obtained by performing point cloud data verification on the optical signal parameters, and a multispectral fusion result obtained by performing multispectral fusion on the optical signal parameters. The optical signal parameters are obtained by mapping an optical code representing the driving intention of the sending vehicle according to the light language coding rules by a sending vehicle. The sending vehicle and the receiving vehicle are respectively determined based on the traffic participant information, the first target vehicle, and the second target vehicle. The first target vehicle and the second target vehicle are located within a preset range.

2. The vehicle control method according to claim 1, characterized in that: The process of controlling the second target vehicle according to the light language semantic label so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the control result of the second target vehicle includes: Controlling the lights of the second target vehicle according to the light language semantic label, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the light control result of the second target vehicle; or, The lights and speed of the second target vehicle are controlled according to the light language semantic label, so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the light and speed control results of the second target vehicle.

3. The vehicle control method according to claim 1, characterized in that: The process of controlling the second target vehicle according to the light language semantic label so that the first target vehicle executes or terminates the driving intention corresponding to the light language semantic label based on the control result of the second target vehicle further includes: Using the vehicle light language code table as an optical coding interaction protocol between a transmitting vehicle and a receiving vehicle; Furthermore, the feedback code or the warning code is parsed by the transmitting vehicle, and the driving intention is executed according to the parsing result of the feedback code, or the driving intention is terminated according to the parsing result of the warning code.

4. The vehicle control method according to claim 3, characterized in that: The process of identifying the optical signal parameters by the receiving vehicle and generating a feedback code or a warning code according to the identification result of the optical signal parameters includes: Capturing vehicle light signals from the optical signal parameters by the receiving vehicle, including extracting vehicle light flashing frequency and vehicle light color; and Performing point cloud data verification on the optical signal parameters by the receiving vehicle, including locating the sending vehicle by using the point cloud data and verifying the motion trajectory of the sending vehicle; Performing multi-spectral fusion from the optical signal parameters by the receiving vehicle, including collecting visible light and near-infrared band images, and filtering ambient light based on band differences; The feedback code or the warning code is generated according to the vehicle light signal capture result, the point cloud data verification result and the multi-spectral fusion result.

5. The vehicle control method according to claim 3 or 4, characterized in that: The method further includes: performing time window competition according to the priorities of the transmitting vehicle and the receiving vehicle, and transmitting a coding instruction according to the time window competition result; wherein the coding instruction includes: the optical signal parameter, the feedback code or the warning code; And / or, within a preset time window, the vehicle with the highest priority among the traffic participant information, the first target vehicle and the second target vehicle is used as the sending vehicle, and the remaining vehicles are used as receiving vehicles, and only the vehicle with the highest priority is allowed to send the optical signal parameters.

6. The vehicle control method according to claim 1, characterized in that: The process of obtaining the vehicle light language code table includes: Define the meaning of light language based on the type of light, the position of the vehicle and the number of times the light flashes; and Determining a vehicle light language encoding rule based on a pre-determined or real-time determined vehicle light language field, a vehicle light language field encoding length, and a vehicle light language field description, and obtaining a complete vehicle light language encoding based on the vehicle light language encoding rule; wherein the vehicle light language field includes a signal type, a flashing mode, a flashing number, a light source direction, a color identifier, a priority, and a preset reserved bit; The light language meaning definition, the vehicle light language field and the vehicle light language complete code are associated to form the vehicle light language code table.

7. The vehicle control method according to claim 1 or 6, characterized in that: The process of controlling the second target vehicle according to the light language semantic label includes: Displaying the light language semantic label through a pre-configured target display in the second target vehicle; wherein the target display is located in the field of view of the driver of the second target vehicle; Based on the light language semantic label and the occlusion information of the first target vehicle, an audio warning is issued to the second target vehicle; and when the distance between the first target vehicle and the second target vehicle is less than a preset distance, the second target vehicle is braked and the brake lights of the second target vehicle are controlled to flash; wherein the occlusion information is identified by an enhanced display device.

8. The vehicle control method according to claim 1, wherein: The process of determining the light language feature of the first target vehicle based on the multimodal data includes: Calculating a headlight brightness change frequency based on pixel changes in the headlight image information to obtain a flashing frequency of the first target vehicle; and Segmenting the headlight image information into red and yellow regions according to a preset color space, and obtaining the headlight color of the first target vehicle based on the red and yellow region segmentation results; wherein the headlight color corresponding to the red region segmentation result is red, and the headlight color corresponding to the yellow region segmentation result is yellow; and, Predicting the light language intention of the first target vehicle based on the vehicle light image information within a preset time period and the vehicle motion trajectory of the first target vehicle within a preset continuous time period; and When the headlights in the headlight image information are blocked, predicting the headlight positions of the first target vehicle using the point cloud data of the first target vehicle; and The three-dimensional data frame of the point cloud data is projected into a two-dimensional image to locate the headlight area of ​​the first target vehicle.

9. A vehicle control system, characterized in that: The system includes: a data acquisition module, configured to acquire multimodal data of a first target vehicle, the multimodal data including headlight image information, vehicle position information, and environmental information, the environmental information including road information and traffic participant information; a light language feature module, configured to determine light language features of the first target vehicle based on the multimodal data, the light language features including light flashing frequency, light color, light area, and light language intent; A vehicle control module, configured to map the light language features into light language semantic labels using a pre-obtained light language coding table, and control a second target vehicle according to the light language semantic labels so that the first target vehicle executes or terminates a driving intention corresponding to the light language semantic label based on a control result of the second target vehicle. The module comprises: identifying optical signal parameters by a receiving vehicle, generating a feedback code or a warning code based on the identification result of the optical signal parameters, and executing the driving intention according to the feedback code, or terminating the driving intention according to the warning code. The identification result of the optical signal parameters includes a light signal capture result obtained by capturing the light signal from the optical signal parameters, a point cloud data verification result obtained by performing point cloud data verification on the optical signal parameters, and a multispectral fusion result obtained by performing multispectral fusion on the optical signal parameters. The optical signal parameters are obtained by mapping an optical code representing the driving intention of the sending vehicle according to the light language coding rules by a sending vehicle. The sending vehicle and the receiving vehicle are respectively determined based on the traffic participant information, the first target vehicle, and the second target vehicle, and the first target vehicle and the second target vehicle are located within a preset range.

10. A computer device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the vehicle control method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Light interaction method, device and equipment and computer readable storage medium

    CN114954215A