Method for external address calling, vehicle, storage medium and program product
Patent Information
- Application Number
- CN202610988034.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-03
AI Technical Summary
[0004]本申请的主要目的在于提供一种发明名称对外喊话方法、车辆、存储介质及计算机程序产品,旨在解决传统车辆预警方法不够准确及时的技术问题
在本申请中,提出一种对外喊话方法,通过将实时采集的车辆周围的摄像头数据输入视觉语言模型,得到视觉语言模型对车辆环境的语义理解结果,基于该语义理解结果确定待喊话目标并生成警示词,进一步的,基于语义理解结果确定车辆对外喊话策略,最后,根据车辆对外喊话策略,采用车载外置扬声器阵列对待喊话目标播放警示词。由此,通过引入视觉语言模型对车辆环境进行理解,从而通过准确的车辆对外喊话策略、待喊话目标和警示词,实现准确和及时的车辆预警。
Smart Images

Figure CN122481609B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to methods for making announcements, vehicles, storage media, and computer program products. Background Technology
[0002] With the development of autonomous driving technology and Advanced Driver Assistance Systems (ADAS), vehicles' ability to perceive their surroundings has significantly improved. Traditional vehicle warning systems mainly rely on audible and visual alarms (such as the "beep" sound of reversing radar and flashing turn signals) to alert pedestrians or vehicles. However, traditional vehicle warning methods have the following limitations: 1. Limited warning information dimension. Traditional alarm sounds can only convey general signals such as "danger" or "please be careful," failing to convey specific intentions or instructions (for example, "Please stop, you are in the blind spot" and "Please pass quickly, I am reversing" have completely different meanings). 2. Lack of scene semantic understanding. Existing perception systems are mostly based on object detection, only recognizing "people" and "vehicles," making it difficult to understand complex dynamic interaction scenarios (such as long-tail scenarios like pedestrians waving, children playing hesitantly on the roadside, and non-motorized vehicles going the wrong way). 3. Passive and delayed interaction. Alarms are mostly triggered when the risk of collision is extremely high, lacking proactive and negotiated interactions based on intention prediction. 4. Rigid and fixed wording. Existing public address systems (if any) typically play pre-recorded fixed audio, which cannot generate targeted natural language prompts based on real-time traffic conditions, making them appear rigid and unintelligent.
[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide an invention called a method for making public announcements, a vehicle, a storage medium, and a computer program product, which aims to solve the technical problem that traditional vehicle warning methods are not accurate and timely enough.
[0005] To achieve the above objectives, this application proposes a method for making public announcements, the method comprising: Real-time data collection from cameras surrounding the vehicle; The camera data is input into a visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment; Based on the semantic understanding results, the target to be addressed is determined and a warning message is generated; Based on the semantic understanding results, a strategy for vehicles to communicate with the outside world is determined; According to the vehicle's external communication strategy, the warning message is played to the target to be addressed using an external vehicle speaker array.
[0006] In one embodiment, the step of determining the vehicle's external communication strategy based on the semantic understanding result includes: Real-time collection of driver status perception data; Based on the state perception data, the driving state of the vehicle driver is determined; Based on the semantic understanding results and the driving status, a vehicle external communication strategy is determined.
[0007] In one embodiment, the step of determining the vehicle's external communication strategy based on the semantic understanding result and the driving state includes: The scene category and hazard level of the vehicle environment are determined based on the semantic understanding results; Based on the scenario category and the danger level, a first strategy for the vehicle to make public announcements is determined; the first strategy is adjusted based on the driving status to obtain a second strategy for the vehicle to make public announcements. Alternatively, a third strategy for the vehicle to make a public announcement can be determined based on the scenario category, the danger level, and the driving status.
[0008] In one embodiment, the step of inputting the camera data into a visual language model to obtain the semantic understanding result of the visual language model of the vehicle environment includes: Based on the semantic understanding results, the scene complexity of the vehicle environment, the estimated collision time between the target to be addressed and the vehicle, and the confidence level of the semantic understanding results output by the visual language model are determined. The danger level of the vehicle environment is determined based on the estimated collision time, the confidence level, the scene complexity, and a first preset judgment rule corresponding to the estimated collision time, the confidence level, and the scene complexity. Alternatively, a risk score can be calculated based on the estimated collision time, the confidence level, the scene complexity, and their respective weighting coefficients; and the hazard level of the vehicle environment can be determined based on the risk score and a second preset judgment rule corresponding to the risk score.
[0009] In one embodiment, the step of playing the warning message to the target using an onboard external speaker array according to the vehicle's external announcement strategy includes: Collect camera data from targets that have been verbally addressed; The camera data of the target that has been called out is input into the visual language model, and the visual language model determines the behavioral changes of the target that has been called out. Based on the behavioral changes of the already addressed target, determine whether playing the warning message to the already addressed target is effective; If playing the warning message to the target that has already been addressed is ineffective, the step of playing the warning message to the target to be addressed using the vehicle's external loudspeaker array, according to the vehicle's external addressing strategy, is repeated.
[0010] In one embodiment, the step of determining whether playing the warning message to the already addressed target is effective based on changes in the target's behavior includes: Based on the reward signal, a new warning word is generated, and / or a new vehicle external communication strategy is determined; wherein the reward signal includes: a positive reward corresponding to the effectiveness of playing the warning word to the target that has been addressed, and a negative reward corresponding to the ineffectiveness of playing the warning word to the target that has been addressed.
[0011] In one embodiment, the step of obtaining the semantic understanding result of the vehicle environment by the visual language model further includes: Receive roadside perception data around the vehicle; The camera data and the roadside perception data are input into the visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment.
[0012] Furthermore, to achieve the above objectives, this application also proposes an external megaphone device, which includes: The first module is used to collect real-time camera data around the vehicle. The second module is used to input the camera data into the visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment. The third module is used to determine the target to be addressed and generate warning words based on the semantic understanding results. The fourth module is used to determine the vehicle's external communication strategy based on the semantic understanding results; The fifth module is used to play the warning words to the target to be addressed using an onboard external speaker array, according to the vehicle's external announcement strategy.
[0013] In addition, to achieve the above objectives, this application also proposes a vehicle comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the external broadcasting method described above.
[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the external broadcasting method described above.
[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the external broadcasting method described above.
[0016] One or more technical solutions proposed in this application have at least the following technical effects: This application proposes a method for issuing public address systems. It involves inputting real-time camera data collected around the vehicle into a visual language model to obtain the model's semantic understanding of the vehicle environment. Based on this semantic understanding, the target to be addressed is identified, and a warning message is generated. Further, based on the semantic understanding, a vehicle public address strategy is determined. Finally, according to the strategy, an onboard external speaker array is used to play the warning message to the target. Thus, by introducing a visual language model to understand the vehicle environment, accurate and timely vehicle warnings are achieved through a precise vehicle public address strategy, target, and warning message. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the first embodiment of the method for making public announcements according to this application. Figure 2 This is a flowchart illustrating the second embodiment of the method for making public announcements according to this application. Figure 3 This is an application diagram illustrating the second embodiment of the method for making public announcements according to this application; Figure 4 This is a flowchart illustrating the third embodiment of the method for making public announcements according to this application. Figure 5 This is a schematic diagram of the module structure of the external loudspeaker device according to an embodiment of this application; Figure 6 This is a schematic diagram of the device structure of the hardware operating environment involved in the external broadcasting method in the embodiments of this application.
[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0023] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or vehicle capable of performing the above functions. The following description uses a vehicle as an example to illustrate this embodiment and the subsequent embodiments.
[0024] Based on this, the embodiments of this application provide a method for making public announcements, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the method for making public announcements according to this application.
[0025] In this embodiment, the method for making external announcements includes steps S10 to S50: Step S10: Collect real-time camera data around the vehicle; In one embodiment, data from multiple cameras surrounding the vehicle is acquired in real time. These include a front-view main camera with resolutions of 1920×1080 at 30fps and a field of view (FOV) of 120°, responsible for monitoring pedestrians and vehicles ahead; four surround-view fisheye cameras with resolutions of 1280×720 at 30fps, covering 360° blind spot monitoring; and a rear-view camera specifically for monitoring vehicles / pedestrians approaching from behind when reversing. Furthermore, the real-time acquired data from the multiple cameras surrounding the vehicle undergoes preprocessing, including distortion correction, image stitching to generate a bird's-eye view (BEV) image, and temporal frame caching (retaining the most recent 2 seconds of historical frames for motion trend analysis).
[0026] Step S20: Input the camera data into the visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment; In one embodiment, preprocessed camera data is input into a Visual Language Model (VLM), and scene understanding is performed using pre-designed prompt templates. This improves the coverage of long-tail scenarios for the public address system, leveraging the generalization ability of the VLM to effectively address long-tail scenarios that traditional rule-based algorithms struggle to cover (such as abnormal behavior under special weather conditions or non-standard traffic participants).
[0027] In one embodiment, the prompt template is as follows: "Analyze the following vehicle camera footage and identify if an emergency traffic scenario exists: 1. Are there pedestrians, cyclists, or other vehicles in the footage? 2. What is the motion state of the target object (stationary / moving / sudden acceleration)? 3. What is the relative distance and approach speed between the target and the vehicle? 4. Is there a potential collision risk (such as crossing the road, sudden appearance of a blind spot)? 5. Hazard level assessment: Low / Medium / High / Emergency. Please output the analysis results in structured JSON format." In one embodiment, the Visual Language Model (VLM) can be selected and optimized. The base model employs a lightweight multimodal large model (such as Qwen-VL, LLaVA-1.5, InternVL, etc.). During domain fine-tuning, in-vehicle scene datasets (such as BDD100K, nuScenes) are used for instruction tuning. For edge deployment, model quantization (INT8) and TensorRT acceleration are used to achieve real-time inference on the in-vehicle Orin chip, achieving inference latency of less than 200ms in practical applications.
[0028] In one feasible implementation, step S20 may include steps A1 to A2: Step A1: Receive roadside sensing data around the vehicle; Step A2: Input the camera data and roadside perception data into the visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment.
[0029] In one embodiment, preprocessed camera data and received roadside perception data around the vehicle are input into a Visual Language Model (VLM) for scene understanding using a pre-designed prompt template. The specific prompt template is similar to the one used in the previous section analyzing only camera data, and will not be described again here.
[0030] In one embodiment, the vehicle receives data from the roadside unit (RSU) via V2X communication (such as a PC5 interface). The roadside perception data may include: location, speed, acceleration, category (vehicle, pedestrian, non-motorized vehicle, etc.), size, direction, etc. for each target; traffic light status such as the current phase (red, green, yellow) and remaining time; traffic events such as construction, accidents, congestion, etc.; and road geometry information such as lane lines, stop lines, etc.
[0031] In one embodiment, the vehicle's own camera data and roadside perception data are aligned in time and space. During time synchronization, GPS time or Network Time Protocol (NTP) is used to align timestamps. During spatial coordinate transformation, the roadside perception data is transformed from the roadside sensor coordinate system to the vehicle coordinate system, or unified to a global coordinate system (such as UTM). Next, the roadside perception data is converted into a natural language description (text) to be input into a visual language model along with the images. For example, a text description might be generated: "A pedestrian is crossing the road 30 meters ahead of the vehicle. A truck is traveling at 60 km / h in the left lane. The traffic light is red with 10 seconds remaining." In one embodiment, the visual language model (VLM) can process both image and text input simultaneously. Images captured by the vehicle's camera and text descriptions converted from roadside perception data are used as input, prompting the VLM with questions such as, "Please describe the current driving scenario and point out potential risks." The VLM outputs a natural language description of the semantic understanding of the current environment, such as, "There is a pedestrian crossing the road ahead, but the traffic light is about to turn green. The pedestrian may not be able to cross in time; it is recommended to slow down and give way." Furthermore, by optimizing the prompt template by incorporating roadside global information, the VLM can understand the vehicle's global environment 1-3 seconds in advance and generate subsequent announcements, addressing the limitations of limited field of view and delayed warnings in traditional vehicle-mounted perception systems.
[0032] In this way, by integrating the vehicle-mounted VLM visual perception with V2X data from roadside cameras, traffic lights, and radar, the cockpit announcement system can not only perceive the vehicle's local field of vision, but also obtain the full field of vision of the roadside (such as blind spots at intersections and vehicles approaching from behind on curves), thus achieving beyond-line-of-sight announcement warnings.
[0033] Step S30: Based on the semantic understanding results, determine the target to be addressed and generate a warning word; In one embodiment, for complex scenarios where the vehicle is located, a lightweight text generation model is used to synthesize the announcement content in real time. The input is scene elements extracted by the visual language model (VLM) (such as target type, location, speed, environmental features, etc.), and the output is a natural language warning, such as: "There is a fast-approaching electric vehicle on the left ahead. Please be careful and avoid it!"
[0034] In one embodiment, key information is extracted from the semantic understanding results output by the Visual Language Model (VLM), including scene understanding, risk analysis, predictive inference, and action suggestions. For example, the semantic understanding results output by the VLM are structured data: {"Scene Understanding": "At an intersection of urban roads, a pedestrian is crossing the road, and an electric vehicle is traveling against the flow of traffic on the right"; "Risk Analysis": ["Pedestrian crossing the road", "Electric vehicle approaching against the flow of traffic"]; "Predictive Inference": "The pedestrian will reach the vehicle in 3 seconds, and the electric vehicle will meet the vehicle in 5 seconds"; "Action Suggestion": "Slow down and yield to the pedestrian, pay attention to the electric vehicle on the right"}. Based on the risk analysis, the targets requiring warnings can be identified: pedestrians and electric vehicles. Furthermore, targeted natural language warning words can be generated, which must be clear, concise, easy to understand, and consistent with the current traffic scenario. Different warning words are generated for each target. Simultaneously, considering the vehicle's own state (e.g., whether it is moving, its speed) and traffic rules, appropriate tone and content are generated. Specifically, for pedestrians: generate polite, clear, and non-starterable warning messages, such as: "Pedestrians, please note that a vehicle is passing by, please wait"; for electric vehicles: generate clear and unambiguous warning messages, such as: "Electric vehicle, you are going against traffic, please keep to the right."
[0035] Step S40: Based on the semantic understanding results, determine the vehicle's external communication strategy; Understandably, semantic understanding results can include the following structured semantic information: 1. Environmental target categories: pedestrians, non-motorized vehicles, motorized vehicles, obstacles, construction workers, people crossing the road, children on the roadside, pets, objects illegally approaching vehicles, etc.; 2. Target spatial attributes: target's position relative to vehicles, distance, speed, trajectory, and danger level; 3. Scene semantics: road scenes (intersections, ramps, residential areas, pedestrian crossings, narrow road encounters), environmental risk scenes (blind spots, backlighting, obstructed vision, pedestrians about to cross); 4. Target behavior semantics: whether the target is approaching a vehicle, whether it is encroaching on the vehicle's path, whether it is ignoring the vehicle, whether it is performing any illegal actions, and the target's attention state; 5. Environmental background semantics: surrounding background noise, density of the surrounding crowd, and whether it is a no-honking zone.
[0036] Based on the structured semantic information included in the above semantic understanding results, the shouting mode (single / loop, shouting duration, shouting interval), shouting start and stop constraints and priority strategies, and corresponding vocal prosody strategies (voice tone and pause rhythm) in the vehicle's external shouting strategy can be further determined, thereby determining the specific vehicle's external shouting strategy.
[0037] Step S50: According to the vehicle's external announcement strategy, a warning message is played to the target using the vehicle's external speaker array.
[0038] In one embodiment, the hardware configuration of the vehicle-mounted external speaker array is as follows: Parametric speakers are integrated into the front, rear, and left and right rearview mirrors, with a beamwidth of ±15° and an effective distance of 30m. The pointing angle of the speakers is calculated based on the position of the target to be addressed, output by the Visual Language Model (VLM) (image coordinates -> world coordinates transformation). The volume of the warning message is dynamically adjusted based on ambient noise (collected via the vehicle-mounted microphone) and the distance to the target, ensuring that the sound pressure level at the target location reaches 75-85dB. Furthermore, when multiple targets are present, a polling or zoned synchronous playback strategy is employed to play the respective warning message to each target.
[0039] This embodiment proposes a method for issuing public address systems. Real-time data from cameras surrounding the vehicle is input into a visual language model to obtain the model's semantic understanding of the vehicle environment. Based on this semantic understanding, the target to be addressed is identified, and a warning message is generated. Further, a vehicle public address strategy is determined based on the semantic understanding. Finally, according to the strategy, an onboard external speaker array is used to play the warning message to the target. Thus, by introducing a visual language model to understand the vehicle environment, accurate and timely vehicle warnings are achieved through a precise vehicle public address strategy, target, and warning message.
[0040] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 Step S40 may also include steps B1 to B3: Step B1: Collect real-time status perception data of the vehicle driver; In one embodiment, the driver's state perception data includes physiological state data, behavioral operation data, and cognitive state data. Physiological state data can be determined through eye tracking or gaze tracking, facial expression analysis, and physiological signals collected by physiological sensors located throughout the vehicle. Normal behavioral operation data such as steering wheel angle, brake / accelerator pedal depth, shifting frequency, and turn signal usage habits can be determined, as well as behavioral operation data such as sharp turns, frequent fine adjustments, and operation delays. Cognitive state data such as attention and workload can be inferred from the aforementioned data.
[0041] Step B2: Determine the driver's driving status based on the state perception data; In one embodiment, taking the driver's facial image acquired in real time via an in-vehicle camera as an example, the process begins with face detection and feature point localization: For each frame, a lightweight face detection model (such as MobileNet-SSD) is used to detect the face, and then a feature point localization model (such as Dlib or a lightweight keypoint detection model) is used to obtain key points such as the eyes, mouth, and head posture. Next, facial features and temporal features are calculated, including eye and mouth features, head posture features, etc., and temporal features including statistical features over a period of time (such as 60 seconds) (such as blinking frequency, eye closure time ratio, number of yawns, number of abnormal head postures, etc.). Then, driving state recognition is performed, which can be achieved using a rule-based thresholding method or a machine learning classifier (such as SVM, random forest, or lightweight neural network) for state classification. For example, the proportion of time the eyes are closed to a specific period of time can be used as a fatigue indicator. Fatigue can also be detected by combining the frequency of yawning and head posture (such as frequent nodding). Whether the driver is paying attention to the road can be determined by head posture and the direction of eye gaze. If the driver looks elsewhere for a long time (such as on a mobile phone, out the window, etc.), it is judged as distraction. This is how distraction detection is achieved.
[0042] Step B3: Based on the semantic understanding results and driving status, determine the vehicle's external communication strategy.
[0043] In one embodiment, the following decision logic is proposed: When the environmental risk is high based on the semantic understanding results, a warning should be given to external traffic participants regardless of the driver's condition. When the driver is in poor condition (e.g., distracted, fatigued) and there is a certain risk in the environment, the warning should be strengthened, and internal reminders to the driver may also be necessary. When the environmental risk is low and the driver is in normal condition, the warning can be omitted to avoid unnecessary interference. Based on this decision logic, environmental risk levels are defined and quantified based on the semantic understanding results, such as low (0), medium (1), and high (2), and driver condition levels are defined and quantified based on the driving condition, such as normal (0), mild distraction / fatigue (1), and severe distraction / fatigue (2). Furthermore, based on the defined and quantified environmental risk levels and driver condition levels, the following rule table 1 can be designed:
[0044] Table 1 In other words, the system extracts environmental risk levels and specific semantic information (e.g., a pedestrian is crossing the road) from the semantic understanding results of the vehicle environment obtained by the Visual Language Model (VLM). It also determines the driver's state level and specific state (e.g., the driver is looking at their phone). Based on a pre-defined rule table corresponding to the semantic understanding results and driving state, it determines the urgency level of the announcement (no announcement, mild, moderate, strong, urgent). Based on the semantic information and the determined urgency level, it determines the specific content of the announcement. Simultaneously, based on the driver's state, it decides whether an internal reminder to the driver is necessary.
[0045] In one feasible implementation, step B3 may include steps B31 to B33: Step B31: Determine the scene category and hazard level of the vehicle environment based on the semantic understanding results; Step B32: Determine the first strategy for vehicle-to-the-surface communication based on the scenario category and hazard level; adjust the first strategy according to the driving status to obtain the second strategy for vehicle-to-the-surface communication. Step B33, or, based on the scenario category, hazard level, and driving status, determine the third strategy for the vehicle to make an external announcement.
[0046] In one embodiment, another method is provided for determining a vehicle's external communication strategy based on semantic understanding results and driving status.
[0047] First, based on the semantic understanding results, the scene category and hazard level of the vehicle environment are determined. At this point, the following rule table 2 is designed:
[0048] Table 2 Then, based on the scenario category and danger level as shown in Table 2 above, the first strategy for the vehicle to communicate with others can be matched. Finally, the first strategy is adjusted according to the driving status to obtain the second strategy for the vehicle to communicate with others. For example, when the driving status is that the driver is distracted (such as looking down at a mobile phone), and the scenario category and danger level correspond to a high-risk scenario outside the vehicle, the priority of the communication is automatically increased, the volume is increased, and a multimodal warning is triggered; if the driving status is that the driver has already taken evasive action, the intensity of the communication is reduced or the communication is paused to avoid invalid interaction.
[0049] In addition to rule table 2 as shown above, rule tables based on scenario category, hazard level, and driving status can also be designed (not described in detail here), so that a third strategy for the vehicle to make external announcements can be determined according to the actual scenario category, hazard level, and driving status.
[0050] In one possible implementation, steps B21 to B23 may be included after step B2: Step B21: Determine the scene complexity of the vehicle environment, the estimated collision time between the target to be addressed and the vehicle, and the confidence level of the semantic understanding results output by the visual language model based on the semantic understanding results. Step B22: Determine the hazard level of the vehicle environment based on the estimated collision time, confidence level, scene complexity, and the first preset judgment rule corresponding to the estimated collision time, confidence level, and scene complexity. Step B23, or, calculate the risk score based on the estimated collision time, confidence level, scene complexity, and their respective weighting coefficients; determine the hazard level of the vehicle environment based on the risk score and the second preset judgment rule corresponding to the risk score.
[0051] Determining the scene complexity of a vehicle environment based on semantic understanding results refers to classifying or quantifying scene complexity based on information such as target type, number of targets, road type, distribution of traffic participants, occlusion, lighting conditions, and road obstacles identified in the semantic understanding results. For example, when there is only a single static target in the scene, no occlusion, and an open road environment, it is judged as low scene complexity; when there are multiple types of dynamic targets such as pedestrians, non-motorized vehicles, and motorized vehicles in the scene, occlusion exists, or it is located in complex road conditions such as intersections or congested sections, it is judged as high scene complexity.
[0052] The target to be addressed is a pedestrian, non-motorized vehicle, or other obstacle identified in the semantic understanding results that poses a collision risk. Combining the vehicle's current speed and trajectory with the location, speed, and direction of the target to be addressed, the estimated time to collision (TTC) between the vehicle and the target is calculated using distance and relative speed.
[0053] The confidence level of the semantic understanding results output by the visual language model is the classification or recognition confidence probability output by the visual language model itself, and its value is usually between 0 and 1. This value characterizes the reliability of the model's recognition and target determination results in the current scene. The higher the confidence level, the more reliable the semantic understanding results; conversely, there is a risk of recognition bias or misjudgment.
[0054] In one embodiment, step B22 can employ a rule-based matching method, directly matching the three parameters obtained in step B21 with preset rules to determine the hazard level. Specifically, a first preset judgment rule is pre-established and stored. This rule is a graded matching logic based on estimated collision time, confidence level, and scene complexity. For example, the following judgment logic can be set: when the estimated collision time is short, the confidence level is high, and the scene complexity is high, it is judged as an extremely high hazard level; when the estimated collision time is relatively short, the confidence level is moderate, and the scene complexity is moderate, it is judged as a high hazard level; when the estimated collision time is relatively long, the confidence level is relatively high, and the scene complexity is relatively low, it is judged as a medium hazard level; and when the estimated collision time is long, the confidence level is low, and the scene complexity is low, it is judged as a low hazard level. Specifically, by substituting the actual parameter values obtained in step B21 into the aforementioned first preset judgment rule for matching, the hazard level corresponding to the current vehicle environment can be output.
[0055] Step B23 is an alternative implementation to step B22. A quantitative risk score is obtained through weighted calculation, and the hazard level is then determined based on the score. Specifically, the following implementation is used: Corresponding weight coefficients are configured for the estimated collision time, confidence level, and scenario complexity. These weight coefficients can be pre-calibrated and stored according to actual driving scenarios, vehicle conditions, and safety requirements. The sum of all weight coefficients can be normalized to 1. According to a preset calculation formula, the three parameters are weighted and calculated with their corresponding weight coefficients to obtain the quantitative risk score. An example calculation formula is: Risk Score = Estimated Collision Time Parameter × Corresponding Weight Coefficient + Confidence Level Parameter × Corresponding Weight Coefficient + Scenario Complexity Parameter × Corresponding Weight Coefficient. Specifically, this calculation formula is: ; In this system, RiskScore represents the risk score, TTC represents the estimated time to collision, VLMconfidence represents the confidence level, Scenecomplexity represents the scene complexity, and α, β, and γ are weighting coefficients, determined through calibration using real vehicle data. A second pre-defined judgment rule is established, which establishes the correspondence between risk score intervals and hazard levels. For example: a risk score in the first high-score interval corresponds to an extremely high hazard level; a risk score in the second high-score interval corresponds to a high hazard level; a risk score in the medium-score interval corresponds to a medium hazard level; and a risk score in the low-score interval corresponds to a low hazard level. The system matches the calculated risk score with the interval belonging to the second pre-defined judgment rule to ultimately determine the hazard level of the vehicle environment.
[0056] In this application, reference is made to Figure 3By inputting real-time camera data around the vehicle into a visual language model, the semantic understanding of the vehicle environment by the visual language model is obtained. Based on the semantic understanding, the target to be addressed is determined and a warning word is generated. At the same time, the driving state of the vehicle driver is determined based on real-time collected state perception data of the vehicle driver. Furthermore, based on the semantic understanding and driving state, the vehicle's external communication strategy is determined. Finally, according to the vehicle's external communication strategy, the warning word is played to the target using an onboard external speaker array.
[0057] Compared to the limitations of existing technologies with their single-dimensional warning information, this method generates accurate warning words based on the semantic understanding of the vehicle environment using a visual language model, thereby conveying specific intentions or instructions. Furthermore, compared to the limitations of existing technologies lacking scene semantic understanding, this method uses a visual language model to semantically understand the vehicle environment and determines the driver's driving state based on the driver's state perception data, thus fully understanding the internal and external environment of the vehicle. Compared to the limitations of existing technologies with passive and delayed interaction, this method determines the vehicle's external communication strategy based on semantic understanding results and driving state, and proactively and promptly plays warning words to the target based on this strategy. Finally, compared to the limitations of existing technologies with fixed and rigid wording, this method generates targeted warning words adapted to the vehicle environment based on semantic understanding results. Therefore, this application breaks through the traditional warning paradigm based on a single signal and static rules, constructing an intelligent warning method with dynamic scene semantic understanding, intention perception, and natural interaction capabilities.
[0058] In the first application scenario, a pedestrian warning is implemented during reversing: a vehicle is reversing in a parking lot, and a child is chasing a ball towards the rear of the car in the blind spot. The rear-view camera captures the image, and the Visual Language Model (VLM) receives the image and the "reversing" status indicator. The VLM identifies "child," "running," "ball," and "high risk," and infers that "the child may not be paying attention to the vehicle," classifying the risk level as "high." A text message is generated: "Kid, there's a car behind you! Stop immediately! Parents, please supervise your child!" A rapid but clear voice message is generated and played back via speakers on both sides of the rear bumper. Simultaneously, the brake lights flash at a high frequency. The effect: the child hears the specific instruction to stop, parents intervene, and an accident is avoided.
[0059] In the second application scenario, a right-turn blind spot warning is implemented at an intersection: a vehicle is turning right at an intersection without traffic lights, and an electric scooter is approaching straight ahead in the A-pillar blind spot. The side-view camera captures the image, and the Visual Language Model (VLM) analyzes it, identifying "electric scooter," "straight ahead," "close distance," and "driver's view obstructed." A text message is generated: "Electric scooter driver on the right, please be aware, I am turning right, please slow down!" This message is broadcast directionally to the right front via a speaker near the right front wheel arch. Effect: The electric scooter driver is informed of the vehicle's intention in advance and slows down to avoid it.
[0060] In the third application scenario, a highway breakdown warning is implemented: a vehicle stops on the highway due to a breakdown, activates its hazard lights, and a vehicle behind is approaching at high speed with signs of distracted driving. A rear-view telephoto camera captures the traffic flow behind, and a Visual Language Model (VLM) identifies "lane departure," "rapidly decreasing distance," and "suspected driver looking down" from the vehicle behind. A high-intensity warning sound and voice message are generated: "Attention vehicles behind! A breakdown has occurred ahead; please change lanes immediately!" This message is broadcast at maximum volume through a loudspeaker at the rear of the vehicle, accompanied by flashing taillights.
[0061] In this embodiment, compared to the limitation of existing technologies with their single-dimensional warning information, accurate warning words are generated through semantic understanding of the vehicle environment using a visual language model, thereby conveying specific intentions or instructions. Compared to the limitation of existing technologies lacking scene semantic understanding, the vehicle environment is semantically understood through a visual language model, and the driver's driving state is determined through the driver's state perception data, thus fully understanding the internal and external environment of the vehicle. Compared to the limitation of existing technologies with passive and delayed interaction, the vehicle's external communication strategy is determined based on semantic understanding results and driving state, and warning words are played proactively and promptly to the target of the communication according to the strategy. Compared to the limitation of existing technologies with fixed and rigid wording, targeted warning words adapted to the vehicle environment are generated through semantic understanding results. Therefore, this application breaks through the traditional warning paradigm based on a single signal and static rules, constructing an intelligent warning method with dynamic scene semantic understanding, intention perception, and natural interaction capabilities.
[0062] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 After step S50, steps C1 to C4 are also included: Step C1: Collect camera data from the target who has been spoken to; Step C2: Input the camera data of the target who has been called into the visual language model, and the visual language model determines the behavioral changes of the target who has been called; Step C3: Based on the changes in the behavior of the target who has been addressed, determine whether playing the warning message to the target is effective; Step C4: If playing the warning message to the target has no effect, repeat step S50.
[0063] In one embodiment, a vehicle-mounted camera monitors the behavioral changes of the target being addressed in real time, and the addressing strategy is dynamically adjusted based on feedback, forming a closed loop of execution-feedback-adjustment, breaking through the linear logic of existing solutions where the message ends after it finishes playing. For example, a visual language model monitors the behavioral changes of the target being addressed in real time to determine whether the message is effective (e.g., "pedestrian stop" = effective, "pedestrian continues crossing" = ineffective). If the message is ineffective, a secondary message strategy is immediately triggered (e.g., increasing volume, changing the message content, or enabling multimodal alerts); if the message is effective, the message intensity is immediately reduced or the message is paused to avoid invalid interaction.
[0064] In one embodiment, after playing a warning message to the target via an onboard external speaker array according to the vehicle's external announcement strategy, the system further executes closed-loop verification of the announcement effect and adaptive re-announcement logic. First, camera data of the announced target is collected. Immediately after playing the warning message to the target, real-time image data and video stream data, including the announced target, are collected using the vehicle's surround-view cameras, front-view cameras, side-view cameras, or dedicated target detection cameras. The collected camera data includes at least the announced target's location information, posture information, motion trajectory information, and surrounding environment information, thus providing complete and reliable visual input for subsequent behavior recognition. The data collection process can share image data sources with the vehicle's existing perception system, such as collecting data from cameras around the vehicle, without requiring additional hardware, ensuring system compatibility and cost-effectiveness.
[0065] Then, the camera data is input into the visual language model to determine the behavioral changes of the target who has been warned. The collected images and video data of the warned target are input into the pre-trained visual language model. This visual language model has multimodal fusion perception capabilities and can perform semantic understanding and classification of the target's behavior. The model extracts and analyzes the temporal and spatial features of consecutive frames to determine whether the warned target has produced corresponding behavioral changes after receiving the warning words, such as: whether it stops the violation, avoids the vehicle, accelerates away, ignores the warning and continues the original behavior, or exhibits emotionally confrontational behavior. Through behavioral semantic annotation and temporal comparison, the model outputs the identification result of whether the target's behavior has a positive response, no response, or negative response.
[0066] Next, the effectiveness of the warning message is determined based on behavioral changes. Based on the target's behavioral changes output by the visual language model, the effectiveness of the warning is evaluated according to preset validity rules: if the target actively avoids vehicles, stops dangerous behavior, obeys traffic rules, or adjusts their behavior according to the warning diagram after receiving the warning message, the warning is deemed effective; if the target ignores the warning, does not change their behavior, or continues to engage in dangerous behavior or obstructs traffic, the warning is deemed invalid. The validity determination logic can be adaptively adjusted according to different application scenarios, such as road avoidance scenarios, lane violation scenarios, and pedestrian crossing scenarios, with corresponding behavioral judgment thresholds and standards set for each.
[0067] Finally, if the verbal warning is ineffective, the warning operation is re-executed. When it is determined that playing the warning message to the already warned target has not achieved the expected effect, a loop execution mechanism is automatically triggered, re-invoking the vehicle's external verbal warning strategy and playing the warning message again to the already warned target (i.e., the current target to be warned) through the vehicle's external speaker array. During the re-warning process, the verbal warning strategy can be further optimized by combining the target's behavioral characteristics, such as adjusting the volume, frequency, and duration of the warning message, or changing the warning content, to improve the warning effect of the second warning until the target produces an effective behavioral response. This forms a closed-loop control logic of "verification-perception-recognition-determination re-verification," improving the reliability and actual intervention effect of the vehicle's intelligent external verbal warning.
[0068] In one possible implementation, step D is included after step C3: Step D: Based on the reward signal, generate a new warning word, and / or determine a new vehicle external communication strategy; wherein the reward signal includes: a positive reward corresponding to the effectiveness of playing the warning word to the target that has been addressed, and a negative reward corresponding to the ineffectiveness of playing the warning word to the target that has been addressed.
[0069] In one embodiment, a reinforcement learning (RL) algorithm is introduced, using the behavioral feedback of the target being addressed as a reward signal. This allows the system to autonomously optimize the content, speed, volume, and timing of the announcements, achieving self-learning and self-optimization of the announcement strategy. The reward signal is defined as follows: a positive reward is given when the target takes an evasive action (such as stopping, slowing down, or changing lanes); a negative reward is given when the target does not respond or takes a dangerous action (such as accelerating across). Based on the reward signal, the system autonomously adjusts the announcement parameters (e.g., retaining the current announcement strategy when there is a positive reward; adjusting the speed, volume, or regenerating the announcement content when there is a negative reward).
[0070] In one embodiment, firstly, corresponding reward signals are constructed and generated. A reward mechanism for evaluating the effectiveness of the announcements is pre-constructed, converting the results of the announcement effectiveness determination into quantifiable reward signals to guide subsequent iterations and updates of the announcement strategy. Specifically, for positive reward signals: when it is determined that playing the warning words to the already announced target is effective, i.e., the target responds with positive behaviors such as avoiding the warning, ceasing dangerous behavior, or following traffic guidance after the warning, a positive reward signal is generated and output. This positive reward indicates that the currently used warning word content and vehicle announcement strategy have a good intervention effect and are worth retaining or strengthening. For negative reward signals: when it is determined that playing the warning words to the already announced target is ineffective, i.e., the target ignores the warning, does not change dangerous behavior, or continues to obstruct vehicle passage, a negative reward signal is generated and output. This negative reward indicates that the current warning word or announcement strategy has failed to achieve the expected warning purpose and needs to be adjusted and optimized. It is understood that the above reward signals can be represented in a numerical form, for example, assigning a positive reward value to effective announcements and a negative reward value to ineffective announcements, thereby providing a quantitative basis for subsequent strategy optimization.
[0071] Then, new warning words and / or new vehicle announcement strategies are generated based on the reward signals. These reward signals are input into a pre-built reinforcement learning optimization model or policy iteration network. The model, with the optimization objective of maximizing the long-term effectiveness of the announcements, adaptively updates the warning words and vehicle announcement strategies. Specifically, generating new warning words involves selecting more targeted, deterrent, or guiding text content from a pre-set warning word library when a negative reward signal is received, or dynamically generating new warning statements using a natural language generation model. This includes adjusting the tone of the warning, simplifying the instructions, and highlighting the dangerous consequences to enhance the persuasiveness of the warning. Determining new vehicle announcement strategies involves optimizing and adjusting the announcement strategy based on the reward signal feedback. This includes, but is not limited to, adjusting the speaker array's sound angle and directionality, increasing the playback volume, increasing the playback frequency, changing the loop mode, switching between single-speaker / multi-speaker collaborative sound modes, or adaptively matching the announcement level based on target distance and scene type. Joint optimization involves simultaneously generating new warning words and determining new announcement strategies to achieve synergistic optimization of content and form, thereby improving the warning success rate in subsequent announcements.
[0072] Through the closed-loop mechanism of "effect judgment - reward feedback - strategy iteration", the vehicle's external announcements have self-learning and adaptive capabilities. In the process of continuous interaction, the warning words and announcement strategies are continuously optimized to adapt to different target types and complex road scenarios, thereby improving the flexibility and effectiveness of intelligent announcements.
[0073] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the method of external communication of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0074] This application also provides a device for making public announcements; please refer to [reference needed]. Figure 5 The external communication device includes: The first module 10 is used to collect real-time camera data around the vehicle; The second module 20 is used to input the camera data into the visual language model to obtain the semantic understanding result of the visual language model of the vehicle environment; The third module 30 is used to determine the target to be addressed and generate warning words based on the semantic understanding results. The fourth module 40 is used to determine the vehicle's external communication strategy based on the semantic understanding results; The fifth module 50 is used to play the warning words to the target to be addressed using an on-board external speaker array, according to the vehicle's external announcement strategy.
[0075] In one embodiment, the fourth module 40 is further configured to: Real-time collection of driver status perception data; Based on the state perception data, the driving state of the vehicle driver is determined; Based on the semantic understanding results and the driving status, a vehicle external communication strategy is determined.
[0076] In one embodiment, the fourth module 40 is further configured to: The scene category and hazard level of the vehicle environment are determined based on the semantic understanding results; Based on the scenario category and the danger level, a first strategy for the vehicle to make public announcements is determined; the first strategy is adjusted based on the driving status to obtain a second strategy for the vehicle to make public announcements. Alternatively, a third strategy for the vehicle to make a public announcement can be determined based on the scenario category, the danger level, and the driving status.
[0077] In one embodiment, the fourth module 40 is further configured to: After the step of inputting the camera data into the visual language model to obtain the semantic understanding result of the visual language model of the vehicle environment: Based on the semantic understanding results, the scene complexity of the vehicle environment, the estimated collision time between the target to be addressed and the vehicle, and the confidence level of the semantic understanding results output by the visual language model are determined. The danger level of the vehicle environment is determined based on the estimated collision time, the confidence level, the scene complexity, and a first preset judgment rule corresponding to the estimated collision time, the confidence level, and the scene complexity. Alternatively, a risk score can be calculated based on the estimated collision time, the confidence level, the scene complexity, and their respective weighting coefficients; and the hazard level of the vehicle environment can be determined based on the risk score and a second preset judgment rule corresponding to the risk score.
[0078] In one embodiment, the external communication device further includes a sixth module for: After the step of playing the warning message to the target using an onboard external speaker array according to the vehicle's external announcement strategy: Collect camera data from targets that have been verbally addressed; The camera data of the target that has been called out is input into the visual language model, and the visual language model determines the behavioral changes of the target that has been called out. Based on the behavioral changes of the already addressed target, determine whether playing the warning message to the already addressed target is effective; If playing the warning message to the target that has already been addressed is ineffective, the step of playing the warning message to the target to be addressed using the vehicle's external loudspeaker array, according to the vehicle's external addressing strategy, is repeated.
[0079] In one embodiment, the sixth module is further configured to: The step of determining whether playing the warning message to the already addressed target is effective based on changes in the target's behavior includes: Based on the reward signal, a new warning word is generated, and / or a new vehicle external communication strategy is determined; wherein the reward signal includes: a positive reward corresponding to the effectiveness of playing the warning word to the target that has been addressed, and a negative reward corresponding to the ineffectiveness of playing the warning word to the target that has been addressed.
[0080] The external loudspeaker device provided in this application, employing the external loudspeaker method described in the above embodiments, can solve the technical problem that traditional vehicle warning methods are not accurate or timely enough. Compared with the prior art, the beneficial effects of the external loudspeaker device provided in this application are the same as those of the external loudspeaker method provided in the above embodiments, and other technical features in the external loudspeaker device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0081] This application provides a vehicle, the vehicle including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the external announcement method in Embodiment 1 above.
[0082] The following is for reference. Figure 6 It shows a structural schematic diagram of a vehicle suitable for implementing the embodiments of this application. Figure 6 The vehicle shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.
[0083] like Figure 6 As shown, the vehicle may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for vehicle operation. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the vehicle to communicate wirelessly or wiredly with other devices to exchange data. Although the diagram shows vehicles with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.
[0084] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0085] The vehicle provided in this application, employing the external announcement method described in the above embodiments, can solve the technical problem that traditional vehicle warning methods are not accurate and timely enough. Compared with the prior art, the beneficial effects of the vehicle provided in this application are the same as those of the external announcement method provided in the above embodiments, and other technical features of the vehicle are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0086] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0087] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0088] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the external broadcasting method in the above embodiments.
[0089] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0090] The aforementioned computer-readable storage medium may be included in the vehicle or may exist independently and not installed in the vehicle.
[0091] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a vehicle, cause the vehicle to: collect real-time camera data around the vehicle; input the camera data into a visual language model to obtain the semantic understanding result of the visual language model of the vehicle environment; determine the target to be addressed and generate a warning word based on the semantic understanding result; determine the vehicle's external communication strategy based on the semantic understanding result; and play the warning word to the target using an onboard external speaker array according to the vehicle's external communication strategy.
[0092] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0094] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0095] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described external announcement method, which can solve the technical problem that traditional vehicle warning methods are not accurate and timely enough. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the external announcement method provided in the above embodiments, and will not be repeated here.
[0096] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for making an external announcement.
[0097] The computer program product provided in this application can solve the technical problem that traditional vehicle warning methods are not accurate and timely enough. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the external announcement method provided in the above embodiments, and will not be repeated here.
[0098] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for making public announcements, characterized in that, The methods of making public announcements include: Real-time data collection from cameras surrounding the vehicle; The camera data is input into a visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment; Based on the semantic understanding results, identify the targets outside the vehicle that need to be warned and generate warning words; Based on the semantic understanding results, a vehicle external communication strategy is determined; the vehicle external communication strategy includes communication mode, communication start and stop constraints and priority strategy, and corresponding vocal prosody strategy. According to the vehicle's external communication strategy, the warning message is played to the target to be addressed using an onboard external speaker array. Following the step of playing the warning message to the target using an onboard external speaker array according to the vehicle's external announcement strategy, the following steps are included: Collect camera data from targets that have been verbally addressed; The camera data of the target who has been called out is input into the visual language model, and the visual language model determines the behavioral changes of the target who has been called out. Based on the behavioral changes of the already addressed target, determine whether playing the warning message to the already addressed target is effective; If playing the warning message to the target that has already been addressed is ineffective, the step of playing the warning message to the target to be addressed using the vehicle's external loudspeaker array, according to the vehicle's external addressing strategy, is repeated.
2. The method for making public announcements as described in claim 1, characterized in that, The step of determining the vehicle's external communication strategy based on the semantic understanding results includes: Real-time collection of driver status perception data; Based on the state perception data, the driving state of the vehicle driver is determined; Based on the semantic understanding results and the driving status, a vehicle external communication strategy is determined.
3. The method for making public announcements as described in claim 2, characterized in that, The step of determining the vehicle's external communication strategy based on the semantic understanding result and the driving state includes: The scene category and hazard level of the vehicle environment are determined based on the semantic understanding results; Based on the scenario category and the danger level, a first strategy for the vehicle to make public announcements is determined; the first strategy is adjusted based on the driving status to obtain a second strategy for the vehicle to make public announcements. Alternatively, a third strategy for the vehicle to make a public announcement can be determined based on the scenario category, the danger level, and the driving status.
4. The method for making public announcements as described in claim 1, characterized in that, The step of inputting the camera data into the visual language model to obtain the semantic understanding result of the visual language model of the vehicle environment includes: Based on the semantic understanding results, the scene complexity of the vehicle environment, the estimated collision time between the target to be addressed and the vehicle, and the confidence level of the semantic understanding results output by the visual language model are determined. The danger level of the vehicle environment is determined based on the estimated collision time, the confidence level, the scene complexity, and a first preset judgment rule corresponding to the estimated collision time, the confidence level, and the scene complexity. Alternatively, a risk score can be calculated based on the estimated collision time, the confidence level, the scene complexity, and their respective weighting coefficients; and the hazard level of the vehicle environment can be determined based on the risk score and a second preset judgment rule corresponding to the risk score.
5. The method for making public announcements as described in claim 1, characterized in that, The step of determining whether playing the warning message to the already addressed target is effective based on changes in the target's behavior includes: Based on the reward signal, a new warning word is generated, and / or a new vehicle external communication strategy is determined; wherein the reward signal includes: a positive reward corresponding to the effectiveness of playing the warning word to the target that has been addressed, and a negative reward corresponding to the ineffectiveness of playing the warning word to the target that has been addressed.
6. The method for making public announcements as described in claim 1, characterized in that, The step of obtaining the semantic understanding result of the vehicle environment by the visual language model further includes: Receive roadside sensing data around the vehicle; The camera data and the roadside perception data are input into the visual language model to obtain the semantic understanding results of the visual language model of the vehicle environment.
7. A vehicle, characterized in that, The vehicle includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method for making public announcements as described in any one of claims 1 to 6.
8. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the external broadcasting method as described in any one of claims 1 to 6.
9. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method for making public announcements as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic driving risk identification model training method, risk assessment method and system
CN121434774A
Vehicle risk early warning method, vehicle, storage medium and computer program product
CN121686410A
Vehicle and control method thereof
US20180118107A1