A method and system for active response of a vehicle machine in a low-speed scenario of a vehicle

By combining lane-level positioning and pedal signal recognition models, low-speed root cause targets are identified and segmented, and active response rules are configured, solving the customized needs of vehicle cameras in low-speed scenarios and realizing intelligent and convenient operation of the vehicle system in low-speed scenarios.

CN121708560BActive Publication Date: 2026-07-03DONGGUAN FUTURE IMAGING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DONGGUAN FUTURE IMAGING TECH CO LTD
Filing Date
2025-12-19
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing vehicle cameras are difficult to adapt to the differentiated needs of specific scenarios in low-speed environments, leading to increased driver workload and low efficiency. The image recognition function is disconnected from the vehicle control function, making it impossible to provide customized solutions for low-speed scenarios.

Method used

By combining lane-level positioning information and pedal signals with existing recognition models, low-speed root cause targets are identified through statistical and segmentation models, and an active response rule base is configured to achieve active response from the vehicle system.

Benefits of technology

It improves the intelligence and convenience of the vehicle's infotainment system in low-speed scenarios, reduces the driver's workload, realizes the synergy between image recognition and vehicle control, and enhances operational efficiency and safety in low-speed scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708560B_ABST
    Figure CN121708560B_ABST
Patent Text Reader

Abstract

This application relates to the field of vehicle infotainment system (VIS) management technology, specifically a method and system for proactive VIS response in low-speed vehicle scenarios. This application uses existing scene recognition models to identify scenes and further filters target scene labels by combining lane-level positioning information and pedal signals. The target scene labels and the original image are input into a segmentation model to extract low-speed root cause targets that are strongly correlated with the target scene labels. Based on the low-speed root cause targets, proactive response rules are extracted from a rule base, and the VIS proactive response is executed. This application combines scene recognition and low-speed root cause target extraction to configure multiple proactive rules, enriching the proactive response function of the VIS in low-speed scenarios and improving the intelligence and convenience of the VIS.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of vehicle infotainment management technology, specifically a vehicle infotainment active response method and system for low-speed vehicle scenarios. Background Technology

[0002] With the deep integration of the automotive industry with artificial intelligence and the Internet of Things (IoT) technologies, intelligent connected vehicles have become a core direction for the transformation and upgrading of the global automotive industry. Among these, in-vehicle cameras, acting as the "eyes" of vehicle environmental perception, provide drivers with rich environmental information through image acquisition and processing technologies. Their applications have expanded from early reversing cameras and dashcams to advanced driver assistance systems (ADAS) such as lane departure warning, adaptive cruise control, and traffic sign recognition, and have even become an indispensable perception unit in autonomous driving systems. For example, the invention patent application CN202310823831 proposes an image optimization method for autonomous driving systems that can improve the efficiency of in-vehicle image debugging and meet the debugging needs of different autonomous driving scenarios.

[0003] Under the current technological framework, the application logic of vehicle cameras largely revolves around "high-speed dynamic driving safety." For example, when a vehicle is traveling at speeds of 30-60 km / h or higher, the camera can capture dynamic targets such as vehicles, pedestrians, and lane markings in real time. Combined with algorithm models, it can achieve functions such as collision risk prediction and lane keeping assist, effectively reducing safety hazards at high speeds. The formation of this technological path is closely related to the early focus of intelligent driving technology on the core objective of "reducing high-speed accidents," and has also led to the optimization of existing vehicle camera hardware parameters (such as frame rate and dynamic range) and software algorithms (such as moving target tracking and real-time obstacle detection) for high-speed dynamic scenarios.

[0004] However, as users' demands for intelligent automotive experiences continue to upgrade, the vehicle's intelligent processing capabilities in low-speed scenarios are gradually becoming a technological pain point. In these scenarios, the driver's core needs have shifted from "safety and risk avoidance" to "efficient operation" and "convenient interaction," but existing in-vehicle camera technology has revealed significant limitations when adapting to such scenarios.

[0005] Meanwhile, the environmental characteristics of low-speed scenarios differ fundamentally from those of high-speed scenarios, and the image acquisition parameters of existing vehicle cameras (such as focal length, resolution, and white balance) are mostly preset based on high-speed scenarios, making it difficult to adapt to the needs of low-speed scenarios. Therefore, existing vehicle camera technology has a rather broad coverage of low-speed scenarios and fails to provide customized solutions for the differentiated needs of specific scenarios. Summary of the Invention

[0006] In view of this, the purpose of this application is to provide a vehicle-mounted active response method and system for low-speed vehicle scenarios, so as to solve the problems in the background art.

[0007] To achieve the above objectives, this application adopts the following technical solution:

[0008] This application discloses a vehicle-mounted active response method for low-speed vehicle scenarios, comprising the following steps:

[0009] When the target vehicle is detected to be in a low-speed driving mode and the duration of the low-speed driving mode exceeds a preset duration threshold, the vehicle camera is controlled to capture scene images; and the lane-level positioning information and pedal signals of the target vehicle are obtained.

[0010] The scene image is classified based on a pre-built recognition model to obtain multiple scene labels and confidence scores for the scene image.

[0011] The first matching degree between the lane-level positioning and multiple scene labels and the second matching degree between the pedal signal and multiple scene labels are calculated based on a pre-built statistical model; and the target scene label is determined based on the confidence degree of multiple scene labels, the first matching degree of multiple scene labels and the second matching degree of multiple scene labels.

[0012] The target scene label and the scene image are input into a pre-built segmentation model to obtain the segmentation result of one or more low-speed root cause targets in the scene image that match the scene label. The segmentation result includes the category label of the low-speed root cause target and the segmentation mask of the low-speed root cause target.

[0013] Based on the type label of the low-speed root cause target, the current response rule is determined according to the pre-configured response rule library, and the vehicle system actively responds based on the current response rule. The response rule library includes multiple active response rules, which are one or a combination of the following: opening the 360-degree surround view, displaying the segmentation mask image of the low-speed root cause target, and popping out the charging port cover.

[0014] This application also provides an in-vehicle active response system for low-speed vehicle scenarios, including:

[0015] The acquisition module is used to control the vehicle-mounted camera to acquire scene images when it detects that the target vehicle is in a low-speed driving mode and the duration of the low-speed driving mode is greater than a preset duration threshold; and to acquire the lane-level positioning information and pedal signal of the target vehicle.

[0016] The scene classification module is used to classify the scene image based on a pre-built recognition model to obtain multiple scene labels and confidence scores of the scene image.

[0017] The scene determination module is used to calculate the first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels based on a pre-built statistical model; and to determine the target scene label based on the confidence degree of multiple scene labels, the first matching degree of multiple scene labels, and the second matching degree of multiple scene labels.

[0018] The segmentation module is used to input the target scene label and the scene image into a pre-built segmentation model to obtain the segmentation result of one or more low-speed root cause targets in the scene image that match the scene label. The segmentation result includes the category label of the low-speed root cause target and the segmentation mask of the low-speed root cause target.

[0019] An active response module is used to determine the current response rule based on the type label of the low-speed root cause target, the pre-configured response rule library, and to execute the vehicle system active response based on the current response rule. The response rule library includes multiple active response rules, which are one or a combination of the following: opening the 360-degree surround view, displaying the segmentation mask image of the low-speed root cause target, and popping out the charging port cover.

[0020] This application also provides an electronic device, including: a processor and a memory;

[0021] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to cause the electronic device to perform the methods described above.

[0022] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.

[0023] The beneficial effects of this application are as follows: This application provides a vehicle-mounted infotainment system (V2X) active response method and system for low-speed vehicle scenarios. This application uses existing scene recognition models to identify scenes and further filters target scene labels by combining lane-level positioning information and pedal signals. The target scene labels and the original image are input into a segmentation model to extract low-speed root cause targets that are strongly correlated with the target scene labels. Active response rules are extracted from a rule base based on the low-speed root cause targets, and the V2X active response is executed. This application combines scene recognition and low-speed root cause target extraction to configure multiple active rules, enriching the active response function of the V2X in low-speed scenarios and improving the intelligence and convenience of the V2X. Attached Figure Description

[0024] The present application will be further described below with reference to the accompanying drawings and embodiments:

[0025] Figure 1 This is a system framework diagram of an active vehicle-mounted system response method for low-speed vehicle scenarios according to an embodiment of this application;

[0026] Figure 2 This is a flowchart illustrating an active vehicle-mounted system response method for low-speed vehicle scenarios in one embodiment of this application;

[0027] Figure 3 This is a structural diagram of a vehicle-mounted active response system for low-speed vehicle scenarios, as shown in one embodiment of this application. Detailed Implementation

[0028] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0029] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application. Therefore, the drawings only show the layers related to this application and are not drawn according to the actual number, shape and size of the layers in the actual implementation. In the actual implementation, the shape, number and proportion of each layer can be arbitrarily changed, and the layer layout may also be more complex.

[0030] Numerous details are explored in the following description to provide a more thorough explanation of embodiments of this application; however, it will be apparent to those skilled in the art that embodiments of this application may be practiced without these specific details.

[0031] Most existing vehicle cameras operate on a "full-time acquisition + general processing" model. This means that regardless of whether the vehicle is traveling at high speed or parked at low speed, the camera continuously acquires images with fixed parameters, lacking an active switching mechanism based on scene characteristics. For example, when a vehicle approaches a highway tollbooth, the driver needs to manually activate the camera's "scanning mode" and adjust the vehicle's position and angle to ensure the camera clearly captures the payment QR code. During this process, the camera cannot predict the arrival of the "payment scenario" based on vehicle speed changes, nor can it automatically optimize focus and exposure parameters to adapt to QR code recognition requirements. This results in a scanning success rate that depends on the driver's skill level, and in extreme cases, multiple attempts may be required, reducing traffic efficiency. Similarly, when a vehicle approaches its destination (such as the entrance to a residential area or a shopping mall parking lot), the driver often needs to visually observe the surrounding environment to confirm the parking location or locate the target building. While the vehicle camera continuously captures surrounding images, it cannot actively recognize the scene characteristic of "approaching the destination," nor can it specifically zoom in on key areas (such as house numbers or parking lot entrance signs). This forces the driver to rely on subjective judgment, failing to leverage the camera's environmental perception value.

[0032] In low-speed scenarios, the driver's operational needs are often deeply intertwined with vehicle control functions. For example, upon recognizing a charging station, the charging port needs to be opened; upon recognizing a narrow road, the 360-degree surround-view imaging assist parking needs to be activated. However, in existing technologies, the image recognition function of the in-vehicle camera and the vehicle control function are "separated": the camera is only responsible for outputting the recognition result (such as "charging station detected"), while performing related operations (such as opening the charging port) still requires the driver to manually complete the operation through physical buttons or the central control screen, creating a break in the "recognition-execution" chain.

[0033] This disconnect not only increases the driver's workload but can also cause inconvenience due to operational delays. For example, in exit toll payment scenarios, drivers need to scan a QR code to pay, which can easily trigger anxiety if there are vehicles waiting behind them.

[0034] To address the aforementioned problems, this application provides the following technical solution.

[0035] Figure 1 This is a system architecture diagram of a vehicle-mounted active response method for low-speed vehicle scenarios according to an embodiment of this application, as shown below. Figure 1 As shown, this application includes a perception module, a recognition module, and an execution module. The perception module mainly acquires scene images, positioning information, and extracts pedal signals from the CAN bus. These are then input into the scene recognition module for scene recognition. The recognized scene labels and scene images are input into the segmentation model to segment low-speed root cause objects. The low-speed root cause objects, combined with the scene labels, extract execution rules from the rule base and hand them over to the execution unit for execution.

[0036] Figure 2 This is a flowchart of a vehicle-mounted active response method for low-speed vehicle scenarios, as described in one embodiment of this application. Figure 2 As shown, the specific process includes:

[0037] S100: When it is detected that the target vehicle is in a low-speed driving mode and the duration of the low-speed driving mode is greater than a preset duration threshold, the vehicle camera is controlled to capture scene images; and the lane-level positioning information and pedal signals of the target vehicle are obtained.

[0038] It is understood that the specific vehicle for implementing this invention can be any intelligent vehicle equipped with a smart camera, including traditional energy vehicles (fuel vehicles) and new energy vehicles (pure electric drive or hybrid electric vehicles). Each intelligent vehicle is equipped with at least one smart camera.

[0039] When the vehicle is traveling at a low speed for more than 30 seconds, it is determined that the vehicle is in the "low-speed driving model" of this application. At this time, the camera is controlled to capture video, and frames are extracted from the video to obtain scene images.

[0040] S200, classify the scene image based on the pre-built recognition model to obtain multiple scene labels of the scene image and the confidence scores of the multiple scene labels;

[0041] This application uses existing recognition models to classify scenarios. Existing recognition models can be localized models that rely on vehicle computing power or cloud models that rely on cloud services and are deployed on cloud servers.

[0042] The probability (confidence level) of various scene labels output by existing general scene recognition models.

[0043] S300, calculate the first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels based on the pre-built statistical model; and determine the target scene label based on the confidence degree of multiple scene labels, the first matching degree of multiple scene labels, and the second matching degree of multiple scene labels;

[0044] Traditional scene labels (such as "exiting the garage") are merely category IDs, lacking a quantitative description of the physical environment (lane type) and driving behavior (pedal signal) within the scene. In complex scenarios, misidentification of the scene is likely to occur. The accuracy of scene recognition determines the accuracy of subsequent active responses. In order to improve the accuracy of scene recognition in conjunction with the recognition model described above, this application constructs a statistical model.

[0045] This application constructs a statistical model based on sample data to characterize the matching degree of different target types and pedal signal features under multiple scene labels. The purpose is to bind scene labels with measurable environmental-behavioral features through the statistical model. The method for constructing the statistical model is as follows:

[0046] (1) Obtain multiple sample data of various scene labels, wherein the sample data includes lane-level positioning data samples and pedal signal samples of the target time period;

[0047] (2) Map the lane-level positioning data samples onto the map to obtain the lane type samples where the vehicle is located. ;

[0048] This mapping is essentially a non-linear coordinate-semantic mapping that transforms machine-encoded lane IDs into interpretable semantic labels (such as "parking lane"), making scene-lane associations quantifiable.

[0049] (3) The pedal signal samples are parsed into a throttle opening sequence. and brake opening / closing sequence The throttle opening sequence Including throttle opening and closing at multiple time sampling points The brake opening / closing sequence Brake opening and closing degree including multiple time sampling points ;

[0050] Temporal segmentation eliminates signal coupling, providing structured input for feature extraction. It also makes driving behavior patterns explicit (e.g., the braking sequence in a "leaving the garage" scene exhibits high-frequency fluctuations). This provides temporal evidence for scene-behavior association, supporting the dynamic scene recognition proposed in this application.

[0051] (4) From the throttle opening sequence Extracting throttle features And from the brake opening / closing sequence Extract braking features And based on the throttle feature and the braking features Constructing pedal signal features The throttle feature The braking features include average opening degree, standard deviation of opening degree, and non-zero percentage. This includes maximum value, continuous pedaling time, and switching frequency;

[0052] The same driving scenario imposes structural constraints on the driver's operations, reflected in the timing patterns, intensity distribution, and dynamic changes of pedal operations. For example:

[0053] Leaving the garage: frequent starts and stops, low-speed crawling → small pulses of accelerator + high-frequency intervention of brake.

[0054] When meeting oncoming traffic on a narrow road: Slow down cautiously and prepare to stop → Keep the brake lightly pressed and the accelerator pedal close to zero.

[0055] High-speed cruising: Throttle is maintained steadily and brakes are almost non-existent.

[0056] Traffic jam following: Frequent switching between accelerator and brake ("pumping the brake").

[0057] Therefore, effective features should be able to capture these pattern differences.

[0058] The aforementioned throttle characteristics focus on intensity, stability, and frequency of use, while brake characteristics focus on continuity, urgency, and interaction with the throttle. Therefore, the extracted pedal signal features... It can effectively distinguish different scene labels and serve as an important basis for determining scene labels.

[0059] (5) For each scene label, calculate the lane type sample for each vehicle. quantity With the number of multiple sample data The ratio of the two values ​​yields the first degree of match between the lane type and the scene label. , ;

[0060] (6) For each scene label, the pedal signal features of multiple samples Clustering was performed to obtain multiple typical pedal feature clusters, and the number of samples in each typical pedal feature cluster was calculated. With the number of multiple sample data The ratio of the two values ​​is used to obtain the second matching degree between typical pedal features and scene labels. , Wherein, the typical pedal feature is the average feature vector of the feature vectors within the typical pedal feature cluster;

[0061] (6) First matching degree based on multiple scene labels, multiple lane types, and multiple lane types Multiple typical pedal features and the second matching degree of multiple typical pedal features Build a statistical model.

[0062] Steps (5)-(7) are used to calculate the specific frequencies in the sample data to reflect the matching degree between different road types, different pedal features and scene labels, and then to build a statistical model.

[0063] After constructing the above statistical model, the process of calculating the first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels based on the pre-constructed statistical model, and determining the target scene label based on the confidence degree of multiple scene labels, the first matching degree of multiple scene labels, and the second matching degree of multiple scene labels includes:

[0064] S310, determine the lane type of the lane-level positioning, and based on the lane type, determine the first matching degree between the lane-level positioning and multiple scene labels from the statistical model. ;

[0065] Specifically, if the statistical model matches the lane type corresponding to multiple scene labels, it can directly output the first matching degree corresponding to the multiple scene labels. .

[0066] S320, extract the pedal features of the pedal signal, calculate the similarity between the pedal features and multiple typical pedal features of multiple scene labels in the statistical model, and take the second matching degree of the target typical pedal feature with the highest similarity in each scene label as the second matching degree of the scene label. .

[0067] The pedal features are essentially feature vectors. Therefore, similarity is calculated using cosine similarity or Euclidean distance. Then, the most similar typical pedal features are extracted from the typical pedal features of each scene label. This yields the second matching degree between the current pedal features and multiple scene labels. .

[0068] S330, confidence level for multiple scene labels First match score and first match score of multiple scene tags and the second matching degree of multiple scene tags We then perform weighted calculations to obtain the overall confidence level for each scene label. The mathematical expression for the overall confidence level is:

[0069]

[0070] In the formula, As the first weight, As the second weight, It is the third weight;

[0071] S340, the comprehensive confidence level The highest scene label is used as the target scene label.

[0072] Finally, the first matching score, second matching score, and recognition confidence score of each scene label are weighted and summed to obtain the comprehensive confidence score, where the first weight is... Second weight and third weight The values ​​can be 0.4, 0.3, and 0.3.

[0073] S400, the target scene label and the scene image are input into a pre-built segmentation model to obtain the segmentation result of one or more low-speed root cause targets in the scene image that match the scene label, wherein the segmentation result includes the category label of the low-speed root cause target and the segmentation mask of the low-speed root cause target;

[0074] In this application, a segmentation model based on scene label guidance is pre-built. The construction process of the segmentation model includes:

[0075] S1, Obtain the sample image set;

[0076] S2, take the sample images from the sample image set The input is fed into a pre-built recognition model to obtain scene labels. The images from the sample image set are then input into a pre-built general segmentation model to obtain multiple targets and their segmentation masks. ;

[0077] S3, Extract each target and scene label based on the statistical model. The matching degree is calculated, and targets with a matching degree less than a set matching degree threshold are removed to obtain associated targets and their segmentation masks. ;

[0078] Traditional segmentation models (such as general segmentation models) treat images as a collection of pixels without scene constraints, leading to false detections of irrelevant targets in specific scenes (such as segmenting "road width" in a "exiting a garage" scene). This solution uses scene labels S as semantic priors. The scene labels provide scene semantics, while the scene image provides pixel-level context. Together, they enable the model to accurately segment targets under semantic guidance.

[0079] S4, based on sample images Scene tags And segmentation masks for multiple related targets Construct training data, wherein the segmentation masks of the multiple associated targets For training labels;

[0080] S5, Input the training labels into the artificial neural network to obtain the prediction mask;

[0081] S6, based on a pre-built loss function Calculate the loss between the training labels and the predicted mask, wherein the loss function is... Including segmentation loss Matching with scenes ;

[0082] The loss function The mathematical expression is:

[0083]

[0084]

[0085]

[0086]

[0087]

[0088] In the formula, For weight parameters, Represents the target set. Indicates the current scene The relevant target set For the target index, These are the pixel coordinates. Image height, Image width, Indicate target The predicted mask pixels, Indicate target The actual segmentation mask pixels, Represents the binary cross-entropy loss. Indicate target The prediction mask, Indicate target The true segmentation mask.

[0089] The loss function described above uses the standard segmentation loss, i.e. Its principle is consistent with existing segmentation models, so it will not be explained in detail here. The innovation of this application lies in the scene matching loss. This application measures scene matching loss by calculating the binary cross-entropy loss between the predicted mask and the actual mask. This is because this application predefines the semantic scope of the current scene (e.g., "exiting the garage" means "the gate and QR code should appear"), and the scene matching loss... The purpose is to measure the model's performance in the scenario. Is the following section correctly segmented? The targets in the system (i.e., "gate" and "QR code"). For example, when... If the target (such as a QR code) exists in the image but the model fails to segment it (predicted mask is 0), or segments irrelevant targets (such as "road width"), then the segmentation has failed, and the corresponding scene matching loss is [not specified]. Outputs a higher loss value.

[0090] S7, The parameters of the artificial neural network are adjusted based on the loss combined with the minimum gradient method;

[0091] S8. Repeat steps S5-S7 until training is complete and the segmentation model is obtained.

[0092] By backpropagating using the minimum gradient method and continuously adjusting the internal parameters of the artificial neural network, a fitted model can be obtained. The fitted segmentation model can effectively output targets that are related to the current scene label.

[0093] After obtaining the above segmentation model, the scene image and target scene label are input into the model to segment the associated targets that cause the vehicle to be traveling at low speed, namely the root cause targets of low speed.

[0094] This application uses scene labels as semantic anchors to upgrade the model from "segmenting all targets" to "accurately segmenting scene-related targets"; pass Achieve pixel-level optimization for small objectives, enabling the model to truly understand "low-speed root causes" (such as deceleration caused by brakes) in automotive-grade scenarios.

[0095] S500, based on the type label of the low-speed root cause target, the pre-configured response rule base, determine the current response rule, and execute the vehicle system active response based on the current response rule. The response rule base includes multiple active response rules, and the active response rule is one or a combination of multiple actions such as opening the 360 ​​surround view, displaying the segmentation mask image of the low-speed root cause target, and popping out the charging port cover.

[0096] Finally, after performing scene recognition and low-speed root cause target localization using the aforementioned scheme, this application utilizes a configured rule base to apply the low-speed root cause targets, thereby enabling proactive response from the vehicle's infotainment system. This application illustrates how the extracted low-speed root cause targets are applied to the vehicle's infotainment system to provide intelligent and convenient in-vehicle functions through the following examples.

[0097] (1) When the low-speed root cause target is the garage exit gate and the payment QR code, the payment QR code image is extracted based on the mask of the payment QR code and projected onto the vehicle display area;

[0098] In the "exiting the garage" scenario, the segmentation model dynamically filters and generates a target mask (containing the garage exit gate and the payment QR code area) using scene labels. The gate forces vehicles to stop (at low speed), and the QR code is the key medium for payment (physical constraint: payment requires scanning the code to trigger the gate to rise).

[0099] Target extraction: Accurately crop the QR code region image based on segmentation mask (avoiding background interference).

[0100] Active projection: The QR code image is projected onto the vehicle's display area (such as the central control screen) in real time, eliminating the need for manual operation by the driver. Drivers can complete the scan and payment process from inside the vehicle without lowering the windows or adjusting their phone's camera angle to face the QR code.

[0101] (2) When the low-speed root cause target is a fork in the road, the fork in the road area image is extracted based on the fork in the road mask and projected onto the vehicle display area;

[0102] Target extraction: Extract the intersection area image (including lane lines and road signs) based on the mask and exclude irrelevant background (such as vehicles and pedestrians).

[0103] Dynamic projection: The bifurcation image is overlaid on the vehicle navigation interface in an AR-enhanced form (such as highlighting the "turn left / go straight" arrow). In addition, it can be combined with high-precision maps to calculate the optimal route in real time (such as "turn left 300 meters before entering the main road").

[0104] The above process takes less than 0.5 seconds from detection to response (traditional navigation requires the driver to manually check the map).

[0105] (3) When the low-speed root cause target is that the lane width is less than the set value, turn on the 360-degree surround view;

[0106] Object detection: Real-time calculation of lane width (based on segmentation mask).

[0107] Automatic activation: When the lane width is less than 2.5 meters, the 360-degree surround view system (panoramic image + side obstacle warning) will automatically activate without driver intervention. The surround view image is overlaid with lane boundaries (red warning lines) in real time, intuitively displaying safe distances.

[0108] It is activated only in low-speed scenarios (speed <15km / h) to avoid interference at high speeds.

[0109] (4) When the low-speed root cause target is a charging pile, the request to open the charging port cover will be projected onto the vehicle display area.

[0110] In charging station scenarios (such as parking lot charging stations), where the vehicle is moving at low speed (requiring it to stop and charge), the location of the charging station is accurately located based on a mask. The command "Open charging port cover" is projected onto the vehicle's infotainment display area (e.g., a "Click to open charging port" button pops up on the central control screen). It automatically opens after the owner confirms.

[0111] This application discloses a vehicle-mounted infotainment system (V2X) active response method for low-speed vehicle scenarios. This method utilizes existing scene recognition models to identify scenes and further filters target scene labels by combining lane-level positioning information and pedal signals. The target scene labels and the original image are input into a segmentation model to extract low-speed root cause targets strongly correlated with the target scene labels. Active response rules are extracted from a rule base based on these low-speed root cause targets, and the V2X active response is executed. This application combines scene recognition and low-speed root cause target extraction to configure multiple active rules, enriching the V2X active response functionality in low-speed scenarios and improving the intelligence and convenience of the vehicle-mounted infotainment system.

[0112] like Figure 3 As shown, this application also provides a vehicle-mounted active response system for low-speed vehicle scenarios, comprising:

[0113] The acquisition module is used to control the vehicle-mounted camera to acquire scene images when it detects that the target vehicle is in a low-speed driving mode and the duration of the low-speed driving mode is greater than a preset duration threshold; and to acquire the lane-level positioning information and pedal signal of the target vehicle.

[0114] The scene classification module is used to classify the scene image based on a pre-built recognition model to obtain multiple scene labels and confidence scores of the scene image.

[0115] The scene determination module is used to calculate the first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels based on a pre-built statistical model; and to determine the target scene label based on the confidence degree of multiple scene labels, the first matching degree of multiple scene labels, and the second matching degree of multiple scene labels.

[0116] The segmentation module is used to input the target scene label and the scene image into a pre-built segmentation model to obtain the segmentation result of one or more low-speed root cause targets in the scene image that match the scene label. The segmentation result includes the category label of the low-speed root cause target and the segmentation mask of the low-speed root cause target.

[0117] An active response module is used to determine the current response rule based on the type label of the low-speed root cause target, the pre-configured response rule library, and to execute the vehicle system active response based on the current response rule. The response rule library includes multiple active response rules, which are one or a combination of the following: opening the 360-degree surround view, displaying the segmentation mask image of the low-speed root cause target, and popping out the charging port cover.

[0118] This application discloses an in-vehicle infotainment system for low-speed vehicle scenarios. The system utilizes existing scene recognition models to identify scenes and further filters target scene labels by combining lane-level positioning information and pedal signals. The target scene labels and the original image are input into a segmentation model to extract low-speed root cause targets strongly correlated with the target scene labels. Based on these low-speed root cause targets, active response rules are extracted from a rule base, and the in-vehicle infotainment system's active response is executed. This application combines scene recognition and low-speed root cause target extraction to configure multiple active rules, enriching the in-vehicle infotainment system's active response capabilities in low-speed scenarios and improving its intelligence and convenience.

[0119] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0120] The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory so that the terminal performs any of the methods in this embodiment.

[0121] As will be understood by those skilled in the art, the computer-readable storage medium described in this embodiment allows for the implementation of all or part of the steps in the above method embodiments by computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0122] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication between them. The memory is used to store computer programs, the communication interface is used to perform communication, and the processor and the transceiver are used to run the computer programs, so that the electronic terminal performs the steps of the above method.

[0123] In this embodiment, the memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.

[0124] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0125] In the above embodiments, although the present application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. The embodiments of the present application are intended to cover all such substitutions, modifications, and variations falling within the broad scope thereof.

[0126] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by this application.

Claims

1. A vehicle-mounted active response method for low-speed vehicle scenarios, characterized in that, Including the following steps: When the target vehicle is detected to be in a low-speed driving mode and the duration of the low-speed driving mode exceeds a preset duration threshold, the vehicle camera is controlled to capture scene images; and the lane-level positioning information and pedal signals of the target vehicle are obtained. The scene image is classified based on a pre-built recognition model to obtain multiple scene labels and confidence scores for the scene image. The first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels are calculated based on a pre-built statistical model. The target scene label is determined based on the confidence score of multiple scene labels, the first matching score of multiple scene labels, and the second matching score of multiple scene labels. The method for constructing the statistical model includes: acquiring multiple sample data for various scene labels, wherein the sample data includes lane-level positioning data samples and pedal signal samples for a target time period; mapping the lane-level positioning data samples onto a map to obtain a sample of the lane type where the vehicle is located. The pedal signal samples were parsed into a throttle opening / closing sequence. and brake opening / closing sequence The throttle opening sequence Including throttle opening and closing at multiple time sampling points The brake opening / closing sequence Brake opening and closing degree including multiple time sampling points From the throttle opening sequence Extracting throttle features And from the brake opening / closing sequence Extract braking features And based on the throttle feature and the braking features Constructing pedal signal features The throttle feature The braking features include average opening degree, standard deviation of opening degree, and non-zero percentage. This includes maximum value, continuous pedaling time, and switching frequency; for each scene label, it calculates the sample of the lane type for each type of vehicle. quantity With the number of multiple sample data The ratio of the two values ​​yields the first degree of match between the lane type and the scene label. , For each scene label, the pedal signal features of multiple samples. Clustering was performed to obtain multiple typical pedal feature clusters, and the number of samples in each typical pedal feature cluster was calculated. With the number of multiple sample data The ratio of the two values ​​is used to obtain the second matching degree between typical pedal features and scene labels. , Wherein, the typical pedal feature is the average feature vector of the feature vectors within the typical pedal feature cluster; based on the first matching degree corresponding to multiple scene labels, multiple lane types, and multiple lane types. Multiple typical pedal features and the second matching degree of multiple typical pedal features Construct a statistical model; The target scene label and the scene image are input into a pre-built segmentation model to obtain the segmentation result of one or more low-speed root cause targets in the scene image that match the scene label. The segmentation result includes the category label of the low-speed root cause target and the segmentation mask of the low-speed root cause target. Based on the type label of the low-speed root cause target and the pre-configured response rule base, the current response rule is determined, and the vehicle system actively responds based on the current response rule. The response rule base includes multiple active response rules, which are one or a combination of the following: opening the 360-degree surround view, displaying the segmentation mask image of the low-speed root cause target, and popping out the charging port cover.

2. The vehicle-mounted active response method for low-speed vehicle scenarios according to claim 1, characterized in that, The first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels are calculated based on a pre-built statistical model, including: The lane type of the lane-level positioning is determined, and based on the lane type, a first matching degree between the lane-level positioning and multiple scene labels is determined from the statistical model. ; Extract the pedal features from the pedal signal, calculate the similarity between the pedal features and multiple typical pedal features of multiple scene labels in the statistical model, and take the second matching degree of the target typical pedal feature with the highest similarity in each scene label as the second matching degree of the scene label. .

3. The vehicle-mounted active response method for low-speed vehicle scenarios according to claim 1, characterized in that, The target scene label is determined based on the confidence score of multiple scene labels, the first matching score of multiple scene labels, and the second matching score of multiple scene labels, including: Confidence of multiple scene labels First match score and first match score of multiple scene tags and the second matching degree of multiple scene tags We then perform weighted calculations to obtain the overall confidence level for each scene label. The mathematical expression for the overall confidence level is: In the formula, As the first weight, As the second weight, It is the third weight; The overall confidence level The highest scene label is used as the target scene label.

4. The vehicle-mounted active response method for low-speed vehicle scenarios according to claim 1, characterized in that, The method for constructing the segmentation model includes: S1, Obtain the sample image set; S2, take the sample images from the sample image set The input is fed into a pre-built recognition model to obtain scene labels. The images from the sample image set are then input into a pre-built general segmentation model to obtain multiple targets and their segmentation masks. ; S3, Extract each target and scene label based on the statistical model. The matching degree is calculated, and targets with a matching degree less than a set matching degree threshold are removed to obtain associated targets and their segmentation masks. ; S4, based on sample images Scene tags And segmentation masks for multiple related targets Construct training data, wherein the segmentation masks of the multiple associated targets For training labels; S5, Input the training labels into the artificial neural network to obtain the prediction mask; S6, based on a pre-built loss function Calculate the loss between the training labels and the predicted mask, wherein the loss function is... Including segmentation loss Matching with scenes ; S7, The parameters of the artificial neural network are adjusted based on the loss combined with the minimum gradient method; S8. Repeat steps S5-S7 until training is complete and the segmentation model is obtained.

5. The vehicle-mounted active response method for low-speed vehicle scenarios according to claim 4, characterized in that, The loss function The mathematical expression is: In the formula, For weight parameters, Represents the target set. Indicates the current scene The relevant target set For the target index, These are the pixel coordinates. Image height, Image width, Indicate target The predicted mask pixels, Indicate target The actual segmentation mask pixels, This represents the binary cross-entropy loss. Indicate target The prediction mask, Indicate target The true segmentation mask.

6. The vehicle-mounted active response method for low-speed vehicle scenarios according to claim 1, characterized in that, The response rule base includes: When the low-speed root cause target is the garage exit gate and the payment QR code, the payment QR code image is extracted based on the mask of the payment QR code and projected onto the vehicle display area; When the low-speed root cause target is a fork in the road, the fork in the road area image is extracted based on the fork in the road mask and projected onto the vehicle display area. When the low-speed root cause target is that the lane width is less than a set value, turn on the 360-degree surround view. When the target of the low-speed root cause is a charging pile, the request to open the charging port cover will be projected onto the vehicle display area.

7. A vehicle-mounted active response system for low-speed vehicle scenarios, characterized in that, include: The acquisition module is used to control the vehicle camera to acquire scene images when it detects that the target vehicle is in a low-speed driving mode and the duration of the low-speed driving mode is greater than a preset duration threshold. And acquire the lane-level positioning information of the target vehicle and the pedal signal of the target vehicle; The scene classification module is used to classify the scene image based on a pre-built recognition model to obtain multiple scene labels and confidence scores of the scene image. The scene determination module is used to calculate the first matching degree between the lane-level positioning and multiple scene labels, and the second matching degree between the pedal signal and multiple scene labels based on a pre-built statistical model. The target scene label is determined based on the confidence score of multiple scene labels, the first matching score of multiple scene labels, and the second matching score of multiple scene labels. The method for constructing the statistical model includes: acquiring multiple sample data for various scene labels, wherein the sample data includes lane-level positioning data samples and pedal signal samples for a target time period; mapping the lane-level positioning data samples onto a map to obtain a sample of the lane type where the vehicle is located. The pedal signal samples were parsed into a throttle opening / closing sequence. and brake opening / closing sequence The throttle opening sequence Including throttle opening and closing at multiple time sampling points The brake opening / closing sequence Brake opening and closing degree including multiple time sampling points From the throttle opening sequence Extracting throttle features And from the brake opening / closing sequence Extract braking features And based on the throttle feature and the braking features Constructing pedal signal features The throttle feature The braking features include average opening degree, standard deviation of opening degree, and non-zero percentage. This includes maximum value, continuous pedaling time, and switching frequency; for each scene label, it calculates the sample of the lane type for each type of vehicle. quantity With the number of multiple sample data The ratio of the two values ​​yields the first degree of match between the lane type and the scene label. , For each scene label, the pedal signal features of multiple samples. Clustering was performed to obtain multiple typical pedal feature clusters, and the number of samples in each typical pedal feature cluster was calculated. With the number of multiple sample data The ratio of the two values ​​is used to obtain the second matching degree between typical pedal features and scene labels. , Wherein, the typical pedal feature is the average feature vector of the feature vectors within the typical pedal feature cluster; based on the first matching degree corresponding to multiple scene labels, multiple lane types, and multiple lane types. Multiple typical pedal features and the second matching degree of multiple typical pedal features Construct a statistical model; The segmentation module is used to input the target scene label and the scene image into a pre-built segmentation model to obtain the segmentation result of one or more low-speed root cause targets in the scene image that match the scene label. The segmentation result includes the category label of the low-speed root cause target and the segmentation mask of the low-speed root cause target. An active response module is used to determine the current response rule based on the type label of the low-speed root cause target and a pre-configured response rule library, and to execute the vehicle system active response based on the current response rule. The response rule library includes multiple active response rules, which are one or a combination of the following: opening the 360-degree surround view, displaying the segmentation mask image of the low-speed root cause target, and popping out the charging port cover.

8. An electronic device, characterized in that, include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to cause the electronic device to perform the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image tuning method and device for automatic driving system, and storage medium

    CN116843574A

  • Driving scene recognition method and device, computer equipment and storage medium

    CN115393818A

  • High beam and low beam automatic switching control method and device based on scene recognition

    CN119018045A