Vehicle control method and system, vehicle, storage medium and program product
By using multi-source sensor data fusion and deep learning algorithms to identify features during the reversing phase and dynamically matching the viewing angle, the problem of distraction in traditional reversing assistance technology is solved, improving the safety and convenience of the reversing process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-03
AI Technical Summary
Traditional reversing assist technology relies on manually switching the view, which leads to a distraction of attention while driving and affects safety.
By fusing data from multiple sensors and using deep learning algorithms, the system accurately identifies features during the reversing phase, dynamically matches the optimal viewing angle based on state machine logic, and automatically switches the viewing angle.
It improves the ease of operation and safety during driving, and reduces the risk of scratches caused by untimely or mismatched perspective switching.
Smart Images

Figure CN121777801A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicles, and more particularly to a vehicle control method and system, a vehicle, a storage medium, and a program product. Background Technology
[0002] When reversing, especially in complex scenarios (such as narrow parking spaces, dense parking lots, and angled parking), users need to frequently observe the surrounding environment of the vehicle to judge the distance of obstacles, the position of parking lines, and the relative relationship between the vehicle and the target parking space.
[0003] Traditional reversing assist technologies often involve manually switching between surround view or a single static view. This operation can distract the user and may lead to accidental pressing or releasing of the brake or accelerator pedals while switching views, affecting driving safety. Summary of the Invention
[0004] This application provides vehicle control methods and systems, vehicles, storage media, and program products to improve safety during driving.
[0005] In a first aspect, embodiments of this application provide a vehicle control method, including:
[0006] Acquire multi-source sensor data, including vehicle status data and environmental perception data;
[0007] The multi-source sensor data is fused using a deep learning algorithm to obtain vehicle reversing stage features, which are used to characterize the current stage of the vehicle during reversing.
[0008] Based on the characteristics of the vehicle reversing phase, the target viewpoint is determined through state machine logic, and the display screen of the target viewpoint is output.
[0009] In one possible implementation, the step of fusing the multi-source sensor data based on a deep learning algorithm to obtain vehicle reversing phase features includes:
[0010] Image features of the environmental perception data are extracted based on a convolutional neural network.
[0011] The time-series characteristics of the vehicle status data were analyzed based on a long short-term memory network.
[0012] Based on the image features and time series features, the characteristics of the vehicle reversing phase are determined.
[0013] In one possible implementation, determining the target viewing angle based on the vehicle's reversing phase characteristics using state machine logic and displaying the screen from that target viewing angle includes:
[0014] Based on the vehicle reversing phase characteristics and the mapping relationship between the reversing phase and the viewing angle, the target viewing angle corresponding to the vehicle reversing phase characteristics is determined.
[0015] The transition screen of the target view is generated based on the interpolation algorithm, and the transition screen and the display screen of the target view are output sequentially.
[0016] In one possible implementation, determining the target viewing angle corresponding to the vehicle reversing phase characteristics based on the vehicle reversing phase characteristics and the mapping relationship between the reversing phase and the viewing angle includes:
[0017] When the vehicle is in the initial positioning stage, the target view is determined to be a front view and a panoramic view, based on the vehicle reversing phase feature characterization.
[0018] When the vehicle is in the initial stage of reversing, the target view is determined to be a side-rear view.
[0019] When the vehicle is in the middle of reversing phase, the target viewing angle is determined to be the rear-view frontal view and the side-view.
[0020] When the vehicle is in the reversing phase, indicating that the vehicle is in the final stage of reversing, the target view is determined to be a rear-view frontal view or a bird's-eye view.
[0021] In one possible implementation, the method further includes:
[0022] The weights of each sensor data point are adjusted based on the confidence level of each sensor data point in the multi-source sensor data.
[0023] And / or, adjust the state machine logic based on the user's historical reversing behavior data.
[0024] In one possible implementation, generating the transition scene from the target perspective using an interpolation algorithm includes at least one of the following:
[0025] Intermediate frames between adjacent viewpoints are generated based on a linear interpolation algorithm;
[0026] Generate non-linear transition images based on Bézier curve interpolation algorithm;
[0027] The transition parameters of the interpolation algorithm are adjusted based on the user's historical reversing behavior data.
[0028] Secondly, embodiments of this application provide a vehicle control system, including: a memory and a processor;
[0029] The memory stores computer-executed instructions;
[0030] The processor executes computer execution instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0031] Thirdly, this application provides a vehicle including the vehicle control system described in the second aspect.
[0032] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0033] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0034] The vehicle control method and system, vehicle, storage medium, and program product provided in this application acquire multi-source sensor data, including vehicle state data and environmental perception data. Based on a deep learning algorithm, the multi-source sensor data is fused to obtain vehicle reversing stage features. Based on these features, a state machine logic determines the target viewing angle and outputs the display screen from that angle. Through multi-source sensor data fusion, reversing stage features (such as initial positioning, early, middle, and late reversing stages) are accurately identified. The state machine logic dynamically matches the optimal viewing angle (such as a side-rear view or a rear-view frontal view) based on these stage features, thereby accurately and actively switching the viewing angle throughout the reversing process. This avoids user distraction caused by manually switching the viewing angle, significantly improving operational convenience and safety, while reducing the risk of scratches caused by untimely or mismatched viewing angle switching. Attached Figure Description
[0035] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0036] Figure 1 A schematic diagram of a scenario for the vehicle control method provided in this application;
[0037] Figure 2 A flowchart illustrating the vehicle control method provided in this application;
[0038] Figure 3 A schematic diagram of the vehicle control device provided in this application;
[0039] Figure 4 A schematic diagram of the structure of the electronic device provided in this application.
[0040] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0041] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0042] As described in the background section, traditional reversing assistance technologies typically rely on a fixed-view rearview mirror or a single static image (such as a rearview view). Users need to manually switch between different views (such as left rearview, right rearview, panoramic view, etc.) and may even need to repeatedly compare the image with the screen using a physical rearview mirror. This operation can distract the user and may also lead to accidental pressing or releasing of the brake or accelerator pedals when switching views, affecting driving safety.
[0043] To address this, this application provides a vehicle control method that uses multi-source sensor data fusion to accurately identify reversing stage characteristics (such as initial positioning, early reversing stage, middle reversing stage, and late reversing stage). The state machine logic dynamically matches the optimal viewing angle (such as side and rear view, and rear front view) based on the stage characteristics, thereby accurately and actively switching the viewing angle throughout the entire reversing process. This avoids user distraction caused by manually switching the viewing angle, significantly improves operational convenience and safety, and reduces the risk of scratches caused by untimely or mismatched viewing angle switching.
[0044] Figure 1 A schematic diagram of a scenario for the vehicle control method provided in this application, such as... Figure 1 As shown, the vehicle control system acquires vehicle status data and environmental perception data collected by sensors such as surround view cameras, ultrasonic radar sensors, vehicle speed sensors, and steering wheel angle sensors. It then calls a deep learning model, inputting the vehicle status data and environmental perception data into the deep learning model. The deep learning model fuses and processes the vehicle status data and environmental perception data to output the vehicle reversing phase characteristics.
[0045] After obtaining the characteristics of the vehicle reversing stage, the vehicle control system can determine the current stage of the vehicle reversing process, determine the target view through state machine logic, and output the display screen of the target view to the display screen, so that the user can observe the required view on the display screen without having to park back and forth, stop and wait to observe, and then continue parking, matching the driver's precise view requirements for each stage of reversing.
[0046] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0047] Figure 2 A flowchart illustrating the vehicle control method provided in this application is shown below. Figure 2 As shown, with the vehicle control system as the executing entity, the method includes:
[0048] S101. Acquire multi-source sensor data, including vehicle status data and environmental perception data.
[0049] Multi-source sensor data refers to vehicle operating status and surrounding environment information collected by multiple sensors, including but not limited to vehicle speed, steering wheel angle, obstacle distance, and surround view images. Vehicle status data includes vehicle speed and steering wheel angle; environmental perception data includes obstacle distance and surround view images.
[0050] For example, sensors may include surround-view cameras, ultrasonic radar, vehicle speed sensors, steering wheel speed sensors, etc. Accordingly, multi-source sensor data may include 360° high-definition real-time images of the vehicle's surroundings captured by the surround-view camera; real-time measurements of the distance between the vehicle and surrounding obstacles by the ultrasonic radar sensor; vehicle speed sensors monitoring vehicle speed to determine reversing status; and steering wheel angle sensors determining the vehicle's turning radius (steering radius) to identify the reversing phase.
[0051] For example, after acquiring multi-source sensor data, the multi-source sensor data can be preprocessed before being input into a deep learning model, thereby providing high-quality, standardized input data for the deep learning model.
[0052] For example, distortion correction and stitching can be performed on surround-view images to form a panoramic view; ultrasonic radar data can be filtered and denoised; and steering wheel angle data can be normalized. Distortion correction and stitching ensure the accuracy of panoramic images; filtering and denoising improve the reliability of ultrasonic radar data; and normalization eliminates the dimensional differences between data from different sensors.
[0053] Distortion correction eliminates camera lens distortion (such as the fisheye effect) through algorithms, for example, using the `fisheye::undistortImage` function in OpenCV. Filtering and denoising removes outliers from ultrasonic radar data using moving averages or Kalman filtering; for example, a triple moving average is used to detect sudden changes in distance. Normalization maps steering wheel angle values to the [0,1] interval for easier processing by subsequent algorithms.
[0054] S102. Based on deep learning algorithms, multi-source sensor data are fused and processed to generate vehicle reversing stage features, which are used to characterize the current stage of the vehicle during reversing.
[0055] Deep learning algorithms refer to a series of mathematical models and computational methods used to solve specific tasks. Examples include Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. CNNs are used to analyze obstacles, parking markings, and vehicle positions in surround-view images; LSTM networks are used to analyze time-series data of steering wheel angles and vehicle positions.
[0056] In some embodiments, image features of environmental perception data are extracted based on convolutional neural networks; time series of vehicle state data are analyzed based on long short-term memory networks; and vehicle reversing phase features are generated based on image features and time series features. Through coordinated processing of convolutional neural networks and long short-term memory networks, deep fusion of image features and time series features is achieved, improving the accuracy of vehicle reversing phase feature recognition.
[0057] Convolutional neural networks (CNNs) are deep learning models that extract local features of images through multiple layers of convolutional kernels. For example, the ResNet-50 architecture can be used to extract edge features of obstacles and road markings from surround-view images. Long Short-Term Memory (LSTM) networks are recurrent neural networks that process time-series data. They can capture the temporal dependencies of vehicle state data, such as the trend of steering wheel angle changes and the identification of features during the reversing phase (such as initial positioning and the initial stage of reversing).
[0058] Convolutional neural networks perform layer-by-layer convolution and pooling operations on environmental perception data to extract image features such as obstacles and road markings; long short-term memory networks perform time-series modeling of vehicle state data to analyze its dynamic change trends. Image features and time-series features are fused together through a feature fusion layer to generate vehicle reversing phase features, which serve as inputs to the state machine logic.
[0059] For example, in a slanted parking scenario, the vehicle control system accurately identifies the position of the parking space line based on a convolutional neural network. At the same time, it analyzes the trend of steering wheel angle changes through a long short-term memory network and jointly determines that it is in the initial stage of reversing, thereby switching to the side and rear view in advance and reducing the recognition error caused by insufficient extraction of a single feature.
[0060] For example, a convolutional neural network may include an input layer, a convolutional layer, an activation layer, a pooling layer, a depthwise separable convolutional layer, a fully connected layer, etc.
[0061] Preprocessing of multi-source sensor data may include scaling the original image (1920×1080) to 640×480, pixel normalization (mapping 0-255 to 0-1), and format conversion (converting RGB to tensors, such as 3×480×640, where 3 is the RGB channel).
[0062] The input layer receives preprocessed image data and passes it to the first convolutional layer of the convolutional neural network. The first convolutional layer can use multiple convolutional kernels (e.g., 32) sliding across the input image. Each kernel learns a specific feature pattern; for example, kernel 1 captures horizontal edges (matching the horizontal edges of lane lines), kernel 2 captures vertical edges (matching the vertical lines of lane lines), and kernel 3 captures diagonal edges (matching the edges of curved lanes). For RGB images, each kernel typically performs a convolution operation with three input channels (R, G, B). With 32 kernels, 32 feature maps can be output, meaning the first convolutional layer can output feature maps with a shape of 32×480×640.
[0063] After the convolution operation, the activation layer (ReLU) of the convolutional neural network filters out invalid noise points (such as the edges of small stones on the road surface) and retains only the effective features of lane lines and vehicle edges.
[0064] Pooling layers can use 2x2 max pooling to downsample the feature map, thereby reducing the size of the feature map (e.g., 480×640) to half its original size (240×320), reducing computation by about 75%. It should be noted that while reducing computation, pooling operations can also preserve the location information and approximate shape of the features.
[0065] Deep separable convolutional layers can include depthwise convolution and pointwise convolution. Each channel (e.g., 32 channels) in deep convolution uses an independent convolutional kernel to extract its own features (e.g., channel 1 extracts lane line edges, channel 2 extracts vehicle contours). Pointwise convolution uses 1×1 convolution to combine and fuse features from multiple channels to generate higher-level features. For example, it can extract continuous edges and road surface textures of straight lanes, i.e., lane area features, as well as vehicle contours and rounded edges of headlights, i.e., features of vehicles ahead.
[0066] After multiple rounds of convolution and pooling, the size of the feature map becomes very small (e.g., 15×20), but the number of channels becomes very large (e.g., 256 or 512). High-level convolutional layers can extract semantic-level features from these low-dimensional, high-channel-count feature maps, such as the position and size of vehicles ahead, the direction of lane lines, and the shape and color of traffic signs.
[0067] Then, the fully connected layer flattens the multidimensional feature map extracted by the high-level convolution into a one-dimensional vector (e.g., a 512-dimensional vector) and outputs it to the vehicle control system.
[0068] For example, the core of a Long Short-Term Memory (LSTM) network consists of three gates: the forget gate, the input gate, and the output gate. The forget gate determines which past states to forget; for example, if a vehicle changes from turning to going straight, the forget gate will forget the previous turning angle and retain the straight-going state. The input gate determines which current states to retain; for example, if the brake opening suddenly increases in the current frame, the input gate will focus on retaining this feature. The output gate determines which remembered states to output; for example, if it remembers five consecutive frames of increased brake opening and decreased vehicle speed, the output gate will output the feature that the vehicle is about to stop.
[0069] Taking vehicle status data, including vehicle speed, steering wheel angle, and brake opening, as an example (10 frames (1 second)):
[0070] Vehicle state data undergoes preprocessing, which may include normalization and format conversion. Normalization involves mapping vehicle speed (0-120) to 0-1 and brake opening (0-100) to 0-1. Format conversion involves converting vehicle state data conforming to one batch, 10 time steps, and 3 features into a three-dimensional tensor format.
[0071] Then, the original 3D features (vehicle speed, turning angle, braking) can be encoded into a 16-dimensional vector, outputting 1×10×16 (batch, time step, embedding dimension), making it easier for the Long Short-Term Memory network to learn temporal associations.
[0072] The Long Short-Term Memory (LSTM) network can be configured with 64 hidden units (adapting to onboard computing power) and input 1×10×16 temporal data. The LSM network processes this data frame by frame. For example, frame 1: remembers the vehicle speed of 60, no braking, and driving straight; frames 2-10: the forget gate gradually forgets the high vehicle speed, while the input gate retains the increase in brake opening and the decrease in vehicle speed; the output gate outputs 1×64 hidden states (e.g., containing the temporal correlation features of decreasing vehicle speed and increasing brake opening for 10 consecutive frames).
[0073] The fully connected layer can compress the 64-dimensional hidden state into a 16-dimensional vector and output it to the vehicle control unit.
[0074] For example, the vehicle reversing phase includes initial positioning, initial reversing phase, middle reversing phase, and final reversing phase.
[0075] S103. Based on the characteristics of the vehicle reversing phase, the target viewpoint is determined through state machine logic, and the display screen of the target viewpoint is output.
[0076] State machine logic refers to control logic based on preset state transition rules, used to dynamically switch system states according to input conditions. Target view refers to the optimal display view matched based on the characteristics of the vehicle during the reversing phase, including but not limited to front view, panoramic view, side and rear view, side view, rear-view front view, BEV (Bird's Eye View), etc.
[0077] A front view refers to an image captured by a camera directly in front of the vehicle (usually above the windshield or at the front grille), showing the scene in front of the vehicle. A panoramic view involves capturing images from multiple cameras (e.g., four cameras: front, rear, left, and right), then stitching and merging them to create a 360-degree panoramic image covering the area around the vehicle. A side-rear view is an image captured by a camera mounted to the left or right rear of the vehicle (usually on either side of the rear bumper), showing the scene to the side of the vehicle. A side view is an image captured by a camera mounted to the left or right front of the vehicle (usually at the A-pillar or rearview mirror), showing the scene to the side of the vehicle. A rear-view frontal view is an image captured by a camera mounted directly behind the vehicle (usually in the center of the rear windshield or at the trunk lid), showing the scene directly behind the vehicle. A BEV view is a perspective transformation, fusion, and stitching of images from different angles (e.g., front, rear, left, and right) using algorithms to generate a top-down view of the vehicle and its surroundings from a high altitude.
[0078] In some embodiments, the state machine logic includes a mapping relationship between the reversing phase and the viewpoint. Accordingly, the target viewpoint corresponding to the vehicle's reversing phase characteristics can be determined based on the vehicle's phase characteristics and the mapping relationship between the reversing phase and the viewpoint; that is, the target viewpoint is determined based on the mapping relationship and the vehicle's current phase. Then, a transition frame for the target viewpoint can be generated based on an interpolation algorithm, and the transition frame and the target viewpoint display frame are output sequentially to ensure a smooth and natural switching process and prevent frequent jumps. The interpolation algorithm is an image processing technique that generates intermediate transition frames using mathematical interpolation methods.
[0079] For example, when switching from a side-view to a frontal view during reversing, the vehicle control system generates an intermediate frame using a linear interpolation algorithm, so that the image smoothly transitions from a side-view to a frontal view, avoiding any sudden changes in the image that could negatively impact the user experience.
[0080] In some examples, during the vehicle reversing phase, the feature characterizes the vehicle in the initial positioning stage, determining the target viewing angles as the front view and the panoramic view, and outputting the display images of the front view and the panoramic view. During the initial stage of reversing, the feature characterizes the vehicle in the side-rear view, determining the target viewing angle, and outputting the display image of the side-rear view. During the middle stage of reversing, the feature characterizes the vehicle in the rear-view frontal view and the side view. During the final stage of reversing, the feature characterizes the vehicle in the rear-view frontal view or the bird's-eye view.
[0081] By intelligently analyzing each stage of vehicle reversing and automatically and dynamically switching the surround view display perspective, the system can match the driver's precise perspective requirements for each stage of reversing.
[0082] For example, when the vehicle speed is greater than or equal to a preset speed threshold (e.g., 3 km / h) and the steering wheel angle is less than a first preset angle threshold (e.g., an angle close to 0), the vehicle is determined to be in the initial positioning stage of the reversing phase. When the vehicle is in reverse (R) gear and the steering wheel angle is greater than or equal to a second preset angle threshold (e.g., 90°), the vehicle is determined to be in the initial reversing stage. When the vehicle is in reverse gear and the steering wheel angle is greater than the first preset angle threshold but less than the second preset angle threshold, the vehicle is determined to be in the later reversing stage.
[0083] For example, insufficient fusion of multi-source sensor data can lead to inaccurate environmental perception and vehicle status, affecting the accuracy of perspective switching. Employing a dynamic weight allocation mechanism, which adjusts the weights of each sensor in real time based on the confidence level of the sensor data (e.g., ultrasonic radar has high confidence for close-range detection, and cameras have high confidence for long-range recognition), can optimize the results of multi-source sensor data fusion. Here, confidence level refers to the reliability index of sensor data, typically based on an assessment of sensor measurement accuracy and environmental adaptability.
[0084] Therefore, in some embodiments, the weights of each sensor data are adjusted according to the confidence level of each sensor data in the multi-source sensors. By dynamically adjusting the weights, the system can adapt to the differences in sensor performance under different scenarios, improve the robustness of environmental perception, and optimize the multi-source data fusion effect.
[0085] For example, before acquiring multi-source sensor data, the vehicle control system calculates the confidence level of each sensor data in real time using a sliding window algorithm and dynamically adjusts the weights based on the confidence level.
[0086] For example, when a vehicle is close to an obstacle, the weight of ultrasonic radar data is increased to improve the accuracy of near-range obstacle detection; when the vehicle is in an open area, the wide-area perception capability of the camera is emphasized, the weight of camera data is increased, and the impact of radar blind spots is reduced.
[0087] For example, attention mechanisms (such as the Transformer architecture) can be introduced into convolutional neural networks and long short-term memory networks to enhance the model's ability to focus on key features (such as obstacles and parking space markings). For instance, when processing panoramic images, convolutional neural networks can dynamically allocate computational resources through self-attention mechanisms to prioritize processing areas related to the reversing phase (such as blind spots and parking space lines); similarly, when analyzing time-series data, attention weights can be used to enhance the sensitivity to changes in steering wheel angle.
[0088] Attention mechanisms can improve the model's efficiency in extracting key information and reduce redundant computation. For example, in the initial stage of reversing, the vehicle control system can more accurately identify subtle changes in the steering wheel angle and predict the required side and rear view in advance; in the final stage of reversing, the vehicle control system can enhance its focus on near-field obstacle data from ultrasonic radar, avoiding delays in view switching caused by environmental interference. Ultimately, this improves the accuracy of recognition during the reversing phase, ensuring a high degree of match between view switching and the user's operational intentions.
[0089] In some examples, generating transition frames for the target viewpoint using interpolation algorithms includes at least one of the following methods: generating intermediate frames between adjacent viewpoints based on a linear interpolation algorithm; or generating non-linear transition frames based on a Bézier curve interpolation algorithm. The availability of multiple interpolation algorithms enhances the flexibility of viewpoint switching.
[0090] Bézier curve interpolation is an interpolation method that generates smooth curves using control points, suitable for generating non-linear transition frames. For example, when switching from a side-rear view to a rear-view view, linear interpolation is used to generate intermediate frames; when switching from a BEV view to a rear-view frame, Bézier curve interpolation is used to generate non-linear transition frames to adapt to the switching requirements of different scenes.
[0091] In some embodiments, the state machine logic, i.e., the perspective switching strategy, is adjusted based on the user's historical reversing behavior data. Dynamic adaptation of user behavior data enhances the personalized adaptability of the perspective switching strategy. For example, for experienced users, the system can speed up perspective switching and reduce operation waiting time; for users with less driving experience, the system can extend the display time of key perspectives to reduce operational errors and improve the overall user experience.
[0092] Among them, the user's historical reversing behavior data refers to the operation data recorded by the user during past reversing processes, including steering wheel turning habits, vehicle speed control preferences, and viewing angle switching frequency. For example, if the user is accustomed to frequently making minor adjustments to the steering wheel in the middle of reversing, the system can extend the display time of the side view accordingly.
[0093] For example, by recording users' historical reversing behavior data (such as steering wheel angle habits, vehicle speed control preferences, and viewing angle switching frequency) through the vehicle control system, a user behavior pattern database can be constructed, and the viewing angle switching strategy can be dynamically adjusted.
[0094] The user behavior pattern database identifies users' personalized needs by analyzing historical operation data. For example, for users with less driving experience, the system can extend the display time of key perspectives (such as side and rear views) to reduce operational errors; for experienced users, the system can speed up perspective switching to improve parking efficiency.
[0095] By dynamically adjusting its strategy, the system can adapt to different users' operating habits, improving the adaptability of human-computer interaction. For example, in the initial stage of reversing, the system predicts the side and rear view requirements in advance based on the user's historical steering wheel angle data and actively switches accordingly; in the final stage of reversing, the system can dynamically adjust the display time of the rear view based on the user's vehicle speed control preferences, avoiding incompatibility caused by over-reliance on fixed rules. Ultimately, this achieves personalized optimization of the view switching strategy, balancing safety and efficiency needs, and improving user satisfaction.
[0096] For example, the weights of each sensor data point can be adjusted based on the user's historical reversing behavior data and the confidence levels of each sensor data point. Furthermore, by combining user behavior data with sensor confidence levels, a more precise weight allocation is achieved.
[0097] For example, for users who prefer to reverse quickly, the system can increase the weight of camera data to improve long-distance environmental perception; for less experienced users, the system can increase the weight of ultrasonic radar data to enhance short-distance obstacle detection.
[0098] In some embodiments, the generation of transition frames for the target viewpoint using an interpolation algorithm includes at least one of the following: generating intermediate frames for adjacent viewpoint frames using a linear interpolation algorithm; generating non-linear transition frames using a Bézier curve interpolation algorithm; and adjusting the transition parameters of the interpolation algorithm using the user's historical reversing behavior data.
[0099] Adjacent viewpoints (such as the rear-view front view and the side-rear view, or the left rear-view and the right rear-view) have visual differences. Directly switching between them can cause abrupt changes and stuttering, affecting the user's judgment of the environment. The core advantage of linear interpolation algorithms is their simplicity, efficiency, and low computational cost. They can quickly generate smooth intermediate frames based on the pixel information (or feature information) of two adjacent viewpoints, solving the fundamental problem of image gaps. For example, when switching camera viewpoints while reversing (such as switching from the reversing camera's front view to the 360-degree panoramic side-rear view), linear interpolation can quickly fill in the intermediate transition frame, ensuring visual smoothness and meeting the real-time requirements of in-vehicle systems.
[0100] Linear interpolation produces a smooth, uniform transition, while in actual visual perception, non-linear transitions (such as slow-to-fast transitions or smooth curve transitions) better align with human visual habits and reproduce spatial geometry during viewpoint changes (such as changes in the spatial position of obstacles around a vehicle). Bézier curves (especially quadratic / cubic Bézier curves) offer smooth, controllable, and adjustable curvature characteristics, generating natural non-linear transitions and resolving the abrupt and visually inaccurate transitions of linear interpolation. For example, when reversing, drivers require a high degree of visual continuity; Bézier curve interpolation allows for smoother viewpoint transitions (such as from a side view to a bird's-eye view), preventing drivers from misjudging obstacle positions due to abrupt linear transitions.
[0101] Different drivers have different reversing habits (e.g., some drivers prefer slow, precise adjustments to the parking space, requiring a slower, more detailed transition; others prefer fast reversing, requiring a simpler, faster transition). Furthermore, different scenarios (such as parallel parking and perpendicular parking) also have different requirements for interpolation transitions. Based on users' historical reversing behavior data (such as reversing speed, viewing angle switching frequency, and parking type preferences), the core parameters of the interpolation algorithm (such as the step size of linear interpolation, the control point position of the Bézier curve, and the number of transition frames) can be dynamically adjusted to make the transition effect more aligned with user habits and actual scenario needs. For example, if historical data reveals that a driver frequently performs parallel parking and prefers slow adjustments, the curvature of the Bézier curve will be automatically increased (making the transition from the side view to the BEV view smoother) and the number of intermediate frames increased (preserving more detail). If the driver is found to prefer fast reversing into a parking space, the number of transition frames will be reduced and the interpolation speed increased to improve efficiency.
[0102] For example, an edge computing module can be deployed in an in-vehicle computing platform to offload image rendering tasks required for viewpoint switching (such as panoramic image stitching and interpolation transitions) to the edge computing unit. Through a distributed computing architecture, CNN feature extraction, LSTM stage recognition, and image rendering tasks are processed in parallel, reducing the computational load on the central processing unit. For instance, GPUs can be used to accelerate panoramic image stitching, and FPGAs can be used to implement hardware-level optimization of interpolation algorithms.
[0103] The introduction of edge computing modules can significantly reduce the latency of perspective switching and improve system real-time performance. For example, at the end of the reversing phase, the system can simultaneously process ultrasonic radar data and render the rear view, avoiding screen stuttering caused by competition for computing resources. In angled parking scenarios, the edge computing unit can generate BEV perspective images in real time, reducing user waiting time. Ultimately, this achieves a dual improvement in the smoothness of perspective switching and response speed.
[0104] For example, a multi-level view switching rule base is constructed, dynamically selecting the applicable switching rules based on the complexity of the reversing scenario (such as narrow parking spaces, angled parking, and parallel parking). For instance, in narrow parking space scenarios, the system prioritizes a combination of the side-rear view and the BEV view; in parallel parking scenarios, the system increases its reliance on the rear-view frontal view. The rule base is matched in real-time using a scene classification algorithm (such as CNN-based parking space type recognition) to ensure the scenario adaptability of the switching strategy.
[0105] A multi-level rule base can adapt to the specific needs of different parking scenarios, improving the accuracy of perspective switching. For example, in parallel parking scenarios, the system can switch to a side-rear view in advance to assist in rear wheel trajectory alignment; in angled parking scenarios, the system can dynamically adjust the perspective switching threshold to avoid false triggers caused by changes in the parking space angle. Ultimately, this achieves scenario-based optimization of perspective switching strategies, reducing the need for user intervention.
[0106] The vehicle control method provided in this application embodiment accurately identifies the characteristics of the reversing stage through multi-source sensor data fusion. The state machine logic dynamically matches the optimal viewing angle based on the stage characteristics, thereby accurately and actively switching the viewing angle throughout the reversing process. This avoids the user's attention being distracted due to manually switching the viewing angle, significantly improving the convenience and safety of operation, while reducing the risk of scratches caused by untimely or mismatched viewing angle switching.
[0107] Figure 3 A schematic diagram of the vehicle control device provided in this application is shown below. Figure 3 As shown, the vehicle control device 10 provided in this embodiment includes:
[0108] The first processing module 11 is used to acquire multi-source sensor data, which includes vehicle status data and environmental perception data.
[0109] The second processing module 12 is used to fuse multi-source sensor data based on deep learning algorithms to obtain vehicle reversing stage features, which are used to characterize the current stage of the vehicle during the reversing process.
[0110] The third processing module 13 is used to determine the target viewpoint based on the characteristics of the vehicle reversing stage through state machine logic, and output the display screen of the target viewpoint.
[0111] Optionally, the second processing module 12 is specifically used to extract image features from environmental perception data based on a convolutional neural network;
[0112] Analysis of time-series characteristics of vehicle status data based on long short-term memory network;
[0113] Based on image features and time series features, the characteristics of the vehicle reversing stage are determined.
[0114] Optionally, the third processing module 13 is specifically used to determine the target viewpoint corresponding to the vehicle reversing stage features based on the vehicle reversing stage features and the mapping relationship between the reversing stage and the viewpoint.
[0115] The system generates a transitional image from the target viewpoint based on an interpolation algorithm, and outputs the transitional image and the target viewpoint display image in sequence.
[0116] Optionally, the third processing module 13 is specifically used to determine the target view as the front view and the panoramic view when the vehicle is in the initial positioning stage during the reversing phase feature characterization.
[0117] When the vehicle is in the initial stage of reversing, the target view is determined as the side-rear view.
[0118] When the vehicle is in the middle of reversing, the target viewpoints are determined as the rear frontal viewpoint and the side viewpoint.
[0119] When the vehicle is in the reversing phase, indicating that it is in the final stage of reversing, the target viewpoint is determined as either a rear-view frontal view or a bird's-eye view.
[0120] Optionally, the first processing module 11 is also used to adjust the weight of each sensor data according to the confidence level of each sensor data in the multi-source sensor data;
[0121] And / or, the third processing module 13 is also used to adjust the state machine logic based on the user's historical reversing behavior data.
[0122] Optionally, the third processing module 13 is specifically used to generate intermediate frames of adjacent viewpoint images based on a linear interpolation algorithm;
[0123] Generate non-linear transition images based on Bézier curve interpolation algorithm;
[0124] The transition parameters of the interpolation algorithm are adjusted based on the user's historical reversing behavior data.
[0125] The vehicle control device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0126] Figure 4 This is a schematic diagram of the vehicle control system provided in this application. Figure 4 As shown, the vehicle control system 50 provided in this embodiment includes at least one processor 501 and a memory 502. Optionally, the electronic device 50 further includes a communication component 503. The processor 501, memory 502, and communication component 503 are connected via a bus.
[0127] In a specific implementation, at least one processor 501 executes computer execution instructions stored in memory 502, causing at least one processor 501 to perform the above-described method.
[0128] The specific implementation process of processor 501 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0129] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0130] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0131] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0132] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0133] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0134] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0135] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0136] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0137] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0138] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0139] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0141] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A vehicle control method, characterized in that, include: Acquire multi-source sensor data, including vehicle status data and environmental perception data; The multi-source sensor data is fused using a deep learning algorithm to obtain vehicle reversing stage features, which are used to characterize the current stage of the vehicle during reversing. Based on the characteristics of the vehicle reversing phase, the target viewpoint is determined through state machine logic, and the display screen of the target viewpoint is output.
2. The method according to claim 1, characterized in that, The process of fusing the multi-source sensor data using a deep learning algorithm to obtain vehicle reversing phase features includes: Image features of the environmental perception data are extracted based on a convolutional neural network. The time-series characteristics of the vehicle status data were analyzed based on a long short-term memory network. Based on the image features and time series features, the characteristics of the vehicle reversing phase are determined.
3. The method according to claim 1, characterized in that, The step of determining the target viewpoint based on the vehicle's reversing phase characteristics using state machine logic and displaying the screen from that target viewpoint includes: Based on the vehicle reversing phase characteristics and the mapping relationship between the reversing phase and the viewing angle, the target viewing angle corresponding to the vehicle reversing phase characteristics is determined. The transition screen of the target view is generated based on the interpolation algorithm, and the transition screen and the display screen of the target view are output sequentially.
4. The method according to claim 3, characterized in that, The step of determining the target viewing angle corresponding to the vehicle reversing phase characteristics based on the vehicle reversing phase characteristics and the mapping relationship between the reversing phase and the viewing angle includes: When the vehicle is in the initial positioning stage, the target view is determined to be a front view and a panoramic view, based on the vehicle reversing phase feature characterization. When the vehicle is in the initial stage of reversing, the target view is determined to be a side-rear view. When the vehicle is in the middle of reversing phase, the target viewing angle is determined to be the rear-view frontal view and the side-view. When the vehicle is in the reversing phase, indicating that the vehicle is in the final stage of reversing, the target view is determined to be a rear-view frontal view or a bird's-eye view.
5. The method according to any one of claims 1-4, characterized in that, The method further includes: The weights of each sensor data point are adjusted based on the confidence level of each sensor data point in the multi-source sensor data. And / or, adjust the state machine logic based on the user's historical reversing behavior data.
6. The method according to claim 3, characterized in that, The generation of the transition scene from the target perspective using an interpolation algorithm includes at least one of the following: Intermediate frames between adjacent viewpoints are generated based on a linear interpolation algorithm; Generate non-linear transition images based on Bézier curve interpolation algorithm; The transition parameters of the interpolation algorithm are adjusted based on the user's historical reversing behavior data.
7. A vehicle control system, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-6.
8. A vehicle comprising the vehicle control system of claim 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-6.