AVP vehicle state dynamic updating method and system based on user attention perception

By dynamically updating the method in the automated valet parking system, acquiring and fusing multi-dimensional data sequences, optimizing communication bandwidth allocation and predicting takeover results, the problems of fixed communication resource allocation and delayed takeover intent perception in existing technologies are solved, thereby improving the system's security and response speed.

CN122135589APending Publication Date: 2026-06-02CHENGDU YIBO INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU YIBO INFORMATION TECH CO LTD
Filing Date
2026-04-24
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In automated valet parking systems, existing technologies suffer from fixed communication resource allocation and delayed perception of takeover intent, leading to network link congestion and response time delays. This makes it difficult to achieve high-fidelity early warnings and accurate judgment of user intervention intentions, increasing the risk of vehicle collisions.

Method used

By acquiring the environmental sequence, attention sequence, scaling sequence, and trajectory sequence during the valet parking process, dwell features and distortion features are extracted, and attention coefficients are obtained by fusing them. The quantization step size of the data stream is dynamically adjusted in conjunction with risk indicators, and the takeover result is predicted using a nonlinear mapping function to optimize communication bandwidth allocation and early braking pre-charging.

Benefits of technology

It achieves high-fidelity presentation of user expectations and risks in environments with limited communication resources, reduces the collision risk caused by response delays, and improves safety and communication resource utilization efficiency in non-line-of-sight valet parking scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135589A_ABST
    Figure CN122135589A_ABST
Patent Text Reader

Abstract

The application discloses a kind of AVP vehicle state dynamic updating method and system based on user attention perception, it is related to electric digital data processing technical field, method includes: obtaining the environment sequence of guest parking stage, the attention sequence of parking parameter, scaling sequence and trajectory sequence;Based on the attention and scaling sequence extraction resident and distortion feature to fuse and obtain attention coefficient, and the edge detection of environment sequence is carried out to obtain risk index;According to risk index and attention coefficient determine demand coefficient, input the quantization step of encoding closed loop dynamic adjustment data stream;Extract the radial component of trajectory sequence to screen takeover benchmark to calculate offset degree to obtain intervention trend, combine nonlinear mapping to fuse intervention trend and demand coefficient to obtain takeover result.The application has the advantages of interactive intention accurate perception, communication bandwidth self-adaptation and pre-safety defense.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing technology, specifically to a method and system for dynamically updating the status of AVP vehicles based on user attention perception. Background Technology

[0002] In the current Automated Valet Parking (AVP), the remote monitoring and status feedback mechanism suffers from fixed communication resource allocation and a lag in the perception of takeover intent.

[0003] Specifically, in operating environments such as underground parking garages with no line of sight or narrow parking spaces with numerous blind spots, existing systems typically rely on fixed-frequency data polling mechanisms or uniform real-time transmission of the entire video stream. This data delivery strategy is prone to network congestion in weak network environments where the communication bandwidth between the vehicle and the cloud is limited. This results in blurry and stuttering images of sudden risks (such as pedestrians or load-bearing pillars in blind spots) due to high compression rates or network latency, making it difficult to achieve on-demand high-fidelity warnings. At the same time, the current intelligent parking technology's judgment logic for user intervention intentions is relatively passive. The system is limited to waiting for the user to press the physical or virtual takeover button on the screen before issuing a braking command to the chassis, without fully considering the user's zoom-in actions on the monitoring interface (field distortion) and the sliding trajectory of the finger approaching the takeover button in the two-dimensional space of the screen. This open-loop monitoring model, which separates objective risks in the physical environment from subjective expectations of users at the monitoring end, results in limited communication bandwidth resources being unable to be dynamically allocated to high-risk monitoring areas. This shortens the response time of the parking system when facing emergencies, making it difficult to complete pre-emptive risk interception and braking response before the user actually touches the takeover button. This increases the risk of vehicle collisions in complex parking environments and affects the reliability of human-machine collaborative interaction. Summary of the Invention

[0004] In view of the technical problems described in the background art, the present invention provides a method and system for dynamic updating of AVP vehicle status based on user attention perception.

[0005] A method for dynamic updating of AVP vehicle status based on user attention perception is proposed. The method includes: acquiring the valet parking process and acquiring the environmental sequence, attention sequence, scaling sequence, and trajectory sequence of parking parameters within the corresponding stage; extracting dwell features and distortion features based on the attention sequence and scaling sequence, and fusing them to obtain the attention coefficient; simultaneously performing edge detection based on the environmental sequence corresponding to the current stage to obtain risk indicators; obtaining the demand coefficient of parking parameters based on the risk indicators and attention coefficient, and inputting the demand coefficient into the encoding closed loop to dynamically adjust the quantization step size of the corresponding data stream; acquiring a takeover benchmark that is not related to the parking parameters in the screen coordinate system, extracting the radial component of the trajectory sequence towards the takeover benchmark, and calculating the offset to obtain the intervention trend; and then fusing the intervention trend and the demand coefficient through a nonlinear mapping function to obtain the takeover result.

[0006] Optionally, based on the attention sequence and the scaling sequence, dwell features and distortion features are extracted and fused to obtain the attention coefficient, including: extracting the difference between the maximum dwell time and the duration threshold in the attention sequence as the first feature; calculating the discrete gradient of adjacent duration values ​​in the attention sequence and taking the algebraic sum of all discrete gradients within the time window to obtain the second feature; obtaining the dwell feature based on the superposition result of the first feature and the second feature; calculating the rate of change of adjacent scaling factors in the scaling sequence and extracting the spatial derivative extremum as the distortion feature; converting the distortion feature into intent gain and multiplying the dwell feature with the intent gain to obtain the attention coefficient.

[0007] Optionally, edge detection is performed based on the environmental sequence corresponding to the current stage to obtain risk indicators, including: calculating the absolute difference between adjacent distance values ​​in the environmental sequence; if any absolute difference is greater than a safety threshold, an environmental step is triggered, and the first value is assigned to the risk indicator; if none of the absolute differences are greater than the safety threshold, no environmental step is triggered, and the second value is assigned to the risk indicator.

[0008] Optionally, the demand coefficient is input into the coding closed loop to dynamically adjust the quantization step size of the corresponding data stream, including: multiplying the demand coefficient by a dimensional constant to convert it into a pure number, and then inputting it into a logarithmic function to obtain the composite gain; constructing a compressed denominator based on the composite gain and the balance coefficient; and generating a non-linearly reduced quantization step size by dividing the base step size by the compressed denominator.

[0009] Optionally, the radial component of the trajectory sequence toward the takeover reference is extracted, and the offset is calculated to obtain the intervention trend. This includes: connecting the historical touch points in the trajectory sequence with the takeover reference in the screen coordinate system to form a reference vector; calculating the displacement vectors of two adjacent historical touch points and projecting the displacement vectors onto the corresponding reference vectors to obtain the radial component; defining the magnitude of the radial component pointing toward the takeover reference as the approach value and the magnitude of the component deviating from the reference as the divergence value; summing the approach value and the divergence value within the time window and using the accumulated result as the intervention trend.

[0010] Optionally, the intervention trend and demand coefficient are fused through a nonlinear mapping function to obtain the takeover result, including: if several pre-trends are greater than or equal to zero, then no takeover demand is directly output as the takeover result; if several pre-trends are less than zero, then the absolute value of the intervention trend, the demand coefficient, the control coefficient, and the dimensional constant are multiplied together to construct a penalty index; the penalty index is input into the probability activation function, and the value of the mapping output is used as the takeover result.

[0011] A dynamic update system for AVP vehicle status based on user attention perception is also provided, including: an acquisition module for acquiring the valet parking process and acquiring the environmental sequence, attention sequence, scaling sequence, and trajectory sequence of parking parameters within the corresponding stage; an evaluation module for extracting dwell features and distortion features based on the attention sequence and scaling sequence, and fusing them to obtain the attention coefficient; simultaneously performing edge detection based on the environmental sequence corresponding to the current stage to obtain risk indicators; a feedback module for obtaining the demand coefficient of parking parameters based on the risk indicators and attention coefficient, and inputting the demand coefficient into the coding closed loop to dynamically adjust the quantization step size of the corresponding data stream; and a prediction module for acquiring a takeover benchmark that is not related to the parking parameters in the screen coordinate system, extracting the radial component of the trajectory sequence to the takeover benchmark, and calculating the offset to obtain the intervention trend; subsequently, the intervention trend and the demand coefficient are fused through a nonlinear mapping function to obtain the takeover result.

[0012] Optionally, the evaluation module is also used to: extract the difference between the maximum dwell time and the duration threshold in the attention sequence as a first feature; calculate the discrete gradient of adjacent duration values ​​in the attention sequence, and perform an algebraic sum of all discrete gradients within the time window to obtain a second feature; obtain the dwell feature based on the superposition result of the first feature and the second feature; calculate the rate of change of adjacent scaling factors in the scaling sequence, and extract the spatial derivative extrema as a distortion feature; convert the distortion feature into intent gain, and multiply the dwell feature by the intent gain to obtain the attention coefficient.

[0013] Optionally, the evaluation module is also used to: calculate the absolute difference between adjacent distance values ​​in the environmental sequence; if any absolute difference is greater than the safety threshold, an environmental step is triggered and the first value is assigned to the risk indicator; if none of the absolute differences are greater than the safety threshold, no environmental step is triggered and the second value is assigned to the risk indicator.

[0014] Optionally, the feedback module is also used to: convert the demand coefficient into a pure number by multiplying it with the dimensional constant, and then input the logarithmic function to obtain the composite gain; construct a compressed denominator based on the composite gain and the balance coefficient; and generate a nonlinearly reduced quantization step size by dividing the base step size by the compressed denominator.

[0015] The beneficial effects of this invention are reflected in: In the entire AVP vehicle status dynamic update method based on user attention perception, firstly, by constructing an attention coefficient that includes temporal dwell gradient and spatial field of view distortion, the user's zooming in on the screen is quantified into a mathematical weight, improving upon the existing method of judging attention solely based on dwell time and achieving perception of user expectations. Furthermore, the quantified user attention is product-coupled with the risk of the physical environment to generate a comprehensive feedback demand coefficient, which is then introduced into the quantization step size adjustment closed loop of cloud source coding. This optimizes the bandwidth allocation mechanism under limited communication resources, enabling the system to respond to risks or increases in user attention. At the same time, the encoder is instructed to reduce the compression ratio of the video stream in a specific area and allocate higher transmission bandwidth to ensure the faithful presentation of risky scenes. Furthermore, by extracting the radial projection component and intervention trend of the user's touch trajectory approaching the takeover virtual control in the two-dimensional coordinate system of the screen, and using nonlinear mapping to integrate it with environmental risks, the takeover probability can be predicted before the user's finger actually touches the takeover button. This provides the chassis domain controller with time for early braking and pre-charging, reduces the collision risk caused by reaction delay, constructs a front-end safety defense closed loop, and improves the safety and communication resource utilization efficiency in non-line-of-sight valet parking scenarios. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0017] Figure 1 This is a schematic diagram illustrating the steps of the AVP vehicle status dynamic update method based on user attention perception of the present invention. Figure 2 This is a schematic diagram of a portion of step S1 in the AVP vehicle status dynamic update method based on user attention perception of the present invention. Figure 3 This is a schematic diagram of a portion of step S2 in the AVP vehicle status dynamic update method based on user attention perception of the present invention. Figure 4 This is a schematic diagram of part of step S3 in the AVP vehicle status dynamic update method based on user attention perception of the present invention. Figure 5 This is a schematic diagram of part of step S4 in the AVP vehicle status dynamic update method based on user attention perception of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0020] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0021] This invention provides a method for dynamically updating the AVP vehicle status based on user attention awareness, such as... Figure 1 As shown, in one embodiment, the method includes: S1. Acquire the target valet parking process, which consists of multiple consecutive automatic valet parking stages, and acquire the environmental sequence, the focus sequence of parking parameters, the scaling sequence, and the trajectory sequence from the previous acquisition cycle before the current acquisition cycle.

[0022] S2. Extract the dwell features in the time dimension and the distortion features in the spatial dimension based on the attention sequence and the scaling sequence, and fuse them to obtain the attention coefficient; at the same time, obtain the automatic valet parking stage to which the parking parameters belong at the current moment, and perform edge detection based on the environmental sequence corresponding to the current stage to obtain risk indicators.

[0023] S3. Obtain the demand coefficient of parking parameters based on risk indicators and attention coefficients, and input the demand coefficient into the coding closed loop to dynamically adjust the quantization step size of the corresponding data stream.

[0024] S4. Obtain a takeover baseline in the screen coordinate system that is not related to parking parameters, extract the radial component of the trajectory sequence to the takeover baseline, and calculate the offset to obtain the intervention trend; then, fuse the intervention trend with the demand coefficient through a nonlinear mapping function to obtain the takeover result.

[0025] In this embodiment, it should be noted that in S1, during the valet parking monitoring process, the first step is to perform the synchronous acquisition of multi-dimensional data sequences and scene initialization. This stage aims to solve the fundamental problem of existing data collection having a single dimension and lacking spatiotemporal alignment.

[0026] Specifically, the clock synchronization between the vehicle and mobile terminal is completed based on the network time server of the cloud platform, and a two-dimensional coordinate system is established on the user terminal screen. During the set data collection cycle of the parking and turning phase, the sensors deployed on the vehicle body detect the distance of the obstacle to the right rear at a fixed frequency, generating an environmental sequence containing distance values ​​of 180, 175, 170 and 140, in centimeters.

[0027] Simultaneously, the terminal interface records the user's observation time of the right rear wheel blind spot image, forming an attention sequence containing values ​​such as 1000, 1200, and 1500, in milliseconds. It also captures the user's two-finger zoom-in screen operation, recording a zoom sequence with a maximum zoom rate of 2.0, and tracks the coordinate movement trajectory sequence of the fingers on the screen. By constructing a comprehensive data matrix covering environmental physical distance, user interface dwell time, field-of-view zoom level, and touch movement trajectory, a correlation mapping between physical space and digital interaction space is established. This parallel acquisition method of multi-source data changes the previous monitoring mode that relied solely on a single sensor or button command, providing a structured data foundation for accurately quantifying the user's psychological state and assessing environmental step risks. This enables intelligent parking to acquire data through two-way monitoring of environmental perception and human-machine interaction.

[0028] In S2, based on the acquired underlying data, the process then moves to feature extraction and risk assessment. This step focuses on addressing the issue that a single-dimensional threshold cannot accurately measure user intervention expectations and environmental abrupt changes. The current stage's duration threshold is set to 1000. The maximum duration of 1500 in the attention sequence is extracted, and the difference is calculated to obtain a first feature of 500. Simultaneously, the differences between adjacent duration values ​​of 1200 and 1000, and 1500 and 1200, are calculated and summed to obtain a second feature of 500. These two are added together to obtain a base time attention score of 1000. To introduce a spatial dimension, a scaling rate of 2.0 is extracted and combined with a time constant of 0.5. A spatial distortion gain is generated through an exponential relationship, amplifying the base time attention score to 3718, thereby quantifying the potential intervention intent hidden when a user rapidly zooms in on the screen.

[0029] Simultaneously, edge detection is performed on the environmental sequence, calculating the absolute differences between adjacent values ​​as 5, 5, and 30. Since the last difference of 30 exceeds the preset safety threshold of 20, a step change in the physical environment is determined, and the value 1 is assigned to the risk indicator. This calculation logic transforms the user's implicit intervention intention into a quantified attention coefficient and converts the approach of obstacles in the physical space into discrete risk markers. This overcomes the limitation of separating risk and attention in existing open-loop monitoring, achieving dual sensitivity to high-risk environments and high-intensity interaction states, and improving the accuracy and responsiveness of safety assessments in complex parking scenarios.

[0030] In S3, after obtaining the quantized risk and attention status, the adaptive adjustment stage of source coding is entered to address the transmission delay and blurring issues of high-value risk images caused by bandwidth equalization in weak network environments. First, the attention coefficient of 3718 is multiplied by the risk index of 1 and the base attention coefficient is added to calculate a demand coefficient of 7436. This coefficient represents the urgency of data delivery to a specific monitored area under sudden environmental changes and high user attention. Subsequently, the cloud platform uses this demand coefficient to adjust the compression rate of the video stream. Using an initial delivery step size of 32 as the numerator, a logarithmic compression denominator is constructed, containing a balance coefficient of 2.0 and a dimension alignment coefficient of 0.001. After non-linear dimensionality reduction using the logarithmic function, the original quantization step size of 32 is dynamically compressed to 6. This significant reduction in the quantization step size indicates that the underlying encoder has reduced the data compression rate of the video stream in this blind spot, thereby allocating higher priority transmission bandwidth to it within the limited communication link.

[0031] This closed-loop control logic changes the existing data distribution strategy of uniform transmission across the entire dataset, achieving adaptive matching between monitoring image quality and network bandwidth resources. In emergency situations, it can ensure low-latency and high-fidelity transmission of images in high-risk areas, allowing users to clearly observe approaching obstacles and improving the effectiveness of remote monitoring information and the utilization rate of communication resources.

[0032] In S4, after ensuring clear transmission of the risky image, intervention intent prediction based on interaction trajectory is performed to overcome the braking command lag problem caused by existing methods relying on physical button triggers. Within the terminal screen coordinate system, using the emergency brake control as a reference, the user's finger swipe trajectory is converted into a radial projection. As the user's finger continuously approaches the brake control, these negative approach values ​​are accumulated within a time window, obtaining an intervention trend with an absolute value of 400. This value objectively reflects the urgency of the finger's convergence in physical space.

[0033] Subsequently, the intervention trend of 400, the demand coefficient of 7436, and the dimensional constant of 0.000001 are multiplied by the control coefficient of 0.8 to construct a penalty index. After mapping with a probabilistic activation function, this index outputs a takeover result of 0.915. This calculation logic indicates that an intervention probability of 0.915 has been predicted based on the trajectory deviation before the user's finger actually touches the brake button on the screen. Relying on this high-confidence pre-prediction, the cloud can issue a brake pre-charge command to the vehicle chassis domain controller in advance without waiting for the final mechanical touch control loop. This pre-defense mechanism compensates for the physiological reaction delay between human visual perception and action execution, buys valuable braking redundancy time for valet parking, reduces the probability of collisions with nearby obstacles, and enhances the human-machine collaborative safety of intelligent parking.

[0034] In summary, the entire AVP vehicle status dynamic update method based on user attention perception firstly quantifies the user's zooming in on the screen into mathematical weights by constructing an attention coefficient that includes temporal dwell gradient and spatial field-of-view distortion. This improves upon existing methods that rely solely on dwell time to judge attention and achieves perception of user expectations. Furthermore, the quantified user attention is coupled with the risk of the physical environment through a product to generate a comprehensive feedback demand coefficient. This coefficient is then introduced into the quantization step size adjustment closed loop of cloud-based source coding, optimizing the bandwidth allocation mechanism under limited communication resources. This allows for the response to risk events or increases in user attention. When the system is activated, the encoder is instructed to reduce the compression ratio of the video stream in a specific area and allocate higher transmission bandwidth to ensure the faithful presentation of risky scenes. Furthermore, by extracting the radial projection component and intervention trend of the user's touch trajectory approaching the takeover virtual control in the two-dimensional coordinate system of the screen, and using nonlinear mapping to integrate it with environmental risks, the probability of takeover can be predicted before the user's finger actually touches the takeover button. This provides the chassis domain controller with time for early braking and pre-charging, reduces the collision risk caused by reaction delay, constructs a front-end safety defense closed loop, and improves the safety and communication resource utilization efficiency in non-line-of-sight valet parking scenarios.

[0035] like Figure 2As shown, in one specific implementation, S1 includes: S11, a synchronization clock and coordinate system. Using the network time protocol server on the cloud platform as a reference, synchronization messages are sent to waiting vehicles and user terminals, setting a basic data collection cycle length. Simultaneously, a screen coordinate system is established on the user terminal device to map terminal interaction behavior.

[0036] S12. Divide the valet parking process into stages. By reading the real-time vehicle speed and turning angle data of the vehicle chassis, the valet parking process is divided into the cruising to find a parking space stage, the parking and turning stage, and the parking space correction stage.

[0037] S13. Construct a multi-dimensional data sequence. Activate the surround-view sensors around the waiting vehicle to detect obstacle distances at a fixed frequency, generating an environmental sequence composed of distance values ​​from the previous collection period. Activate the interactive tracking interface within the user terminal, targeting the data displayed on the screen. Parking parameters are used to record user time data for the area to form an attention sequence; zoom data is recorded when the user performs zoom operations to form a zoom sequence; and the coordinate movement path of the touch point in the screen coordinate system is recorded to form a trajectory sequence.

[0038] In this embodiment, it should be noted that in S11, during the automatic valet parking monitoring process, a synchronization clock and coordinate system are first established. This step aims to resolve the spatiotemporal misalignment issue between the vehicle's physical sensors, the cloud monitoring platform, and the user terminal. Using the network time protocol server on the cloud platform as a reference, synchronization messages are sent to the waiting vehicle and the user terminal, setting a basic acquisition cycle length to ensure that all data frames processed subsequently have a unified timestamp. Simultaneously, an independent two-dimensional pixel coordinate system is established on the user terminal device, with the upper left corner of the screen as the origin, the rightward direction as the positive horizontal axis, and the downward direction as the positive vertical axis. This alignment of the spatiotemporal reference provides a benchmark for accurately capturing the user's finger swipe trajectory on the screen and the corresponding changes in the external environment. By eliminating timing errors between distributed systems, this step ensures the accuracy of subsequent multi-dimensional data sequence fusion calculations within the same time slice, laying the foundation for low-latency closed-loop feedback and accurate intent prediction.

[0039] In S12, following the establishment of time and space benchmarks, the valet parking phase is divided. Its core purpose is to match differentiated risk assessment thresholds based on the vehicle's current motion state. The onboard computing unit of the waiting vehicle reads real-time vehicle speed data and steering wheel angle data from the chassis, using these underlying kinematic parameters to divide the entire valet parking process into a cruising to find a parking space phase, a parking maneuvering phase, and a parking space correction phase. Taking the parking maneuvering phase as an example, the available physical space around the vehicle is drastically compressed, increasing the sensitivity to lateral obstacles. By dividing the continuous parking process into discrete phases with clear physical semantics, the corresponding environmental safety thresholds and interaction dwell time thresholds can be dynamically invoked for different phases in subsequent data processing. This division logic optimizes the fixed setting of a single global threshold, enhances the algorithm's adaptability to complex parking environments, and improves the scenario fit and decision-making targeting of risk warnings.

[0040] In S13, after confirming that the current stage is the parking maneuvering phase, S13 is executed to construct a multi-dimensional data sequence. This aims to address the technical problem that a single sensor cannot comprehensively characterize the overall risks of a parking scenario. The surround-view sensor on the right rear side of the vehicle waiting to park is activated to detect obstacle distances at a fixed frequency, generating data such as... (The sentence is incomplete and requires more context to translate accurately). The environmental sequence is measured in centimeters. Simultaneously, the interactive tracking interface within the user terminal is activated, recording the user's temporal data for the "right rear wheel blind spot video stream" parking parameter on the screen, forming data such as... The attention sequence in milliseconds; simultaneously recording the maximum zoom rate when the user performs a two-finger zoom operation. The system records the scaling sequence and the trajectory sequence formed by the movement of the touch point coordinates. This mechanism, which synchronously collects vehicle-side physical space ranging and end-user digital interaction behavior, constructs a heterogeneous data matrix that includes changes in the objective environment and the subject's psychological state, providing underlying data support for subsequent quantification of risk and interaction intent.

[0041] It should also be noted that the acquisition period and the fixed frequency of the sensor in S1 are determined based on a comprehensive evaluation of a limited number of parking physical tests and the capacity of the communication link. Specifically, multiple (e.g., 300) automatic parking experiments at typical vehicle speeds (e.g., 5 km / h) are conducted in a closed test field. The minimum sampling rate required to completely reconstruct the relative motion trajectory of obstacles is evaluated using the Nyquist sampling theorem. Simultaneously, the average data packet transmission delay of the vehicle-to-cloud network under weak network conditions is considered for constraint optimization. For example, test data shows that when the vehicle is traveling at the parking speed limit, the theoretical minimum ranging frequency needs to reach 10Hz (i.e., a period of 100 milliseconds) to identify the displacement of a small obstacle with a width of 10 cm. Considering the average delay of 20 milliseconds due to network jitter, the system amplifies the sampling margin. Through a trade-off between computing power and communication bandwidth, the fixed frequency of the sensor is ultimately determined to be 20Hz, i.e., a single acquisition period is preset to 50 milliseconds, to ensure the fidelity of the underlying data sequence and the real-time nature of the warning.

[0042] like Figure 3 As shown, in one specific implementation, S2 includes: S21, extracting basic features of parking. Extract the sequence of interest and set a duration threshold suitable for the current parking stage. Traverse the parking duration values ​​in the sequence; if there is data exceeding the duration threshold, extract the difference between the maximum parking duration value and the duration threshold as the first feature; if not, set the first feature to zero.

[0043] The specific value of the duration threshold is not arbitrarily set, but is obtained based on statistical analysis of historical user eye movement and interaction test data from a large number of real valet parking scenarios. Specifically, the system pre-collects user screen observation data from a limited number of (e.g., 500) parking maneuvers, extracting the average observation duration before any risk intervention and the average gaze duration before any risk intervention. Subsequently, a Gaussian mixture model is used to perform cluster analysis on these two duration distributions to find the critical point where their probability density functions intersect, and a certain safety margin is added to this as the final duration threshold. For example, statistical analysis of historical data shows that the average duration of normal saccades is 600 milliseconds, while the average duration of gazes that generate anxiety expectations is 1300 milliseconds. The model calculates that the distinguishing threshold between the two is 950 milliseconds. After adding a 50-millisecond margin, the duration threshold for the current parking stage can be scientifically determined to be 1000 milliseconds, thus ensuring the objectivity and accuracy of the extracted dwell features.

[0044] S22. Extracting Dwell Gradient Features. Calculate the difference between the dwell time value of the next acquisition point and the dwell time value of the previous point in the sequence, defining it as the discrete gradient. Calculate the algebraic sum of all discrete gradients within the time window to obtain the second feature. If the second feature is less than zero, add the first feature to zero to obtain the base time-of-interest value; if the second feature is not less than zero, add the first feature to the second feature to obtain the base time-of-interest value. The combination of the above values ​​constitutes the dwell feature.

[0045] The length of the time window was determined based on a limited number of historical interaction data points related to human short-term visual memory and touch continuity. Specifically, a limited batch (e.g., 800 groups) of simulated emergency intervention operation logs from users on in-vehicle or mobile devices were collected. The average duration of the complete psychophysical action chain—from the moment the user's anxious gaze to the point where their finger continuously slid towards the brake control—was extracted. A distribution fitting algorithm was then used to find time cutoff points covering over 95% of the continuous operational intentions.

[0046] S23. Obtaining the attention coefficient by fusing spatial distortion features. The scaling sequence is retrieved, the rate of change of scaling factor between adjacent sampling points is calculated, and the extreme values ​​of its spatial derivatives are extracted as distortion features. Subsequently, the dwell features and distortion features are fused to obtain the corrected attention coefficient. Its mathematical expression is:

[0047] In the formula, each character is defined as follows: For the first time after merging spatial distortion The attention factor for each parking parameter is measured in milliseconds. The dwell characteristics are calculated based on S21 and S22, with the physical unit being milliseconds; It is a natural constant; The distortion features extracted from the scaling sequence represent the scaling rate, in reciprocals of seconds. This is a preset time constant, in seconds, used as the reciprocal unit of time to compensate for distortion features.

[0048] Wherein, time constant The value is determined by fitting and calibrating historical user zooming operations with actual psychological urgency levels. Specifically, during the development phase, testers were invited to face sudden situations of varying urgency in a simulator, simultaneously recording their maximum zooming rate sequence when zooming the screen with two fingers, and their subsequent psychological urgency scores (0-10 points) from a questionnaire assessment. Then, the least squares method was used to perform exponential curve fitting between the collected finite group (e.g., 300 groups) of zooming rates and their corresponding subjective urgency scores, solving for the exponential factor that minimizes the sum of squared errors, and establishing it as the time constant. For example, after collecting a batch of samples with maximum zooming rates of 1.5, 2.0, and 2.5 corresponding to urgency scores of 5, 7, and 9, the fitting calculation yielded the following results: When the value is 0.5, the output trend of the exponential term best matches the nonlinear growth gradient of the psychological score. Therefore, 0.5 is set as the preset time constant to ensure the mathematical rationality of feature fusion.

[0049] S24. Perform edge detection on the environmental sequence. Calculate the absolute difference in distance values ​​between adjacent obstacles in the environmental sequence. Compare this absolute difference with the preset safety threshold corresponding to the current stage.

[0050] The safety threshold is not an empirical blind test value, but is calculated based on historical test data of the vehicle's chassis dynamic limits and sensor signal-to-noise ratio. Specifically, the calculation process involves first conducting a limited number (e.g., 200) of obstacle approach braking experiments at different vehicle speeds in a closed test area, recording the maximum physical braking distance from receiving the braking command to complete stop. Simultaneously, the extreme value of the background noise fluctuation of this type of surround-view sensor in a static, stable environment is statistically analyzed. The maximum braking distance is added to the sensor's background noise extreme value, and then multiplied by a redundancy coefficient of 1.2 to obtain the safety threshold for that stage. For example, if historical calibration data shows that the maximum braking distance at a typical parking speed is 12 cm, and the sensor's maximum inherent ranging noise error is 4.6 cm, the system multiplies the sum of these two values ​​by 1.2, resulting in (12 + 4.6) × 1.2 = 19.92 cm, and then rounds up to 20 cm as the preset safety threshold. This accurately defines environmental steps while eliminating normal noise fluctuations.

[0051] S25. Determine the environmental step state. If any absolute difference is greater than the safety threshold, an environmental step is triggered, and the first value is assigned to the risk indicator.

[0052] S26. Determine the steady-state environmental condition. If all absolute differences are not greater than the safety threshold, the environmental step is not triggered, and the second value is assigned to the risk indicator.

[0053] In this embodiment, it should be noted that in S21, based on the multidimensional data obtained above, basic features of dwell time are extracted. This step is used to initially filter out normal visual saccades of the user and extract excessive dwell behaviors that represent anxiety. This is done within the attention sequence. , and The value is in milliseconds, and the duration threshold for the adaptation and redirection phase is set to [value missing]. Milliseconds. By traversing the sequence, the maximum dwell time was found. The maximum dwell time exceeded the set duration threshold in milliseconds. The difference between this maximum dwell time and the duration threshold is extracted. The first feature is the millisecond. If none of the values ​​in the sequence exceed the threshold, this feature is set to zero. This setting shifts the benchmark for attention assessment from absolute duration to relative over-limit duration, eliminating data redundancy from routine monitoring. It allows computing resources to focus on potential risk points that cause users to stare for extended periods, providing a numerical basis for quantifying users' basic psychological expectations.

[0054] In S22, to compensate for the limitation of static thresholds in detecting the gradual accumulation of user anxiety, S22 is executed to extract dwell gradient features. The difference between adjacent acquisition points in the attention sequence is calculated and defined as the discrete gradient, i.e. milliseconds and Milliseconds. The algebraic sum of all discrete gradients within the time window is used to obtain the second feature. Milliseconds. Since this second feature is greater than zero, it indicates that the user's attention to this area is continuously increasing, thus the first feature... With the second feature Add them together to get the base time focus value. Milliseconds. Through the computational logic of differentiation and integration, the dynamic growth trend of user gaze duration was successfully captured. This approach can identify anxious observation behaviors that are frequent and gradually lengthening, even if they do not exceed the timeout limit individually, thus improving the smoothness and sensitivity of the attention assessment model in the time dimension.

[0055] In S23, given that a single time dimension cannot fully reflect the urgency of user intervention on the touchscreen device, S23 integrates spatial distortion features to obtain the final attention coefficient. Specifically, in the interaction intent quantification stage of automated valet parking monitoring, to address the technical problem that a single time dimension cannot accurately characterize the instantaneous spatial focus expectation of users due to limited field of vision, a mathematical expression is used... The local scaling operation of a 2D screen is converted into a computable intention gain multiplier.

[0056] In this expression, The attention coefficient represents the fused spatial distortion, and its physical unit is milliseconds. It is the basic time attention extracted from the user's gaze duration, which serves as the base data for evaluation; The physical significance of introducing the exponential function, which is a natural constant, is that the user's rapid zooming action with two fingers usually represents a nonlinear surge in the expected intervention. Linear multiplication cannot truly reflect this rapid intervention impulse, while the exponential term can give appropriate weight sensitivity to the instantaneous distortion rate. The spatial derivative extrema extracted from the scaling sequence represent the maximum scaling rate of the screen, expressed in reciprocals of seconds. This is the preset scaling sensitivity time constant, in seconds. Its existence is not only used to adjust the steepness of the exponential curve, but also, at the physical dimension level, related to... The reciprocal units of time are multiplied to cancel each other out, ensuring that the power of the exponent is a dimensionless pure number, satisfying rigorous mathematical logic; the constant 1 in parentheses serves as a baseline, ensuring that the basic time focus is preserved without being excessively attenuated in the absence of scaling or with a very small scaling factor.

[0057] Based on specific data from underground parking garage entry scenarios, when the base time attention... At milliseconds, a user was detected rapidly zooming in due to poor visibility, and the maximum zoom rate was recorded. Combined with a preset time constant seconds, the calculation process is expanded to Milliseconds. Through this nonlinear fusion calculation, the original 1000 milliseconds of the normal time dwell time is amplified to 3718 milliseconds by a physical amplification action with an amplitude of 2.0. This establishes a collaborative mapping mechanism between spatial interaction and time dwell time in the underlying mathematical architecture. Furthermore, in practical terms, it assigns numerical weights to the target monitoring screen that match the intensity of user attention, providing a discriminative input basis for subsequent adaptive scheduling of communication bandwidth.

[0058] In S24, while assessing the user's subjective intent, edge detection is performed on the environmental sequence. The aim is to identify abrupt changes posing a collision threat from continuous ranging data. This involves retrieving the environmental sequence... , , and The values ​​in centimeters are used to calculate the absolute differences in distance to adjacent obstacles, respectively. , and Centimeters. Then, these absolute differences are compared with the preset values ​​for the current warehousing stage. The system compares data against centimeter-level safety thresholds. This high-frequency differential comparison effectively filters out subtle distance fluctuations caused by slight vehicle movement or inherent sensor noise, focusing only on data points where the first-order spatial derivative changes drastically. This allows for the detection of scenarios such as sudden approach of load-bearing pillars or dynamic pedestrians in complex garage environments, providing a direct quantitative benchmark for assessing the threat level of the physical environment.

[0059] In S25, when the differential comparison in S24 meets the warning condition, the environmental step state is determined and risk indicators are assigned values. This is due to the last absolute difference in the environmental sequence. centimeters larger than the preset A centimeter-level safety threshold is used to determine if a step change has occurred in the physical environment. At this point, an environmental step flag is triggered, and a preset dimensionless Boolean constant is set. The value is assigned to the risk indicator. This logic distills the complex changes in physical spatial distance into a concise binary judgment result. It strips away the specific absolute distance values, retaining the core information of whether a risk has occurred. By discretizing continuous physical distance signals into explicit risk Boolean values, it eliminates dimensional interference for the subsequent mathematical fusion of multi-source heterogeneous data, ensuring that a definite and interference-resistant trigger signal can be output to the upper-level control logic when an obstacle approaches.

[0060] In S26, to ensure stability under normal security conditions and avoid the consumption of communication resources by redundant data processing, S26 is set to handle the steady-state environment. If the absolute difference between all adjacent distance values ​​calculated in S24 is not greater than... A safety threshold of centimeters implies that the physical environment surrounding the vehicle is in a stable and predictable state of change. Under this condition, an environmental step flag is not triggered, and a preset dimensionless Boolean constant is applied. Assign a value to the risk indicator. This decision branch handles situations where sensor data fluctuates normally within a safe range. Assign a value to the risk indicator. This ensures that environmental parameters are zeroed out during subsequent fusion computing, maintaining data feedback frequency at the basic cruising level. This not only reduces unnecessary communication load between the cloud and the vehicle but also guarantees the logical integrity of the monitoring and feedback mechanism under different risk levels.

[0061] like Figure 4 As shown, in one specific implementation, S3 includes: S31, quantifying the demand coefficient. The attention coefficient is used as the base value, and the product of the attention coefficient and the risk indicator is calculated as the environmental added value. The base value and the environmental added value are added to obtain the demand coefficient for parking parameters.

[0062] S32. Adaptive Adjustment of Quantization Step Size. The demand coefficient is incorporated into the encoder's quantization parameter calculation module to dynamically generate the target quantization step size. Its mathematical expression is:

[0063] In the formula, each character is defined as follows: The quantization step size is assigned to the target of the current corresponding data stream and is a dimensionless integer. The basic quantization step size is a dimensionless integer. This is a preset dimensionless equilibrium coefficient; It is the natural logarithm function; The demand coefficient obtained for S31, in milliseconds; is the dimensional alignment factor, in milliseconds (reciprocal). This formula forces a smaller quantization step size, reducing the compression ratio and achieving bandwidth tilting when feedback requirements increase.

[0064] Among them, the balance coefficient Alignment coefficient with dimensions Its value is based on the throughput limit test data of the cloud encoder in a weak network environment. By setting different bandwidth limits in the network simulator, a limited number of simulated video streams with different demand coefficients are continuously input (e.g., 400 times), and the relationship between the quantization step size and the peak signal-to-noise ratio and latency of the final image is monitored. The system uses a genetic algorithm to perform multi-objective optimization iteration on these two coefficients with "minimum bandwidth usage and no stuttering" as the fitness function. For example, test data shows that when the input demand coefficients are 2000, 5000, and 8000 milliseconds, the system expects the quantization step size to smoothly decrease from 32 to about 20, 10, and 5, respectively. Through algorithmic evolution calculations, it was found that when Setting it to 0.001 effectively maps the independent variable to the single-digit range, and When the slope of the logarithmic curve is controlled by a value of 2.0, it best fits the desired decay curve mentioned above, thus establishing the specific values ​​of the two.

[0065] In this embodiment, it should be noted that in S31, after obtaining the attention coefficient representing subjective expectations and the risk indicator representing objective threats, S31 is executed to quantify the demand coefficient, in order to solve the monitoring lag problem caused by the separation of risk and attention. The attention coefficient in milliseconds is used as the base value, and its correlation with risk indicators is calculated. The product of these factors yields the environmental added value. Milliseconds. The base value is added to the environmental added value to obtain the demand coefficient for parking parameters. Milliseconds. This computational logic, combining addition and multiplication, has a clear physical orientation: when the environment is stable, the risk indicator is zero, and the demand coefficient equals the user's daily level of attention; while when the environment experiences a sudden change, the risk indicator becomes one, and the demand coefficient doubles based on the user's level of attention. This design mathematically couples and amplifies physical changes and psychological focus, quantifying the real urgency of the current surveillance video stream in the network transmission queue.

[0066] In S32, after obtaining the high demand coefficient, the quantization step size is adaptively adjusted to resolve the blurring and stuttering issues caused by uniform transmission of high-value, high-risk images in weak network environments. Specifically, after obtaining the high demand coefficient representing user attention and sudden environmental changes, S32 aims to solve the delay and blurring problem caused by uniform transmission of risk frames in monitoring images in weak network environments, through an expression... An adaptive adjustment closed loop for the quantization step size of the source coding was constructed.

[0067] in the formula The target quantization step size is a dimensionless integer. In video coding standards, the smaller this value, the lower the compression rate and the larger the transmission bandwidth required. This serves as the base quantization step size distributed from the cloud network layer. Furthermore, the computational logic employs division and a logarithmic function. The reason for the nesting is that communication bandwidth resources are limited. If a linear mapping is used, when the demand coefficient surges, the quantization step size will quickly decay to zero, causing encoder abnormalities. The natural logarithm function has the mathematical property that the slope gradually flattens as the independent variable increases, which can dynamically compress a wide range of demand coefficients to a reasonable range. This ensures that the quantization step size shrinks non-linearly when responding to high-risk demands, but remains within the range of legal positive integers, thus avoiding a single video channel occupying the entire network. It is a dimensionless bandwidth resource balancing coefficient used to control the sensitivity of quantization step size reduction; The demand coefficient is obtained in milliseconds. The unit of measurement alignment coefficient is the reciprocal of milliseconds. Multiplying it by the demand coefficient eliminates the time dimension and satisfies the rule that the independent variable of the logarithmic function is a pure number. The constant 1 in parentheses ensures that the logarithmic term is zero when the demand is extremely small, and the target quantization step size smoothly falls back to the base step size.

[0068] Demand coefficient calculated based on scenario data milliseconds, initial base quantization step size Balance coefficient Alignment coefficient Substituting into the formula, we get The output quantization step size is rounded down to 6. This calculation allows the encoder to compress the quantization step size from 32 to 6 when facing the condition of a load-bearing column approaching. By reducing the quantization loss of high-frequency components in the image, a higher priority transmission channel is allocated to the video stream in the blind spot of this road, realizing a physical adaptive match between the clarity of the monitoring image and the dynamic environmental risks.

[0069] like Figure 5 As shown, in one specific embodiment, S4 includes: S41, constructing a reference vector and performing radial projection. In the screen coordinate system, the center pixel of the takeover reference is calibrated. The coordinates of historical touch points in the trajectory sequence are extracted. The historical touch points are connected to the takeover reference to form a reference vector sequence. The actual displacement vectors of two adjacent historical touch points are calculated, and the displacement vectors are projected onto the corresponding reference vectors to obtain the radial components.

[0070] S42. Calculate the intervention trend. If the projection points to the control benchmark, its magnitude is defined as the approach value; if it deviates, it is defined as the divergence value. The approach value and the divergence value within the time window are summed, and the accumulated result is the intervention trend.

[0071] S43. Calculate the takeover results. If the intervention trend is less than zero, fuse the intervention trend with the demand coefficient using a mapping function. Its mathematical expression is:

[0072] In the formula, each character is defined as follows: The takeover result is represented by a takeover probability value, with the domain being... ; It is a natural constant; The absolute value of the intervention trend, in pixels; This is the demand factor, expressed in milliseconds. These are dimensionless control coefficients; is a dimensional constant, with units of . , used to offset the physical dimensions generated by the product term.

[0073] Among them, control coefficient With dimensional constant Its value also follows a data-driven logic based on real interactive feedback. A positive sample set is constructed by retrieving the absolute value of the intervention trend and the demand coefficient in the second before the user finally presses the physical emergency brake button from a limited number (e.g., 600) of real valet parking history records; simultaneously, a negative sample set is constructed by extracting data from normal, unattended parking scans. Then, a logistic regression algorithm is used, and maximum likelihood estimation is employed to train and fit the probability activation function. The dimensional constant is... First, a specific constant is set based on the objective criteria of aligning physical dimensions to compensate for the units. For example, when training a model using historical data, it is found that when... When dynamically iterated to 0.8, the logistic regression model generally exhibited a probability of overtaking positive samples with a takeover probability greater than 0.85 and a probability of overtaking negative samples with a takeover probability strictly less than 0.2. Its intent classification accuracy and recall reached an optimal level of over 96%. Therefore, [the model was]... The value is fixed at 0.8 to ensure the scientific validity and error tolerance of the predicted takeover results.

[0074] S44. Filter out non-intervention behaviors. If the intervention trend is greater than or equal to zero, it indicates that the trajectory is deviating from the takeover baseline, and the prediction result of no takeover need is directly output.

[0075] In this embodiment, it should be noted that in S41, after ensuring clear transmission of the risk image, a reference vector is constructed and radially projected to solve the problem that disordered screen touch trajectories are difficult to interpret for specific operational intentions. The center of the "emergency brake" control is calibrated as the takeover reference within the screen coordinate system. Historical touch points are extracted from the trajectory sequence and connected to the takeover reference to form a radial reference vector sequence. Subsequently, the actual displacement vectors of adjacent touch points are calculated and projected onto the corresponding radial reference vectors using the vector dot product rule to obtain the radial components. This mathematical calculation method of vector projection effectively filters out invalid displacements generated when the user's finger slides or wanders horizontally on the screen, extracting effective movement components that directly point to or deviate from the direction of the brake control. This step establishes a precise one-dimensional evaluation system pointing to potential risk intervention targets in a two-dimensional disordered interactive coordinate system, improving the directionality of intent recognition.

[0076] In S42, the intervention trend is calculated based on the acquired radial projection components. The purpose is to confirm the coherence and authenticity of the finger approach behavior by accumulating historical actions. The direction of the projection components is evaluated; if it points towards the control reference, it is defined as a negative approach value; if it deviates, it is defined as a positive divergence value. These values ​​within the sliding time window are summed to obtain the absolute value of the intervention trend. The value is a pixel, and the sign is negative. This algebraic summation logic allows positive divergence values ​​generated by short-term retreat or hesitation to partially offset the convergent values, thus filtering out large, one-time displacements caused by accidental touches. By calculating the cumulative offset over a period of time, the intervention trend quantifies the urgency with which the user's finger converges towards the brake button in physical space. This provides a stable and reliable data characterization for determining whether the user has a strong motivation to terminate the intervention.

[0077] In S43, after confirming a clear approaching trend of the finger, the final takeover result is calculated to address the lag in braking commands caused by relying on physical button touch. Specifically, after ensuring the clear delivery of key environmental data, S43 aims to predict takeover behavior in advance by analyzing the user's physical sliding trajectory in the two-dimensional space of the screen, overcoming the engineering defect of delayed braking response caused by relying on mechanical button triggers, and employing mathematical expressions. The screen interaction trajectory was mapped non-linearly with the underlying environmental risks.

[0078] In this expression, This represents the probability value of the target takeover obtained, and its domain is strictly constrained by the functional properties to the interval between zero and one, conforming to the definition paradigm of probability theory; the core driving term of the formula is the natural constant. The negative exponent part is composed of the product of multiple physical quantities, among which... The absolute value of the calculated intervention trend represents the cumulative pixel displacement as the user's finger approaches the emergency takeover control, in pixels. The demand coefficient obtained earlier represents a quantitative rating of the current environmental risk level, measured in milliseconds.

[0079] Furthermore, multiplying the absolute value of the displacement trend with the demand coefficient mathematically constructs a risk coupling mechanism with dual confirmation. This means that only when the physical environment risk is high and the subjective intervention gesture is urgent will the core product of the exponential term increase, thereby triggering a higher probability of takeover. This effectively filters out false triggers caused by users' fingers unintentionally sliding across the takeover area in a safe environment. To control the shape coefficient, the steepness of the probability curve at the threshold boundary is determined; It is a dimensional transformation constant, whose physical unit is the product of the reciprocal of the pixel and the reciprocal of the millisecond. It is used to completely cancel out the dimensions of the complex and ensure the physical validity of the exponential operation.

[0080] Substitute scenario data and intervene in the absolute value of the trend. Pixels, demand factor Milliseconds, controlling the shape factor Dimensional transformation constant Substituting into the takeover probability formula, we first calculate the core product of the exponential part as follows: Then, by solving the overall probability function, we can obtain... This computational logic indicates that before the user's finger actually presses the brake control, a 91.5% probability of intervention has been calculated based on the convergence strength of the trajectory and the severity of the environment. This allows the cloud to issue a pre-charge command to the chassis in advance, thus mitigating the physiological delay in human action execution through a pre-positioned safety loop.

[0081] In S44, to ensure the overall operational stability of parking and prevent false intentions from interfering with normal driving, non-intervention behaviors are filtered out. If the intervention trend calculated in S42 is greater than or equal to zero, it indicates that the user's trajectory within the time window is generally divergent, or that they are simply idly browsing within a safe distance of the takeover control. In this case, nonlinear probability calculations are not initiated; instead, the prediction result indicating no takeover need is directly output. This logical branch eliminates the computational resource consumption and false alarms caused by users dragging maps, switching perspectives, or unintentionally touching edges while viewing the surrounding monitoring environment. By setting a zero-value filtering boundary, sensitivity to intervention intentions is maintained while suppressing overreactions, ensuring the smoothness of the automated valet parking process and the rigor of chassis control commands.

[0082] This invention also provides an AVP vehicle status dynamic update system based on user attention perception, the system comprising: The acquisition module is used to acquire the valet parking process and to acquire the environmental sequence, the attention sequence of parking parameters, the zoom sequence, and the trajectory sequence within the corresponding stage. The evaluation module is used to extract dwell features and distortion features based on the attention sequence and the scaling sequence, and fuse them to obtain the attention coefficient; at the same time, it performs edge detection based on the environmental sequence corresponding to the current stage to obtain risk indicators. The feedback module is used to obtain the demand coefficient of parking parameters based on risk indicators and attention coefficients, and input the demand coefficient into the coding closed loop to dynamically adjust the quantization step size of the corresponding data stream. The prediction module is used to obtain a takeover baseline that is not related to parking parameters in the screen coordinate system, extract the radial component of the trajectory sequence to the takeover baseline, and calculate the offset to obtain the intervention trend; then, the intervention trend and the demand coefficient are fused through a nonlinear mapping function to obtain the takeover result.

[0083] In one specific implementation, the evaluation module is further configured to: extract the difference between the maximum dwell time and the duration threshold in the attention sequence as a first feature; calculate the discrete gradient of adjacent duration values ​​in the attention sequence, and perform an algebraic sum of all discrete gradients within the time window to obtain a second feature; obtain the dwell feature based on the superposition result of the first feature and the second feature; calculate the rate of change of adjacent scaling factors in the scaling sequence, and extract the spatial derivative extrema as a distortion feature; convert the distortion feature into an intent gain, and multiply the dwell feature by the intent gain to obtain the attention coefficient.

[0084] In one specific implementation, the evaluation module is further configured to: calculate the absolute difference between adjacent distance values ​​in the environmental sequence; if any absolute difference is greater than a safety threshold, an environmental step is triggered, and a first value is assigned to the risk indicator; if none of the absolute differences are greater than the safety threshold, an environmental step is not triggered, and a second value is assigned to the risk indicator.

[0085] In one specific implementation, the feedback module is also used to: convert the demand coefficient into a pure number by multiplying it with a dimensional constant, and then input the logarithmic function to obtain the composite gain; construct a compressed denominator based on the composite gain and the balance coefficient; and generate a nonlinearly reduced quantization step size by dividing the base step size by the compressed denominator.

[0086] To enable those skilled in the art to fully understand and implement the technical solutions described in this specification, the following section, using a specific application scenario, provides a detailed deduction and data analysis of the entire process of the AVP vehicle status dynamic update method and system based on user attention perception.

[0087] Assume a vehicle is in an underground parking garage and performing the parking maneuver in step S1, with its right rear wheel approaching a load-bearing pillar. During the current data acquisition period, ultrasonic sensors deployed on the vehicle continuously detect the distance to the obstacle at the right rear, generating an environmental sequence. The unit is centimeters. At this moment, the user is monitoring the parking process via their mobile phone screen. The system, through the interactive tracking interface, records that the user showed a high level of attention to the parking parameter "right rear wheel blind spot video stream." The attention sequence for this area is extracted as follows: The unit is milliseconds; at the same time, due to poor visibility, the user performed a two-finger zoom operation on the screen, which was recorded as the corresponding zoom sequence; and the user's finger then slid within the screen coordinate system, and the system simultaneously recorded the corresponding trajectory sequence.

[0088] The system then proceeds to S2 to quantify the user's psychological expectations. First, S21 and S22 are executed to extract dwell characteristics, and a duration threshold for the current inbound redirection phase is set. Milliseconds. Traversing the sequence of interest. The maximum stay duration is If the threshold is exceeded, the first feature is extracted. Milliseconds. Then, the discrete gradients of adjacent sampling points in the sequence are calculated, respectively. and Within the time window, the algebraic sum is used to obtain the second feature. Due to the second feature The system adds the first feature to the second feature to obtain the basic time-based attention score. Milliseconds. This result quantifies, from an underlying logical perspective, the objective fact that users not only have longer viewing times but also continuously increasing levels of attention.

[0089] To elevate the 2D screen operation to a higher dimension, the system executes S23, fusing spatial distortion features. The scaling sequence is retrieved, and calculations show that the extreme value of the spatial derivative (i.e., the screen scaling rate) for the user's two-finger zoom is... The unit is The system's preset scaling sensitivity time constant Seconds. Substitute the above physical quantities into the formula. The calculation process is as follows: Milliseconds. Calculation results show that, accompanied by The sudden zoom-in gesture instantly increased the system's attention rating for the video stream from [previous value]. leap to It accurately captured the user's high-priority spatial focus intervention intentions.

[0090] While extracting user intent, the system simultaneously executes steps S24 to S26 to capture objective steps in the physical environment. (This is for the environmental sequence.) Calculate the absolute difference between adjacent values ​​as follows: Centimeters. The preset safety threshold for the current warehousing stage is [value missing]. centimeters. Due to the absolute difference within the last sampling period. This indicates that the load-bearing column has entered a high-risk collision boundary. The system determines that the first-order spatial derivative has crossed the safety threshold interface, triggering an environmental step flag, which sets the preset Boolean constant. The value is assigned to the risk indicator. After entering S31, the system calculates the demand coefficient. The formula logic is the base value plus the product of the base value and the risk indicator, that is... millisecond.

[0091] After obtaining the high-level demand coefficient, the system executes S32 for adaptive adjustment of the source coding. The initial basic quantization step size is then distributed from the cloud network layer. Preset balance coefficient Dimensional alignment coefficient Substitute into the formula Perform the calculation: The system rounds the result down to obtain the target quantization step size. Quantization step size from sudden drop This means that the cloud encoder significantly reduced the data compression rate of the "right rear wheel blind spot video stream" and occupied the communication bandwidth with the highest priority, which quickly reduced the compression distortion rate of the video image on the user terminal screen and completed the closed loop from logical calculation to physical bandwidth allocation.

[0092] Once the image becomes clear, the user perceives the distance as too close and panics, initiating a finger swipe towards the "emergency brake" control at the bottom of the screen. The system immediately executes steps S41 and S42, constructing a radial reference vector within the screen coordinate system, using the center of the "emergency brake" control as the intervention reference. By tracking the user's finger trajectory sequence, the projection of the actual displacement onto the radial reference vector is calculated. As the finger continuously approaches the brake control, the projection component manifests as a negative approach value towards the reference. Within the current swipe time window, the magnitudes of these negative approach values ​​are summed to obtain the absolute value of the intervention trend. Pixels. This value objectively represents the amount of physical displacement of a user's finger within the screen space as it converges towards a potentially dangerous key.

[0093] Since the intervention trend is negative, the system eventually switches to the nonlinear takeover probability mapping of S43. The system's pre-calibrated control coefficients... dimensional constant Substitute the relevant data into the takeover probability formula. First, the core product of the exponential part is calculated: Then, the probability function is solved: Calculations show that the system predicts a high risk of injury before the user's finger actually touches the emergency brake button. The system can pre-charge the brakes to the chassis domain controller without waiting for the final mechanical touch, thus eliminating the risk of physical collision due to human reaction delays.

[0094] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.

[0095] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.

[0096] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for dynamically updating the status of AVP vehicles based on user attention perception, characterized in that the method... include: Acquire the valet parking process and obtain the environmental sequence, the attention sequence of parking parameters, the zoom sequence, and the trajectory sequence within the corresponding stage; Based on the attention sequence and the scaling sequence, dwell features and distortion features are extracted and fused to obtain the attention coefficient; at the same time, edge detection is performed based on the environmental sequence corresponding to the current stage to obtain risk indicators. Based on the risk indicators and the attention coefficient, the demand coefficient of the parking parameters is obtained, and the demand coefficient is input into the coding closed loop to dynamically adjust the quantization step size of the corresponding data stream. Within the screen coordinate system, a takeover reference that is not related to the parking parameters is obtained, the radial component of the trajectory sequence toward the takeover reference is extracted, and the offset is calculated to obtain the intervention trend. The intervention trend and the demand coefficient are then fused using a nonlinear mapping function to obtain the takeover result.

2. The method for dynamic updating of AVP vehicle status based on user attention perception according to claim 1, characterized in that, The step of extracting dwell features and distortion features based on the attention sequence and the scaling sequence, and fusing them to obtain the attention coefficient, includes: The difference between the maximum dwell time and the duration threshold in the sequence of interest is extracted as the first feature; Calculate the discrete gradients of adjacent duration values ​​in the sequence of interest, and take the algebraic sum of all the discrete gradients within the time window to obtain the second feature; The dwelling feature is obtained based on the superposition result of the first feature and the second feature; Calculate the rate of change of adjacent scaling factors in the scaling sequence, and extract the spatial derivative extrema as the distortion feature; The distortion feature is converted into an intent gain, and the dwell feature is multiplied by the intent gain to obtain the attention coefficient.

3. The method for dynamic updating of AVP vehicle status based on user attention perception according to claim 1, characterized in that, The step of performing edge detection based on the environmental sequence corresponding to the current stage to obtain risk indicators includes: Calculate the absolute difference between adjacent distance values ​​in the environmental sequence; If any of the absolute differences is greater than the safety threshold, an environmental step is triggered, and the first value is assigned to the risk indicator. If none of the absolute differences are greater than the safety threshold, an environmental step is not triggered, and the second value is assigned to the risk indicator.

4. The method for dynamic updating of AVP vehicle status based on user attention perception according to claim 1, characterized in that, The step of inputting the demand coefficient into the coding closed loop and dynamically adjusting the quantization step size of the corresponding data stream includes: After multiplying the demand coefficient by the dimensional constant to convert it into a pure number, the composite gain is obtained by inputting it into a logarithmic function. Construct a compressed denominator based on the aforementioned composite gain and balance coefficient; The quantization step size is generated by dividing the base step size by the compression denominator, resulting in a nonlinearly reduced quantization step size.

5. The method for dynamic updating of AVP vehicle status based on user attention perception according to claim 1, characterized in that, The step of extracting the radial component of the trajectory sequence toward the takeover reference and calculating the offset to obtain the intervention trend includes: Within the screen coordinate system, a reference vector is formed by connecting the historical touch points in the trajectory sequence with the takeover reference. Calculate the displacement vector of two adjacent historical touch points, and project the displacement vector onto the corresponding reference vector to obtain the radial component; The modulus of the radial component pointing towards the nozzle reference is defined as the approaching value, and the modulus of the component moving away from it is defined as the diverging value. The approach value and the divergence value within the time window are summed, and the accumulated result is used as the intervention trend.

6. The method for dynamic updating of AVP vehicle status based on user attention perception according to claim 5, characterized in that, The step of fusing the intervention trend with the demand coefficient through a nonlinear mapping function to obtain the takeover result includes: If the intervention trend is greater than or equal to zero, then no takeover requirement is directly output as the takeover result; If the intervention trend is less than zero, then the penalty index is constructed by multiplying the absolute value of the intervention trend, the demand coefficient, the control coefficient, and the dimensional constant together. The penalty index is input into the probability activation function, and the value of the mapping output is used as the takeover result.

7. An AVP vehicle status dynamic update system based on user attention perception, characterized in that, include: The acquisition module is used to acquire the valet parking process and to acquire the environmental sequence, the attention sequence of parking parameters, the zoom sequence, and the trajectory sequence within the corresponding stage. The evaluation module is used to extract dwell features and distortion features based on the attention sequence and the scaling sequence, and fuse them to obtain the attention coefficient; at the same time, it performs edge detection based on the environmental sequence corresponding to the current stage to obtain risk indicators. The feedback module is used to obtain the demand coefficient of the parking parameters based on the risk indicators and the attention coefficient, and input the demand coefficient into the coding closed loop to dynamically adjust the quantization step size of the corresponding data stream. The prediction module is used to obtain a takeover reference that is not related to the parking parameters in the screen coordinate system, extract the radial component of the trajectory sequence to the takeover reference, and calculate the offset to obtain the intervention trend. The intervention trend and the demand coefficient are then fused using a nonlinear mapping function to obtain the takeover result.

8. The AVP vehicle status dynamic update system based on user attention perception according to claim 7, characterized in that, The evaluation module is also used for: The difference between the maximum dwell time and the duration threshold in the sequence of interest is extracted as the first feature; Calculate the discrete gradients of adjacent duration values ​​in the sequence of interest, and take the algebraic sum of all the discrete gradients within the time window to obtain the second feature; The dwelling feature is obtained based on the superposition result of the first feature and the second feature; Calculate the rate of change of adjacent scaling factors in the scaling sequence, and extract the spatial derivative extrema as the distortion feature; The distortion feature is converted into an intent gain, and the dwell feature is multiplied by the intent gain to obtain the attention coefficient.

9. The AVP vehicle status dynamic update system based on user attention perception according to claim 7, characterized in that, The evaluation module is also used for: Calculate the absolute difference between adjacent distance values ​​in the environmental sequence; If any of the absolute differences is greater than the safety threshold, an environmental step is triggered, and the first value is assigned to the risk indicator. If none of the absolute differences are greater than the safety threshold, an environmental step is not triggered, and the second value is assigned to the risk indicator.

10. The AVP vehicle status dynamic update system based on user attention perception according to claim 7, characterized in that, The feedback module is also used for: After multiplying the demand coefficient by the dimensional constant to convert it into a pure number, the composite gain is obtained by inputting it into a logarithmic function. Construct a compressed denominator based on the aforementioned composite gain and balance coefficient; The quantization step size is generated by dividing the base step size by the compression denominator, resulting in a nonlinearly reduced quantization step size.