Vehicle parking-out control method and device, vehicle and computer readable medium

CN122474054BActive Publication Date: 2026-09-08CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610936610.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-08
Estimated Expiration
2046-06-26

AI Technical Summary

Technical Problem

[0005]本申请提供了一种车辆泊出控制方法、装置、车辆及计算机可读介质,以解决嘈杂环境语音指令识别精度低的技术问题

Benefits of technology

本申请提供了一种车辆泊出控制方法,包括:获取目标车辆的环境声音信号和车外用户发出的语音信号;对语音信号进行环境自适应语音处理后识别出用户意图,其中,环境自适应语音处理包括利用环境声音信号的环境表征来抑制语音信号中的环境噪声,从而对语音信号的声学特征进行校准,并在校准后生成识别文本,识别文本用于识别用户意图;在用户意图为泊出意图的情况下,根据语音信号携带的生物特征和活体特征对发出语音信号的用户进行身份校验,并在校验通过的情况下,根据用户位置确定车辆的泊出方向;根据泊出方向,控制车辆执行泊出操作。本申请通过环境自适应语音处理方式,依托环境表征校准语音声学特征,有效抑制环境噪声干扰,提升复杂场景下语音识别与意图解析的准确度,解决了嘈杂环境语音指令识别精度低的技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122474054B_ABST
    Figure CN122474054B_ABST
Patent Text Reader

Abstract

The application relates to a vehicle parking-out control method and device, a vehicle and a computer readable medium. The method comprises the following steps: acquiring an environmental sound signal of a target vehicle and a voice signal emitted by a user outside the vehicle; recognizing a user intention after performing environmental self-adaptive voice processing on the voice signal, wherein the environmental self-adaptive voice processing comprises suppressing environmental noise in the voice signal by using an environmental representation of the environmental sound signal, thereby calibrating the acoustic characteristics of the voice signal, and generating a recognition text after the calibration, the recognition text being used to recognize the user intention; in the case that the user intention is a parking-out intention, performing identity verification on the user emitting the voice signal according to biological characteristics and living characteristics carried by the voice signal, and in the case that the verification is passed, determining a parking-out direction of the vehicle according to the position of the user; and controlling the vehicle to perform a parking-out operation according to the parking-out direction. The application relies on the environmental representation to suppress environmental noise interference, and solves the technical problem of low voice recognition accuracy in a noisy environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent vehicle control technology, and in particular to a vehicle parking control method, device, vehicle, and computer-readable medium. Background Technology

[0002] With the continuous development of intelligent vehicle technology, automatic parking has become a standard feature in mainstream models. To address the inconvenience of in-vehicle operation and remote control of parking via mobile phone, the industry has gradually introduced external voice control automatic parking solutions. These solutions rely on voice interaction to trigger vehicle parking from outside the vehicle, further enhancing the intelligence and convenience of vehicle use.

[0003] Currently, although some vehicles on the market have achieved the function of automatic parking and automatic parking exit controlled by external voice, such voice control systems generally have obvious defects. For example, in scenarios where vehicles are often in open-air parking lots, underground garages and other scenarios with frequent traffic, the outdoor environment is noisy and complex. The existing voice recognition solutions have insufficient anti-interference ability, which directly leads to low accuracy of voice command recognition. It is very easy to have command parsing errors, function mis-triggering or response failure, which seriously affects the normal use of the external voice parking exit function.

[0004] There is currently no effective solution to the problem of low accuracy in voice command recognition in noisy environments. Summary of the Invention

[0005] This application provides a vehicle parking control method, device, vehicle, and computer-readable medium to solve the technical problem of low accuracy in voice command recognition in noisy environments.

[0006] According to one aspect of the embodiments of this application, this application provides a vehicle parking exit control method, including: acquiring an ambient sound signal of a target vehicle and a voice signal emitted by a user outside the vehicle; recognizing the user's intention after performing environmental adaptive voice processing on the voice signal, wherein the environmental adaptive voice processing includes using the environmental representation of the ambient sound signal to suppress environmental noise in the voice signal, thereby calibrating the acoustic features of the voice signal, and generating recognition text after calibration, the recognition text being used to recognize the user's intention; if the user's intention is a parking exit intention, verifying the identity of the user who emitted the voice signal based on the biometric and liveness features carried by the voice signal, and if the verification is successful, determining the parking exit direction of the vehicle based on the user's location; and controlling the vehicle to perform a parking exit operation according to the parking exit direction.

[0007] Optionally, the speech signal undergoes environment-adaptive speech processing, including the following steps performed by a sound environment coding thread: extracting acoustic features of the environmental sound signal; mapping the acoustic features of the environmental sound signal to a query key vector using a query key generator; calculating the similarity between the query key vector and each known environmental representation in the environmental representation lookup table using a pre-established environmental representation lookup table, and using this similarity as the attention weight of each known environmental representation; and weighting and summing each known environmental representation according to its respective attention weight to obtain the environmental representation of the environmental sound signal.

[0008] Optionally, the environment-adaptive speech processing of the speech signal further includes performing the following steps through a speech recognition thread: downsampling the speech signal to obtain the acoustic features of the speech signal; fusing the acoustic features of the speech signal with the environmental representation of the environmental sound signal to obtain fused features; generating a real-valued mask based on the fused features using a pre-trained adaptive calibrator, and using the real-valued mask to calibrate the acoustic features of the speech signal to suppress noise, to obtain calibrated features of the acoustic features of the speech signal; and inputting the calibrated features into an end-to-end speech recognition model to obtain the recognized text output by the end-to-end speech recognition model.

[0009] Optionally, before performing environment-adaptive speech processing on the speech signal, the method further includes: performing end-to-end co-training on all model components in the environment-adaptive speech processing to co-optimize model parameters for environment representation extraction, acoustic feature calibration, and speech recognition.

[0010] Optionally, the identity verification of the user emitting the voice signal is performed based on the biometric and liveness features carried by the voice signal, including: extracting voiceprint features and at least one liveness feature from the voice signal; and confirming that the user has passed the vehicle owner identity verification when the similarity between the voiceprint feature and the preset authorized voiceprint is greater than or equal to a first preset threshold, and the similarity between the liveness feature and the preset liveness feature is greater than or equal to a second preset threshold, wherein the preset authorized voiceprint is the voiceprint feature pre-set by the authorized vehicle owner stored locally in the vehicle, and the preset liveness feature is the liveness feature pre-set by the authorized vehicle owner stored locally in the vehicle.

[0011] Optionally, before determining the parking direction of the vehicle based on the user's location, the method further includes determining the user's location in the following manner: obtaining candidate locations for the user, wherein the candidate locations include at least one of the following: a first location of the user obtained based on sound source localization, a second location of the user obtained using an ultra-wideband (UWB) positioning module, a third location of the user obtained based on Bluetooth-assisted positioning, and a fourth location of the user obtained based on camera acquisition; the ultra-wideband (UWB) positioning module includes positioning base stations deployed at four locations: front left, front right, rear left, and rear right of the vehicle, and the positioning base stations are used for bidirectional ranging and positioning with the UWB tag embedded in the vehicle key carried by the user; assigning weights to each candidate location based on the confidence level of the data used to calculate each candidate location, wherein the confidence level is positively correlated with the data quality of each data point; and performing weighted fusion on each candidate location to obtain the final user location.

[0012] Optionally, the parking direction of the vehicle is determined based on the user's location, including: calculating the azimuth and distance of the final user's position relative to the front of the vehicle, and identifying the parking space type where the vehicle is located; if the parking space type is a parallel parking space and the azimuth is within a preset acute angle range, determining forward parking as the parking direction; if the parking space type is a parallel parking space and the azimuth is within a preset obtuse angle range, determining backward parking as the parking direction; if the parking space type is a perpendicular parking space and the azimuth is within a preset acute angle range, determining forward lateral parking as the parking direction; if the parking space type is a perpendicular parking space and the azimuth is within a preset obtuse angle range, determining backward lateral parking as the parking direction; if the parking space type is an angled parking space, determining the direction in which the user can directly get into the vehicle after parking as the parking direction; if there are obstacles in all directions or the distance is less than a preset safe distance, pausing the parking operation and prompting the user to adjust the position.

[0013] According to another aspect of the embodiments of this application, this application provides a vehicle parking control device, including: a voice acquisition module, used to acquire environmental sound signals of a target vehicle and voice signals emitted by a user outside the vehicle; a voice processing and semantic recognition module, used to recognize the user's intention after performing environmental adaptive voice processing on the voice signal, wherein the environmental adaptive voice processing includes using the environmental representation of the environmental sound signal to suppress environmental noise in the voice signal, thereby calibrating the acoustic features of the voice signal, and generating recognition text after calibration, the recognition text being used to recognize the user's intention; an identity verification and direction decision module, used to verify the identity of the user who emitted the voice signal based on the biometrics and liveness features carried by the voice signal when the user's intention is to park, and to determine the parking direction of the vehicle based on the user's location if the verification is successful; and an execution module, used to control the vehicle to perform a parking operation according to the parking direction.

[0014] According to another aspect of the embodiments of this application, this application provides a vehicle, including a memory, a processor, a communication interface and a communication bus. The memory stores a computer program that can run on the processor. The memory and the processor communicate with each other through the communication bus and the communication interface. When the processor executes the computer program, it implements the steps of the above method.

[0015] According to another aspect of the embodiments of this application, this application also provides a computer-readable medium having processor-executable non-volatile program code that causes the processor to perform the above-described method.

[0016] Compared with related technologies, the technical solutions provided in this application have the following advantages: This application provides a vehicle parking control method, comprising: acquiring ambient sound signals of a target vehicle and voice signals emitted by a user outside the vehicle; recognizing the user's intent after performing environment-adaptive speech processing on the voice signals, wherein the environment-adaptive speech processing includes using the environmental representation of the ambient sound signals to suppress environmental noise in the voice signals, thereby calibrating the acoustic features of the voice signals, and generating recognition text after calibration, which is used to recognize the user's intent; when the user's intent is to park, verifying the identity of the user who emitted the voice signals based on the biometric and liveness features carried by the voice signals, and determining the parking direction of the vehicle based on the user's location if the verification is successful; and controlling the vehicle to perform a parking operation based on the parking direction. This application, through environment-adaptive speech processing, relies on environmental representation to calibrate the acoustic features of the voice, effectively suppressing environmental noise interference, improving the accuracy of speech recognition and intent parsing in complex scenarios, and solving the technical problem of low accuracy in voice command recognition in noisy environments. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0019] Figure 1 This is a schematic flowchart of an optional vehicle parking control method according to an embodiment of this application; Figure 2 This is a schematic flowchart of an optional environment-adaptive speech processing method according to an embodiment of this application; Figure 3This is a schematic diagram of an optional adaptive calibrator processing flow according to an embodiment of this application; Figure 4 This is a block diagram of an optional vehicle parking control device according to an embodiment of this application; Figure 5 This is a schematic diagram of an optional vehicle structure provided for an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustration and has no specific meaning in itself. Therefore, "module" and "part" may be used interchangeably.

[0022] To address the problems mentioned in the background art, according to one aspect of the embodiments of this application, an embodiment of a vehicle parking exit control method is provided, such as... Figure 1 As shown, the method may include the following steps: Step S102: Acquire the ambient sound signal of the target vehicle and the voice signal emitted by the user outside the vehicle; Step S104: After performing environment-adaptive speech processing on the speech signal, the user's intention is identified. The environment-adaptive speech processing includes using the environmental representation of the environmental sound signal to suppress environmental noise in the speech signal, thereby calibrating the acoustic features of the speech signal, and generating recognition text after calibration. The recognition text is used to identify the user's intention. Step S106: If the user's intention is to park out, the user who sent the voice signal is identified based on the biometric and liveness features carried in the voice signal. If the verification is successful, the parking direction of the vehicle is determined based on the user's location. Step S108: Control the vehicle to perform a parking operation according to the parking direction.

[0023] In this embodiment, the ambient sound signal refers to the ambient noise of the external environment where the vehicle is located. The speech signal refers to the voice sound wave signal emitted by the user outside the vehicle, including user semantic commands, voiceprint features, and liveness features such as breathing and intonation. Environment-adaptive speech processing refers to a speech processing mechanism designed for noisy outdoor scenarios such as garages and open-air parking lots, relying on dual parallel threads to complete environmental feature extraction and speech signal optimization. The dual parallel threads include a sound environment encoding thread and a speech recognition thread. The sound environment encoding thread is an independently running program thread specifically used to collect and analyze the ambient sound around the vehicle, generating feature data that characterizes the environmental noise. The speech recognition thread is a program thread running in parallel with the sound environment encoding thread, using environmental representation to correct the original speech signal and complete speech-to-text and intent recognition. Environmental representation refers to the low-dimensional feature vector obtained by encoding the time-frequency features of the ambient noise around the vehicle, which can accurately reflect the noise type, intensity, and distribution pattern of the current environment. Biometric features can be voiceprints, and liveness features refer to features possessed only by real human voices, including breathing rhythm, intonation fluctuations, and vocal vibrato, used to distinguish real human speech from recorded or synthesized speech. Identity verification refers to the process by which the vehicle verifies the legitimacy of the voice sender based on locally stored information. Optimal parking exit direction refers to the vehicle's exit direction determined by considering the user's location, parking space layout, and the distribution of surrounding obstacles, while also balancing driving safety, passenger safety, and user habits.

[0024] In step S102, the vehicle is in a normal parking state with the engine off, locks engaged, and the automatic parking system in standby mode. The high-sensitivity onboard microphone array deployed around the vehicle remains in standby mode, monitoring external sound signals. When an authorized owner or other user issues a natural voice command (including various colloquial parking-related expressions such as "The sides of the car are too narrow, it's inconvenient for me to get in," "The car is too crowded, move it out," "I need to get in, move the car," etc.) within the vehicle's effective interaction range, the microphone array captures the analog sound wave signal propagating through the air in real time. The onboard audio circuit then performs analog-to-digital conversion, transforming the continuous analog sound waves into a standardized digital voice signal. After conversion, the onboard voice module performs preliminary filtering on the original digital voice signal, eliminating transient noise while fully preserving the semantic information, voiceprint information, and biometric and liveness characteristics corresponding to human voice production. The processed complete voice signal is then synchronously transmitted to the voice processing unit for voice signal acquisition and preliminary preprocessing. The onboard microphone array also continuously or at preset intervals collects ambient noise from outside the vehicle.

[0025] In step S104, as Figure 2 As shown, the system starts a sound environment encoding thread and a speech recognition thread. The two threads run independently and in parallel without blocking each other, enabling real-time updates of environmental features and synchronous speech processing.

[0026] The sound environment encoding thread continuously collects ambient sounds around the vehicle, performs feature extraction and encoding operations on the ambient sounds, and combines them with a pre-established environmental representation lookup table to complete feature matching and fusion, ultimately generating an environmental representation that fully reflects the current noise state. This environmental representation is then synchronously pushed to the speech recognition thread. The environmental representation is automatically updated periodically according to changes in external noise, always matching the real-time environmental state, and continuously pushing the latest environmental representation data to the speech recognition thread. The process of establishing the environmental representation lookup table is as follows: A large number of real parking scene noise samples are pre-collected, covering various typical noisy environments such as open-air parking lots, underground garages, roadside traffic, shopping mall pedestrian flow, and equipment noise; all noise samples are uniformly subjected to frame segmentation, windowing, and spectral feature extraction operations to obtain the basic acoustic feature sequences for each scene; an unsupervised clustering algorithm is used to cluster all acoustic feature sequences, with each cluster corresponding to a standardized environmental noise; the cluster center vector of each cluster is taken as the standard environmental representation for that scene, and all standard environmental representations are stored one-to-one with scene labels to obtain the environmental representation lookup table.

[0027] The speech recognition thread receives the original speech signal and the environmental representation output by the sound environment encoding thread. It uses the environmental representation to perform targeted calibration on the acoustic features of the original speech signal to cancel out the interference of environmental noise on the effective speech.

[0028] The calibration method can be as follows: Based on the fusion features, generate a frame-by-frame, frequency-by-frequency real-valued noise reduction mask, with the mask value ranging from 0 to 1: a mask approaching 0 corresponds to the noise-dominant frequency point, and signal attenuation suppression is performed; a mask approaching 1 corresponds to the effective human voice frequency point, and the signal amplitude is preserved and slightly enhanced; multiply the mask element-wise with the original speech acoustic features to complete the directional attenuation of the noise component; then combine the noise spectrum offset recorded in the environmental characterization for compensation correction to offset the human voice spectrum distortion caused by noise; finally, eliminate the inter-frame mask value abrupt change through smoothing filtering to obtain calibrated speech acoustic features without environmental noise interference.

[0029] After feature calibration, the calibrated acoustic features are input into the speech recognition model, which then converts them into standard recognition text. Finally, semantic parsing is performed on the recognition text to determine the actual operational intent corresponding to the user's voice, thus completing user intent recognition.

[0030] In step S106, when the semantic parsing result determines that the user's intention is to park the vehicle, the system extracts the biometric and liveness features carried in the voice from the collected voice signal, and performs identity verification by combining the voiceprint information in the voice signal. The system compares the extracted voiceprint and liveness features with the authorized vehicle owner feature data stored locally in the vehicle, and determines whether the user's identity is legitimate by combining the comparison results of the two features. If the identity verification result is unsuccessful, the system terminates the parking process; if the identity verification is successful, the system starts the position detection module to obtain the current position information of the user relative to the vehicle, and at the same time retrieves the parking space information and obstacle distribution information collected by the environmental perception module, fuses and analyzes the multiple types of data, and automatically calculates and determines the parking direction of the vehicle according to the preset decision logic.

[0031] In step S108, the vehicle control unit receives the optimal parking direction command and sends control signals to the vehicle chassis, steering, power and other actuators to drive the vehicle to smoothly complete the automatic parking action in the determined parking direction. Throughout the process, the vehicle sensing system monitors the surrounding environment in real time to ensure driving safety.

[0032] Through steps S102 to S108, this application uses an environment-adaptive speech processing method to calibrate speech acoustic features based on environmental representation, effectively suppressing environmental noise interference, improving the accuracy of speech recognition and intent parsing in complex scenarios, and solving the technical problem of low accuracy in speech command recognition in noisy environments.

[0033] In an optional embodiment, environment-adaptive speech processing of the speech signal includes performing the following steps via a sound environment coding thread: Step 1: Extract the acoustic features of the ambient sound signal; Step 2: Map the acoustic features of the ambient sound signal into a query key vector using a query key generator; Step 3: Using a pre-established environment representation lookup table, calculate the similarity between the query key vector and each known environment representation in the environment representation lookup table, and use this similarity as the attention weight for each known environment representation. Step 4: Weight and sum the known environmental representations according to their respective attention weights to obtain the environmental representation of the environmental sound signal.

[0034] In this embodiment, ambient sound signals refer to various background sounds outside the vehicle other than the user's voice, including ambient noise such as pedestrian noise, traffic noise, and equipment operation noise. The query key generator is a functional module built into the sound environment encoding thread, responsible for converting ambient acoustic features into standardized vector data. The query key vector is a feature vector obtained by the query key generator and is used to perform similarity matching calculations with pre-stored environmental features. The environmental representation lookup table refers to a pre-established database on the vehicle's local machine, storing standard environmental representation data for various typical scenarios, covering different types and intensities of ambient noise. The attention weights of each known environmental representation are used to characterize the degree of matching between the current environment and each standard environmental feature in the lookup table; the higher the weight value, the higher the feature similarity between the two.

[0035] In this embodiment of the application, after the vehicle starts the environment adaptive voice processing function and runs the dual threads in parallel, such as Figure 2 As shown, the sound environment encoding thread independently executes the entire process of environment acquisition, feature conversion, and matching fusion.

[0036] Specifically, the sound environment encoding thread adopts a long-period continuous sampling mode, with a sampling duration much longer than the duration of a single user voice message, ensuring complete capture of the overall time-frequency patterns of environmental noise. The sampling range covers the entire space in front of, behind, to the sides of the vehicle, with no sampling blind spots; the sampling process distinguishes between static continuous noise and dynamic sudden noise, completely recording the original information such as noise intensity changes and frequency distribution, and finally acquiring continuous and complete external environmental sound signals, which are then transmitted to the back-end feature extraction unit.

[0037] The system performs frame segmentation and windowing preprocessing on the acquired raw environmental sound signals, splitting the continuous audio signal into multiple short audio frames; for each audio frame, it extracts multi-dimensional acoustic features such as frequency, energy, and spectral distribution, and combines them to form an acoustic feature sequence corresponding to the environmental sound.

[0038] The system inputs the extracted acoustic feature sequence into the query key generator. The generator compresses the dimensions and unifies the format of the multidimensional acoustic features according to the preset feature mapping rules and vector encoding rules, and converts the discrete acoustic feature sequence into a query key vector with fixed dimensions and a unified format.

[0039] The preset feature mapping rule can be as follows: extract multi-dimensional acoustic features based on a fixed 20ms audio frame, divide each frame's acoustic features into three sub-features: Mel spectrum component, energy component, and temporal fluctuation component; normalize each sub-feature to the [0,1] interval to eliminate numerical deviations caused by different noise intensities; concatenate the three normalized sub-features according to a fixed sorting rule of spectrum priority, energy priority, and temporal supplementation to form a single-frame ordered feature array; use a multi-layer fully connected network as the mapping base layer, set three hidden layers to compress the feature dimension layer by layer, and combine ReLU activation function and batch normalization processing after each layer, finally mapping the variable-length acoustic feature sequence into an intermediate feature vector with fixed dimension and standardized numerical distribution, thus completing the feature mapping.

[0040] The preset vector encoding rules can be as follows: a floating-point vector with a fixed dimension of 128 is used as the unified output format, and all intermediate feature vectors obtained by mapping are normalized in dimension through a learnable linear encoding layer; during the encoding process, a positional encoding embedding mechanism is used to superimpose the temporal information of the audio frame onto the end of the feature vector, preserving the temporal features of noise changing over time; after encoding, L2 regularization is performed on the 128-dimensional overall vector, constraining the vector magnitude to 1, eliminating the interference of noise energy differences in different scenes on subsequent similarity calculations; the normalized vector after regularization is the query key vector that can be directly used for matching operations.

[0041] The system retrieves the vehicle's locally stored environment representation lookup table, iterates through each pre-stored known environment representation (i.e., the registered environment representation) in the lookup table, and simultaneously performs a similarity calculation between the previously generated query key vector and each known environment representation to quantify the matching degree between the current environment and each standard environment. The similarity calculation can be performed using the cosine similarity algorithm to complete the vector matching calculation. The calculation formula is as follows: Sim(A,B)= ; Where A is the query key vector generated in real time, and B is a single known environment representation vector in the lookup table; all dimensions in the vector participate in the dot product operation synchronously, and the output result range is [-1,1]; the negative similarity of the output is uniformly set to 0, and only the effective matching score in the range of 0~1 is retained, which intuitively represents the degree of matching between the current environment and the standard noise scene.

[0042] The process of establishing the environmental representation lookup table is as follows: A large number of real parking scene noise samples are pre-collected, covering various typical noisy environments such as open-air parking lots, underground garages, roadside traffic, shopping mall pedestrian flow, and equipment noise; framing, windowing, and spectral feature extraction operations are uniformly performed on all noise samples to obtain the basic acoustic feature sequences for each scene; an unsupervised clustering algorithm is used to cluster all acoustic feature sequences, with each cluster corresponding to a type of standardized environmental noise; the cluster center vector of each cluster is taken as the standard environmental representation for that scene, and all standard environmental representations are stored in a one-to-one correspondence with scene labels to obtain the environmental representation lookup table.

[0043] The matching degree values ​​are then normalized and converted into attention weights for each known environmental representation. The weights directly reflect the degree to which the standard environmental representation represents the current real environment.

[0044] The normalization method is as follows: collect the original cosine similarity scores S1, S2, ..., Sn of all known environmental representations in a single retrieval; sum the total score Ssum = S1 + S2 + ... + Sn; and assign an attention weight Wi = Si / Ssum to each representation, which converts the original cosine similarity scores into values ​​between 0 and 1, thus completing the normalization. After normalization, the sum of all weights is always equal to 1. The larger the weight value, the higher the matching degree between the standard environmental representation and the current real noise, and the higher the proportion allocated in the subsequent weighted summation.

[0045] Finally, the system uses the calculated attention weights as weighting coefficients to perform a weighted summation operation on all known environmental representations in the environmental representation lookup table. The higher the weight of a standard representation, the larger its proportion in the final result, thus achieving adaptive fusion of multiple standard features. After the operation is complete, a unique environmental representation matching the real-time noise state of the current sound environment is output, and this environmental representation is pushed in real-time to the parallel-running speech recognition thread, providing a basis for subsequent speech feature calibration and noise reduction processing. Simultaneously, the sound environment encoding thread cyclically executes all the above steps, periodically updating the environmental representation to adapt to dynamically changing external noise.

[0046] This application uses a standardized process of environmental sound acquisition, feature mapping, weight calculation and weighted fusion to accurately generate environmental representations that fit the real-time scene, providing reliable data support for speech noise reduction.

[0047] In an optional embodiment, the environment-adaptive speech processing of the speech signal further includes performing the following steps via a speech recognition thread: Step 1: Downsample the speech signal to obtain its acoustic features; Step 2: Fuse the acoustic features of the speech signal with the environmental representation of the ambient sound signal to obtain the fused features; Step 3: Use the pre-trained adaptive calibrator to generate a real-valued mask based on the fusion features, and use the real-valued mask to calibrate the acoustic features of the speech signal to suppress noise, thereby obtaining the calibrated features of the acoustic features of the speech signal. Step 4: Input the calibration features into the end-to-end speech recognition model to obtain the recognized text output by the end-to-end speech recognition model.

[0048] In this embodiment, downsampling refers to the operation of dimensionality reduction and simplification of speech acoustic features, reducing data computation and improving processing efficiency while retaining core speech information. Fusion features refer to composite features obtained by concatenating and dimensionally fusing the original speech acoustic features with environmental representations, containing both speech information and environmental noise information. The adaptive calibrator is a built-in AI module that generates a noise mask based on the fusion features, achieving accurate noise reduction and correction of speech features. The real-valued mask is a numerical matrix generated by the adaptive calibrator, used to mark noise regions and effective speech regions in the speech features, achieving targeted noise suppression. Calibration features refer to the pure speech acoustic features that have undergone mask denoising and noise suppression processing, eliminating environmental interference. The end-to-end speech recognition model can adopt a U2++ model that integrates CTC and attention encoder-decoder architectures; inputting calibration features directly outputs text content.

[0049] In this embodiment of the application, after the sound environment encoding thread completes the environment representation output, such as Figure 2 , Figure 3 As shown, the speech recognition thread performs downsampling, feature fusion, adaptive calibration, and speech recognition operations sequentially based on the original speech signal and environmental representation.

[0050] Specifically, such as Figure 2 , Figure 3 As shown, the speech recognition thread receives the acquired raw speech signal, inputs it into the adaptive calibrator, and performs frame-by-frame processing and downsampling on the entire speech signal. The downsampling process simplifies the speech data dimensions according to preset rules, eliminating redundant secondary features while fully preserving core acoustic features related to semantics, voiceprints, and liveness, reducing the computational load without sacrificing recognition accuracy. After downsampling, the acoustic features of the simplified speech signal are obtained.

[0051] The pre-training process of the adaptive calibrator is as follows: Noisy human speech samples are collected, including superimposed noise from garages of different car models, conversations with passersby, vehicle horns, etc., and simultaneously paired with clean, noise-free baseline speech. Each sample is labeled with the standard environmental representation, voiceprint, and live physiological features corresponding to the scene, and divided into training, validation, and test sets in a 7:2:1 ratio. The adaptive calibrator incorporates downsampling, feature fusion, mask generation, and speech enhancement output branches, with each branch weight initialized using Xavier. The mask output layer uses the Sigmoid activation function, outputting a real-valued mask in the 0-1 range. The training input is a noisy speech frame plus the corresponding environmental representation, and the output is a denoised and enhanced speech frame. The loss function uses the Mel-spectrum MSE loss of clean and enhanced speech plus the speech recognition-assisted cross-entropy loss. The Adam optimizer is used for backpropagation iterative updates of all branch parameters, and the denoising signal-to-noise ratio and recognition accuracy are calculated on the validation set after each iteration. When the signal-to-noise ratio of the test set improves by ≥12dB, the speech recognition accuracy improves by ≥15%, and there are no significant fluctuations for 10 consecutive rounds, the iteration stops, all network weights and bias parameters of the adaptive calibrator are frozen, pre-training is completed and burned to the vehicle's local storage unit.

[0052] Preset rules for simplifying speech data dimensions can be: time-domain simplification rules, such as retaining the three core time-domain features of speech fundamental frequency, formants, and respiratory fluctuations, while removing high-frequency weak noise time-domain components that have no significant semantic or liveness information; frequency-domain simplification rules, such as retaining only the effective frequency band of human voice from 300Hz to 3400Hz, filtering out low-frequency vehicle vibration noise below 300Hz and high-frequency environmental screeching noise above 3400Hz; dimensionality compression rules, such as using principal component analysis (PCA) to reduce the dimensionality of the retained time-frequency joint features, retaining more than 95% of the cumulative variance contribution rate, and discarding secondary feature dimensions with a variance proportion of less than 5%; and sampling simplification rules, such as downsampling human speech from an original sampling rate of 16kHz to 8kHz, simultaneously using anti-aliasing filtering, halving the data volume without losing key features for voiceprint and liveness recognition.

[0053] The adaptive calibrator retrieves the environmental representation of the ambient sound signal generated by the ambient sound encoding thread and fuses the acoustic features of the speech signal with the environmental representation. The fusion process follows preset data splicing rules, associating and binding speech features with noise features to generate a fused feature that simultaneously contains valid speech information and ambient noise information, enabling the subsequent calibration module to accurately distinguish between speech and noise regions. The preset data splicing rules can be as follows: a dual-channel parallel splicing architecture is used. The first channel is the downsampled speech acoustic feature matrix (dimension N×D1, where N is the number of audio frames and D1 is the speech feature dimension); the second channel is a fixed vector of the real-time output environmental representation (dimension 1×D2). The environmental representation vector is copied and extended along the frame dimension to generate an N×D2 matrix matching the number of speech feature frames, achieving temporal dimension alignment. After alignment, the dual-channel matrix is ​​directly spliced ​​along the feature dimension to obtain a fused feature matrix of dimension N×(D1+D2). After splicing, a learnable fusion weight layer is added to adaptively adjust the numerical proportions of the speech features and environmental representation, completing the data fusion.

[0054] The adaptive calibrator analyzes the distribution and intensity range of noise based on the environmental noise information carried in the fused features, and automatically generates a real-valued mask for the corresponding speech feature dimension. This real-valued mask is then used to perform region-by-region correction on the original speech acoustic features, suppressing and weakening noise regions marked by the mask, while preserving and enhancing effective speech regions. This process thoroughly filters out interference from environmental noise, ultimately yielding the calibrated features after noise removal.

[0055] The adaptive calibrator inputs the noise-reduced calibration features into an end-to-end speech recognition model that integrates CTC and attention encoder-decoder architectures. The model decodes the calibration features, performs word matching, and sentence reconstruction, generating fluent and accurate recognition text according to natural language rules. This recognition text fully restores the semantic content of the user's original speech. Based on this recognition text, the system completes subsequent user intent analysis, thereby advancing processes such as identity verification and parking decision-making. Natural language rules specifically include word segmentation rules, semantic extraction rules, and sentence normalization rules. Word segmentation rules employ a word segmentation dictionary tailored to in-vehicle spoken language scenarios, incorporating parking-related vocabulary (moving the car, driving out, getting in the car, reversing, parallel exit, etc.), modal particles, and colloquial filler words; prioritizing the segmentation of consecutive numbers, locative words, and reduplicated words; filtering out meaningless filler words such as "trouble," "a little," and "help me," while retaining core action nouns and locative nouns. Semantic extraction rules refer to establishing a keyword matching library for parking intentions, setting weights for core trigger words, and prioritizing the extraction of core requests such as "drive out," "move out," "move the car," and "get in the car." It distinguishes between three semantic branches: simple inquiries, parking requests, and termination operations, eliminating irrelevant chatter. Sentence regularization rules refer to standardizing and reorganizing incomplete or disordered sentences into a unified standard sentence format: "User requests vehicle parking." Simultaneously, it filters repetitive semantic fragments, merges consecutive identical parking requests, and outputs concise and standardized recognition text to the intent parsing module.

[0056] This application leverages downsampling to improve computational efficiency, achieves precise noise reduction through feature fusion and adaptive mask calibration, and combines a high-performance end-to-end speech recognition model to complete speech-to-text conversion. After layer-by-layer optimization, it significantly improves the accuracy and stability of speech recognition in noisy environments, effectively solving the problems of speech recognition failure and parsing errors caused by outdoor noise.

[0057] In an optional embodiment, before performing environment-adaptive speech processing on the speech signal, the method further includes: End-to-end co-training is performed on all model components in environment-adaptive speech processing to collaboratively optimize model parameters for environment representation extraction, acoustic feature calibration, and speech recognition.

[0058] In this embodiment, end-to-end collaborative training refers to a training method that jointly trains three types of model components: environmental representation extraction, acoustic feature calibration, and speech recognition, rather than training a single module individually. The aforementioned model components include all AI functional modules, including the sound environment coding-related model, the adaptive calibrator, and the end-to-end speech recognition model.

[0059] In this embodiment, the training objects are determined based on all functional modules under the environment adaptive speech processing architecture, including the environment feature extraction model, query key generator, and feature weighted fusion model corresponding to the sound environment encoding thread, as well as the downsampling module, adaptive calibrator, and end-to-end speech recognition model corresponding to the speech recognition thread.

[0060] Massive amounts of audio data were collected across various parking scenarios, noise levels, and vocal characteristics, and divided into training, validation, and test sets. The dataset covers typical parking scenarios such as underground garages, open-air parking lots, and roadside parking, and includes real human voices with different accents, speaking speeds, and timbres, as well as various environmental noise samples. An end-to-end joint training framework was adopted, linking environmental representation extraction, acoustic feature calibration, and speech recognition into a complete training chain. During training, forward propagation sequentially executes the entire process of environmental encoding, feature fusion, mask calibration, and speech recognition, obtaining the error value between the recognition result and the standard text. During backpropagation, the error is propagated back layer by layer to each model component, synchronously adjusting the internal model parameters of all components. The training process does not optimize any single module individually, but rather aims to optimize the overall speech recognition accuracy and noise reduction effect, adjusting all parameters in a coordinated manner to achieve collaborative adaptation across multiple modules.

[0061] The training and validation process is repeated cyclically to continuously iterate model parameters, thereby reducing recognition errors and improving noise suppression capabilities. Once the model's overall performance on the test dataset reaches a preset standard, training is stopped, and all optimized model parameters are solidified and burned into the vehicle's local storage unit.

[0062] This application uses end-to-end collaborative training to jointly optimize the parameters of the entire model, enabling deep adaptation of modules such as environment coding, feature calibration, and speech recognition, thereby improving the overall collaboration capability and comprehensive performance of the entire speech processing system and further enhancing the speech recognition effect in noisy environments.

[0063] In an optional embodiment, the user emitting the voice signal is verified based on the biometric and liveness characteristics carried by the voice signal, including: Step 1: Extract the voiceprint features and at least one liveness feature from the speech signal; Step 2: If the similarity between the voiceprint feature and the preset authorized voiceprint is greater than or equal to the first preset threshold, and the similarity between the liveness feature and the preset liveness feature is greater than or equal to the second preset threshold, confirm that the user has passed the vehicle owner identity verification. The preset authorized voiceprint is the voiceprint feature pre-set by the authorized vehicle owner stored locally in the vehicle, and the preset liveness feature is the liveness feature pre-set by the authorized vehicle owner stored locally in the vehicle.

[0064] In this embodiment, voiceprint features are unique acoustic features formed by human voice timbre, vocal frequency, etc. Each user possesses unique voiceprint information, which is the core basis for identity recognition. Preset authorized voiceprints refer to the voiceprint data of authorized vehicle owners collected and stored locally in advance, serving as the benchmark sample for voiceprint comparison. The first preset threshold refers to the voiceprint similarity judgment standard set by the system, used to distinguish between authorized and unauthorized users. Preset liveness features refer to the liveness features of authorized vehicle owners collected and stored locally in advance, serving as the benchmark sample for liveness verification. The second preset threshold refers to the liveness feature similarity judgment standard set by the system, used to distinguish between genuine user voice and forged signals such as AI-synthesized voice and electronically recorded voice-changing voice.

[0065] In this embodiment of the application, after semantic recognition determines that the intention to leave is executed, dual security verification is completed by combining voiceprint features and liveness features. Specifically, the system retrieves the complete speech signal after noise reduction and calibration, and performs two types of feature extraction operations respectively. That is, based on the timbre, fundamental frequency, formants and other information of the speech, voiceprint features are extracted, and based on the physiological laws of human vocalization, liveness features such as breathing rhythm, intonation fluctuations, and natural vibrato are extracted from the speech waveform.

[0066] The system extracts real-time voiceprint features and performs a global similarity comparison with preset authorized voiceprints stored offline in the vehicle. The comparison process uses a feature matching algorithm to calculate the overlap between the two sets of voiceprint data, obtaining a voiceprint similarity value. This value is then compared with a preset similarity threshold. Simultaneously, the system verifies the extracted liveness features, matching the real-time liveness features with pre-stored real-person voice samples (i.e., preset liveness features) to determine whether the current speech possesses the physiological characteristics of a real human voice, thus filtering out forgeries such as electronic recordings, AI-synthesized speech, and voice-changing speech. Liveness verification is considered successful only if the similarity between the liveness feature and the preset liveness feature is greater than or equal to a second preset threshold.

[0067] The system combines voiceprint comparison results with liveness detection results for a comprehensive judgment. If the voiceprint feature similarity is greater than or equal to a first preset threshold, and the liveness detection passes (i.e., the similarity between the liveness feature and the preset liveness feature is greater than or equal to a second preset threshold), the current speaker is determined to be the authorized vehicle owner, the overall identity verification passes, and the system proceeds to the subsequent parking direction decision process. If the voiceprint similarity is lower than the preset threshold, or the liveness detection fails, the identity verification fails, the system announces an identity verification failure message through the vehicle's external speaker, and simultaneously terminates the current parking interaction process, leaving the vehicle in its original state.

[0068] This application employs a dual verification mechanism combining voiceprint and liveness detection. It relies on voiceprint to accurately match authorized vehicle owners while using liveness detection to prevent malicious forgery methods such as recording and synthesized speech. This significantly enhances the security and anti-hacking capabilities of identity verification, preventing unauthorized personnel from accidentally triggering vehicle parking.

[0069] In an optional embodiment, before determining the parking direction of the vehicle based on the user's location, the method further includes determining the user's location in the following manner: Step 1: Obtain the user's candidate location. The candidate location includes at least one of the following: the user's first location obtained based on sound source localization, the user's second location obtained using an ultra-wideband (UWB) positioning module, the user's third location obtained based on Bluetooth-assisted positioning, and the user's fourth location obtained based on camera acquisition. The ultra-wideband (UWB) positioning module includes positioning base stations located at four positions: front left, front right, rear left, and rear right of the vehicle. The positioning base stations are used for bidirectional ranging and positioning with the UWB tag built into the vehicle key carried by the user. Step 2: Assign weights to each candidate position based on the confidence level of the data used to calculate each candidate position, where the confidence level is positively correlated with the data quality of each data point; Step 3: Perform weighted fusion on each candidate location to obtain the final user location.

[0070] In this embodiment, candidate location refers to user location data calculated using different positioning methods, with each method corresponding to an independent candidate location. Sound source localization is a method that infers the location of the speaker based on the reception time difference of an external microphone array. UWB positioning utilizes ultra-wideband positioning technology, relying on bidirectional ranging between the vehicle's positioning base station and the UWB tag embedded in the car key to achieve high-precision spatial positioning. Bluetooth-assisted positioning uses Bluetooth signal strength and transmission latency to achieve auxiliary positioning, used as supplementary positioning when UWB signals are blocked. Camera visual positioning relies on the vehicle's surround-view cameras to identify human body shapes and combines this with visual algorithms to calculate the user's location. Confidence score is used to evaluate the reliability and accuracy of a single location data point; higher data quality results in a higher confidence score. Data quality includes, but is not limited to, image clarity, signal strength, and voice clarity.

[0071] This embodiment describes a multimodal fusion positioning process for user location, which is executed after identity verification and before the decision on the departure direction, integrating multiple positioning methods to improve location accuracy and reliability.

[0072] In this embodiment, the system simultaneously activates four types of independent positioning units to calculate the user's location in parallel and generate multiple candidate locations. Specifically, it can employ sound source localization technology, utilizing a distributed microphone array on the vehicle body to calculate the location and distance of the speaker based on the time and intensity differences of the voice signals received by different microphones, generating a first location; it can employ UWB positioning technology, using four UWB positioning base stations deployed around the vehicle to conduct bidirectional ranging and communication with the UWB tag in the user's car key, calculating a high-precision location based on the ultra-wideband signal, generating a second location; it can employ Bluetooth-assisted positioning, using real-time communication between the car key and the vehicle's Bluetooth module to calculate the relative position based on Bluetooth signal strength and transmission delay, generating a third location as backup positioning data in occluded scenarios; and it can employ camera visual positioning, using surround-view cameras to capture images outside the vehicle, using a human body recognition algorithm to locate the user's figure, and combining this with a visual ranging algorithm to calculate the location, generating a fourth location.

[0073] The system evaluates the acquisition environment, signal status, and data integrity of each candidate location group to determine the confidence level of the single location data. For example, if the UWB signal is unobstructed and the data is stable, the second location corresponding to UWB has the highest confidence level. If there are walls or vehicles obstructing the UWB signal, causing it to weaken, the UWB confidence level is reduced, while the confidence levels of Bluetooth positioning and visual positioning are increased. At the same time, the system dynamically adjusts the confidence levels of sound source positioning and visual positioning based on environmental noise and image clarity, and finally assigns a corresponding weight coefficient to each candidate location.

[0074] Using the assigned confidence level as a weighting coefficient, a weighted fusion calculation is performed on the orientation and distance data of all candidate locations. Location data with higher confidence levels have a larger proportion in the final result, thereby mitigating the bias caused by low-quality positioning data. After the fusion calculation is completed, a unique and accurate end-user location is output, including complete spatial information such as the user's precise azimuth angle and straight-line distance relative to the vehicle. This location data is then transmitted to the intelligent driving domain controller for subsequent parking direction decisions.

[0075] This application integrates multiple positioning methods to achieve multimodal complementarity, and through dynamic weighted fusion of confidence levels, it balances the positioning accuracy in normal scenarios with the positioning continuity in occluded scenarios, effectively improving the accuracy of user location detection and environmental adaptability, and providing reliable location data for berthing direction decision-making.

[0076] In an optional embodiment, determining the parking direction of the vehicle based on the user's location includes: Step 1: Calculate the azimuth and distance of the end user's location relative to the front of the vehicle, and identify the parking space type where the vehicle is located; Step 2a: If the parking space type is a parallel parking space and the azimuth angle is within the preset acute angle range, determine that parking forward is the parking direction; Step 2b: If the parking space type is a parallel parking space and the azimuth angle is within the preset obtuse angle range, determine that parking backwards is the parking direction. Step 2c: If the parking space type is a perpendicular parking space and the azimuth angle is within the preset acute angle range, determine the parking direction as parking out to the side forward. Step 2d: If the parking space type is a perpendicular parking space and the azimuth angle is within the preset obtuse angle range, determine the parking direction as parking to the rear side. Step 2e: When the parking space type is a slanted parking space, determine the direction from which the user can directly get into the car after parking as the parking direction; In step 2f, if there are obstacles in all directions or the distance is less than the preset safe distance, the parking operation is paused and the user is prompted to adjust the position.

[0077] In this embodiment, the azimuth angle of the end user's position relative to the front of the vehicle refers to the angle formed by the user's position and the centerline of the front of the vehicle, with the front of the vehicle as the reference, used to determine the user's relative orientation. Parking space type refers to the parking space configuration, categorized into parallel, perpendicular, and angled parking spaces. The preset acute angle range refers to the azimuth angle interval set by the system, representing the area where the user is located on the front side of the vehicle, for example, 0°±30°. The preset obtuse angle range refers to the azimuth angle interval set by the system, representing the area where the user is located on the rear side of the vehicle, for example, 180°±30°. The preset safety distance refers to the minimum safe distance threshold between the vehicle and the user; a distance less than this poses a collision risk, for example, 0.3m.

[0078] This embodiment describes the decision-making process for parking direction, which is executed based on the fused user location, parking space, and obstacle data. It is the core decision-making step for controlling vehicle parking.

[0079] In this embodiment, the system receives the end user's location, calculates the user's azimuth angle relative to the front of the vehicle and the straight-line distance between the user and the vehicle; at the same time, it retrieves images and radar data collected by the environmental perception module, identifies the type of parking space the vehicle is currently in (parallel parking space / perpendicular parking space / angled parking space), and simultaneously acquires the distribution data of all obstacles around the vehicle.

[0080] The system categorizes parking spaces by type and executes differentiated decision-making logic based on the user's azimuth angle range. Specifically, for parallel parking scenarios, if the user's azimuth angle is within a preset acute angle range (i.e., the user is at the front of the car), to avoid collisions when parking backward, the system determines that parking forward is the appropriate direction; if the user's azimuth angle is within a preset obtuse angle range (i.e., the user is at the rear of the car), then parking backward is the appropriate direction. For perpendicular parking scenarios, if the user's azimuth angle is within a preset acute angle range, considering the distribution of obstacles behind, the system determines that parking forward to the side is the appropriate direction; if the user's azimuth angle is within a preset obtuse angle range, then parking backward to the side is the appropriate direction. For angled parking scenarios, considering both the parking space's tilt angle and the user's standing position, the system prioritizes the driving direction that allows the user to directly board the vehicle after parking, as the appropriate parking direction.

[0081] In this embodiment, if obstacles are detected in all parking directions, preventing the vehicle from exiting normally; or if the distance between the vehicle and the pedestrian is less than a preset safe distance, posing a risk of personal collision, all parking processes are immediately suspended. The user is then prompted via external voice and light signals to adjust their position. If there is no safety risk or omnidirectional obstacles, the appropriate parking direction determined above is locked, a formal parking control command is generated, and sent to the vehicle execution module to initiate the automatic parking action.

[0082] This application combines multi-dimensional data such as parking space type, user location, safety distance, and obstacle distribution to intelligently determine the parking direction, eliminating the need for manual direction selection. This not only matches user habits but also avoids the risk of collisions between people and vehicles and vehicle scrapes from the source, significantly improving the intelligence level and operational safety of automatic parking.

[0083] In an optional embodiment, the method further includes: During the process of controlling the vehicle to perform a parking operation, parking status prompts are broadcast through the vehicle's external speakers; The turn signals indicate the optimal parking direction, and the brake lights are continuously illuminated to alert surrounding vehicles and pedestrians. The docking status information is pushed synchronously through the linked mobile application.

[0084] This embodiment describes a multi-channel status feedback process during vehicle parking, which runs synchronously throughout the entire parking operation and serves as an auxiliary interactive element of the main method. Specifically: As the vehicle begins its parking maneuver, the system activates external speakers positioned at the front and rear of the vehicle to broadcast voice announcements. Depending on the current process stage, the system sequentially broadcasts the identity verification result, parking direction, operating status, and any abnormality alerts. Upon completion of parking, it broadcasts a message indicating that parking is finished and passengers can board, allowing users outside the vehicle to monitor the vehicle's status in real time.

[0085] The system synchronously controls the vehicle's lighting system. For example, according to the determined optimal parking direction, the corresponding turn signal is illuminated to clearly indicate the direction the vehicle is about to travel. Throughout the parking maneuver, the vehicle's brake lights remain constantly on, warning surrounding vehicles and pedestrians to take evasive action. After the parking maneuver is completed, the lights automatically return to normal.

[0086] For users who have already bound their vehicles, the system synchronously pushes the status data of each operational node to the user's mobile application via the vehicle communication module throughout the entire parking process. Users can remotely view the entire process information, including parking start-up, operation, and completion, without needing to be near the vehicle; this feedback method only provides status display, does not participate in vehicle control, and does not rely on the mobile communication network to complete the parking operation.

[0087] This application provides simultaneous feedback on parking status through three channels: external voice prompts, vehicle lights, and a mobile app. This enables multi-dimensional human-computer interaction and external warnings, enhancing users' awareness of vehicle status and effectively alerting nearby pedestrians and vehicles, thereby improving the safety and interactive experience of the parking process.

[0088] In an optional embodiment, the method further includes: The voice signal is analyzed by performing fuzzy semantic understanding through a built-in lightweight edge AI agent model to parse the parking intention. The lightweight edge AI agent model does not require a fixed wake word or fixed sentence structure, supports self-learning based on user history interaction data, and adapts to the accents and speech speeds of different car owners.

[0089] In this embodiment, the lightweight edge AI agent resides locally in the vehicle and remains in standby mode throughout. Once the external voice module collects the user's voice signal and completes noise reduction calibration, the recognized text is directly input into the agent. No fixed wake-up word or command phrase needs to be preset; any natural spoken language related to parking can trigger parsing. The agent incorporates a fuzzy semantic understanding algorithm to segment and semantically decompose colloquial, fragmented, and inconsistent natural language statements, removing redundant colloquial words and extracting the core intent. Regardless of the user's expression, the core intent of "the vehicle needs to park" can be accurately identified.

[0090] The system continuously records each user's voice content, accent features, and speech rate characteristics, forming a user-specific interaction database. The intelligent agent periodically retrieves this database for local self-learning, continuously optimizing the parsing logic and gradually adapting it to the user's unique accent, speaking speed, and commonly used colloquial expressions. When encountering speech with low recognition confidence, it combines historical interaction data to assist in the judgment, reducing the probability of recognition failure.

[0091] After completing the fuzzy semantic parsing, the agent outputs the standardized berthing intention command to the downstream module to advance the identity verification, berthing decision and other processes; if the semantics cannot be recognized, a voice prompt is triggered to guide the user to rephrase the command.

[0092] This application relies on a lightweight AI agent on the edge to achieve fuzzy semantic parsing without fixed commands, adapting to users' natural spoken language interaction; combined with local self-learning capabilities, it adapts to the accents and speaking habits of different users, getting rid of the format limitations of traditional voice commands, and greatly improving the naturalness and ease of use of external voice interaction.

[0093] According to another aspect of the embodiments of this application, a vehicle parking exit control system is provided, comprising: The voice module is used to receive voice signals from users outside the vehicle and interpret their intentions through a built-in intelligent agent understanding model. The intelligent cockpit controller is used to process voice signals, extract voiceprint features and compare them to complete the vehicle owner's identity verification. The vehicle owner location module is used to accurately determine the vehicle owner's position and distance relative to the vehicle. The intelligent driving domain controller, as the core of system decision-making, integrates multi-source data, determines the parking direction, and controls the parking process; An environmental perception module is used for identifying surrounding obstacles; The vehicle execution module is used to receive instructions from the intelligent driving domain controller, complete steering, acceleration / deceleration, and gear shifting operations, and achieve autonomous parking. The feedback and notification module is used to synchronize the parking status with the car owner through multiple channels.

[0094] In this embodiment, the external microphone group of the voice module collects the user's voice signal, performs analog-to-digital conversion and preliminary noise reduction, and simultaneously undertakes subsequent voice broadcasting through an external speaker, achieving integrated voice input and output. The processed voice data is transmitted to the intelligent cockpit controller. The intelligent cockpit controller receives the voice data, extracts voiceprint features and liveness features, and retrieves the local authorized feature library to complete dual identity verification; the verification result and voice parsing intent are simultaneously sent to the intelligent driving domain controller. After successful identity verification, the vehicle owner location module activates the multimodal positioning function, integrating UWB, Bluetooth, visual, and sound source positioning data to calculate the user's precise location and upload the location data to the intelligent driving domain controller in real time. The environmental perception module uses LiDAR, 4D millimeter-wave radar, and surround-view cameras to scan the parking space shape and obstacle distribution from all directions, transmitting environmental data and parking space data to the intelligent driving domain controller. As the core of the system, the intelligent driving domain controller integrates four types of data: identity information, user location, parking space, and obstacles, determines the optimal parking direction and driving path according to preset logic, generates complete control commands, and sends them to the vehicle execution module. The vehicle execution module receives control commands and coordinates the steering, power, braking, and gear shifting mechanisms to automatically park according to the planned path. During the driving process, it receives obstacle monitoring data from the environmental perception module in real time to enable emergency braking and hazard avoidance. Throughout the entire parking cycle, the feedback and prompting module simultaneously activates the external speakers, vehicle lights, and mobile app to push real-time operating status, abnormal prompts, and completion notifications to the user and surrounding entities, ensuring the entire system operates in a closed loop.

[0095] This application utilizes an environment-adaptive speech processing method, relying on environmental representation to calibrate speech acoustic features, effectively suppressing environmental noise interference, improving the accuracy of speech recognition and intent parsing in complex scenarios, and solving the technical problem of low accuracy in speech command recognition in noisy environments.

[0096] According to another aspect of the embodiments of this application, such as Figure 4 As shown, a vehicle parking control device is provided, comprising: The voice acquisition module 401 is used to acquire the ambient sound signals of the target vehicle and the voice signals emitted by the user outside the vehicle. The speech processing and semantic recognition module 403 is used to recognize the user's intention after performing environment-adaptive speech processing on the speech signal. The environment-adaptive speech processing includes using the environmental representation of the environmental sound signal to suppress environmental noise in the speech signal, thereby calibrating the acoustic features of the speech signal, and generating recognition text after calibration. The recognition text is used to recognize the user's intention. The identity verification and direction decision module 405 is used to verify the identity of the user who sent the voice signal based on the biometric and liveness characteristics carried by the voice signal when the user's intention is to park out. If the verification is successful, the parking direction of the vehicle is determined based on the user's location. The execution module 407 is used to control the vehicle to perform a parking operation according to the parking direction.

[0097] It should be noted that the voice acquisition module 401 in this embodiment can be used to execute step S102 in this application embodiment, the voice processing and semantic recognition module 403 in this embodiment can be used to execute step S104 in this application embodiment, the identity verification and direction decision module 405 in this embodiment can be used to execute step S106 in this application embodiment, and the execution module 407 in this embodiment can be used to execute step S108 in this application embodiment.

[0098] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules, as part of a device, can operate in environments such as... Figure 1 The hardware environment shown can be implemented either through software or through hardware.

[0099] Optionally, the speech processing and semantic recognition module is specifically used for: extracting acoustic features of the ambient sound signal; mapping the acoustic features of the ambient sound signal into a query key vector through a query key generator; calculating the similarity between the query key vector and each known environmental representation in the environmental representation lookup table using a pre-established environmental representation lookup table, and using this similarity as the attention weight of each known environmental representation; and weighting and summing each known environmental representation according to its respective attention weight to obtain the environmental representation of the ambient sound signal.

[0100] Optionally, the speech processing and semantic recognition module is further configured to: downsample the speech signal to obtain the acoustic features of the speech signal; fuse the acoustic features of the speech signal with the environmental representation of the environmental sound signal to obtain fused features; generate a real-valued mask based on the fused features using a pre-trained adaptive calibrator, and use the real-valued mask to calibrate the acoustic features of the speech signal to suppress noise, thereby obtaining calibrated features of the acoustic features of the speech signal; and input the calibrated features into an end-to-end speech recognition model to obtain the recognized text output by the end-to-end speech recognition model.

[0101] Optionally, the speech processing and semantic recognition module also includes an end-to-end training module, specifically used for: end-to-end collaborative training of all model components in environment-adaptive speech processing to collaboratively optimize model parameters for environment representation extraction, acoustic feature calibration, and speech recognition.

[0102] Optionally, the identity verification and direction decision module is specifically used to: extract voiceprint features and at least one liveness feature from the voice signal; and confirm that the user has passed the vehicle owner identity verification when the similarity between the voiceprint feature and the preset authorized voiceprint is greater than or equal to a first preset threshold, and the similarity between the liveness feature and the preset liveness feature is greater than or equal to a second preset threshold. The preset authorized voiceprint is the voiceprint feature pre-set by the authorized vehicle owner stored locally in the vehicle, and the preset liveness feature is the liveness feature pre-set by the authorized vehicle owner stored locally in the vehicle.

[0103] Optionally, the identity verification and direction decision module further includes a vehicle owner positioning unit, specifically used for: obtaining candidate locations of the user, wherein the candidate locations include at least one of the following: a first location of the user obtained based on sound source localization, a second location of the user obtained using an ultra-wideband (UWB) positioning module, a third location of the user obtained based on Bluetooth-assisted positioning, and a fourth location of the user obtained based on camera acquisition; the ultra-wideband (UWB) positioning module includes positioning base stations deployed at four locations on the front left, front right, rear left, and rear right of the vehicle, and the positioning base stations are used for bidirectional ranging and positioning with the UWB tag embedded in the vehicle key carried by the user; assigning weights to each candidate location based on the confidence level of the data used to calculate each candidate location, wherein the confidence level is positively correlated with the data quality of each data point; and performing weighted fusion of each candidate location to obtain the final user location.

[0104] Optionally, the identity verification and direction decision module is further configured to: calculate the azimuth and distance of the end user's position relative to the front of the vehicle, and identify the parking space type where the vehicle is located; if the parking space type is a parallel parking space and the azimuth is within a preset acute angle range, determine that parking forward is the parking direction; if the parking space type is a parallel parking space and the azimuth is within a preset obtuse angle range, determine that parking backward is the parking direction; if the parking space type is a perpendicular parking space and the azimuth is within a preset acute angle range, determine that parking forward to the side is the parking direction; if the parking space type is a perpendicular parking space and the azimuth is within a preset obtuse angle range, determine that parking backward to the side is the parking direction; if the parking space type is an angled parking space, determine that the direction in which the user can directly get into the vehicle after parking is the parking direction; if there are obstacles in all directions or the distance is less than a preset safe distance, suspend the parking operation and prompt the user to adjust the position.

[0105] According to another aspect of the embodiments of this application, this application provides a vehicle, such as Figure 5As shown, the system includes a memory 501, a processor 503, a communication interface 505, and a communication bus 507. The memory 501 stores a computer program that can run on the processor 503. The memory 501 and the processor 503 communicate through the communication interface 505 and the communication bus 507. When the processor 503 executes the computer program, it implements the steps of the above method.

[0106] The memory and processor in the aforementioned vehicle communicate with each other via a communication bus and communication interface. The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0107] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0108] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0109] According to another aspect of the embodiments of this application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of any of the above embodiments.

[0110] Optionally, in embodiments of this application, the computer-readable medium is configured to store program code for the processor to perform the following steps: Acquire ambient sound signals from the target vehicle and voice signals from users outside the vehicle; After performing environment-adaptive speech processing on the speech signal, the user's intention is identified. The environment-adaptive speech processing includes using the environmental representation of the environmental sound signal to suppress environmental noise in the speech signal, thereby calibrating the acoustic features of the speech signal, and generating recognition text after calibration. The recognition text is used to identify the user's intention. If the user's intention is to park out, the user who sent the voice signal is identified based on the biometric and liveness characteristics carried in the voice signal. If the verification is successful, the parking direction of the vehicle is determined based on the user's location. Control the vehicle to perform the parking operation according to the parking direction.

[0111] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0112] In specific implementation, the embodiments of this application can be referred to the above embodiments and have corresponding technical effects.

[0113] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0114] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0115] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0116] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0117] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0118] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0119] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0120] If the aforementioned function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks. It should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0121] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A vehicle parking exit control method, characterized in that, include: Acquire ambient sound signals from the target vehicle and voice signals from users outside the vehicle; After performing environment-adaptive speech processing on the speech signal, the user's intention is identified. The environment-adaptive speech processing includes using the environmental representation of the environmental sound signal to suppress environmental noise in the speech signal, thereby calibrating the acoustic features of the speech signal, and generating recognition text after calibration. The recognition text is used to identify the user's intention. If the user's intention is to park, the user who sent the voice signal is identified based on the biometric and liveness characteristics carried by the voice signal. If the verification is successful, the parking direction of the vehicle is determined based on the user's location. According to the parking direction, control the vehicle to perform a parking operation; The environmental adaptive speech processing of the speech signal includes the following steps performed by a sound environment coding thread: extracting acoustic features of the environmental sound signal; mapping the acoustic features of the environmental sound signal into a query key vector using a query key generator; calculating the similarity between the query key vector and each known environmental representation in the environmental representation lookup table using a pre-established environmental representation lookup table, and using this similarity as the attention weight of each known environmental representation; and weighting and summing each known environmental representation according to its respective attention weight to obtain the environmental representation of the environmental sound signal.

2. The method according to claim 1, characterized in that, The environmental adaptive speech processing of the speech signal further includes performing the following steps through a speech recognition thread: The speech signal is downsampled to obtain its acoustic features; The acoustic features of the speech signal are fused with the environmental representation of the ambient sound signal to obtain fused features; A real-valued mask is generated based on the fusion features using a pre-trained adaptive calibrator, and the acoustic features of the speech signal are calibrated using the real-valued mask to suppress noise, thereby obtaining the calibrated features of the acoustic features of the speech signal. The calibration features are input into the end-to-end speech recognition model to obtain the recognized text output by the end-to-end speech recognition model.

3. The method according to any one of claims 1 to 2, characterized in that, Before performing environment-adaptive speech processing on the speech signal, the method further includes: All model components in the environment-adaptive speech processing are trained end-to-end to collaboratively optimize model parameters for environment representation extraction, acoustic feature calibration, and speech recognition.

4. The method according to claim 1, characterized in that, The step of verifying the identity of the user who emitted the voice signal based on the biometric and liveness characteristics carried by the voice signal includes: Extract the voiceprint features and at least one liveness feature from the speech signal; If the similarity between the voiceprint feature and the preset authorized voiceprint is greater than or equal to a first preset threshold, and the similarity between the liveness feature and the preset liveness feature is greater than or equal to a second preset threshold, the user is confirmed to have passed the vehicle owner identity verification. The preset authorized voiceprint is the voiceprint feature pre-set by the authorized vehicle owner and stored locally in the vehicle. The preset liveness feature is the liveness feature pre-set by the authorized vehicle owner and stored locally in the vehicle.

5. The method according to claim 1, characterized in that, Before determining the parking direction of the vehicle based on the user's location, the method further includes determining the user's location in the following manner: The candidate locations of the user are obtained, wherein the candidate locations include at least one of the following: a first location of the user obtained based on sound source localization, a second location of the user obtained using an ultra-wideband (UWB) positioning module, a third location of the user obtained based on Bluetooth-assisted positioning, and a fourth location of the user obtained based on camera acquisition. The ultra-wideband (UWB) positioning module includes positioning base stations arranged at four positions: front left, front right, rear left, and rear right of the vehicle. The positioning base stations are used to perform bidirectional ranging and positioning with the UWB tag built into the vehicle key carried by the user. Weights are assigned to each candidate position based on the confidence level of the data used to calculate each candidate position, wherein the confidence level is positively correlated with the data quality of each data point; The candidate locations are weighted and fused to obtain the final user location.

6. The method according to claim 5, characterized in that, The step of determining the parking direction of the vehicle based on the user's location includes: Calculate the azimuth angle and distance of the end user's location relative to the front of the vehicle, and identify the parking space type of the parking space where the vehicle is located; When the parking space type is a parallel parking space and the azimuth angle is within a preset acute angle range, the parking direction is determined to be forward parking. When the parking space type is a parallel parking space and the azimuth angle is within a preset obtuse angle range, the parking direction is determined to be parking backwards; When the parking space type is a perpendicular parking space and the azimuth angle is within the preset acute angle range, the parking direction is determined to be parking forward and to the side. When the parking space type is a perpendicular parking space and the azimuth angle is within the preset obtuse angle range, the parking direction is determined to be parking to the rear side. When the parking space type is an angled parking space, the direction in which the user can directly get into the car after parking is determined as the parking direction; If there are obstacles in all directions or the distance is less than the preset safe distance, the parking operation will be paused and the user will be prompted to adjust the position.

7. A vehicle parking exit control device for implementing the vehicle parking exit control method as described in any one of claims 1 to 6, characterized in that, include: The voice acquisition module is used to acquire ambient sound signals of the target vehicle and voice signals emitted by users outside the vehicle. The speech processing and semantic recognition module is used to recognize the user's intent after performing environment-adaptive speech processing on the speech signal. The environment-adaptive speech processing includes using the environmental representation of the environmental sound signal to suppress environmental noise in the speech signal, thereby calibrating the acoustic features of the speech signal, and generating recognition text after calibration. The recognition text is used to recognize the user's intent. The identity verification and direction decision module is used to verify the identity of the user who sent the voice signal based on the biometric and liveness features carried by the voice signal when the user's intention is to park out, and to determine the parking direction of the vehicle based on the user's location if the verification is successful. The execution module is used to control the vehicle to perform a parking operation according to the parking direction.

8. A vehicle comprising a memory, a processor, a communication interface, and a communication bus, wherein the memory stores a computer program executable on the processor, and the memory and the processor communicate via the communication bus and the communication interface, characterized in that... When the processor executes the computer program, it implements the vehicle parking control method according to any one of claims 1 to 6.

9. A computer-readable medium having processor-executable non-volatile program code, characterized in that, The program code causes the processor to execute the vehicle parking control method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Automobile voice command acquiring and processing system and method

    CN103928026A

  • Voice control parking system and method, vehicle and storage medium

    CN115910056A