Voiceprint recognition method and system for vehicle-mounted child safety seat
By simultaneously acquiring microphone audio and radio frequency signals in a car child safety seat, and utilizing a multi-task deep neural network model with feature suppression and prosodic pattern analysis, the problem of misjudgment when children imitate speech in electromagnetic interference environments is solved, ensuring the safety and reliability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies in voiceprint recognition systems for in-vehicle child safety seats fail to effectively address the issue of misjudgment when children imitate adult commands in specific electromagnetic interference environments such as highway service areas, leading to safety risks.
By simultaneously acquiring the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals in preset frequency bands, interference markers are generated. Then, a multi-task deep neural network model is used for feature suppression and prosodic pattern analysis to distinguish between adult and child voice commands, generate voiceprint judgment scores, and generate control authorization commands only when the score is higher than a preset compliance threshold.
This effectively prevents children's imitation of voice from being misinterpreted as adult commands in specific electromagnetic interference environments, ensuring children's safety and improving the system's reliability and safety in severe electromagnetic environments.
Smart Images

Figure CN122050367A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice recognition technology, and more specifically, to a voiceprint recognition method and system for vehicle child safety seats. Background Technology
[0002] With the development of smart cockpit technology, voice interaction has become an important means to improve the ease of operation of in-vehicle equipment. This includes introducing voice interaction into the control of in-vehicle child safety seats. Users can adjust the seat's heating, ventilation, or angle adjustment through voice commands without having to turn around or operate manually. This can provide convenience for caregivers while driving and reduce interference with the driver's attention.
[0003] Considering that children may accidentally trigger the above functions out of curiosity or unconsciously imitate their parents' pronunciation, in order to ensure children's safety, some existing technical solutions integrate a voiceprint verification subsystem into the voice control link of the voice interaction system. After receiving a voice command, the verification subsystem extracts the speaker's voiceprint features and classifies them for adults and children, aiming to ensure that control is limited to adults based on biometrics.
[0004] Since the voiceprint verification subsystem technology is not fully mature, existing technologies have attempted to train more interference-resistant models through data augmentation. For example, Chinese patent application CN120183406A discloses a speech recognition method that superimposes single-person speech samples with different types of interference sounds such as music, TV background noise, and tire noise. Positive samples are constructed for the same speaker as the single-person speech sample, and negative samples are constructed for different speakers. These samples are then used to train the speech recognition model in two stages. This method improves the recognition accuracy of the model in general noise environments by enriching the types of noise in the training samples.
[0005] While the aforementioned methods can improve speech recognition accuracy, they primarily focus on enhancing the model's resistance to broad random noise. The interference sounds in their training data are all general background noise. Although they can effectively isolate background noise interference, in real-world in-vehicle scenarios, especially in environments with specific electromagnetic interference such as highway service areas, high-power FM radio signals from service areas may couple into the equipment circuitry, generating deterministic harmonic interference that falls into the speech frequency band. This interference differs from ordinary noise; it systematically distorts the spectral characteristics of the voiceprint. When children imitate adult commands in this environment, their speech will exhibit a mixed feature contaminated by interference. Existing models trained on general noise are not optimized for this type of interference coupled with imitation behavior, and their voiceprint classifiers may fail, leading to misjudging children's speech as adult commands.
[0006] More seriously, since the vehicle may be parked at a service area at this time, the driver may not be aware of the situation due to needing a rest. This misauthorization could lead to the child unintentionally triggering the seat function, potentially causing safety risks. Existing technologies, including the aforementioned published patent applications, have not recognized the safety hazards in this specific scenario, nor have they provided any targeted solutions. Summary of the Invention
[0007] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution:
[0008] In a first aspect, the present invention discloses a voiceprint recognition method for a vehicle-mounted child safety seat, comprising:
[0009] Simultaneously acquire the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals in preset frequency bands;
[0010] When the intensity of a spatial radio frequency signal exceeds a preset intensity threshold within any acquisition time window, an interference marker is generated for the microphone audio signal acquired within the same acquisition time window.
[0011] The microphone audio signal with interference tags is input into the pre-trained voiceprint recognition model;
[0012] The voiceprint recognition model was trained using a training set containing the following samples: positive samples are pure instructional speech from adults, and negative samples are synthetic interference speech generated by coupling pure instructional speech imitated by children with simulated harmonic interference.
[0013] The voiceprint recognition model is configured to perform:
[0014] Based on interference markers, feature suppression processing for harmonic interference related to preset frequency bands in microphone audio signals is activated.
[0015] The microphone audio signal after feature suppression processing is analyzed to generate a voiceprint determination score;
[0016] The final control authorization command is generated only when the voiceprint judgment score is higher than the preset compliance threshold.
[0017] Furthermore, the voiceprint recognition model is a multi-task deep neural network, which includes a shared feature extraction layer and an instruction compliance judgment task branch; the shared feature extraction layer is used to extract basic acoustic features from the input signal; the instruction compliance judgment task branch is used to output a voiceprint judgment score.
[0018] Furthermore, the training methods for voiceprint recognition models include:
[0019] Obtain clean, adult-grade instruction speech as positive sample speech;
[0020] Obtain the pure instructional speech imitated by children as the negative sample base speech;
[0021] The pure instructional speech imitated by children is coupled in the time and frequency domain with harmonic interference noise of simulated preset frequency radio broadcast signals to generate synthetic interference speech as negative sample speech.
[0022] A training set is constructed based on positive and negative sample speech, where training labels are used to distinguish whether the speech is a valid adult authorization instruction;
[0023] The voiceprint recognition model is trained using a training set.
[0024] Furthermore, methods for generating synthetic interference speech include:
[0025] The interference noise generated by the pure instructional speech imitated by the child and the characteristics of the simulated preset frequency band radio broadcast signal is linearly superimposed in the time and frequency domain. The spectral energy of the interference noise is concentrated in the frequency positions corresponding to the harmonics of each preset frequency band in the 300-3400Hz speech baseband.
[0026] Furthermore, the methods for feature suppression processing include:
[0027] Based on the interference markers, a frequency domain mask is dynamically generated. The frequency domain mask is used to attenuate or zero the energy of the microphone audio signal in the frequency domain corresponding to the harmonic interference frequency range.
[0028] Furthermore, the preset frequency band is the FM broadcast band from 87MHz to 108MHz; the preset intensity threshold is dynamically calibrated based on the average electromagnetic environmental noise inside the vehicle when the vehicle is ignited.
[0029] Furthermore, the process of generating a voiceprint judgment score by the voiceprint recognition model also includes prosodic pattern analysis. Specifically, prosodic pattern analysis includes: calculating the fundamental frequency profile stability index of the microphone audio signal, analyzing the statistical regularity of the accent distribution, and comparing the index and regularity with the pre-stored prosodic template of adult self-spoken speech.
[0030] Furthermore, after generating the voiceprint determination score and before the determination step, a confidence correction step is also included:
[0031] When interference markers are present, the voiceprint determination score is adjusted downward based on the magnitude by which the spatial radio frequency signal strength exceeds a preset strength threshold.
[0032] Furthermore, the method for simultaneously acquiring the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals in a preset frequency band includes:
[0033] The microphone audio signal and the spatial radio frequency signal strength are collected by physically isolated microphone preamplifier circuit and radio frequency strength detection circuit respectively, and the signal sampling clocks of the microphone preamplifier circuit and the radio frequency strength detection circuit are kept synchronized by the same clock source.
[0034] Secondly, this invention discloses a voiceprint recognition system for a vehicle-mounted child safety seat, used to implement the aforementioned voiceprint recognition method for a vehicle-mounted child safety seat. The voiceprint recognition system for a vehicle-mounted child safety seat includes: a synchronous acquisition and marking module, a voiceprint recognition module, and a voiceprint arbitration module.
[0035] The synchronous acquisition and tagging module is used to synchronously acquire the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals, and generate an interference tag for the corresponding microphone audio signal when the spatial radio frequency signal strength exceeds a preset strength threshold.
[0036] The voiceprint recognition module is a pre-trained voiceprint recognition model used to receive microphone audio signals with interference tags, perform feature suppression processing based on interference tags, and output a voiceprint judgment score.
[0037] The voiceprint arbitration module is used to receive voiceprint judgment scores and generate a final control authorization command when the voiceprint judgment score is higher than a preset compliance threshold.
[0038] Compared with related technologies, the present invention has the following beneficial effects:
[0039] This invention breaks away from the conventional approach of processing only acoustic signals, creatively introducing synchronous monitoring of specific radio frequency signals in space. Upon detecting a broadcast signal exceeding a threshold, it immediately adds a precise timestamp to the voice frame at the same moment. This timestamp acts as a targeted navigation signal. Based on this precise navigation signal, a voiceprint recognition model is used for feature suppression, specifically weakening narrow frequency bands in the speech spectrum corresponding to known broadcast harmonics. This eliminates interference factors while preserving the original speech to the maximum extent. Furthermore, because the voiceprint recognition model is not trained using conventional clean speech or generalized noise data, but rather couples the child's imitated commands with simulated specific broadcast harmonic interference to create negative samples for training, it is highly effective in distinguishing between genuine adult commands disguised by specific interference and children's imitated commands. Ultimately, the model can determine whether to wake up the voice interaction system and issue control authorization commands based on the output voiceprint score, effectively avoiding misjudging interfered children's imitated voices as compliant adult commands during voiceprint verification, thus ensuring the safety of children alone in the car seat.
[0040] This invention utilizes dynamically generated frequency domain masks for feature suppression. Based on the threat indicated by the interference marker, a mask template that precisely corresponds to the current interference spectrum is calculated and generated in real time. This template can selectively attenuate or reduce the microphone audio signal to zero in the specific narrowband range where the interference is most severe in the frequency domain, while preserving the complete information of other uncontaminated frequency bands of the speech to the maximum extent, thereby improving the possibility of reliable discrimination of this information in severe electromagnetic environments.
[0041] This invention captures unnatural features such as rhythmic breaks, rhythmic imbalances, or harsh intonation commonly found in children's imitations of pronunciation through prosodic pattern analysis. Even if interference blurs the timbre features, the system can still effectively identify unnatural imitations of children's pronunciations based on these deep-seated behavioral pattern differences that are not easily distorted by electromagnetic interference. This reduces the probability of misjudging interfered children's imitations as compliant adult instructions. Attached Figure Description
[0042] Figure 1 A flowchart illustrating the steps of the voiceprint recognition method for a vehicle-mounted child safety seat provided by the present invention;
[0043] Figure 2 This is a flowchart illustrating the data processing in the voiceprint recognition system for a vehicle-mounted child safety seat provided by the present invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Example 1
[0046] Please see Figure 1 As shown, this embodiment provides a voiceprint recognition method for in-vehicle child safety seats, which can be applied to the voice interaction system of child safety seats. The method includes steps one to six.
[0047] Step 1: Simultaneously acquire the spatial radio frequency signal strength of the microphone audio signal and the radio broadcast signal of the preset frequency band.
[0048] Step 2: When the spatial radio frequency signal strength exceeds the preset strength threshold within any acquisition time window, an interference mark is generated for the microphone audio signal acquired within the same acquisition time window.
[0049] Since step one mainly introduces the simple preliminary data acquisition, this section will describe and explain the part about generating interference markers in conjunction with step two.
[0050] For example, the preset frequency band is the FM broadcast band from 87MHz to 108MHz; the preset intensity threshold is dynamically calibrated based on the average electromagnetic ambient noise inside the vehicle when the vehicle is ignited.
[0051] The microphone audio signal refers to the raw analog electroacoustic signal collected by a high signal-to-noise ratio omnidirectional microphone positioned on the side or in front of the child safety seat headrest. In this embodiment, the microphone's frequency response range covers the 300Hz-3400Hz voice baseband, and its output is digitized via a 24-bit high-precision analog-to-digital converter (ADC), with a sampling frequency set to 16kHz to meet the needs of voice processing.
[0052] In this embodiment, the preset frequency band can be the FM broadcast band from 87MHz to 108MHz. This frequency band is chosen because in environments such as highway service areas and urban commercial areas, the high-power broadcast signals in this frequency band are known and common sources of strong electromagnetic interference. Therefore, it is preset to specifically monitor this frequency band. In other embodiments, the preset frequency band can also be the aviation radio band (usually 118-137MHz) to cope with special electromagnetic environments such as near airports.
[0053] The spatial radio frequency signal strength refers to the field strength of radio waves propagating in space within a preset frequency band, measured in decibels per milliwatt (dBm). In this embodiment, it can be measured using a dedicated radio frequency strength detection circuit. This circuit includes a broadband antenna tuned to 87-108MHz, a low-noise amplifier, and a radio frequency power detection chip, which converts the detected radio frequency signal strength into a DC voltage signal, which is then read by an ADC.
[0054] The acquisition time window refers to the basic time unit for signal processing by the system. To ensure strict time alignment between acoustic and radio frequency signals, in this embodiment, a fixed-length time frame can be used as the acquisition time window, for example, each frame is 32 milliseconds long, corresponding to 512 audio sampling points. The radio frequency signal strength is also sampled once within each synchronization time window and the average value is calculated.
[0055] An exemplary method for synchronously acquiring the spatial radio frequency signal strength of a microphone audio signal and a radio broadcast signal in a preset frequency band includes: acquiring the microphone audio signal and the spatial radio frequency signal strength using a physically isolated microphone preamplifier circuit and a radio frequency strength detection circuit, respectively, wherein the signal sampling clocks of the microphone preamplifier circuit and the radio frequency strength detection circuit are kept synchronized by the same clock source.
[0056] Specifically, the audio acquisition path can be microphone → preamplifier (with anti-aliasing filter) → audio ADC.
[0057] The radio frequency detection path can be: broadband antenna → 87-108MHz bandpass filter → low noise amplifier → radio frequency power detector → radio frequency ADC.
[0058] The sampling clocks of both ADCs can be generated by dividing the same high-stability crystal oscillator, ensuring that the sampling start time and sampling interval of both are completely synchronized.
[0059] Within each acquisition time window, the system controller can read a frame of 512-point audio sampling data AF[t] from the audio ADC buffer, and simultaneously read multiple RF intensity values RS sampled within that time window from the RF ADC, calculating their average value RS[t], where t is the time window index. Both are synchronized via a hardware clock to ensure that AF[t] and RS[t] describe the physical phenomena within the same time period.
[0060] Then, the average radio frequency intensity RS[t] of the current time window is compared with a preset intensity threshold Th. Th is not a fixed value. For example, during the vehicle power-on initialization phase, the system can continuously collect RS for 10 seconds while the engine is idling and there is no broadcast in the vehicle, calculate its mean μ and standard deviation σ, and then dynamically set Th=μ+3σ. This method can adapt to the background noise of different vehicle electrical systems.
[0061] The judgment logic can be set as follows: if RS[t]>Th, then it is determined that the audio signal in the current time window is subjected to strong broadcast signal coupling interference. The system generates a binary interference flag FI[t]=1 and binds this flag to the corresponding AF[t]. If it does not exceed the limit, then FI[t]=0, that is, it is marked as no interference.
[0062] It is worth mentioning that the system can format the results of the synchronization processing and output them as input to subsequent modules. The output is a data volume containing an audio frame sequence AS and an interference marker sequence FS. The audio frame sequence AS represents a list containing T consecutive audio frames, i.e., AS=[AF[1], AF[2],...,AF[T]], where each AF[t] corresponds to 512 sampling data points in a sampling time window. The interference marker sequence FS represents a binary list of the same length as AS and strictly aligned, i.e., FS=[FI[1],FI[2],...,FI[T]]. When FI[t]=1, it indicates that AF[t] is interfered with; when it is 0, it indicates that it is not interfered with.
[0063] Steps one and two mainly introduce the preliminary acquisition and marking steps. Unlike the software timestamp alignment method, this implementation adopts a hardware synchronization scheme, which better ensures the spatiotemporal accuracy of interference marking from the source. After setting the intensity threshold detection, the subsequent complex processing flow is only activated when a specific interference is detected. When there is no interference, the system can run in a low-power mode, which can improve energy efficiency.
[0064] Step 3: Input the microphone audio signal with interference markers into the pre-trained voiceprint recognition model.
[0065] The voiceprint recognition model was trained using a training set containing the following samples: positive samples are pure adult instruction speech, and negative samples are synthetic interference speech generated by coupling pure instruction speech imitated by children with simulated harmonic interference.
[0066] For example, the voiceprint recognition model can be a multi-task deep neural network, which includes a shared feature extraction layer and an instruction compliance judgment task branch; the shared feature extraction layer is used to extract basic acoustic features from the input signal; the instruction compliance judgment task branch is used to output a voiceprint judgment score.
[0067] In this embodiment, a multi-task deep neural network can be constructed as a voiceprint recognition model. Its core design goal is to simultaneously extract effective features from potentially interfered speech and complete the judgment of instruction compliance.
[0068] Specifically, the input layer of the model can receive a fixed-length audio frame sequence with the shape [T,F], where T is the number of time steps (e.g., 32ms per frame, 31 frames in total, approximately 1 second), and F is the frequency domain feature dimension (e.g., 80-dimensional Mel spectrum), i.e., T=31, F=80. At the same time, the model receives an interference tag sequence synchronized with the audio frame sequence.
[0069] The core structure of the shared feature extraction layer can be a 4-layer one-dimensional convolutional neural network (1D-CNN) as the shared bottom layer, with batch normalization and ReLU activation function following each convolutional layer.
[0070] For adjusting its layers and parameters, the kernel width can be 5, 3, 3, 3, respectively, and the number of channels can be 64, 128, 256, 256, respectively. Pooling layers with a stride of 2 are interspersed after the 2nd and 4th convolutional layers to compress the temporal dimension and expand the receptive field.
[0071] Then there's the connection relationship. The original [T,F] features are transformed into a high-level feature tensor with temporal relationships after passing through this shared CNN. The tensor has a shape of [T',C] (T' is the reduced time step, and C=256 is the number of channels). This shared layer ensures that the model learns robust basic acoustic representations from the disturbed speech.
[0072] The structure of the instruction compliance judgment task branch can be to receive the output [T', C] of the shared feature extraction layer.
[0073] Its hierarchy begins with a bidirectional long short-term memory network (Bi-LSTM) with 128 hidden units to model global temporal dependencies. The hidden state at the final time step of the Bi-LSTM is then passed through a fully connected layer, and finally connected to a fully connected layer with an output dimension of 1, using the sigmoid activation function.
[0074] The output of this branch is a scalar value between 0 and 1, which is the voiceprint determination score. It directly corresponds to the model output in step two. The closer the score is to 1, the more confident the model is that the voice is a valid adult authorized instruction; the closer it is to 0, the more likely it is an instruction imitated by a child or an invalid instruction under interference.
[0075] Furthermore, the training method for the voiceprint recognition model includes: acquiring pure adult instruction speech as positive sample speech; acquiring pure instruction speech imitated by children as negative sample base speech; coupling the pure instruction speech imitated by children with harmonic interference noise of simulated preset frequency radio broadcast signals in the time and frequency domain to generate synthetic interference speech as negative sample speech; constructing a training set based on positive sample speech and negative sample speech, wherein training labels are used to distinguish whether the speech is a valid adult authorized instruction.
[0076] Specifically, for the construction of the training data, the positive sample speech was as follows: at least 50 different adults (half male and half female) read out all the preset control commands (such as "turn on heating" and "adjust the temperature") in clean speech. Each command was recorded multiple times, forming approximately 10,000 positive samples. Conventional speech enhancement (volume normalization and slight noise addition) was performed to increase diversity.
[0077] The negative sample base speech is specifically: collecting the speech of at least 30 children aged 3-8 imitating the reading of the same instructions, forming about 6,000 base negative samples.
[0078] For the generation of synthetic interference speech, i.e., negative samples:
[0079] For example, a method for generating synthetic interference speech includes: linearly superimposing in the time-frequency domain a child’s imitation of pure instructional speech with interference noise generated by simulating the characteristics of a radio broadcast signal in a preset frequency band, wherein the spectral energy of the interference noise is concentrated at the frequency positions corresponding to the harmonics of the preset frequency band within the 300-3400Hz speech baseband.
[0080] Specifically, for each child's imitated speech, the potential harmonics are calculated based on the current simulated preset frequency band. For example, if the current simulated preset frequency band is 98MHz, then the second harmonic of 98MHz is 196MHz (exceeding the speech frequency band). However, the difference frequency and intermodulation interference generated by the nonlinearity of the circuit will fall into the speech frequency band. An interference noise spectrum can be generated using an empirical formula. .
[0081] The generated interference noise spectrum Inverse Fourier transform is performed to obtain time-domain noise. Then, according to the specified signal-to-interference ratio (e.g., 0dB, meaning the interference equals the speech energy), the time-domain signal of the child's imitated speech is compared with the signal-to-interference ratio (SIR). Linear superposition is performed to generate synthetic interference speech. ,in The scaling factor α is calculated based on the target signal-to-interference ratio preset during data synthesis to ensure that the ratio of interference noise intensity to original speech intensity meets the requirements of the simulation scenario. The final negative sample set consists of both clean child speech and synthesized interference child speech.
[0082] Finally, the labels were set: positive samples were labeled 1, and all negative samples (pure and synthetic) were labeled 0.
[0083] The next step is the model training process. First, set the loss function; binary cross-entropy loss can be used. This is the standard loss function for binary classification tasks, and its calculation formula is:
[0084] ;
[0085] Where y is the true label (0 or 1) and p is the score predicted by the model.
[0086] Then the Adam optimizer can be used with the following parameters: initial learning rate lr=0.001, weight decay Wd=1e-4 (i.e. L2 regularization to prevent overfitting).
[0087] For the training strategy, the batch size can be set to bs=32, the training run for 100 epochs, and a learning rate cosine annealing strategy can be used, with the learning rate decaying from the initial value to 0 in each epoch according to the cosine function. When the performance no longer improves on the validation set (early stopping method), the optimal model parameters are saved.
[0088] In the implementation of step three, by using precisely simulated coupled interference data for training, the model inherently possesses the ability to resist specific interference, which is better than the model trained on general noise data. Furthermore, the combination of model scoring and post-processing correction based on radio frequency strength constitutes a double insurance at the software algorithm level. Even in the case of occasional misjudgment by the model, the final correction can be made through physical signal strength, thereby improving the overall security of the system.
[0089] Step 4 (executed by the voiceprint recognition model): Based on the interference marker, activate the feature suppression processing of harmonic interference related to the preset frequency band in the microphone audio signal; analyze the microphone audio signal after feature suppression processing to generate a voiceprint judgment score.
[0090] Step four describes the steps of inputting the audio signal with interference tags into the trained voiceprint recognition model for real-time inference.
[0091] Specifically, the system receives the output data body from Embodiment 1, which is the audio frame sequence AS and its corresponding interference tag sequence FS. In the typical configuration of this embodiment, T=31, that is, the sequence length is 31 frames, corresponding to about 1 second of audio (31 frames × 32 milliseconds / frame ≈ 1 second).
[0092] Then comes feature calculation. For each audio frame AF[t] in AS, its 80-dimensional Mel frequency cepstral coefficients are calculated as the basic acoustic feature Xm[t]. After processing all frames, the feature representation of the entire sequence is obtained, which is a feature matrix Xm with shape [31, 80].
[0093] Then comes the feature suppression processing step based on the interference marker. For example, the feature suppression processing method includes: dynamically generating a frequency domain mask based on the interference marker. The frequency domain mask is used to attenuate or zero the energy of the microphone audio signal in the frequency domain corresponding to the harmonic interference frequency range.
[0094] Specifically, for judging the suppression condition, the first step is to check the FS. If more than 50% of the entire sequence is marked as "1", then the current speech segment is determined to be in a strong interference environment, and the module corresponding to feature suppression can be activated.
[0095] Then, a frequency domain mask is generated and applied. Since the interference noise has a concentrated spectrum, and on the Mel scale, these center frequencies correspond to specific Mel band indices Mi, a binary mask matrix Mask with the same dimensions as Xm can be dynamically generated. In the Mask matrix, all columns belonging to the [Mi-2, Mi+2] frequency band range (i.e., the interference band and its adjacent bands) are set to 0, indicating complete suppression, while the remaining frequency band columns are set to 1, indicating preservation.
[0096] Next, suppression is applied by multiplying the feature matrix element by element: Xs = Xm⊙Mask. Here, ⊙ represents the Hadamard product, which is equivalent to "zeroing" the energy of the most disturbed frequency band at the feature level, preventing contaminated features from entering subsequent deep networks.
[0097] The suppressed feature matrix Xs is then input into the constructed multi-task deep neural network for shared feature extraction: Xs passes through 4 layers of 1D-CNN to extract high-level temporal features; compliance judgment: the high-level features pass through Bi-LSTM and fully connected layers, and finally the Sigmoid output layer generates a scalar value pr, which is the original voiceprint judgment score, in the range (0, 1).
[0098] Furthermore, the process of generating a voiceprint judgment score by the voiceprint recognition model also includes prosodic pattern analysis. Specifically, prosodic pattern analysis includes: calculating the fundamental frequency profile stability index of the microphone audio signal, analyzing the statistical regularity of the accent distribution, and comparing the index and regularity with the pre-stored prosodic template of adult self-spoken speech.
[0099] The prosodic template for adult spontaneous speech can be constructed through offline analysis of batches of adult natural speech data. The specific method involves collecting a large amount of natural, continuous speech from adults in various contexts, extracting prosodic features such as fundamental frequency contours and energy stresses, and then forming a template through statistical modeling (e.g., calculating the range of typical features or establishing a reference distribution). Once generated, this template is pre-stored in the system as fixed reference data and used in real-time analysis to compare the prosodic features of the current speech, thereby determining whether it conforms to the adult natural pronunciation pattern. This process is standardized offline processing, independent of the end user, ensuring the consistency of the discrimination benchmark and the out-of-the-box usability of the system.
[0100] Specifically, in this embodiment, prosodic pattern analysis is not an independent external algorithm, but rather internalized within the model training and inference. Since the Bi-LSTM layer used in the voiceprint recognition model is essentially a powerful temporal modeler, during training, the model can learn through comparative study of numerous "adult instructions" and "children's imitation instructions," enabling the hidden states of the Bi-LSTM to automatically learn to encode prosodic features such as intonation patterns, speech rate stability, and stress rhythm. When a speech segment is input, the Bi-LSTM's processing of the sequence is essentially the process of encoding and analyzing its prosodic features.
[0101] Therefore, the PR score output by the model is a comprehensive discrimination result that integrates voiceprint features and deep prosodic features. For example, a speech with a slightly immature timbre but smooth and stable prosodic features may receive a high score because its prosodic features are close to those of an adult; conversely, a speech with a timbre that is close to that of an adult after being distorted by interference, but with discontinuous prosodic features and abnormal rhythm, will receive a low score due to abnormal prosodic coding.
[0102] Since prosody reflects a speaker's language habits, rhythm control, and fluency of pronunciation—which are higher-level neuromuscular control patterns—children are very likely to slip up when imitating adult instructions. This extended analysis can quantitatively assess indicators such as the stability of the fundamental frequency of speech and the regularity of stress distribution, and compare them with the prosody of mature adult spontaneous speech. It can keenly capture unnatural features such as prosodic breaks, rhythmic imbalances, or stiff intonation that are common in children's pronunciation imitations.
[0103] Therefore, even if interference blurs the timbre characteristics, the system can still effectively identify children's unnatural imitations based on these deep-seated behavioral pattern differences that are not easily distorted by electromagnetic interference, thereby reducing the probability of misjudging interfered children's imitations as compliant adult instructions.
[0104] In the implementation of step four, the voiceprint recognition model integrates feature layer suppression (active purification of input) and model layer discrimination, and outputs a comprehensive voiceprint judgment score. This score integrates multi-dimensional analysis results such as voiceprint and prosody, and therefore has high reliability.
[0105] Step five: When interference markers are present, the voiceprint determination score is adjusted downwards based on the magnitude by which the spatial radio frequency signal strength exceeds a preset strength threshold. This step occurs after generating the voiceprint determination score and before the judgment step.
[0106] Specifically, since this step is a post-processing step to improve decision security, when the system confirms that the current voice segment is interfered with (i.e., the FS contains interference markers), this correction logic will be activated. The core purpose of the correction is that when the intensity of the detected radio interference is greater, the system should take a more conservative approach to the original model score, that is, it is more inclined to lower the score in order to reduce the risk of misauthorization.
[0107] In this embodiment, the system can read the maximum radio frequency interference intensity value monitored in the current time period and compare it with a dynamic threshold. Based on the magnitude of the interference exceeding the threshold, an attenuation factor is calculated using a preset, monotonically increasing function. This attenuation factor is applied to the model's original score pr, resulting in a lower and more cautious final score PF. This process ensures that the system's authorization threshold automatically increases in environments with stronger physical interference.
[0108] Step six: Make a judgment based on the voiceprint determination score. Only when the voiceprint determination score is higher than the preset compliance threshold will the final control authorization instruction be generated.
[0109] Specifically, the compliance threshold must first be set. After the model training is completed, a validation set independent of the training set can be used to comprehensively evaluate the model performance. This validation set also consists of labeled adult clean instruction speech and synthetic interference speech. The system can input all samples of the validation set into the model to obtain a series of voiceprint judgment scores. By analyzing the distribution of these scores on positive and negative samples, the optimal threshold can be determined.
[0110] Furthermore, receiver operating characteristic curves can be plotted, and with the goal of maximizing overall discrimination performance (e.g., selecting points with equal error rates) or more strictly controlling the rate of false acceptance by children, a fixed score value can be ultimately selected as a preset compliance threshold. For example, if safety is the highest priority, the threshold can be set at the highest score percentile on the validation set that achieves zero false acceptance by children.
[0111] This approach ensures that compliance thresholds are matched with the capabilities of the model, enabling the final decision to achieve a quantitatively validated optimal balance between security and usability.
[0112] Once the compliance threshold is set, when the final score PF is received, it will be compared with the compliance threshold. Only when the PF is greater than the threshold will the interactive function of the seat be activated by a command. If it is less than the threshold, the command will be rejected. Then, the system can combine the voice recognition results to generate specific corresponding control authorization commands and send them to the controller of the child safety seat through the vehicle CAN bus or private serial protocol to execute the corresponding comfort function adjustments.
[0113] In the implementation of steps five and six, by analyzing the comprehensive voiceprint judgment score, if strong interference is detected, the score can be automatically corrected downward according to the interference intensity, making the decision more conservative. Only when the score is higher than a strict safety threshold will the seat's interactive function be activated to generate and execute commands. This decision-level correction (conservative adjustment based on physical signals) can further improve the credibility of the commands.
[0114] It is worth mentioning that the entire reasoning and decision-making process, from feature suppression and model inference to score correction, can be completed within 100 milliseconds on embedded hardware (such as ARM-Cortex-A series processors), which well meets the real-time requirements of in-vehicle interaction and improves its practicality.
[0115] In summary, this solution breaks away from the conventional approach of processing only acoustic signals. It creatively introduces synchronous monitoring of specific radio frequency signals in space. Once a broadcast signal exceeding a threshold is detected, a precise timestamp is immediately added to the speech frame at the same moment. This timestamp acts as a targeted navigation signal. Based on this precise navigation signal, feature suppression is performed through a voiceprint recognition model, thereby specifically weakening the narrow frequency bands in the speech spectrum corresponding to known broadcast harmonics. This eliminates interference factors while preserving the original speech to the maximum extent. Furthermore, since the voiceprint recognition model is not trained using conventional clean speech or generalized noise data, but rather couples the child's imitated commands with simulated specific broadcast harmonic interference to create negative samples for training, it is highly effective in distinguishing between genuine adult commands disguised by specific interference and children's imitated commands. Finally, the model can determine whether to wake up the voice interaction system and issue control authorization commands based on the output voiceprint judgment score, effectively avoiding misjudging interfered children's imitated voices as compliant adult commands during voiceprint verification, thus ensuring the safety of children alone in the seat.
[0116] Furthermore, feature suppression is achieved by dynamically generating frequency domain masks. Based on the threat indicated by the interference marker, a mask template that precisely corresponds to the current interference spectrum is calculated and generated in real time. This template can selectively attenuate or reduce the microphone audio signal to zero in the specific narrowband range where the interference is most severe in the frequency domain, while preserving the complete information of other uncontaminated frequency bands of speech to the maximum extent, thus improving the possibility of reliable discrimination of this information in severe electromagnetic environments.
[0117] Furthermore, by analyzing prosodic patterns, the system captures unnatural features commonly found in children's pronunciation imitations, such as prosodic breaks, rhythmic imbalances, or harsh intonations. Even if interference blurs the timbre characteristics, the system can still effectively identify children's unnatural pronunciation imitations based on these deep-seated behavioral pattern differences that are not easily distorted by electromagnetic interference. This reduces the probability of misjudging interfered children's imitations as compliant adult instructions.
[0118] Example 2
[0119] Please see Figure 2 As shown, this embodiment provides a voiceprint recognition system for a vehicle-mounted child safety seat, used to implement the voiceprint recognition method for a vehicle-mounted child safety seat disclosed in Embodiment 1. This system can be loaded as a subsystem onto the voice interaction system of the seat. The voiceprint recognition system for a vehicle-mounted child safety seat includes: a synchronous acquisition and marking module, a voiceprint recognition module, and a voiceprint arbitration module.
[0120] The synchronous acquisition and tagging module is configured to synchronously acquire the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals, and generate an interference tag for the corresponding microphone audio signal when the spatial radio frequency signal strength exceeds a preset strength threshold;
[0121] The voiceprint recognition module is configured to receive microphone audio signals with interference tags, perform feature suppression processing based on interference tags, and output a voiceprint determination score.
[0122] The voiceprint arbitration module is configured to receive voiceprint judgment scores and generate a final control authorization command when the voiceprint judgment score is higher than a preset compliance threshold.
[0123] Since this system uses the voiceprint recognition method for in-vehicle child safety seats in Example 1, it has the same effect, which will not be repeated here.
[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0125] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A voiceprint recognition method for a vehicle-mounted child safety seat, characterized in that, include: Simultaneously acquire the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals in preset frequency bands; When the intensity of a spatial radio frequency signal exceeds a preset intensity threshold within any acquisition time window, an interference marker is generated for the microphone audio signal acquired within the same acquisition time window. The microphone audio signal with interference tags is input into the pre-trained voiceprint recognition model; The voiceprint recognition model was trained using a training set containing the following samples: positive samples are pure instructional speech from adults, and negative samples are synthetic interference speech generated by coupling pure instructional speech imitated by children with simulated harmonic interference. The voiceprint recognition model is configured to perform: Based on interference markers, feature suppression processing for harmonic interference related to preset frequency bands in microphone audio signals is activated. The microphone audio signal after feature suppression processing is analyzed to generate a voiceprint determination score; The final control authorization command is generated only when the voiceprint judgment score is higher than the preset compliance threshold.
2. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, The voiceprint recognition model is a multi-task deep neural network, which includes a shared feature extraction layer and an instruction compliance judgment task branch; The shared feature extraction layer is used to extract basic acoustic features from the input signal; the instruction compliance discrimination task branch is used to output the voiceprint judgment score.
3. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, Training methods for voiceprint recognition models include: Obtain clean, adult-grade instruction speech as positive sample speech; Obtain the pure instructional speech imitated by children as the negative sample base speech; The pure instructional speech imitated by children is coupled in the time and frequency domain with harmonic interference noise of simulated preset frequency radio broadcast signals, and the resulting synthetic interference speech is used as negative sample speech. A training set is constructed based on positive and negative sample speech, where training labels are used to distinguish whether the speech is a valid adult authorization instruction; The voiceprint recognition model is trained using a training set.
4. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 3, characterized in that, Methods for generating synthetic interference speech include: The interference noise generated by the pure instructional speech imitated by the child and the characteristics of the simulated preset frequency band radio broadcast signal is linearly superimposed in the time and frequency domain. The spectral energy of the interference noise is concentrated in the frequency positions corresponding to the harmonics of each preset frequency band in the 300-3400Hz speech baseband.
5. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, The methods for feature suppression processing include: Based on the interference markers, a frequency domain mask is dynamically generated. The frequency domain mask is used to attenuate or zero the energy of the microphone audio signal in the frequency domain corresponding to the harmonic interference frequency range.
6. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, The preset frequency band is the FM broadcast band from 87MHz to 108MHz; the preset intensity threshold is dynamically calibrated based on the average electromagnetic environmental noise inside the vehicle when the vehicle is ignited.
7. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, The process of generating a voiceprint judgment score by the voiceprint recognition model also includes prosodic pattern analysis. Specifically, prosodic pattern analysis includes: calculating the fundamental frequency profile stability index of the microphone audio signal, analyzing the statistical law of stress distribution, and comparing the index and law with the pre-stored prosodic template of adult self-spoken speech.
8. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, After generating the voiceprint assessment score and before the assessment step, a confidence correction step is also included: When interference markers are present, the voiceprint determination score is adjusted downward based on the magnitude by which the spatial radio frequency signal strength exceeds a preset strength threshold.
9. The voiceprint recognition method for a vehicle-mounted child safety seat according to claim 1, characterized in that, Methods for simultaneously acquiring the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals in a preset frequency band include: The microphone audio signal and the spatial radio frequency signal strength are collected by physically isolated microphone preamplifier circuit and radio frequency strength detection circuit respectively, and the signal sampling clocks of the microphone preamplifier circuit and the radio frequency strength detection circuit are kept synchronized by the same clock source.
10. A voiceprint recognition system for a vehicle-mounted child safety seat, used to implement the voiceprint recognition method for a vehicle-mounted child safety seat as described in any one of claims 1 to 8, characterized in that, The system includes: The synchronous acquisition and tagging module is used to synchronously acquire the spatial radio frequency signal strength of microphone audio signals and radio broadcast signals, and generate an interference tag for the corresponding microphone audio signal when the spatial radio frequency signal strength exceeds a preset strength threshold; The voiceprint recognition module is a pre-trained voiceprint recognition model used to receive microphone audio signals with interference tags, perform feature suppression processing based on interference tags, and output a voiceprint judgment score. The voiceprint arbitration module is used to receive voiceprint judgment scores and generate a final control authorization command when the voiceprint judgment score is higher than a preset compliance threshold.