Intelligent Voice Interaction Method and System for 5G Communication Devices Based on Deep Learning

Through deep learning technology, the noise filtering threshold and dual encryption are dynamically adjusted on 5G communication devices, which solves the problems of voice signal recognition errors and resource overload, and achieves high accuracy and smooth voice interaction, reducing latency and lag.

CN120126463BActive Publication Date: 2025-07-22SHENZHEN BESNEL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510578703.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-22
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In dynamic and complex scenarios, it is difficult for the prior art to accurately distinguish the vocal characteristics in the voice signal from the acceleration fluctuation signals caused by device movement, resulting in an increase in the voiceprint recognition error rate and a deviation in semantic understanding. The fixed noise weight allocation mechanism can easily cause local processing resources to be overloaded or cloud communication delay exceeding the standard.

Method used

The intelligent voice interaction method of 5G communication devices based on deep learning is adopted, and the joint input signal is generated by obtaining voice data and accelerometer data, and the frequency band noise filtering threshold is dynamically adjusted, and the dual encryption technology of voiceprint characteristics and device unique identifiers is combined to intelligently allocate cloud or local processing resources to achieve synchronization of multimodal output.

Benefits of technology

It improves the accuracy of voice command recognition, reduces the risk of key leakage, improves user operation fluency, and reduces the phenomenon of lag caused by resource preemption, and realizes millisecond synchronization between voice and visual interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126463B_ABST
    Figure CN120126463B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of voice interaction, and relates to an intelligent voice interaction method and system for 5G communication devices based on deep learning, including: generating a joint input signal containing voice frequency domain information and motion state; outputting a denoised voice segment and its corresponding environmental interference level parameter; generating a basic voiceprint ID containing the user's biometric identifier and a dynamic key segment; calculating a cloud processing priority score according to the environmental interference level parameter and the remaining battery power value of the device, and when the score exceeds a preset threshold, sending a voice processing request containing the basic voiceprint ID to the edge computing node, otherwise triggering a local semantic recognition process; receiving the semantic recognition result from the cloud or locally, and generating a multi-round dialogue response instruction in combination with the context parameters in the user's historical interaction data. The present invention solves the problem that the fixed noise weight distribution mechanism is likely to cause local processing resource overload or cloud communication delay exceeding the standard.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of voice interaction, and relates to an intelligent voice interaction method and system for 5G communication devices based on deep learning. Background Art

[0002] In the voice interaction technology for mechanical signal measurement in a noisy environment, the core problem lies in the insufficient separation accuracy of signal features in a dynamic and complex scenario, resulting in the vulnerability of effective voice signals to interference from device motion noise and external environmental vibrations. When the device is in a high-speed movement or high-frequency mechanical vibration environment, existing methods are difficult to accurately distinguish the human voice features in the voice signal from the acceleration fluctuation signals caused by the device's own movement, leading to an increase in the error rate of voiceprint recognition and semantic understanding deviation.

[0003] Traditional solutions mainly adopt an acceleration signal-assisted noise reduction mechanism. By pre-setting a motion noise database to match the time-domain energy characteristics of the device's accelerometer, static filtering parameters are formulated to suppress interference in specific frequency bands. The device movement state is judged by monitoring the peak value of the three-axis acceleration. When the acceleration exceeds the preset threshold, the noise reduction process is started, and a band filter is used to eliminate motion-related noise.

[0004] Based on the above problems, traditional methods are prone to signal misjudgment in sudden mechanical vibration or complex noise scenarios. The fixed noise weight allocation mechanism ignores the dynamic correlation between device power and environmental interference, easily causing local processing resource overload or exceeding the standard of cloud communication delay, and weakening the comprehensive response ability of the system. Summary of the Invention

[0005] To solve the above problems, the present invention provides an intelligent voice interaction method and system for 5G communication devices based on deep learning.

[0006] In the first aspect, the present invention provides an intelligent voice interaction method for 5G communication devices based on deep learning, adopting the following technical solutions:

[0007] The intelligent voice interaction method for 5G communication devices based on deep learning includes the following steps:

[0008] S1. Obtain the original voice data, call the accelerometer data of the device, and generate a joint input signal containing voice frequency domain information and motion state;

[0009] S2. Perform real-time processing on the joint input signal, and output the noise-reduced voice segment and its corresponding environmental interference level parameter;

[0010] S3. Decompose the voiceprint features in the noise-reduced voice segment, and generate a basic voiceprint ID and a dynamic key segment containing the user's biometric identifier;

[0011] S4. Calculate the cloud processing priority score based on the environmental interference level parameter and the remaining battery level of the device. When the score exceeds the preset threshold, send a voice processing request containing the basic voiceprint ID to the edge computing node; otherwise, trigger the local semantic recognition process.

[0012] S5. Receive the semantic recognition results from the cloud or local, and generate a multi-round dialogue response instruction by combining the context parameters in the user's historical interaction data.

[0013] S6. Dynamically allocate rendering resources for the response instruction according to the idle computing power of the target device's GPU, preferentially use the voice synthesis engine to output the result, and synchronously inject the visual interaction interface when the screen wake-up state is detected.

[0014] In a further solution of the present invention, step S1 includes the following steps:

[0015] Obtain continuous raw voice data collected by the microphone of the target device, and synchronously activate the accelerometer sensor to collect the three-axis acceleration values of the device.

[0016] Extract the voice frequency features of each time slice from the framed processing result of the raw voice data, and calculate the time domain energy value of the device by combining the three-axis acceleration values.

[0017] Align the timestamps of the voice frequency features and the time domain energy values to generate a joint input signal containing voice frequency domain information and motion state parameters.

[0018] In a further solution of the present invention, step S2 includes the following steps:

[0019] Based on the voice frequency features and device motion state parameters in the joint input signal, dynamically divide the dynamic interference threshold of the voice frequency band.

[0020] Adjust the band gain coefficient of the adaptive filtering model according to the dynamic interference threshold, perform noise reduction processing on the original voice waveform, and output the noise-reduced voice segment.

[0021] Calculate the energy difference between the original voice segment and the noise-reduced voice segment to generate an environmental interference level parameter for quantifying the degree of noise pollution.

[0022] In a further solution of the present invention, step S3 includes the following steps:

[0023] Extract the fundamental frequency features and harmonic components in the noise-reduced voice segment to generate a dynamic voiceprint hash value containing user biometric features.

[0024] Perform an encryption operation based on the dynamic voiceprint hash value and the unique hardware identifier of the device to generate a dynamic key segment.

[0025] Use the pre - set public key to encrypt the dynamic key and the environmental interference level parameter. The pre - set public key is the asymmetric encryption public key pre - allocated by the cloud to the device. After generating the authentication parameter, upload it to the cloud.

[0026] A further solution of the present invention for generating a dynamic key segment includes the following steps:

[0027] The dynamic key segment is generated by the exclusive - OR operation combination of the device MAC address and the voiceprint hash value, further enhancing security. The formula for key generation is

[0028] ;

[0029] where, is the generated 256 - bit dynamic key, is the physical address of the device network adapter, which is the unique hardware identifier of the device, is the hash function, is the bit - wise exclusive - OR operation, is the dynamic voiceprint hash value.

[0030] A further solution of the present invention for step S4 includes the following steps:

[0031] Calculate the cloud processing priority score through linear weighting according to the environmental interference level parameter and the remaining battery power value of the device;

[0032] If the priority score exceeds the preset threshold, fragment and transmit the voice processing request containing the dynamic key segment to the edge computing node;

[0033] If the priority score does not reach the threshold, trigger the local semantic recognition engine to parse the noise - reduced voice segment.

[0034] A further solution of the present invention for step S5 includes the following steps:

[0035] Extract the intent identifier and entity word vector from the semantic recognition result;

[0036] Retrieve the historical interaction data, match the context - related context parameters, generate the context constraint conditions for multi - turn conversations, and the conversation logic restriction rules dynamically constructed based on the historical interaction data are used to constrain the semantic scope of the current response;

[0037] Adjust the semantic priority of the response instruction according to the dependency weight, and generate a response instruction containing multi - turn conversation logic.

[0038] A further solution of the present invention for step S6 includes the following steps:

[0039] Real - time monitor the GPU rendering queue load rate of the target device, and dynamically allocate the time windows for voice synthesis and graphics rendering based on the idle computing power;

[0040] Within the time window, preferentially call the speech synthesis engine to stream-process response instructions, and cache the visualization interface metadata in a low-power queue;

[0041] When a screen wake-up signal is detected, immediately render the interactive interface based on the metadata and synchronize the timestamp with the speech stream.

[0042] A further solution of the present invention is to dynamically allocate time windows for speech synthesis and graphics rendering based on idle computing power, including the following steps:

[0043] According to the ratio of the current number of rendering tasks of the GPU to the maximum number of concurrent tasks, determine the upper limit of the duration of the allocation window, and calculate the periodic allocation window of the idle computing power, satisfying the following formula,

[0044] ;

[0045] Wherein, represents the finally allocated window duration; represents the maximum number of concurrent tasks of the GPU; represents the device reference time unit; represents the current number of rendering tasks.

[0046] In a second aspect, the present invention provides an intelligent voice interaction system for 5G communication devices based on deep learning, adopting the following technical solutions:

[0047] A combined input signal generation module for generating a combined input signal containing voice frequency domain information and motion state;

[0048] A voice segment noise reduction module for real-time processing of the combined input signal and outputting the noise-reduced voice segment and its corresponding environmental interference level parameter;

[0049] An identity authentication module for decomposing the voiceprint features in the noise-reduced voice segment and generating a basic voiceprint ID and a dynamic key segment containing the user's biometric identifier;

[0050] A distributed processing routing module for calculating the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device;

[0051] A multi-round dialogue management module for receiving semantic recognition results from the cloud or locally, and generating multi-round dialogue response instructions in combination with the context parameters in the user's historical interaction data;

[0052] A multi-modal output scheduling module for preferentially using the speech synthesis engine to output results and synchronously injecting a visual interactive interface when the screen wake-up state is detected.

[0053] In summary, the present invention includes the following beneficial technical effects:

[0054] 1. By jointly analyzing the device motion state and voice signals, the band noise filtering threshold is dynamically adjusted to eliminate environmental interference while retaining the biometric details of the user's voiceprint. Combining the accelerometer data, the device motion intensity is sensed in real time, thereby dynamically optimizing the noise reduction algorithm parameters, improving the voice command recognition accuracy in high-noise scenarios such as cycling and in-vehicle by more than 40%, and reducing the voice distortion problem caused by excessive noise reduction at the same time;

[0055] 2. A dual-encryption technology using voiceprint biometrics and hardware identifiers is adopted to construct a dynamic security mechanism, effectively resisting identity forgery and replay attacks; the exclusive OR operation and encrypted transmission are performed on the voiceprint hash value generated in real time and the unique device identifier. Compared with the traditional fixed key scheme, this method reduces the risk of key leakage by more than 75%, and the dynamically generated key fragments for each interaction can adapt to the encryption intensity requirements in different noise environments;

[0056] 3. A dynamic scheduling strategy for multi-modal output is introduced, and the interaction experience is improved through GPU computing power monitoring and audio-visual synchronization mechanism. The processing window is divided according to the real-time rendering load, giving priority to ensuring the smoothness of speech synthesis, and caching the interface data to meet the immediate rendering requirements when the screen is awakened. This method breaks through the limitations of the traditional single output mode, realizes millisecond-level synchronization between voice broadcast and visual interface switching, improves the operation fluency of users in multi-task concurrent scenarios by more than 50%, and avoids the card phenomenon caused by resource contention at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. The drawings are used to provide a further understanding of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0058] Figure 1 Flow chart showing the method of intelligent voice interaction for 5G communication devices based on deep learning.

[0059] Figure 2 Structure diagram showing the intelligent voice interaction system for 5G communication devices based on deep learning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0061] The following will make a preferred and detailed description of the present invention in conjunction with the attached Figure 1 - Figure 2 drawings.

[0062] Referring to the attached Figure 1 drawings, the present invention proposes an intelligent voice interaction method for 5G communication devices based on deep learning, including the following steps:

[0063] S1. Obtain the original voice data, call the accelerometer data of the device, and generate a combined input signal containing voice frequency domain information and motion state;

[0064] S2. Perform real-time processing on the combined input signal, and output the noise-reduced voice segment and its corresponding environmental interference level parameter;

[0065] S3. Decompose the voiceprint features in the noise-reduced voice segment to generate a basic voiceprint ID containing the user's biometric identifier and a dynamic key segment;

[0066] S4. Calculate the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device. When the score exceeds the preset threshold, send a voice processing request containing the basic voiceprint ID to the edge computing node; otherwise, trigger the local semantic recognition process;

[0067] S5. Receive the semantic recognition result from the cloud or locally, and generate a multi-round dialogue response instruction in combination with the context parameter in the user's historical interaction data;

[0068] S6. Dynamically allocate rendering resources for the response instruction according to the idle computing power of the target device's GPU, preferentially use the voice synthesis engine to output the result, and synchronously inject the visual interaction interface when the screen wake-up state is detected.

[0069] In one embodiment of the present invention, step S1 includes the following steps:

[0070] S11. Collect the original voice data of the human body;

[0071] Specifically, the original voice data emitted by the human body is continuously collected through the built-in microphone of the target device; the original voice data is the human voice signal collected by the built-in microphone of the target device without noise reduction processing, including the speaking content and environmental noise. The original voice data collected by the microphone is divided into frames, and the frequency characteristics of each time slice in the voice waveform are extracted. The original voice data collected by the microphone is a continuous time domain signal, which needs to be divided into several fixed-length voice segments through frame processing;

[0072] Among them, by frame segmentation, the continuous speech stream is divided into segments with equal time windows, which is convenient for analyzing the local characteristics of non-stationary signals segment by segment. Frequency characteristics represent the energy distribution parameters in the frequency domain of speech signals, for example: the Mel frequency cepstrum coefficients generated by Fourier transform reflect the speech components in different frequency bands.

[0073] For example, assuming that the user says the original voice of "turn on the lights", the typical frame length is 20-40 milliseconds, for example, 25 milliseconds per frame, and each frame signal is subjected to fast Fourier transform to generate a frequency feature vector containing 128-dimensional MFCC.

[0074] S12, obtaining accelerometer data of the device;

[0075] Specifically, within 0.1 seconds after the collection of the original human voice data begins, the accelerometer sensor of the target device is synchronously activated to obtain the accelerometer data of the target device in three-dimensional space; the accelerometer data, the motion parameters measured by the built-in inertial sensor of the target device, including the three-axis acceleration value and the angular displacement, calculates the time domain energy value of the current motion state of the target device, which satisfies the following formula:

[0076] ;

[0077] in, , , Indicates the three-axis acceleration value, that is, the instantaneous acceleration data of the target device in three-dimensional space; Indicates the time-domain energy value of the device, which represents the movement intensity of the target device. The higher the energy value, the more intense the device movement.

[0078] S13, generating a joint input signal including speech frequency domain information and motion state;

[0079] Specifically, the timestamps of the speech frequency features and the time domain energy values are aligned and encapsulated into a multi-dimensional data packet containing speech frequency domain information and motion status as a joint input signal; the joint input signal contains the speech frequency features and the time domain energy values of the device motion status.

[0080] In one embodiment of the present invention, step S2 comprises the following steps:

[0081] Based on the voice frequency characteristics of the combined input signal and the device motion state parameters, divide the dynamic interference threshold of the voice frequency band. The dynamic interference threshold is the upper limit of the frequency band noise tolerance dynamically adjusted based on the real-time motion state and historical noise data, and is used to distinguish voice and interference;

[0082] Among them, the voice frequency band represents a specific frequency range divided in the voice signal spectrum;

[0083] The calculation of the dynamic interference threshold depends on the statistical relationship between the motion state and historical noise, and satisfies the following formula,

[0084] ;

[0085] Among them, represents the th dynamic interference threshold of the current frame in the frequency band, represents the motion state weight coefficient, represents the device time-domain energy value, represents the historical noise weight coefficient, represents the th historical noise mean value of the current frame in the frequency band, is calculated by combining historical data, and are determined by laboratory tests.

[0086] Exemplarily, in a cycling scenario, the high-frequency vibration detected by the device , combined with the average noise in the 500 - 1000 Hz frequency band in the historical data . Assuming , , then the dynamic interference threshold is ;

[0087] If the actual noise in the th frequency band exceeds the dynamic interference threshold , it is determined that the th frequency band is interference and is filtered out.

[0088] Use a lightweight noise reduction model to perform adaptive filtering on the voice waveform, and adjust the filtering gain coefficient according to the dynamic interference threshold. The lightweight noise reduction model adopts an architecture that combines a convolutional neural network and an adaptive filter, uses the multi-dimensional data packet of the combined input signal as the input, and performs adaptive filtering on the voice waveform as the output;

[0089] The filtering gain coefficient is the scaling ratio of the signal intensity of each frequency band in the filtering process, retains the voice component and attenuates the noise, and is adjusted according to the ratio of the dynamic interference threshold to the current frequency band noise, and satisfies the following formula,

[0090] ;

[0091] Among them, is the filtering gain coefficient of the th frequency band of the current frame, is the dynamic interference threshold of the th frequency band of the current frame, is the noise energy value of the th frequency band of the current frame.

[0092] When the noise of the current frequency band is lower than the dynamic interference threshold, the filtering gain coefficient decreases linearly with the noise of the current frequency band; when the noise of the current frequency band exceeds the dynamic interference threshold, the filtering gain coefficient forcibly attenuates 90% of the signal energy.

[0093] Exemplarily, if the noise energy value of the noise of a certain frequency band , the dynamic interference threshold of this frequency band , then , the noise of the signal in this frequency band is suppressed by 78% of the signal energy; if the noise energy value of the noise of a certain frequency band , the noise of the current frequency band exceeds the dynamic interference threshold, and the noise of the signal in this frequency band is suppressed by 90% of the signal energy.

[0094] The speech signal after adaptive filtering processing is reconstructed into a time-domain speech signal through inverse Fourier transform and spliced into continuous noise-reduced speech segments; the noise-reduced speech segments, the clear speech signal waveform after frequency-domain filtering and signal reconstruction. Calculate the total energy difference between the original speech segment and the noise-reduced speech segment, generate an environmental interference level parameter, which satisfies the following formula,

[0095] ;

[0096] Among them, represents the energy value of the original speech segment; represents the energy value of the noise-reduced speech segment; represents the environmental interference level parameter, and the higher the value, the more serious the noise pollution.

[0097] The environmental interference level parameter represents an index that quantifies the ratio of noise to speech energy, is used to evaluate the difficulty of noise reduction, calculates the total energy difference between the original speech segment and the noise-reduced speech segment, maps the result to low, medium, and high interference levels, and is used to guide the decision-making of subsequent speech recognition. For example, in a high-interference environment, a more robust recognition algorithm is preferentially called;

[0098] Combined with the calculation of historical data, set the first threshold of the environmental interference level parameter to , and the second threshold of the environmental interference level parameter to ;

[0099] Low interference level, corresponding to , indicating that the environmental noise is negligible, and the system enables a low-latency lightweight noise reduction model;

[0100] Medium interference level, corresponding to , it is necessary to start the enhanced noise suppression mode and extend the voice endpoint detection window length;

[0101] High interference level, corresponding to , force the trigger of multi-device collaborative verification or require the user to confirm sensitive operations again.

[0102] The environmental interference level parameter endows the system with the ability to dynamically adapt to the noise environment;

[0103] Exemplarily, in a quiet office, the low interference level can retain the high-frequency details of the voice to improve the accuracy of voiceprint recognition; in a factory mechanical noise scenario, the high interference level will guide the system to filter redundant frequency bands to reduce the risk of misrecognition;

[0104] The user performs voice payment outdoors in windy and rainy weather, , set and , the system determines it as a medium interference level, automatically enhances the anti-noise weight of the voice feature extraction algorithm, and increases the audio sampling duration to 2 seconds to ensure the voiceprint matching accuracy of the "Pay 100 yuan" instruction.

[0105] The beneficial effect of this verification example: In strenuous exercise or high-noise environments, it improves the signal-to-noise ratio of voice signals, reduces voice distortion caused by excessive noise reduction, and provides cleaner input data for subsequent voiceprint feature extraction.

[0106] In one embodiment of the present invention, step S3 includes the following steps:

[0107] S31. Extract the fundamental frequency features and harmonic components of the user's voiceprint from the noise-reduced voice segment to generate a dynamic voiceprint hash value;

[0108] Specifically, after short-time analysis of the noise-reduced voice segment, extract the fundamental frequency features and harmonic components of the user's voiceprint and calculate the amplitude distribution of each harmonic component; the fundamental frequency features of the user's voiceprint and the first three-order harmonic components form a voiceprint feature vector, and the voiceprint feature vector is processed through a hash function to generate a unique dynamic voiceprint hash value,

[0109] ;

[0110] Among them, represents a 128-bit hash value, represents the fundamental frequency mean value, are the amplitudes of the first three-order harmonic components respectively;

[0111] Among them, the fundamental frequency feature represents the fundamental frequency of vocal cord vibration in the voice signal; the dynamic voiceprint hash value represents a unique identifier generated based on real-time voiceprint features, supporting identity binding and verification across sessions.

[0112] Exemplarily, after the user says "turn on the air conditioner", the amplitudes of the first three harmonic components are set 、 、 , and after being processed by the SHA-256 hash function, a hash value of a fixed length is generated 。

[0113] S32. Generate a dynamic key based on the device hardware unique identifier and the dynamic voiceprint hash value;

[0114] Specifically, the dynamic key is generated by combining the device MAC address and the voiceprint hash value through an exclusive OR operation, further enhancing security. The formula for key generation is

[0115] ;

[0116] Among them, is the generated 256-bit dynamic key, is the physical address of the device network adapter (device hardware unique identifier), is the hash function, is the bitwise exclusive OR operation, is the dynamic voiceprint hash value.

[0117] Exemplarily, the MAC address of a certain 5G communication device is 00-B0-D0-63-C2-26, and the voiceprint hash value is a3f5...7c1d. After the exclusive OR operation and hash processing, a dynamic key is generated for encrypting subsequent operation instructions.

[0118] S33. Encrypt the dynamic key and the environmental interference level parameter using the preset public key. The preset public key is an asymmetric encryption public key pre-allocated by the cloud to the device. After generating the identity verification parameter, upload it to the cloud;

[0119] Specifically, the device pre-stores the RSA public key published by the cloud, performs asymmetric encryption on the dynamic key and the environmental interference level parameter, and generates an identity verification parameter. The identity verification parameter represents an encrypted data packet containing the dynamic key and the environmental interference level, and is used for the cloud to verify legality and operation permissions after reverse decryption. The encryption formula is

[0120] ;

[0121] Among them, represents the encrypted identity verification parameter, represents the environmental interference level parameter, Indicates a public key encryption operation, is the generated 256-bit dynamic key, Indicates data splicing.

[0122] Exemplarily, a dynamic key is preset in advance , interference level parameter , and after being encrypted by the public key, an authentication parameter is generated , and the original data can be decrypted and restored after being uploaded to the cloud.

[0123] Beneficial effects of this verification example: The dual binding of voiceprint biometrics and device hardware constructs a dynamic security mechanism, solving the risks of identity forgery and replay attacks in traditional voice interactions; its core principle lies in the real-time generation of an encryption key containing voiceprint characteristics and the unique identifier of the device, and combining asymmetric encryption to synchronously encode noise environment factors into the verification parameter, enabling the cloud to accurately identify legitimate users and eliminate malicious interference.

[0124] In one embodiment of the present invention, step S4 includes the following steps:

[0125] Map the environmental interference level parameter to an interference weight coefficient, and quantify the remaining battery value of the device into an available computing resource value. The available computing resource value is obtained through the mapping relationship between the remaining battery value and the preset power consumption model, and satisfies the following formula,

[0126] ;

[0127] Wherein, is the current remaining battery value of the device, is the energy consumption per unit time of local semantic recognition, represents the expected time required for voice processing, represents the available computing resource value, and this formula converts the battery power into the resource capacity for sustainable computing, restricting overloaded local processing;

[0128] Among them, the interference weight coefficient converts the interference level parameter into a priority determination weight through a piecewise function. When the interference level parameter ≥0.7, the weight is 0.9, 0.3≤ <0.7, the weight is 0.6, and when the interference level parameter <0.3, the weight is 0.2, aiming to strengthen the dependence on cloud processing in a high-interference environment.

[0129] Calculate the cloud processing priority score through the linear weighting of the interference weight coefficient and the available computing resource value, and satisfy the following formula,

[0130] ;

[0131] Wherein, Represents the cloud processing priority score, where a higher score indicates a higher priority for cloud processing; Represents the interference weight coefficient; Represents the available computing resource value; Represents the dynamic adjustment coefficient.

[0132] Exemplarily, set the interference weight coefficient (interference level parameter When ≥ 0.7, the weight is ); , by default take 0.7, and reduce it to 0.5 when the device is in a moving state to reduce the energy consumption risk of cloud communication; ;

[0133] It can be calculated that: .

[0134] When the priority score exceeds the threshold of the priority score, trigger cloud processing and perform cloud sharding transmission, and send the request containing the basic voiceprint ID in shards to multiple computing units of the edge node; otherwise, call the local decoding module to start semantic recognition;

[0135] Among them, for sharding transmission, the voice processing request is split into multiple data blocks and transmitted to different edge computing units for parallel processing. Each shard contains the basic voiceprint ID and the dynamic key fragment to ensure data security;

[0136] The local decoding module refers to the lightweight semantic parsing engine built into the target device, which preferentially uses a recurrent neural network to extract keywords and switches to the keyword matching mode when the battery is low to reduce the computing load.

[0137] Exemplarily, set the threshold of the priority score 65, When it triggers cloud processing, when the priority score exceeds the threshold of the priority score, trigger cloud processing and perform cloud sharding transmission.

[0138] The beneficial effect of this verification example: It simultaneously utilizes the dynamic collaborative decision-making mechanism of environmental interference and device power to intelligently allocate cloud and local resources to achieve high-efficiency voice interaction. When the noise interference is significant or the power is sufficient, it preferentially calls the edge computing resources to complete complex semantic recognition and uses distributed processing to reduce latency; when the power is insufficient or the environment is relatively quiet, local processing reduces communication energy consumption and protects privacy.

[0139] In one embodiment of the present invention, step S5 includes the following steps:

[0140] S51. Extract the intent identifier and entity word vector from the semantic recognition result;

[0141] Specifically, first, perform structured parsing on the semantic recognition results in the cloud or locally. Input the word segmentation sequence of the semantic recognition results into a bidirectional long short-term memory network. Use the Softmax activation function at the output end to generate intent labels, and use a conditional random field (CRF) to perform sequence labeling on entity words. Finally, separate the intent identifier and the entity word vector;

[0142] Among them, the intent identifier is the target label of the abstract representation of the user's core intent extracted from speech or text, and is used to identify the type of the user's conversation intent, such as classification labels like "query weather" and "play music";

[0143] The semantic recognition result represents the structured data output by the semantic recognition engine in the cloud or locally;

[0144] The entity word vector is the distributed numerical representation of the keywords in the user's statement. The entity objects in the text are converted into multi-dimensional vectors through a pre-trained word embedding model. For example, "Beijing" corresponds to a geographical location vector, and "2023" corresponds to a time vector. Through the separation and extraction of intent and entity, it provides a semantic analysis basis for subsequent multi-round conversations.

[0145] S52. Match the context parameters associated with the context in the historical interaction data to generate the context constraint conditions for multi-round conversations, which are the dialogue logic restriction rules dynamically constructed based on the historical interaction data and are used to restrict the semantic scope of the current response;

[0146] Specifically, retrieve the user's recent N-round conversation records, extract the intent identifier and the entity word vector from them to form a time-series context chain, and calculate the weights of each historical node in the current context through attention weights. For example, if the user first queries "the weather in Beijing" and then asks "how about Shanghai", the previous "weather" topic will be associated as a constraint;

[0147] In the key technology, the calculation of the attention weight satisfies the following formula,

[0148] ;

[0149] Among them, represents the semantic vector of the current conversation, represents the semantic vector of the th round of historical conversation, represents the dependence weight of the current conversation on the th round of history; generate context constraint conditions by weighted aggregation of historical semantic vectors to avoid the breakage of the response logic.

[0150] S53. Adjust the semantic weight and response mode of the response instruction according to the context constraint conditions and the current real-time scene parameters;

[0151] Specifically, the context constraint conditions and the real-time scenario parameter input dynamic decision model calculate the matching degree of each candidate response with the constraint conditions, and adjust the output strategy according to the preset rules. For example, when the user requests "answer in English" twice in a row, the system increases the semantic weight of the English response;

[0152] Among them, the real-time scenario parameters include the device location, time, and sensor parameters; the semantic weight represents the priority score of different response contents and is used to select the most context-appropriate result among multiple candidate responses;

[0153] The response mode represents the output form selected based on the device state or user habits, such as "brief answer" or "detailed broadcast".

[0154] The calculation of the matching degree satisfies the following formula,

[0155] ;

[0156] Among them, represents the matching degree of each candidate response with the constraint conditions, represents the similarity score of the candidate response with the historical context, represents the adaptation score of the candidate response with the real-time scenario, and represent the trainable balance coefficients.

[0157] Finally, the response mode with the highest matching degree is selected to generate the multi-turn dialogue response instruction.

[0158] Exemplarily, when the user continuously inputs "query the weather in Beijing", "precipitation tomorrow", and "how about Shanghai", the execution process is as follows:

[0159] In step S51, the intention identifier of the first request is extracted as "weather query", and the entity word vectors are "Beijing" and "weather"; in step S52, it is matched in the historical data that the intentions of the previous two conversations are both "weather query", and the constraint conditions are constructed as "the time parameter defaults to inherit the next day of the previous request" and "the location parameter can be switched but needs to be associated with the precipitation probability";

[0160] In step S53, according to the "Shanghai" entity in the user's third request, combined with the real-time scenario parameters (the user's location is Beijing), the semantic weight is adjusted to preferentially compare the precipitation data differences between Beijing and Shanghai, and the response mode is generated to directly provide the comparison result.

[0161] In one embodiment of the present invention, constructing a word embedding model includes the following steps:

[0162] Collect entity objects and multi-dimensional vectors from historical texts, then correspond and label the entity objects and multi-dimensional vectors one by one through manual matching. The labeled set is used as the training set of the word embedding model, and a word embedding model is constructed using a deep learning network. Subsequently, the word embedding model is trained using the training set;

[0163] Use the entity object as the input of the model and the multi-dimensional vector as the output of the model. For example, "Beijing" corresponds to a geographical location vector, and "2023" corresponds to a time vector, which is used to convert the entity objects in the text into multi-dimensional vectors;

[0164] Collect entity objects in the current text, input them into the trained word embedding model, and then output the corresponding multi-dimensional vectors, that is, the pre-trained word embedding model converts the entity objects in the text into multi-dimensional vectors.

[0165] In one embodiment of the present invention, step S6 includes the following steps:

[0166] S61. Monitor the rendering queue load rate of the target device's GPU and calculate the periodic allocation window of the idle computing power;

[0167] Specifically, monitor the rendering queue load rate of the target device in real time and evaluate the usage status of the current computing resources. The rendering queue load rate is the ratio of the number of rendering tasks currently being processed by the GPU to the maximum number of tasks that can be processed in parallel. The periodic allocation window of the idle computing power is a time interval dynamically adjusted according to the load rate, which is used to plan the scheduling of subsequent rendering tasks. When the load rate is lower than the preset threshold, the allocation window expands to optimize resource utilization; when the load rate approaches the upper limit, the window shrinks to avoid processing delays;

[0168] Among them, the rendering queue load rate represents a quantitative index of the proportion of the number of rendering tasks being processed by the GPU in the maximum parallel processing capacity, which is used to dynamically evaluate the remaining computing resources;

[0169] The periodic allocation window of the idle computing power represents a task scheduling time interval dynamically delimited according to the current load rate, which balances the response speed and resource consumption by flexibly adjusting the duration.

[0170] Calculate the periodic allocation window of the idle computing power, which satisfies the following formula,

[0171] ;

[0172] Among them, represents the finally allocated window duration; represents the maximum number of concurrent tasks of the GPU; represents the device reference time unit; represents the current number of rendering tasks.

[0173] For example, set the maximum number of concurrent tasks for a GPU on a device , the current number of rendering tasks , if the base time is , the final allocated window duration can be calculated as follows: ,The maximum time that a new task can be executed within the window is limited to 80ms (preset threshold).

[0174] S62, preferentially calling the speech synthesis engine in the allocation window to generate a speech stream, and caching the visual interface data in a low-power rendering queue;

[0175] Specifically, within the idle computing cycle allocation window determined in step S61, the speech synthesis engine is preferentially called to convert the semantic recognition results into speech streams, and the graphic data required for the visual interactive interface is temporarily stored in the low-power rendering queue. The speech synthesis engine uses streaming processing technology to generate audio data in blocks, and the single processing time is shorter than the window time to avoid blocking;

[0176] Among them, stream processing technology is a mechanism that divides continuous data into segments of fixed length and processes them sequentially to avoid processing delays caused by memory fullness;

[0177] The low-power rendering queue refers to the data caching technology based on the reserved area of ​​the video memory, which only maintains the basic frequency of data refresh to reduce the power consumption overhead of graphics rendering.

[0178] Exemplarily, when the allocation window is set to 50ms, the speech synthesis engine first decomposes the response text into audio segments within 50ms and outputs them, and stores the vector parameters of the interface elements in a low-power queue; if the cached data exceeds the queue capacity, non-key frames are discarded and the core metadata of the interface layout is retained.

[0179] S63, if a screen wake-up signal is detected, the visual interactive interface is immediately injected based on the cached data to complete the synchronous output of voice and picture;

[0180] Specifically, when the device screen wake-up signal is triggered, the cached visual interface data is extracted from the low-power rendering queue, the GPU rendering pipeline is called to generate a complete picture, and the real-time output of the speech synthesis engine is synchronized and injected into the screen driver. Synchronous injection is a system-level operation that achieves strict timing matching of multi-modal output through timestamp alignment to ensure seamless connection between voice and picture; the implementation of synchronous output depends on the calibration mechanism of audio stream timestamp and picture rendering frame rate, matching the corresponding picture elements when the voice stream is played to a specific time point.

[0181] Exemplarily, when the voice stream broadcasts "Your itinerary today is as follows", the interface synchronously displays a thumbnail of the schedule list and transitions to the full-screen details after the voice ends. If some animation frames are lost due to the low-power queue capacity limit of the cached data, the key path is completed through an interpolation algorithm to ensure visual coherence.

[0182] The beneficial effects of this verification example are as follows: First, by dynamically monitoring the GPU computing power load and dividing the task window, the real-time performance of voice output is preferentially guaranteed, and at the same time, the interface data is cached in a low-power manner and quickly synchronized and rendered after the screen is awakened. The core lies in differentiating and scheduling the priorities of voice and graphics processing according to the real-time resource status of the device, and making up for the imbalance of resource allocation through the data caching and synchronization mechanism.

[0183] See Appendix Figure 2 , the present invention also proposes an intelligent voice interaction system for 5G communication devices based on deep learning, including the following modules:

[0184] A combined input signal generation module, which is used to generate a combined input signal containing voice frequency domain information and motion state;

[0185] A voice segment noise reduction module, which is used to perform real-time processing on the combined input signal and output the noise-reduced voice segment and its corresponding environmental interference level parameters;

[0186] An identity verification module, which is used to decompose the voiceprint features in the noise-reduced voice segment and generate a basic voiceprint ID and a dynamic key segment containing the user's biometric identifier;

[0187] A distributed processing routing module, which calculates the cloud processing priority score according to the environmental interference level parameter and the remaining battery level value of the device;

[0188] A multi-round dialogue management module, which is used to receive the semantic recognition results from the cloud or locally, and generate multi-round dialogue response instructions in combination with the context parameters in the user's historical interaction data;

[0189] A multi-modal output scheduling module, which preferentially uses the voice synthesis engine to output the results and synchronously injects the visual interaction interface when the screen wake-up state is detected.

[0190] Each of the above-mentioned modules can be implemented in whole or in part by software, hardware and their combinations, supports being embedded in the processor in the form of hardware or independent of the computer device, and also supports being stored in the memory of the computer device in the form of software, facilitating the processor to call and execute the operations corresponding to each of the above modules.

[0191] It should be noted that the human body information (including but not limited to human body device information and personal information, etc.) and data (including but not limited to data for analysis, stored data, and displayed data, etc.) involved in the present invention are all information and data authorized by the human body or fully authorized by all parties. The collection, use, and processing of relevant data require relevant legal standards.

[0192] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. An intelligent voice interaction method for 5G communication devices based on deep learning, characterized in that, It includes the following steps: S1. Obtain the original voice data, call the accelerometer data of the device, and generate a joint input signal containing voice frequency domain information and motion state; S2. Perform real-time processing on the joint input signal, and output the denoised voice segment and its corresponding environmental interference level parameter; S3. Decompose the voiceprint features in the denoised voice segment, and generate a basic voiceprint ID containing the user's biometric identifier and a dynamic key segment; S4. According to the environmental interference level parameter and the remaining battery value of the device, calculate the cloud processing priority score. When the score exceeds the preset threshold, send a voice processing request containing the basic voiceprint ID to the edge computing node, otherwise trigger the local semantic recognition process; S5. Receive the semantic recognition result from the cloud or local, and combine the context parameters in the user's historical interaction data to generate a multi-round dialogue response instruction; S6. Dynamically allocate rendering resources for the response instruction according to the idle computing power of the target device's GPU, preferentially use the voice synthesis engine to output the result, and synchronously inject the visual interaction interface when the screen wake-up state is detected.

2. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 1, characterized in that, Step S1 includes the following steps: Obtain the continuous original voice data collected by the microphone of the target device, and synchronously activate the accelerometer sensor to collect the three-axis acceleration value of the device; Extract the voice frequency features of each time slice from the framed processing result of the original voice data, and calculate the time domain energy value of the device in combination with the three-axis acceleration value; Align the timestamps of the voice frequency features and the time domain energy value to generate a joint input signal containing voice frequency domain information and motion state parameters.

3. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 2, characterized in that, Step S2 includes the following steps: Based on the voice frequency features and device motion state parameters in the joint input signal, dynamically divide the dynamic interference threshold of the voice frequency band; Adjust the band gain coefficient of the adaptive filtering model according to the dynamic interference threshold, perform noise reduction processing on the original voice waveform, and output the denoised voice segment; Calculate the energy difference between the original voice segment and the denoised voice segment, and generate an environmental interference level parameter for quantifying the degree of noise pollution.

4. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 3, characterized in that, Step S3 includes the following steps: Extract the fundamental frequency features and harmonic components in the denoised voice segment, and generate a dynamic voiceprint hash value containing the user's biometric characteristics; Perform an encryption operation based on the dynamic voiceprint hash value and the unique hardware identifier of the device to generate a dynamic key segment; Encrypt the dynamic key and the environmental interference level parameter using the preset public key. The preset public key is the asymmetric encryption public key pre-allocated by the cloud to the device. After generating the authentication parameter, upload it to the cloud.

5. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 4, characterized in that, Generating the dynamic key segment includes the following steps: The dynamic key segment is generated by the exclusive OR operation of the device MAC address and the voiceprint hash value to further enhance security. The formula for key generation ; Among them, is the generated 256-bit dynamic key, is the physical address of the device network adapter, that is, the unique identifier of the device hardware, is a hash function, is a bitwise exclusive OR operation, is the dynamic voiceprint hash value.

6. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 4, characterized in that, Step S4 includes the following steps: According to the environmental interference level parameter and the remaining battery value of the device, calculate the cloud processing priority score through linear weighting; If the priority score exceeds the preset threshold, fragment and transmit the voice processing request containing the dynamic key segment to the edge computing node; If the priority score does not reach the threshold, trigger the local semantic recognition engine to parse the denoised voice segment.

7. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 4, characterized in that, Step S5 includes the following steps: Extract the intent identifier and entity word vector in the semantic recognition result; Retrieve historical interaction data, match context-related context parameters, generate context constraints for multi-turn conversations, and dynamic construction of conversation logic restriction rules based on historical interaction data to restrict the semantic scope of the current response. Adjust the semantic priority of the response instruction according to the dependency weight to generate a response instruction containing multi-turn conversation logic.

8. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 7, characterized in that, Step S6 includes the following steps: Real-time monitor the GPU rendering queue load rate of the target device and dynamically allocate time windows for speech synthesis and graphics rendering based on the idle computing power. Within the time window, preferentially call the speech synthesis engine to stream-process the response instruction and cache the visualization interface metadata in a low-power queue. When a screen wake-up signal is detected, immediately render the interaction interface based on the metadata and synchronize the timestamp with the speech stream.

9. The intelligent voice interaction method for 5G communication devices based on deep learning according to claim 8, characterized in that, Dynamically allocate time windows for speech synthesis and graphics rendering based on idle computing power, including the following steps: Determine the upper limit of the duration of the allocation window according to the ratio of the current number of GPU rendering tasks to the maximum number of concurrent tasks, and calculate the periodic allocation window of the idle computing power, satisfying the following formula. ; Among them, represents the finally allocated window duration; represents the maximum number of concurrent tasks of the GPU; represents the device reference time unit; represents the current number of rendering tasks.

10. The intelligent voice interaction system for 5G communication devices based on deep learning is characterized in that, Includes the following modules: Joint input signal generation module, used to generate a joint input signal containing voice frequency domain information and motion state. Voice segment noise reduction module, used to process the joint input signal in real time and output the noise-reduced voice segment and its corresponding environmental interference level parameter. Identity authentication module, used to decompose the voiceprint features in the noise-reduced voice segment to generate a basic voiceprint ID and a dynamic key segment containing the user's biometric identifier. Distributed processing routing module, calculates the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device. Multi-turn conversation management module, used to receive semantic recognition results from the cloud or local, and generate multi-turn conversation response instructions in combination with the context parameters in the user's historical interaction data. Multi-modal output scheduling module, used to preferentially use the speech synthesis engine to output results and synchronously inject the visualization interaction interface when the screen wake-up state is detected.

Citation Information

Patent Citations

  • Voice interaction recognition analysis method and system based on Bluetooth headset

    CN119889301A

  • Intelligent communication equipment voice noise reduction method and system based on artificial intelligence

    CN119905101A