5G communication equipment intelligent voice interaction method and system based on deep learning

By adopting a deep learning-based intelligent voice interaction method on 5G communication devices, combining voice and acceleration data for real-time noise reduction and feature extraction, the accuracy and security of speech recognition in high-noise environments are solved, and higher recognition accuracy and user experience are achieved.

CN120126463AActive Publication Date: 2025-06-10SHENZHEN BESNEL TECH CO LTD

Patent Information

Application Number
CN202510578703.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-10
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

In dynamic and complex scenarios, in the noise environment of mechanical signal measurement, existing voice interaction technologies are difficult to accurately distinguish the vocal characteristics in the voice signal from the acceleration fluctuation signals caused by device movement, resulting in an increase in the error rate of voiceprint recognition and a deviation in semantic understanding.

Method used

The intelligent voice interaction method of 5G communication devices based on deep learning is adopted. By obtaining original voice data and accelerometer data, generating joint input signals, real-time noise reduction processing, extracting voiceprint features, generating dynamic key fragments, and calculating cloud processing priority based on environmental interference level and device power, and dynamically allocating speech synthesis and graphic rendering resources.

Benefits of technology

In high-noise scenarios, improve the accuracy of voice command recognition by more than 40%, reduce the problem of voice distortion, reduce the risk of identity forgery and playback attacks, improve user operation fluency by more than 50%, and reduce the phenomenon of lag caused by resource seizure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126463A_ABST
    Figure CN120126463A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of voice interaction, and relates to a 5G communication equipment intelligent voice interaction method and system based on deep learning, and the method comprises the steps: generating a joint input signal containing voice frequency domain information and a motion state; outputting the de-noised voice segments and the corresponding environmental interference level parameters; generating a basic voiceprint ID containing a user biological feature identifier and a dynamic key fragment; calculating a cloud processing priority score according to the environmental interference level parameter and the equipment residual electric quantity value, when the score exceeds a preset threshold value, sending a voice processing request containing a basic voiceprint ID to an edge computing node, and otherwise, triggering a local semantic recognition process; and receiving a semantic recognition result from a cloud end or a local end, and generating a multi-round dialogue response instruction in combination with the context parameters in the historical interaction data of the user. According to the invention, the problem that local processing resource overload or cloud communication delay exceeding is easily caused by a fixed noise weight distribution mechanism is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of voice interaction, and relates to an intelligent voice interaction method and system for 5G communication devices based on deep learning. Background Art

[0002] In the voice interaction technology for mechanical signal measurement in a noisy environment, the core problem lies in the insufficient accuracy of signal feature separation in a dynamic and complex scenario, resulting in the susceptibility of effective voice signals to interference from device motion noise and external environmental vibrations. When the device is in a high-speed moving or high-frequency mechanical vibration environment, existing methods are difficult to accurately distinguish the human voice features in the voice signal from the acceleration fluctuation signals caused by the device's own motion, leading to an increase in the error rate of voiceprint recognition and semantic understanding deviation.

[0003] Traditional solutions mainly adopt an acceleration signal-assisted noise reduction mechanism. By pre-setting a motion noise database to match the time-domain energy characteristics of the device's accelerometer, static filtering parameters are formulated to suppress interference in specific frequency bands. The device's movement state is judged by monitoring the peak values of the three-axis acceleration. When the acceleration exceeds the preset threshold, the noise reduction process is started, and a frequency band filter is used to eliminate the motion-related noise.

[0004] Based on the above problems, traditional methods are prone to signal misjudgment in sudden mechanical vibration or complex noise scenarios. The fixed noise weight allocation mechanism ignores the dynamic correlation between the device's power and environmental interference, easily causing local processing resource overload or excessive cloud communication delay, and weakening the comprehensive response ability of the system. Summary of the Invention

[0005] To solve the above problems, the present invention provides an intelligent voice interaction method and system for 5G communication devices based on deep learning.

[0006] In a first aspect, the present invention provides an intelligent voice interaction method for 5G communication devices based on deep learning, adopting the following technical solutions: The intelligent voice interaction method for 5G communication devices based on deep learning includes the following steps: S1. Obtain the original voice data, call the accelerometer data of the device, and generate a joint input signal containing voice frequency domain information and motion state; S2. Perform real-time processing on the joint input signal, and output the noise-reduced voice segment and its corresponding environmental interference level parameter; S3. Decompose the voiceprint features in the noise-reduced voice segment, and generate a basic voiceprint ID containing the user's biometric identifier and a dynamic key segment; S4. Calculate the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device. When the score exceeds the preset threshold, send a voice processing request containing the basic voiceprint ID to the edge computing node, otherwise trigger the local semantic recognition process; S5. Receive semantic recognition results from the cloud or locally, combine context parameters in the user's historical interaction data, and generate multi-turn dialogue response instructions; S6. Dynamically allocate rendering resources for response instructions according to the idle computing power of the target device's GPU, preferentially use the speech synthesis engine to output results, and synchronously inject into the visual interaction interface when the screen wake-up state is detected.

[0007] In a further solution of the present invention, step S1 includes the following steps: Obtain continuous raw speech data collected by the microphone of the target device, and synchronously activate the accelerometer sensor to collect the three-axis acceleration values of the device; Extract the speech frequency features of each time slice from the framed processing result of the raw speech data, and calculate the time-domain energy value of the device in combination with the three-axis acceleration values; Align the timestamps of the speech frequency features and the time-domain energy values to generate a joint input signal containing speech frequency domain information and motion state parameters.

[0008] In a further solution of the present invention, step S2 includes the following steps: Based on the speech frequency features and device motion state parameters in the joint input signal, dynamically divide the dynamic interference threshold of the speech frequency band; Adjust the band gain coefficient of the adaptive filtering model according to the dynamic interference threshold, perform noise reduction processing on the original speech waveform, and output the noise-reduced speech segment; Calculate the energy difference between the original speech segment and the noise-reduced speech segment, and generate an environmental interference level parameter for quantifying the degree of noise pollution.

[0009] In a further solution of the present invention, step S3 includes the following steps: Extract the fundamental frequency features and harmonic components in the noise-reduced speech segment to generate a dynamic voiceprint hash value containing user biometric features; Perform an encryption operation based on the dynamic voiceprint hash value and the device hardware unique identifier to generate a dynamic key segment; Use the preset public key to encrypt the dynamic key and the environmental interference level parameter. The preset public key is an asymmetric encryption public key pre-allocated by the cloud to the device, and upload the generated authentication parameter to the cloud after generating it.

[0010] In a further solution of the present invention, generating the dynamic key segment includes the following steps: The dynamic key segment is generated by combining the device MAC address and the voiceprint hash value through an exclusive OR operation, which further enhances security. The formula for key generation is ;

[0011] Wherein, is the generated 256-bit dynamic key. The physical address of the device network adapter is the unique identifier of the device hardware. is a hash function. is a bitwise exclusive OR operation. is the dynamic voiceprint hash value.

[0012] In a further solution of the present invention, step S4 includes the following steps: According to the environmental interference level parameter and the remaining battery power value of the device, calculate the cloud processing priority score through linear weighting. If the priority score exceeds the preset threshold, fragment and transmit the voice processing request containing the dynamic key segment to the edge computing node. If the priority score does not reach the threshold, trigger the local semantic recognition engine to parse the noise-reduced voice segment.

[0013] In a further solution of the present invention, step S5 includes the following steps: Extract the intent identifier and entity word vector from the semantic recognition result. Retrieve the historical interaction data, match the context-related context parameters, generate the context constraint conditions for the multi-round dialogue, and the dialogue logic restriction rules dynamically constructed based on the historical interaction data are used to restrict the semantic scope of the current response. Adjust the semantic priority of the response instruction according to the dependency weight, and generate a response instruction containing the multi-round dialogue logic.

[0014] In a further solution of the present invention, step S6 includes the following steps: Real-time monitor the GPU rendering queue load rate of the target device, and dynamically allocate the time windows for voice synthesis and graphics rendering based on the idle computing power. Within the time window, preferentially call the voice synthesis engine to stream-process the response instruction, and cache the visualization interface metadata in a low-power queue. When detecting the screen wake-up signal, immediately render the interaction interface based on the metadata and synchronize the timestamp with the voice stream.

[0015] In a further solution of the present invention, dynamically allocating the time windows for voice synthesis and graphics rendering based on the idle computing power includes the following steps: According to the ratio of the current number of GPU rendering tasks to the maximum number of concurrent tasks, determine the upper limit of the duration of the allocation window, and calculate the periodic allocation window of the idle computing power, satisfying the following formula: ;

[0016] Wherein, represents the finally allocated window duration; represents the maximum number of concurrent tasks of the GPU; represents the device reference time unit. Indicates the current number of rendering tasks.

[0017] In a second aspect, the present invention provides an intelligent voice interaction system for 5G communication devices based on deep learning, adopting the following technical solutions: A combined input signal generation module, used to generate a combined input signal containing voice frequency domain information and motion state; A voice segment noise reduction module, used to perform real-time processing on the combined input signal and output the noise-reduced voice segment and its corresponding environmental interference level parameter; An identity authentication module, used to decompose the voiceprint features in the noise-reduced voice segment and generate a basic voiceprint ID and a dynamic key segment containing the user's biometric identifier; A distributed processing routing module, which calculates the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device; A multi-round dialogue management module, used to receive the semantic recognition results from the cloud or locally, and generate multi-round dialogue response instructions by combining the context parameters in the user's historical interaction data; A multi-modal output scheduling module, which preferentially uses the speech synthesis engine to output results and synchronously injects the visual interaction interface when detecting the screen wake-up state.

[0018] In summary, the present invention includes the following beneficial technical effects: 1. Through the combined analysis of the device motion state and the voice signal, the frequency band noise filtering threshold is dynamically adjusted, eliminating environmental interference while retaining the biometric details of the user's voiceprint. Combining the accelerometer data to perceive the device motion intensity in real time, thereby dynamically optimizing the noise reduction algorithm parameters, improving the voice command recognition accuracy by more than 40% in high-noise scenarios such as cycling and in-vehicle, and at the same time reducing the voice distortion problem caused by excessive noise reduction; 2. Adopting a dual encryption technology of voiceprint biometrics and hardware identifiers to construct a dynamic security mechanism, effectively resisting identity forgery and replay attacks; performing exclusive OR operation and encrypted transmission on the real-time generated voiceprint hash value and the unique device identifier. Compared with the traditional fixed key scheme, this method reduces the risk of key leakage by more than 75%, and the dynamic key segment generated for each interaction can adapt to the encryption strength requirements in different noise environments; 3. Introducing a dynamic scheduling strategy for multi-modal output, improving the interaction experience through GPU computing power monitoring and audio-visual synchronization mechanism. Dividing the processing window according to the real-time rendering load, giving priority to ensuring the smoothness of speech synthesis, and caching the interface data to meet the immediate rendering requirements when the screen wakes up. This method breaks through the limitations of the traditional single output mode, achieving millisecond-level synchronization between voice announcements and visual interface switching, improving the operation fluency of users in multi-task concurrent scenarios by more than 50%, and at the same time avoiding the jamming phenomenon caused by resource contention. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. The drawings are used to provide a further understanding of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 Flow schematic diagram of the intelligent voice interaction method for 5G communication devices based on deep learning is disclosed.

[0021] Figure 2 Structural schematic diagram of the intelligent voice interaction system for 5G communication devices based on deep learning is disclosed. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0023] The following will make a preferred and detailed description of the present invention in combination with the attached Figure 1 - Figure 2 The present invention is preferably described in detail as follows.

[0024] Referring to the attached Figure 1 , the present invention proposes an intelligent voice interaction method for 5G communication devices based on deep learning, including the following steps: S1. Obtain the original voice data, call the accelerometer data of the device, and generate a joint input signal including voice frequency domain information and motion state; S2. Perform real-time processing on the joint input signal, and output the denoised voice segment and its corresponding environmental interference level parameter; S3. Decompose the voiceprint features in the denoised voice segment, and generate a basic voiceprint ID and a dynamic key segment including the user's biometric identifier; S4. Calculate the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device. When the score exceeds the preset threshold, send a voice processing request including the basic voiceprint ID to the edge computing node, otherwise trigger the local semantic recognition process; S5. Receive the semantic recognition result from the cloud or locally, and generate a multi-round dialogue response instruction in combination with the context parameters in the user's historical interaction data; S6. Dynamically allocate rendering resources for response instructions based on the idle computing power of the target device GPU, give priority to using the speech synthesis engine to output results, and synchronously inject the visual interactive interface when the screen wake-up state is detected.

[0025] In one embodiment of the present invention, step S1 comprises the following steps: S11, collecting original voice data of human body; Specifically, the original voice data emitted by the human body is continuously collected through the built-in microphone of the target device; the original voice data is the human voice signal collected by the built-in microphone of the target device without noise reduction processing, including the speaking content and environmental noise. The original voice data collected by the microphone is divided into frames, and the frequency characteristics of each time slice in the voice waveform are extracted. The original voice data collected by the microphone is a continuous time domain signal, which needs to be divided into several fixed-length voice segments through frame processing; Among them, by frame segmentation, the continuous speech stream is divided into segments with equal time windows, which is convenient for analyzing the local characteristics of non-stationary signals segment by segment. Frequency characteristics represent the energy distribution parameters in the frequency domain of speech signals, for example: the Mel frequency cepstrum coefficients generated by Fourier transform reflect the speech components in different frequency bands.

[0026] For example, assuming that the user says the original voice of "turn on the lights", the typical frame length is 20-40 milliseconds, for example, 25 milliseconds per frame, and each frame signal is subjected to fast Fourier transform to generate a frequency feature vector containing 128-dimensional MFCC.

[0027] S12, obtaining accelerometer data of the device; Specifically, within 0.1 seconds after the collection of the original human voice data begins, the accelerometer sensor of the target device is synchronously activated to obtain the accelerometer data of the target device in three-dimensional space; the accelerometer data, the motion parameters measured by the built-in inertial sensor of the target device, including the three-axis acceleration value and the angular displacement, calculates the time domain energy value of the current motion state of the target device, which satisfies the following formula: ;

[0028] in, , , Represents the three-axis acceleration value, that is, the instantaneous acceleration data of the target device in three-dimensional space; Indicates the time-domain energy value of the device, which represents the movement intensity of the target device. The higher the energy value, the more intense the device movement.

[0029] S13, generating a joint input signal including speech frequency domain information and motion state; Specifically, the speech frequency features are aligned with the timestamps of the time-domain energy values and encapsulated into a multi-dimensional data packet containing speech frequency-domain information and motion states as a combined input signal; the combined input signal includes the speech frequency features and the time-domain energy values of the device motion states.

[0030] In one embodiment of the present invention, step S2 includes the following steps: Based on the speech frequency features of the combined input signal and the device motion state parameters, the dynamic interference thresholds of the speech frequency bands are divided. The dynamic interference threshold is the upper limit of the frequency band noise tolerance dynamically adjusted based on the real-time motion state and historical noise data, and is used to distinguish speech and interference; Among them, the speech frequency band refers to a specific frequency range divided in the speech signal spectrum; The calculation of the dynamic interference threshold depends on the statistical relationship between the motion state and historical noise, and satisfies the following formula ;

[0031] Among them, represents the dynamic interference threshold of the th frequency band of the current frame, represents the motion state weight coefficient, represents the device time-domain energy value, represents the historical noise weight coefficient, represents the th historical noise mean value of the frequency band of the current frame, is calculated by combining historical data, and are determined by laboratory tests.

[0032] Exemplarily, in a cycling scenario, the high-frequency vibration detected by the device , combined with the average noise in the 500 - 1000 Hz frequency band in historical data . Assuming , , then the dynamic interference threshold is ; If the actual noise in the nd frequency band exceeds the dynamic interference threshold , it is determined that the nd frequency band is interference and is filtered out.

[0033] Use a lightweight noise reduction model to perform adaptive filtering on the speech waveform and adjust the filtering gain coefficient according to the dynamic interference threshold. The lightweight noise reduction model adopts an architecture combining a convolutional neural network and an adaptive filter, uses the multi-dimensional data packet of the combined input signal as the input, and performs adaptive filtering on the speech waveform as the output; The filtering gain coefficient is the scaling ratio of the signal intensity of each frequency band during the filtering process, which retains the speech components and attenuates the noise. It is adjusted according to the ratio of the dynamic interference threshold to the current band noise and satisfies the following formula: ;

[0034] where, is the filtering gain coefficient of the th frequency band of the current frame, is the dynamic interference threshold of the th frequency band of the current frame, is the noise energy value of the th frequency band of the current frame.

[0035] When the current band noise is lower than the dynamic interference threshold, the filtering gain coefficient decreases linearly with the current band noise; when the current band noise exceeds the dynamic interference threshold, the filtering gain coefficient forcibly attenuates 90% of the signal energy.

[0036] Exemplarily, if the noise energy value of a certain band noise , and the dynamic interference threshold of this band , then , and 78% of the signal energy of the noise of the signal in this band is suppressed; if the noise energy value of a certain band noise , and the current band noise exceeds the dynamic interference threshold, 90% of the signal energy of the noise of the signal in this band is suppressed.

[0037] The speech signal after adaptive filtering processing is reconstructed into a time-domain speech signal through inverse Fourier transform and spliced into a continuous noise-reduced speech segment; the noise-reduced speech segment, the clear speech signal waveform after frequency-domain filtering and signal reconstruction. Calculate the total energy difference between the original speech segment and the noise-reduced speech segment to generate an environmental interference level parameter, which satisfies the following formula: ;

[0038] where, represents the energy value of the original speech segment; represents the energy value of the noise-reduced speech segment; represents the environmental interference level parameter, and the higher the value, the more serious the noise pollution.

[0039] The environmental interference level parameter represents an index that quantifies the ratio of noise to speech energy and is used to evaluate the difficulty of noise reduction. By calculating the total energy difference between the original speech segment and the noise-reduced speech segment, the result is mapped to low, medium, and high interference levels, which is used to guide the decision-making of subsequent speech recognition. For example, in a high-interference environment, a more robust recognition algorithm is preferentially called; Combined with the calculation of historical data, the first threshold of the environmental interference level parameter is set to , the second threshold of the environmental interference level parameter is ; Low interference level, corresponding to , indicating that the environmental noise can be ignored, and the system enables a low-latency lightweight noise reduction model; Medium interference level, corresponding to , it is necessary to start the enhanced noise suppression mode and extend the voice endpoint detection window length; High interference level, corresponding to , force the trigger of multi-device collaborative verification or require the user to confirm sensitive operations twice.

[0040] The environmental interference level parameter endows the system with the ability to dynamically adapt to the noise environment; Exemplarily, in a quiet office, the low interference level can retain the high-frequency details of the voice to improve the accuracy of voiceprint recognition; in a factory mechanical noise scenario, the high interference level will guide the system to filter redundant frequency bands to reduce the risk of misrecognition; The user performs voice payment outdoors in windy and rainy weather, , set and , the system determines it as a medium interference level, automatically enhances the anti-noise weight of the voice feature extraction algorithm, and increases the audio sampling duration to 2 seconds to ensure the voiceprint matching accuracy of the "Pay 100 yuan" instruction.

[0041] The beneficial effects of this verification example: In strenuous exercise or high-noise environments, it improves the signal-to-noise ratio of the voice signal, reduces voice distortion caused by excessive noise reduction, and provides purer input data for subsequent voiceprint feature extraction.

[0042] In one embodiment of the present invention, step S3 includes the following steps: S31. Extract the fundamental frequency features and harmonic components of the user's voiceprint from the noise-reduced voice segment to generate a dynamic voiceprint hash value; Specifically, after short-time analysis of the noise-reduced voice segment, extract the fundamental frequency features and harmonic components of the user's voiceprint and calculate the amplitude distribution of each harmonic component; the fundamental frequency features of the user's voiceprint and the first three harmonic components form a voiceprint feature vector, and the voiceprint feature vector is processed through a hash function to generate a unique dynamic voiceprint hash value, ;

[0043] Among them, represents a 128-bit hash value, represents the fundamental frequency mean, are the amplitudes of the first three harmonic components respectively; Among them, the fundamental frequency feature represents the basic frequency of vocal cord vibration in the voice signal; the dynamic voiceprint hash value represents a unique identifier generated based on real-time voiceprint features, supporting cross-session identity binding and verification.

[0044] Exemplarily, after the user says "turn on the air conditioner", the amplitudes of the first three harmonic components are set 、 、 , and after being processed by the SHA-256 hash function, a hash value of a fixed length is generated 。

[0045] S32. Generate a dynamic key based on the device hardware unique identifier and the dynamic voiceprint hash value; Specifically, the dynamic key is generated by combining the device MAC address and the voiceprint hash value through an exclusive OR operation, further enhancing security. The formula for key generation is ;

[0046] Among them, is the generated 256-bit dynamic key, is the physical address of the device network adapter (device hardware unique identifier), is the hash function, is the bitwise exclusive OR operation, is the dynamic voiceprint hash value.

[0047] Exemplarily, the MAC address of a certain 5G communication device is 00-B0-D0-63-C2-26, and the voiceprint hash value is a3f5...7c1d. After the exclusive OR operation and hash processing, a dynamic key is generated for encrypting subsequent operation instructions.

[0048] S33. Encrypt the dynamic key and the environmental interference level parameter using the pre-set public key. The pre-set public key is the asymmetric encryption public key pre-allocated by the cloud to the device, and after generating the identity verification parameter, upload it to the cloud; Specifically, the device pre-stores the RSA public key published by the cloud, and performs asymmetric encryption on the dynamic key and the environmental interference level parameter to generate the identity verification parameter. The identity verification parameter represents an encrypted data packet containing the dynamic key and the environmental interference level, and is used for the cloud to verify the legality and operation authority after reverse decryption. The encryption formula is ;

[0049] Among them, represents the encrypted identity verification parameter, represents the environmental interference level parameter, represents the public key encryption operation, is the generated 256-bit dynamic key, represents data splicing.

[0050] Exemplarily, the dynamic key is pre-set, and the interference level parameter After being encrypted by the public key, the authentication parameters are generated. After being uploaded to the cloud, the original data can be decrypted and restored.

[0051] The beneficial effects of this verification example: the dual binding of voiceprint biometrics and device hardware constructs a dynamic security mechanism, solving the risks of identity forgery and replay attacks in traditional voice interactions; its core principle lies in the real-time generation of an encryption key containing voiceprint characteristics and the unique device identifier, and combining asymmetric encryption to synchronously encode the noise environment factors into the verification parameters, enabling the cloud to accurately identify legitimate users and exclude malicious interference.

[0052] In one embodiment of the present invention, step S4 includes the following steps: Map the environmental interference level parameter to an interference weight coefficient, and quantify the remaining battery value of the device into an available computing resource value. The available computing resource value is obtained through the mapping relationship between the remaining battery value and the preset power consumption model, and satisfies the following formula, ;

[0053] Where, is the current remaining battery value of the device, is the energy consumption per unit time of local semantic recognition, represents the expected time required for voice processing, represents the available computing resource value. This formula converts the battery power into the resource capacity for sustainable computing, restricting overloaded local processing; Among them, the interference weight coefficient converts the interference level parameter into a priority determination weight through a piecewise function. When the interference level parameter ≥0.7, the weight is 0.9, 0.3≤ <0.7, the weight is 0.6, and when the interference level parameter <0.3, the weight is 0.2, aiming to strengthen the dependence on cloud processing in a high-interference environment.

[0054] Calculate the cloud processing priority score through the linear weighting of the interference weight coefficient and the available computing resource value, and satisfy the following formula, ;

[0055] Where, represents the cloud processing priority score. The higher this score, the higher the priority of cloud processing; represents the interference weight coefficient; represents the available computing resource value; represents the dynamic adjustment coefficient.

[0056] Exemplarily, set the interference weight coefficient (interference level parameter When it is ≥ 0.7, the weight is ); , the default value is 0.7, and when the device is in a moving state, it is reduced to 0.5 to reduce the energy consumption risk of cloud communication; ; It can be calculated that: .

[0057] When the priority score exceeds the threshold of the priority score, cloud processing is triggered and cloud sharding transmission is executed. The request containing the basic voiceprint ID is sharded and sent to multiple computing units of the edge node; otherwise, the local decoding module is called to start semantic recognition; Among them, for sharding transmission, the voice processing request is split into multiple data blocks and transmitted to different edge computing units for parallel processing. Each shard contains the basic voiceprint ID and the dynamic key fragment to ensure data security; The local decoding module refers to the lightweight semantic parsing engine built into the target device. It preferentially uses a recurrent neural network to extract keywords and switches to the keyword matching mode when the battery power is low to reduce the computing load.

[0058] Exemplarily, set the threshold of the priority score 65, When it is, cloud processing is triggered. When the priority score exceeds the threshold of the priority score, cloud processing is triggered and cloud sharding transmission is executed.

[0059] The beneficial effects of this verification example: By simultaneously using the dynamic collaborative decision-making mechanism of environmental interference and device power, intelligent allocation of cloud and local resources is realized to achieve high-efficiency voice interaction. When the noise interference is significant or the power is sufficient, edge computing resources are preferentially called to complete complex semantic recognition, and distributed processing is used to reduce latency; when the power is insufficient or the environment is relatively quiet, local processing reduces communication energy consumption and protects privacy.

[0060] In one embodiment of the present invention, step S5 includes the following steps: S51. Extract the intent identifier and entity word vector from the semantic recognition result; Specifically, first, the semantic recognition result from the cloud or local is structurally parsed. The word segmentation sequence of the semantic recognition result is input into a bidirectional long short-term memory network, and the Softmax activation function is used at the output end to generate intent labels. The conditional random field CRF is used to perform sequence labeling on the entity words, and finally, the intent identifier and entity word vector are separated; Among them, the intent identifier is the target label of the abstract representation of the user's core intent extracted from speech or text, used to identify the type of the user's conversation intent, such as classification labels like "query weather" and "play music"; Semantic recognition result, which represents the structured data output by the cloud or local semantic recognition engine; Entity word vector, which is the distributed numerical representation of keywords in the user's statement. The entity objects in the text are converted into multi-dimensional vectors through a pre-trained word embedding model. For example, "Beijing" corresponds to a geographical location vector, and "2023" corresponds to a time vector. By separating and extracting the intent and entity, it provides a semantic analysis basis for subsequent multi-turn conversations.

[0061] S52. Context parameters matching the context association in the historical interaction data are used to generate context constraint conditions for multi-turn conversations, which are dialogue logic restriction rules dynamically constructed based on historical interaction data and are used to restrict the semantic scope of the current response; Specifically, retrieve the user's recent N-round conversation records, extract the intent identifiers and entity word vectors from them to form a temporal context chain, and calculate the weights of each historical node in the current context through attention weights. For example, if the user first queries "the weather in Beijing" and then asks "how about Shanghai", the previous "weather" theme will be associated as a constraint; In the key technology, the calculation of the attention weight satisfies the following formula, ;

[0062] Among them, represents the semantic vector of the current conversation, represents the semantic vector of the nth round of historical conversation, represents the dependence weight of the current conversation on the nth round of history; the context constraint conditions are generated by weighted aggregation of historical semantic vectors to avoid the breakage of the response logic.

[0063] S53. According to the context constraint conditions and the current real-time scene parameters, adjust the semantic weight and response mode of the response instruction; Specifically, the context constraint conditions and real-time scene parameters are input into a dynamic decision-making model to calculate the matching degree of each candidate response with the constraint conditions, and the output strategy is adjusted according to the preset rules. For example, when the user requests "answer in English" twice in a row, the system increases the semantic weight of the English response; Among them, the real-time scene parameters include device location, time, and sensor parameters; the semantic weight represents the priority score of different response contents and is used to select the most context-compliant result among multiple candidate responses; The response mode represents the output form selected based on the device state or user habits, such as "brief answer" or "detailed broadcast".

[0064] The calculation of the matching degree satisfies the following formula, ;

[0065] Among them, represents the matching degree between each candidate response and the constraint conditions, represents the similarity score between the candidate response and the historical context, represents the adaptation score between the candidate response and the real-time scenario, and represents a trainable balance coefficient.

[0066] Finally, select the response mode with the highest matching degree to generate a multi-turn dialogue response instruction.

[0067] Exemplarily, the user continuously inputs "Query the weather in Beijing", "Precipitation tomorrow", "What about Shanghai?", and the execution process is as follows: In step S51, extract the intent identifier of the first request as "weather query", and the entity word vectors as "Beijing" and "weather"; in step S52, match that the intents of the previous two conversations in the historical data are both "weather query", and construct the constraint conditions "The time parameter defaults to inherit the next day of the previous request" and "The location parameter can be switched but needs to be associated with the precipitation probability"; In step S53, according to the "Shanghai" entity in the user's third request, combined with the real-time scenario parameters (the user's location is Beijing), adjust the semantic weight to give priority to comparing the precipitation data differences between Beijing and Shanghai, and generate a response mode to directly provide the comparison result.

[0068] In one embodiment of the present invention, constructing a word embedding model includes the following steps: Collect entity objects and multi-dimensional vectors in historical texts, then correspond and label the entity objects and multi-dimensional vectors one by one through manual matching, and use the labeled set as the training set of the word embedding model. Use a deep learning network to construct the word embedding model, and then use the training set to train the word embedding model; Use the entity object as the input of the model and the multi-dimensional vector as the output of the model. For example, "Beijing" corresponds to the geographical location vector, and "2023" corresponds to the time vector, so as to convert the entity objects in the text into multi-dimensional vectors; Collect entity objects in the current text, input them into the trained word embedding model, and then output the corresponding multi-dimensional vectors, that is, the pre-trained word embedding model converts the entity objects in the text into multi-dimensional vectors.

[0069] In one embodiment of the present invention, step S6 includes the following steps: S61. Monitor the rendering queue load rate of the target device GPU and calculate the periodic allocation window of the idle computing power; Specifically, the GPU rendering queue load rate of the target device is monitored in real time to evaluate the usage status of the current computing resources. The rendering queue load rate is the ratio of the number of rendering tasks currently being processed by the GPU to the maximum number of tasks that can be processed in parallel. The idle computing power cycle allocation window is a time interval dynamically adjusted according to the load rate, which is used to plan the scheduling of subsequent rendering tasks. When the load rate is lower than the preset threshold, the allocation window is extended to optimize resource utilization; when the load rate approaches the upper limit, the window is narrowed to avoid processing delays; Among them, the rendering queue load rate is a quantitative indicator representing the proportion of the number of rendering tasks being processed by the GPU to the maximum parallel processing capacity, which is used to dynamically evaluate the remaining computing resources; The idle computing power cycle allocation window represents a task scheduling time interval dynamically delimited according to the current load rate, and balances the response speed and resource consumption by flexibly adjusting the duration.

[0070] Calculate the cycle allocation window of the idle computing power, which satisfies the following formula, ;

[0071] Among them, represents the finally allocated window duration; represents the maximum number of concurrent tasks of the GPU; represents the device reference time unit; represents the current number of rendering tasks.

[0072] Exemplarily, set the maximum number of concurrent tasks of the GPU of a certain device , the current number of rendering tasks , if the reference time , the finally allocated window duration, it can be calculated that: , the maximum time for executing new tasks within the window is limited to 80ms (preset threshold).

[0073] S62. Prioritize the invocation of the speech synthesis engine within the allocation window to generate a speech stream, and cache the visualization interface data in the low-power rendering queue; Specifically, within the idle computing power cycle allocation window determined in step S61, prioritize the invocation of the speech synthesis engine to convert the semantic recognition result into a speech stream, and at the same time temporarily store the graphic data required by the visualization interaction interface in the low-power rendering queue. The speech synthesis engine uses streaming processing technology to generate audio data in blocks, and the single processing duration is shorter than the window duration to avoid blocking; Among them, the streaming processing technology is a mechanism that divides continuous data into segments with a fixed duration and processes them sequentially, avoiding processing delays caused by the memory being full; The low-power rendering queue represents a data caching technology based on the video memory reservation area, which only maintains the data refresh at the base frequency to reduce the power consumption overhead of graphics rendering.

[0074] Exemplarily, when the allocation window is set to 50 ms, the speech synthesis engine first decomposes the response text into audio segments within 50 ms for output, and stores the vector parameters of the interface elements in the low-power queue; if the cached data exceeds the queue capacity, non-critical frames are discarded, and the core metadata of the interface layout is retained.

[0075] S63. When a screen wake-up signal is detected, inject the visual interaction interface immediately based on the cached data to complete the synchronous output of voice and picture; Specifically, when the device screen wake-up signal is triggered, the cached visual interface data is extracted from the low-power rendering queue, and the GPU rendering pipeline is called to generate a complete picture, which is synchronously injected into the screen driver together with the real-time output of the speech synthesis engine. Synchronous injection is a system-level operation that realizes strict timing matching of multi-modal output through timestamp alignment, ensuring seamless connection between voice and picture; the realization of synchronous output depends on the audio stream timestamp and the picture rendering frame rate calibration mechanism, and the corresponding picture elements are matched at a specific time point when the voice stream is played.

[0076] Exemplarily, when the voice stream broadcasts "Your schedule today is as follows", the interface synchronously displays the thumbnail of the schedule list and transitions to the full-screen details after the voice ends. If some animation frames are lost due to the capacity limit of the low-power queue in the cached data, the key path is completed through the interpolation algorithm to ensure visual coherence.

[0077] The beneficial effects of this verification example: First, by dynamically monitoring the GPU computing power load and dividing the task window, the real-time performance of voice output is guaranteed first, and at the same time, the interface data is cached in a low-power manner and quickly synchronously rendered after the screen wakes up. Its core lies in differentiating the scheduling priorities of voice and graphics processing according to the real-time resource status of the device, and making up for the imbalance of resource allocation through the data caching and synchronization mechanism.

[0078] See Appendix Figure 2 , the present invention also proposes an intelligent voice interaction system for 5G communication devices based on deep learning, including the following modules: A combined input signal generation module, which is used to generate a combined input signal containing voice frequency domain information and motion state; A voice segment noise reduction module, which is used to perform real-time processing on the combined input signal and output the noise-reduced voice segment and its corresponding environmental interference level parameters; An identity verification module, which is used to decompose the voiceprint features in the noise-reduced voice segment and generate a basic voiceprint ID and a dynamic key segment containing the user's biometric identifier; A distributed processing routing module, which calculates the cloud processing priority score according to the environmental interference level parameter and the remaining battery level of the device; The multi-turn dialogue management module is used to receive the semantic recognition results from the cloud or local, and generate multi-turn dialogue response instructions by combining the context parameters in the user's historical interaction data; The multi-modal output scheduling module is used to preferentially use the speech synthesis engine to output results, and synchronously inject the visual interaction interface when the screen wake-up state is detected.

[0079] Each of the above-mentioned modules can be implemented in whole or in part by software, hardware, and their combinations, supporting hardware form to be embedded in or independent of the processor in the computer device, and also supporting software form to be stored in the memory of the computer device, facilitating the processor to call and execute the operations corresponding to each of the above modules.

[0080] It should be noted that the human body information (including but not limited to human body device information and personal information, etc.) and data (including but not limited to data for analysis, stored data, and displayed data, etc.) involved in the present invention are all information and data authorized by the human body or fully authorized by all parties. The collection, use, and processing of relevant data require relevant legal standards.

[0081] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A 5G communication device intelligent voice interaction method based on deep learning, characterized in that: The following steps are involved: S1, obtaining original voice data, calling the accelerometer data of the device, and generating a joint input signal including voice frequency domain information and motion state; S2, processing the combined input signal in real time, and outputting the noise-reduced speech segment and its corresponding environmental interference level parameter; S3, decomposing the voiceprint features in the noise-reduced speech segment to generate a basic voiceprint ID and a dynamic key segment containing a user's biometric identifier; S4. Calculate the cloud processing priority score based on the environmental interference level parameter and the remaining power value of the device. When the score exceeds the preset threshold, send a voice processing request containing the basic voiceprint ID to the edge computing node, otherwise trigger the local semantic recognition process; S5, receiving the semantic recognition results from the cloud or local, combining the context parameters in the user's historical interaction data, and generating multi-round dialogue response instructions; S6. Dynamically allocate rendering resources for response instructions based on the idle computing power of the target device GPU, give priority to using the speech synthesis engine to output results, and synchronously inject the visual interactive interface when the screen wake-up state is detected.

2. The 5G communication device intelligent voice interaction method based on deep learning according to claim 1 is characterized in that: Step S1 includes the following steps: Obtain the continuous original voice data collected by the microphone of the target device, and synchronously activate the accelerometer sensor to collect the three-axis acceleration value of the device; The frame processing result of the original voice data extracts the voice frequency characteristics of each time slice, and calculates the time domain energy value of the device in combination with the three-axis acceleration value; The timestamps of the speech frequency features and the time domain energy values ​​are aligned to generate a joint input signal containing the speech frequency domain information and motion state parameters.

3. The 5G communication device intelligent voice interaction method based on deep learning according to claim 2 is characterized in that: Step S2 includes the following steps: Dynamically divide the dynamic interference threshold of the speech frequency band based on the speech frequency characteristics and device motion state parameters in the joint input signal; The frequency band gain coefficient of the adaptive filtering model is adjusted according to the dynamic interference threshold, the original speech waveform is subjected to noise reduction processing, and the noise-reduced speech segment is output; The energy difference between the original speech segment and the noise-reduced speech segment is calculated to generate an environmental interference level parameter that quantifies the degree of noise pollution.

4. The 5G communication device intelligent voice interaction method based on deep learning according to claim 3 is characterized in that: Step S3 includes the following steps: Extract the fundamental frequency features and harmonic components in the noise-reduced speech segment and generate a dynamic voiceprint hash value containing the user's biometric features; Perform encryption operations based on the dynamic voiceprint hash value and the device hardware unique identifier to generate a dynamic key fragment; The dynamic key and the environmental interference level parameter are encrypted using the preset public key, which is an asymmetric encryption public key pre-assigned to the device by the cloud. The authentication parameters are generated and uploaded to the cloud.

5. The 5G communication device intelligent voice interaction method based on deep learning according to claim 4 is characterized in that: Generating a dynamic key fragment includes the following steps: The dynamic key fragment is generated by combining the device MAC address and the voiceprint hash value through an XOR operation to further enhance security. The formula for key generation is: ; in, is a 256-bit dynamic key generated, The physical address of the device's network adapter is the unique hardware identifier of the device. is a hash function, is a bitwise XOR operation, It is the dynamic voiceprint hash value.

6. The 5G communication device intelligent voice interaction method based on deep learning according to claim 4 is characterized in that: Step S4 includes the following steps: The cloud processing priority score is calculated by linear weighting based on the environmental interference level parameters and the remaining power value of the device; If the priority score exceeds the preset threshold, the voice processing request fragment containing the dynamic key fragment is transmitted to the edge computing node; If the priority score does not reach the threshold, the local semantic recognition engine is triggered to parse the noise-reduced speech segment.

7. The 5G communication device intelligent voice interaction method based on deep learning according to claim 4 is characterized in that: Step S5 includes the following steps: Extract intent identifiers and entity word vectors from semantic recognition results; Retrieve historical interaction data, match contextual parameters associated with the context, generate contextual constraints for multiple rounds of dialogue, and dynamically build dialogue logic restriction rules based on historical interaction data to constrain the semantic scope of the current response; The semantic priority of the answer instruction is adjusted according to the dependency weight, and a response instruction containing multi-round dialogue logic is generated.

8. The 5G communication device intelligent voice interaction method based on deep learning according to claim 7 is characterized in that: Step S6 includes the following steps: Monitor the GPU rendering queue load rate of the target device in real time, and dynamically allocate time windows for speech synthesis and graphics rendering based on idle computing power; Prioritize calling the speech synthesis engine to stream process the response instruction within the time window, and cache the visual interface metadata in a low-power queue; When a screen wake-up signal is detected, the interactive interface is rendered instantly based on the metadata and timestamps are synchronized with the voice stream.

9. The 5G communication device intelligent voice interaction method based on deep learning according to claim 8, characterized in that: Dynamically allocating time windows for speech synthesis and graphics rendering based on idle computing power includes the following steps: According to the ratio of the current number of GPU rendering tasks to the maximum number of concurrent tasks, the upper limit of the allocation window is determined, and the period allocation window of idle computing power is calculated to meet the following formula: ; in, Indicates the final allocated window duration; Indicates the maximum number of concurrent tasks on the GPU; Indicates the device base time unit; Indicates the current number of rendering tasks.

10. The 5G communication equipment intelligent voice interaction system based on deep learning is characterized by: Includes the following modules: A joint input signal generating module, used to generate a joint input signal including speech frequency domain information and motion state; A speech segment noise reduction module is used to process the combined input signal in real time and output the noise-reduced speech segment and its corresponding environmental interference level parameter; The authentication module is used to decompose the voiceprint features in the noise-reduced speech segment and generate a basic voiceprint ID and a dynamic key segment containing the user's biometric identification; The distributed processing routing module calculates the cloud processing priority score based on the environmental interference level parameters and the remaining power value of the device; The multi-round dialogue management module is used to receive the semantic recognition results from the cloud or local, and generate multi-round dialogue response instructions based on the context parameters in the user's historical interaction data; The multimodal output scheduling module is used to give priority to the output results of the speech synthesis engine and simultaneously inject the visual interactive interface when the screen wake-up state is detected.

Citation Information

Patent Citations

  • Voice interaction recognition analysis method and system based on Bluetooth headset

    CN119889301A

  • Intelligent communication equipment voice noise reduction method and system based on artificial intelligence

    CN119905101A

  • User biological feature authentication method and system

    US20170118205A1

Cited By

  • Data identification method and system based on neural network model and application

    CN120452432A

  • Intelligent pickup and speech recognition system based on multi-modal fusion

    CN120954408A

  • Calculation power distribution method and system for cloud computing platform to support edge collaboration

    CN122395201A

  • A cloud computing platform supports edge coordination computing power allocation method and system

    CN122395201B