Target voice regulation and control method, device and equipment based on intelligent glasses
By obtaining audio and portrait information in smart glasses, determining the target speaker, and dynamically adjusting voice characteristics based on user preference data, the problem of difficulty in personalizing smart glasses is solved, and a clear and natural voice interaction experience is achieved in different scenarios.
Patent Information
- Application Number
- CN202510990661.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-08-26
AI Technical Summary
The existing speech processing technology of smart glasses is difficult to differentiate sound optimization and adjustment based on the auditory habits and scene needs of different users, making it difficult for users to obtain an auditory experience that meets their own needs, limiting personalized capabilities.
By obtaining audio and portrait information, determining the target speaker, and dynamically adjusting the voice characteristics based on the mapping relationship between the noise level and the user's preference data, using the amplifier on the smart glasses to adjust the frequency band to achieve personalized sound optimization.
Accurately capture the voice of the target spokesperson, improve the accuracy and interaction efficiency of speech recognition, provide personalized voice auditory experience, adapt to changes in different environments and voice content, and enhance product usage performance.
Smart Images

Figure CN120544552A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart wearable technology, and in particular to a method, device and equipment for target voice control based on smart glasses. Background Art
[0002] Smart glasses, as a new type of wearable device that integrates optical, electronic and artificial intelligence technologies, are widely used in scenarios such as conversations and lecture interpretation. Users can use smart glasses to achieve voice interaction and cross-border communication.
[0003] Currently, smart glasses' voice processing technology mostly uses general-purpose speech recognition and enhancement algorithms. They collect ambient audio through an array of receivers, use beamforming technology to capture the sound source, and then combine it with acoustic models to achieve speech recognition and noise reduction.
[0004] However, different users have different preferences and sensitivities for sound. For example, music lovers may prefer to enhance low-frequency sound effects, while hearing-impaired users may prefer to emphasize the human voice frequency band. Some users prefer warm and mellow timbre, while others prefer clear and transparent sound quality. However, existing technologies generally focus on standardizing the processing of voice signals, making it difficult to optimize and adjust sound based on the listening habits and scenario requirements of different users. This makes it difficult for users to obtain a listening experience that suits their needs, greatly limiting the personalized capabilities of smart glasses in voice interaction scenarios. Summary of the Invention
[0005] The embodiments of the present application provide a target voice control method, device and equipment based on smart glasses, which are used to solve the following technical problems: the existing technology usually focuses on the standardized processing of voice signals, and it is difficult to perform differentiated sound optimization and adjustment according to the auditory habits and scene requirements of different users, resulting in it being difficult for users to obtain an auditory experience that meets their own needs, which greatly limits the personalization capabilities of smart glasses in voice interaction scenarios.
[0006] The embodiments of this application adopt the following technical solutions: The present application provides a method for target voice control based on smart glasses. The method comprises the following steps: obtaining audio information of a speaker in a current scene and dividing the audio information into multiple audio segments based on audio features; obtaining portrait information of a person in the current scene and associating the portrait information with corresponding audio segments to determine a target speaker; when the noise level of the current scene reaches a preset noise level, determining the user's sound preference information corresponding to the current scene based on a mapping relationship between the noise level and smart glasses usage preference data; analyzing the target speaker's voice and responding to a dynamic sound adjustment strategy based on the sound features obtained after the analysis to optimize the target speaker's voice; and adjusting the frequency band of the optimized sound through an amplifier provided on the smart glasses based on the sound preference information to achieve target voice control.
[0007] In one implementation of the present application, portrait information in the current scene is obtained, and the portrait information is associated with the corresponding audio clip to determine the target speaker, specifically including: when it is determined to be a two-person scene, directional sound reception is started, and the sound reception direction is determined based on the direction corresponding to the portrait information and the sound source distance corresponding to the audio clip, and the audio clip is associated with the portrait information to determine the target speaker; when it is determined to be a multi-person scene, omnidirectional sound reception is started, and based on the portrait information of the person whose lip shape changes, the time when the sound appears, the direction corresponding to each portrait information, and the sound source distance corresponding to each audio clip, each portrait information is associated with each audio clip, and the person whose lip shape changes is taken as the target speaker.
[0008] In one implementation of the present application, when it is determined to be a multi-person scene, the method also includes: responding to a preset scheme startup request, matching the audio clips collected in the current multi-person scene with preset data in a database; if there is preset data matching the audio clip in the database, determining the portrait information associated with the audio clip in the database to determine the target speaker based on the portrait information; if there is no preset data matching the audio clip in the database, collecting the portrait information in the current scene and the audio information corresponding to each portrait information, and establishing a correspondence between the portrait information and the audio information, and storing the portrait information and the audio information in the database based on the correspondence.
[0009] In one implementation of the present application, before determining the sound preference information corresponding to the user in the current scenario based on the mapping relationship between the noise level and the smart glasses usage preference data, the method also includes: classifying the historical usage data corresponding to the smart glasses based on different sound scenarios; wherein the historical usage data at least includes historical noise level data, historical volume adjustment data and historical sound quality equalization parameter adjustment data; annotating the classified historical usage data with noise level labels; based on different sound scenarios, extracting sound adjustment features for the historical usage data corresponding to different noise level labels; obtaining user sound preference data based on the sound adjustment features and the collaborative filtering algorithm; mapping different sound scenarios, different noise levels and sound preference data to construct a smart glasses usage preference data table.
[0010] In one implementation of the present application, the voice of the target speaker is analyzed, specifically including: determining the sound wave frequency distribution characteristics and waveform characteristics based on the audio clip corresponding to the target speaker, and obtaining the timbre characteristics based on the sound wave frequency distribution characteristics and waveform characteristics; obtaining the output vocabulary and syllable change corresponding to the target speaker within a preset time period, comparing the output vocabulary and syllable change with preset variable thresholds respectively, and obtaining the speaking speed characteristics based on the comparison results; determining the slope between adjacent frames of the fundamental frequency curve based on the audio clip corresponding to the target speaker, and determining the peak and trough data corresponding to the fundamental frequency curve, and obtaining the intonation characteristics based on the slope and the peak and trough data.
[0011] In one implementation of the present application, a dynamic sound adjustment strategy is responded to based on the sound features obtained after analysis to optimize the voice of the target speaker, specifically including: when there is a deviation between the timbre features and the target timbre features, the gain of different frequency bands of the audio clip is adjusted according to the deviation; when there is a deviation between the speaking rate features and the target speaking rate features, the audio clip is denoised and the syllables of the denoised audio clip are extended through a preset sound signal processing algorithm; when there is a deviation between the intonation features and the target intonation features, the fundamental frequency change trajectory is extracted based on the audio clip, the fundamental frequency change trajectory is compared with the target intonation template through dynamic time warping, and the segments whose comparison errors do not meet the preset conditions are adjusted.
[0012] In one implementation of the present application, an audio clip is denoised by a preset sound signal processing algorithm, and the syllables of the denoised audio clip are extended, specifically including: denoising and reconstructing the audio clip by wavelet transform, and performing signal-to-noise ratio enhancement on the audio clip by a spectral subtraction algorithm in speech enhancement; performing semantic analysis and sentiment analysis on the speech recognition text corresponding to the processed audio clip, and determining key semantics in the speech recognition text based on the analysis results; determining a syllable extension coefficient based on the relationship between the deviation and the preset ratio, and syllable extension of the key semantics based on the extension coefficient.
[0013] In one implementation of the present application, based on the sound preference information, the frequency band of the optimized sound is adjusted by the power amplifier set on the smart glasses to achieve target voice control, specifically including: dividing the horizontal space into multiple sector-shaped areas with the smart glasses as the center of the circle; wherein each sector-shaped area corresponds to a set of power amplifier parameters; matching the sound source position corresponding to the target speaker with the multiple sector-shaped areas, determining the sector-shaped area to be adjusted, and determining the initial power amplifier configuration corresponding to the sector-shaped area to be adjusted; based on the sound preference information, adjusting the initial power amplifier configuration to perform frequency band enhancement on the target speaker's voice, and to perform frequency band attenuation on the non-target speaker's voice.
[0014] An embodiment of the present application provides a target voice control device based on smart glasses, comprising: a division unit, which obtains audio information of a speaker in a current scene and divides the audio information into multiple audio segments according to audio features; an association unit, which obtains portrait information of a person in the current scene and associates the portrait information with corresponding audio segments to determine a target speaker; a parsing unit, which determines, when the noise level of the current scene reaches a preset noise level, the user's corresponding sound preference information in the current scene based on a mapping relationship between the noise level and smart glasses usage preference data; an optimization unit, which analyzes the voice of the target speaker and responds to a dynamic sound adjustment strategy based on the sound features obtained after the analysis to optimize the voice of the target speaker; and an adjustment unit, which adjusts the frequency band of the optimized sound through an amplifier provided on the smart glasses based on the sound preference information to achieve target voice control.
[0015] An embodiment of the present application provides a target voice control device based on smart glasses, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to: obtain audio information of a speaker in a current scene, and divide the audio information into multiple audio segments according to audio features; obtain portrait information of a person in the current scene, and associate the portrait information with corresponding audio segments to determine a target speaker; when the noise level of the current scene reaches a preset noise level, determine the user's corresponding sound preference information in the current scene based on a mapping relationship between the noise level and smart glasses usage preference data; analyze the voice of the target speaker, and respond to a dynamic sound adjustment strategy based on the sound features obtained after the analysis to optimize the voice of the target speaker; based on the sound preference information, adjust the frequency band of the optimized sound through an amplifier provided on the smart glasses to achieve target voice control.
[0016] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: by associating audio clips with portrait information to determine the target speaker, the target speaker's voice can be accurately captured to avoid interference from other sounds, thereby improving voice recognition accuracy and interaction efficiency. Secondly, the target speaker's voice is analyzed and a dynamic adjustment strategy is responded to, and personalized sound quality optimization is performed for different target speakers to ensure clear and natural sound in different environments or when the voice content changes. The embodiments of the present application also determine the sound preference information in different scenarios and under different noises based on the user's historical usage data, and perform frequency band adjustment based on the sound preference information to achieve target voice control, so that the smart glasses can provide more intelligent and humane services. Users can obtain an ideal voice and hearing experience in various scenarios without having to manually adjust the settings frequently, thereby enhancing product performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments described in the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings: Figure 1 A schematic diagram of the architecture of target voice control based on smart glasses provided in an embodiment of the present application; Figure 2 A flow chart of a target voice control method based on smart glasses provided in an embodiment of the present application; Figure 3A flow chart of a method for determining a target speaker in different scenarios provided by an embodiment of the present application; Figure 4 A flow chart of a method for determining sound preference information provided in an embodiment of the present application; Figure 5 A flow chart of a method for dynamic sound adjustment strategy provided in an embodiment of the present application; Figure 6 A schematic diagram of a target voice control device based on smart glasses provided in an embodiment of the present application; Figure 7 A schematic structural diagram of a target voice control device based on smart glasses provided in an embodiment of the present application.
[0018] Description of reference numerals: 101: Smart glasses, 102: Database, 103: Glasses applications, 104: Communication network, 105: Cloud. DETAILED DESCRIPTION
[0019] The embodiments of the present application provide a method, apparatus, and device for target voice control based on smart glasses.
[0020] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0021] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0022] The target voice control method based on smart glasses in the embodiment of the present application is mainly used in scenarios such as two-person and multi-person conversations and lecture interpretation. In a two-person conversation scenario, both parties wear glasses. When the surrounding environment is noisy, the directional microphone on the glasses end is activated through scene judgment and algorithmic voiceprint recognition, and only the voice of the user wearing glasses is focused and collected, ensuring that the conversation between the two parties is not disturbed by the noise of the environment. In a multi-person scenario, only the listener wears glasses. When facing a multi-person communication and entering a interpretation scenario, the device uses an omnidirectional microphone to collect all surrounding sounds, identifies the content of different people's speech through algorithmic voiceprints, and accurately identifies and associates the audio clip of the target speaker with the portrait information. The embodiment of the present application also performs personalized optimization of the target voice according to the environmental noise status and the user's personalized needs, while dynamically enhancing the target voice and suppressing interference, and using spatially distributed power amplifiers to achieve efficient and directional sound output, thereby providing users with a clear and natural voice interaction experience in a complex environment.
[0023] Figure 1 A schematic diagram of the architecture of target voice control based on smart glasses provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the target speech control method based on smart glasses 101 can be applied to Figure 1 In the illustrated environment, the application environment may include smart glasses 101, a database 102, a glasses application 103, a communication network 104, and a cloud 105. The smart glasses 101, worn by the user, acquire audio data and portrait data from the current scene. The glasses application 103 analyzes the acquired audio data and portrait data and selects the target speaker. The target speaker's audio information is sent to the cloud 105 via the communication network 104. The cloud 105 analyzes the audio characteristics of the audio information to implement a personalized optimization strategy for the current speaker's voice. The cloud 105 also determines a gain adjustment strategy for the power amplifier based on the current user's voice preferences. The personalized sound optimization strategy and gain adjustment strategy are then sent to the smart glasses 101 and glasses application 103 via the communication network 104. The sound is adjusted by the power amplifier and glasses application 103 on the smart glasses 101. The glasses application 103 displays the target speaker's speech on the smart glasses 101 and translates the speech according to the user's needs.
[0024] It should be noted that the smart glasses in the embodiments of this application integrate a high-precision camera, an IMU sensor, an edge computing chip, a directional microphone and an omnidirectional microphone that can be routed according to the scene, a directional amplifier and an omnidirectional amplifier that can be used according to the scene, a firmware-based voiceprint recognition algorithm, and a directional radio algorithm. The cloud server in the embodiments of this application is used to store voiceprint data, portrait data, and lip shape data.
[0025] Figure 2A flow chart of a target voice control method based on smart glasses provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the target voice control method based on smart glasses includes the following steps: S201: Acquire audio information of a speaker in a current scene, and divide the audio information into multiple audio segments according to audio features.
[0026] In one implementation of the present application, the audio receiving device on the smart glasses receives the audio of the speaker and identifies the timbre, voiceprint information and distance of the sound source; based on the timbre, voiceprint information and distance of the sound source, the audio is segmented to obtain various audio segments.
[0027] Specifically, the smart glasses' audio receivers collect ambient sound, converting the acoustic signals into electrical signals, which are then converted to audio through analog-to-digital conversion. The collected audio is analyzed to construct a voiceprint feature vector for each speaker. Using the smart glasses' array of receivers, the distance between the sound source and the smart glasses is determined based on factors such as the time difference and intensity difference between the sound reaching different receivers. Based on the similarity of timbre and voiceprint features, as well as changes in the distance to the sound source, the demarcation points between different speakers or different speech segments of the same speaker are determined, and the audio is segmented into multiple segments, each corresponding to the speech content of a single speaker.
[0028] S202: Acquire portrait information of a person in the current scene, and associate the portrait information with a corresponding audio clip to determine a target speaker.
[0029] In one implementation of the present application, if a two-person scene is determined, directional sound reception is initiated. The sound reception direction is determined based on the direction corresponding to the portrait information and the distance from the sound source corresponding to the audio clip. The audio clip is then associated with the portrait information to determine the target speaker. If a multi-person scene is determined, omnidirectional sound reception is initiated. Based on the portrait information of the person experiencing lip changes, the time of sound onset, the direction corresponding to each portrait information, and the distance from the sound source corresponding to each audio clip, each portrait information is associated with each audio clip. The person experiencing lip changes is then identified as the target speaker.
[0030] The method for determining the target speaker is as follows: Figure 3 As shown, Figure 3 A flow chart of a method for determining a target speaker in different scenarios provided by an embodiment of the present application, comprising Figure 3 It can be seen that the embodiment of the present application first determines the number of people in the current application scenario, and responds to different target speaker determination methods based on different number of people.
[0031] Specifically, when determining that the current scene is a two-person conversation, the smart glasses worn by the user activate the directional microphone to obtain audio data and portrait information in the current scene. The firmware-side voiceprint recognition algorithm and directional sound reception algorithm in the smart glasses begin to operate. The sound reception direction is determined based on the orientation of the face in the image, or based on the orientation of the face in the image and the distance to the sound source in the audio clip. The acquired audio data and portrait data are analyzed to determine the target speaker. If the user does not need translation, the target speaker's speech content is displayed on the smart glasses. If the user does need translation, the target speaker's speech content is displayed on the smart glasses and translated.
[0032] Specifically, in a multi-person conversation scenario, without implementing a pre-set data solution, the smart glasses worn by the user activate an omnidirectional microphone to capture audio data and portrait data in the current scene. The captured audio and portrait data are analyzed, and the audio clips are associated with faces based on factors such as the face's orientation and lip shape changes, as well as the distance from the sound source in the audio clip and the time of the voice's appearance. When the lip shape change time matches the time of the voice's appearance, and the face's orientation matches the distance from the sound source, the audio clip is associated with the face.
[0033] For example, the time interval from "closed mouth" to "O shape" is t1-t2, which corresponds to the time interval t1'-t2' of the "oh" sound in the audio. If the time difference |(t2-t1)-(t2'-t1')| is less than 80ms, the match is considered. The image of the person whose mouth shape changes is used as the target speaker, and the target speaker is selected.
[0034] The portrait data and audio data are sent to the cloud. The cloud matches the received audio data with the lip data and sends the matching results to the glasses application. The smart glasses lock the target speaker and display and translate the target speaker's speech content according to user needs.
[0035] In one implementation of the present application, in response to a request to initiate a preset scenario, an audio clip collected in a current multi-person scene is matched with preset data in a database. If the database contains preset data matching the audio clip, portrait information associated with the audio clip is determined in the database, and the target speaker is identified based on the portrait information. If the database does not contain data matching the audio clip, the portrait information in the current scene and the audio information corresponding to each portrait information are collected, and a correspondence between the portrait information and the audio information is established. Based on the correspondence, the portrait information and the audio information are stored in the database.
[0036] Specifically, in a multi-person conversation scenario, if there is a preset data plan, the collected audio data will be sent to the cloud. The cloud will match the received audio data with the pre-stored portraits in the database and send the matching results to the glasses application. The smart glasses will lock the target speaker and display and translate the target speaker's speech content according to user needs.
[0037] If there is no pre-set data solution, the audio data and portrait data collected in the current scene are sent to the cloud for matching and storage. When the cloud receives the audio information later, it matches it with the stored information and sends the matching results to the glasses application. The smart glasses then display and translate the target speaker's speech according to the user's needs.
[0038] In the embodiment of the present application, directional sound reception is activated in a two-person scenario, and the sound reception direction and the target speaker are determined based on the portrait direction and the distance from the sound source, thereby reducing interference and improving efficiency; omnidirectional sound reception is used in a multi-person scenario, integrating information such as the portrait direction, lip shape changes, the distance from the sound source, and the time when the sound appears to achieve accurate matching. The two modes are switched on demand, saving hardware resources and reducing the error rate through the complementarity of multimodal information. Secondly, by matching audio clips with pre-set data in the database, historical data can be quickly used to lock the target speaker, reducing the amount of real-time calculations, and quickly locating frequent speakers in repetitive scenarios to improve recognition efficiency; on the other hand, for newly appearing audio clips, portrait and audio information are collected in a timely manner and a corresponding relationship is established for storage, realizing dynamic expansion of the database and enabling it to have self-learning capabilities. As the usage time increases, it can adapt to more complex scenarios and personnel combinations, and continuously optimize the recognition accuracy of the target speaker.
[0039] S203: When the noise level of the current scene reaches a preset noise level, based on a mapping relationship between the noise level and the smart glasses usage preference data, determine the user's sound preference information corresponding to the current scene.
[0040] In one implementation of the present application, different users have different preferences for different types of sounds. For example, some users may be more sensitive to high-pitched sounds and less sensitive to low-pitched sounds. Smart glasses can perform targeted frequency enhancement on the target speaker's voice based on the user's auditory preferences for sounds of different frequencies, while attenuating the frequency band of the non-target speaker's voice to enhance the user's auditory experience. Figure 4 A flow chart of a method for determining sound preference information provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the method for determining sound preference information includes the following steps S401-S405: S401: Classify historical usage data corresponding to the smart glasses based on different sound scenes.
[0041] In the embodiments of this application, historical usage data includes the user's corresponding historical volume adjustment data and historical sound quality equalization parameter adjustment data. These data record the user's past sound adjustment operations in different situations. The scenarios in the embodiments of this application can be: a small meeting room, a noisy cafe, a bustling street, etc. By associating historical usage data with specific usage scenarios, it is possible to analyze user sound preferences based on different scenarios.
[0042] S402: Label the classified historical usage data with noise level labels.
[0043] After completing the scene classification, the noise level label is annotated for each scene's historical data. The labeling system can be divided into categories such as "low noise," "medium noise," and "high noise" according to noise intensity. For example, below 40 decibels is low noise, and 40-70 decibels is medium noise.
[0044] S403 : Based on different sound scenes, extract sound adjustment features from historical usage data corresponding to different noise level labels.
[0045] In different sound scenarios, based on different noise levels, users' sound adjustment methods and habits will vary. For example, in a noisy cafe, if the noise level is high, users may turn up the volume to combat the ambient noise. Based on the combination of scenario and noise level, representative sound adjustment features are extracted from historical volume adjustment data and sound quality equalization parameter adjustment data. These sound adjustment features include the volume change amplitude, the specific numerical range of sound quality parameter adjustments, and the order in which different parameters are adjusted.
[0046] S404: Obtain user voice preference data based on voice adjustment features and collaborative filtering algorithm.
[0047] The collaborative filtering algorithm used in the embodiments of this application is based on the principle of predicting the behavior of a target user using the behavioral data of similar users. The extracted voice adjustment features are used as user behavior data. By calculating the feature similarity of different users in similar sound scenes and similar noise levels, other users with similar voice adjustment habits as the target user are found. The voice preference information of these similar users in different sound scenes and noise levels is then used, combined with the target user's own historical characteristics, to predict the target user's voice preferences in different sound scenes and noise levels.
[0048] S405: Map different sound scenes, different noise levels, and sound preference data to construct a smart glasses usage preference data table.
[0049] Different sound scenes and noise level labels are mapped to user sound preference data in three dimensions to form a structured preference data table. This table can store the corresponding relationship between "scene-noise-preference parameter", such as "outdoor scene-high noise-volume 80% + vocal frequency band gain 5dB".
[0050] Furthermore, after the smart glasses usage preference data table is constructed, when the current scene and noise level are identified, the matching preference parameters in the smart glasses usage preference data table can be quickly called to automatically adjust the volume and sound quality.
[0051] S204: Analyze the target speaker's voice, and respond to a dynamic voice adjustment strategy based on the voice characteristics obtained after the analysis to optimize the target speaker's voice.
[0052] In one implementation of the present application, based on the audio clip corresponding to the target speaker, the sound wave frequency distribution characteristics and waveform characteristics are determined, and the timbre characteristics are obtained based on the sound wave frequency distribution characteristics and waveform characteristics. The output vocabulary and syllable changes corresponding to the target speaker within a preset time period are obtained, and the output vocabulary and syllable changes are respectively compared with preset variable thresholds, and the speaking rate characteristics are obtained based on the comparison results. Based on the audio clip corresponding to the target speaker, the slope between adjacent frames of the fundamental frequency curve is determined, as well as the peak and trough data corresponding to the fundamental frequency curve. Based on the slope and peak and trough data, the intonation characteristics are obtained.
[0053] Specifically, the target speaker's voice is subjected to feature analysis, including timbre features, speaking speed features, and pitch features.
[0054] When analyzing and extracting timbre features, the sound signal is digitally analyzed based on the audio clip corresponding to the target speaker. The frequency distribution of sound waves refers to the energy distribution of different frequency components in the sound. Different sound-producing bodies produce different frequency combinations, such as male voices with more low-frequency components and female voices with more prominent high-frequency components. Waveform features reflect the temporal changes of the sound signal. Using spectrum analysis signal processing technology, the specific parameters of the sound wave frequency distribution and waveform are parsed from the audio clip to obtain timbre feature data that characterizes the target speaker.
[0055] When analyzing and extracting speech rate features, the target speaker's speech content is obtained within a preset time period, such as the past 30 seconds. Using speech recognition technology, the audio is converted into text, and the output vocabulary and syllable variation are then calculated. Syllable variation refers to the increase or decrease in the number of syllables. These two data points are compared with preset variable thresholds. If the output vocabulary is greater than the threshold and the syllable variation is stable, the speaking rate is fast; if the vocabulary is low and the syllable variation is small, the speaking rate is slow. Through quantitative comparison, the corresponding speaking rate features of the target speaker are obtained.
[0056] When analyzing and extracting intonation features, fundamental frequency (FFM) analysis is performed on audio clips of the target speaker to generate a FFM curve that changes over time. The slope of the FFM curve between adjacent frames is calculated. The slope reflects the speed of pitch change, with a larger slope indicating a more dramatic pitch change. The corresponding peak and trough data of the FFM curve are also determined, with peaks representing the highest points in pitch and troughs representing the lowest points. By combining the slope of the FFM curve between adjacent frames and the peak and trough data, the target speaker's intonation is determined to be flat or fluctuating, and whether there are specific patterns of rising and falling intonation, thereby generating intonation features.
[0057] In one implementation of the present application, the sound enhancement strategy is dynamically adjusted based on the timbre characteristics, speech speed characteristics, and pitch characteristics. For example, for a target speaker with a thick timbre, the gain of the mid- and low-frequency bands is appropriately increased to make the sound fuller and more powerful; for a target speaker with a crisp timbre, while maintaining the clarity of the high frequencies, the mid-frequency bands are appropriately enhanced to enhance the thickness of the sound; for a target speaker with a faster speech speed, the sound coherence is optimized and enhanced, the characteristics of the speech syllables are highlighted to ensure that each syllable is clearly distinguishable, and the syllable time can be extended to ensure the integrity and comprehension of the information, so that the listener can better identify each syllable, reduce auditory fatigue, and improve the comprehension efficiency during long-term listening; for a target speaker with a slower speech speed, due to the characteristics of the voice itself or factors such as age, the high-frequency components are insufficient, and the high-frequency part of the sound signal can be increased to make the sound brighter and clearer, increase the penetration of the sound, and help capture the details of the voice.
[0058] Figure 5 This is a flow chart of a method for dynamic sound adjustment strategy provided in an embodiment of the present application. The method for dynamic sound adjustment strategy includes the following steps S501-S503: S501 : When there is a deviation between the timbre feature and the target timbre feature, perform gain adjustment on different frequency bands of the audio segment according to the deviation.
[0059] If the actual timbre characteristics extracted from an audio clip differ from the pre-set target timbre characteristics, the timbre of the current audio does not meet expectations. Based on the difference between the two, the gain of different frequency bands in the audio clip is adjusted. Because sound is composed of multiple frequency bands, the energy distribution of each band determines the timbre characteristics.
[0060] For example, if the target timbre requires a fuller low-frequency effect, but the current audio lacks low-frequency, the gain of the low-frequency band of the current audio will be specifically increased to enhance the sound energy of this frequency band; if the high-frequency part is too sharp, the gain of the high-frequency band will be reduced to make the timbre softer.
[0061] S502: When there is a deviation between the speech rate feature and the target speech rate feature, denoise the audio segment using a preset sound signal processing algorithm, and extend the syllables of the denoised audio segment.
[0062] In one implementation of the present application, an audio clip is de-noised and reconstructed using a wavelet transform, and the signal-to-noise ratio of the audio clip is enhanced using a spectral subtraction algorithm used in speech enhancement. Semantic and sentiment analysis is performed on the speech recognition text corresponding to the processed audio clip, and key semantics are identified in the speech recognition text based on the analysis results. A syllable lengthening coefficient is determined based on the relationship between the deviation and a preset ratio, and syllable lengthening is performed on the key semantics based on the lengthening coefficient.
[0063] Specifically, noise in audio often has specific frequency and time distribution characteristics that are different from those of real speech signals. The wavelet transform can decompose audio signals into multiple subbands, in which the noise and speech components are expressed differently. By analyzing the characteristics of the signals in each subband, the subbands containing noise are processed to suppress or remove the noise components. The processed subbands are then reconstructed to obtain a noise-free audio signal, making the speech clearer and purer. Secondly, the audio clips are framed and the spectrum of each frame is calculated. The spectrum of each frame is subtracted from the noise spectrum. The enhanced spectrum is converted back to a time domain signal through an inverse Fourier transform, thereby improving the signal-to-noise ratio of the audio.
[0064] Secondly, speech recognition is performed on the processed audio clip and converted into text. Semantic analysis and sentiment analysis are performed on the speech recognition text. Based on the results of semantic analysis and sentiment analysis, key semantics are determined in the speech recognition text. According to the deviations detected during the audio processing, such as speech speed deviation, speech clarity deviation, etc., and the preset proportional relationship, the syllable extension coefficient is determined. The preset proportional relationship in the embodiment of the present application reflects the correspondence between the degree of deviation and the required degree of syllable extension. For example, the greater the deviation, the greater the corresponding syllable extension coefficient. Based on the determined extension coefficient, the syllable extension operation is performed on the key semantic part. By adjusting the duration of each syllable in the key semantics, important information is highlighted, making it easier for the audience to capture the core content. At the same time, according to the results of sentiment analysis, the expression rhythm of the speech can be adjusted to enhance the effect of emotional transmission.
[0065] S503. When there is a deviation between the intonation feature and the target intonation feature, extract the fundamental frequency variation trajectory based on the audio segment, compare the fundamental frequency variation trajectory with the target intonation template through dynamic time warping, and adjust the segments where the comparison error does not meet the preset conditions.
[0066] If the intonation characteristics of an audio clip deviate from the target intonation characteristics, the fundamental frequency variation trajectory of the audio clip is extracted. This fundamental frequency variation trajectory reflects the fluctuations in pitch during speech. Next, the dynamic time warping algorithm is used to compare the extracted fundamental frequency variation trajectory with the target intonation template. The dynamic time warping algorithm flexibly aligns the two sequences on the timeline, addressing time differences caused by varying speaking rates.
[0067] Furthermore, during the comparison process, if the comparison error in certain segments exceeds a preset limit, indicating that the intonation of these segments does not match the target, targeted adjustments are made to these problematic segments. For example, by changing the slope of the fundamental frequency curve, adjusting the position and amplitude of peaks and troughs, etc., intonation deviations are corrected to make the audio intonation more closely match the target intonation template.
[0068] The embodiments of the present application can distinguish different speakers by analyzing timbre features for identity recognition or voice personalization services; by comparing speech rate features, it can grasp the speaking rhythm and assist in optimizing the fluency of voice interaction; based on intonation features, it can identify the tone of voice and emotions, and enhance the emotional perception of human-computer interaction. Adjusting the frequency band gain based on timbre deviation can flexibly change the sound texture and adapt to personalized needs; first removing noise and then adjusting the syllable duration based on speech rate deviation can not only ensure audio quality, but also achieve natural speech rate changes to avoid mechanical feeling; by comparing the intonation fundamental frequency trajectory through dynamic time regularization, the intonation fluctuations can be accurately corrected to simulate voice expressions of different emotions or styles. The three work together to fully optimize voice features.
[0069] S205. Based on the sound preference information, the frequency band of the optimized sound is adjusted by the power amplifier provided on the smart glasses to achieve target voice control.
[0070] In one implementation of this application, the horizontal space is divided into multiple sector-shaped areas with the smart glasses as the center; each sector-shaped area corresponds to a set of amplifier parameters. The sound source location corresponding to the target speaker is matched with the multiple sector-shaped areas to determine the sector-shaped area to be adjusted, and the initial amplifier configuration corresponding to the sector-shaped area to be adjusted is determined. Based on the sound preference information, the initial amplifier configuration is adjusted to enhance the frequency band of the target speaker's voice and attenuate the frequency band of the non-target speakers' voices.
[0071] Specifically, the horizontal space, centered on the wearer of the smart glasses, is evenly divided into multiple sector-shaped areas, each corresponding to a set of amplifier parameters. These parameters, including volume gain, frequency equalization, and phase adjustment, are used to control the amplification and playback quality of sound signals in different frequency bands. For example, the horizontal 360° space can be divided into eight 45° sectors, each with a pre-set parameter combination suitable for sound playback in that direction, allowing sounds from different directions to be processed differently.
[0072] Furthermore, the smart glasses' built-in sound array and sound localization algorithm detect and locate the target speaker's voice in real time. This sound source location is matched against the multiple sector-shaped areas, and the sector corresponding to the target speaker's voice is determined. This sector is then used as the sector to be adjusted. Simultaneously, the pre-set initial amplifier configuration for the sector to be adjusted is obtained, serving as the basis for subsequent sound adjustments.
[0073] Furthermore, for the sector-shaped area to be adjusted where the target speaker is located, the energy of their voice in specific frequency bands is enhanced based on the voice preference information. For example, if the user prefers a clear human voice, the mid-frequency portion is enhanced to make the target speaker's voice more prominent and clear. For the sector-shaped area where non-target speakers are located, their voices are subjected to frequency band attenuation processing, reducing the volume of sounds in these directions or suppressing certain frequency bands to reduce interference from ambient noise and irrelevant sounds. Through differentiated power amplifier adjustment, the embodiments of the present application create a personalized listening experience for users that focuses on the target sound and weakens background interference.
[0074] Figure 6 This is a schematic diagram of a target voice control device based on smart glasses provided in an embodiment of the present application. Figure 6 As shown, the target voice control device based on smart glasses includes: A segmentation unit, which obtains audio information of a speaker in a current scene and divides the audio information into multiple audio segments according to audio features; an associating unit, which obtains portrait information of a person in the current scene and associates the portrait information with the corresponding audio clip to determine a target speaker; an analysis unit, which determines, when the noise level of the current scene reaches a preset noise level, the user's sound preference information corresponding to the current scene based on a mapping relationship between the noise level and the smart glasses usage preference data; an optimization unit that analyzes the target speaker's voice and responds to a dynamic voice adjustment strategy based on the voice characteristics obtained after the analysis to optimize the target speaker's voice; The adjustment unit adjusts the frequency band of the optimized sound through the power amplifier provided on the smart glasses based on the sound preference information to achieve target voice control.
[0075] In one implementation of the present application, obtaining portrait information of a person in the current scene and associating the portrait information with the corresponding audio clip to determine the target speaker specifically includes: If it is determined that the scene is a two-person scene, directional sound reception is started, and the sound reception direction is determined based on the direction corresponding to the portrait information and the distance from the sound source corresponding to the audio clip, and the audio clip is associated with the portrait information to determine the target speaker; When it is determined to be a multi-person scene, omnidirectional sound reception is started. Based on the portrait information of the person whose lip shape changes, the time when the sound appears, the direction corresponding to each portrait information, and the sound source distance corresponding to each audio segment, each portrait information is associated with each audio segment, and the person whose lip shape changes is identified as the target speaker.
[0076] In one implementation of the present application, when it is determined to be a multi-person scenario, the following is further included: Respond to the preset scheme start request and match the audio clips collected in the current multi-person scene with the preset data in the database; If there is preset data matching the audio clip in the database, determining portrait information associated with the audio clip in the database to determine the target speaker based on the portrait information; If there is no preset data matching the audio clip in the database, the portrait information in the current scene and the audio information corresponding to each portrait information are collected, and a correspondence between the portrait information and the audio information is established, and the portrait information and the audio information are stored in the database based on the correspondence.
[0077] In one implementation of the present application, before determining the sound preference information corresponding to the user in the current scene based on the mapping relationship between the noise level and the smart glasses usage preference data, the method further includes: Classifying historical usage data corresponding to the smart glasses based on different sound scenarios; wherein the historical usage data includes at least historical noise level data, historical volume adjustment data, and historical sound quality equalization parameter adjustment data; labeling the classified historical usage data with noise level labels; Based on different sound scenes, extracting sound adjustment features from the historical usage data corresponding to different noise level labels; Obtaining user sound preference data based on the sound adjustment features and the collaborative filtering algorithm; Different sound scenes, different noise levels and the sound preference data are mapped to construct a smart glasses usage preference data table.
[0078] In one implementation of the present application, analyzing the target speaker's voice specifically includes: Determining, based on the audio segment corresponding to the target speaker, a sound wave frequency distribution feature and a waveform feature, and obtaining a timbre feature based on the sound wave frequency distribution feature and the waveform feature; Obtaining the output vocabulary and syllable variation of the target speaker within a preset time period, comparing the output vocabulary and syllable variation with preset variable thresholds, and obtaining a speech rate feature based on the comparison results; Based on the audio segment corresponding to the target speaker, the slope of the fundamental frequency curve between adjacent frames is determined, as well as the peak and trough data corresponding to the fundamental frequency curve. The intonation feature is obtained based on the slope and the peak and trough data.
[0079] In one implementation of the present application, a dynamic sound adjustment strategy is implemented based on the sound characteristics obtained after analysis to optimize the target speaker's voice, specifically including: In the event that there is a deviation between the timbre characteristic and the target timbre characteristic, performing gain adjustment on different frequency bands of the audio segment according to the deviation; In the event that there is a deviation between the speech rate feature and the target speech rate feature, denoising the audio segment using a preset sound signal processing algorithm, and performing syllable extension on the denoised audio segment; In the case where there is a deviation between the intonation feature and the target intonation feature, the fundamental frequency change trajectory is extracted based on the audio segment, the fundamental frequency change trajectory is compared with the target intonation template through dynamic time warping, and the segments where the comparison error does not meet the preset conditions are adjusted.
[0080] In one implementation of the present application, the audio segment is denoised using a preset sound signal processing algorithm, and the syllables of the denoised audio segment are extended, specifically including: Performing denoising and reconstruction processing on the audio segment by wavelet transform, and performing signal-to-noise ratio enhancement processing on the audio segment by spectral subtraction algorithm in speech enhancement; Performing semantic analysis and sentiment analysis on the speech recognition text corresponding to the processed audio clip, and determining key semantics in the speech recognition text based on the analysis results; A syllable extension coefficient is determined according to the relationship between the deviation and the preset ratio, and the syllable of the key semantics is extended based on the extension coefficient.
[0081] In one implementation of the present application, based on the sound preference information, the frequency band of the optimized sound is adjusted by the power amplifier provided on the smart glasses to achieve target voice control, specifically including: Taking the smart glasses as the center, the horizontal space is divided into a plurality of sector-shaped areas; wherein each sector-shaped area corresponds to a set of power amplifier parameters; Matching the sound source position corresponding to the target speaker with the plurality of sector-shaped areas to determine the sector-shaped area to be adjusted, and determining the initial power amplifier configuration corresponding to the sector-shaped area to be adjusted; Based on the voice preference information, the initial power amplifier configuration is adjusted to enhance the frequency band of the target speaker's voice and to attenuate the frequency band of the non-target speaker's voice.
[0082] Figure 7 This is a schematic diagram of the structure of a target voice control device based on smart glasses provided in an embodiment of the present application. Figure 7 As shown, a target voice control device based on smart glasses includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any of the above-mentioned target voice control methods based on smart glasses.
[0083] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0084] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. It will be apparent to those skilled in the art that various modifications and variations may be made to the embodiments of the present application. However, such modifications or substitutions do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A target voice control method based on smart glasses, characterized in that: The method comprises: Acquire audio information of a speaker in a current scene, and divide the audio information into multiple audio segments according to audio features; Acquire portrait information of a person in the current scene, and associate the portrait information with the corresponding audio clip to determine a target speaker; When the noise level of the current scene reaches a preset noise level, determining the user's sound preference information corresponding to the current scene based on a mapping relationship between the noise level and the smart glasses usage preference data; Analyzing the target speaker's voice, and responding to a dynamic voice adjustment strategy based on the voice characteristics obtained after the analysis to optimize the target speaker's voice; Based on the sound preference information, the frequency band of the optimized sound is adjusted by the power amplifier set on the smart glasses to achieve target voice control.
2. The target speech control method based on smart glasses according to claim 1, characterized in that: The acquiring of portrait information in the current scene and associating the portrait information with the corresponding audio segment to determine the target speaker specifically includes: If it is determined that the scene is a two-person scene, directional sound reception is started, and the sound reception direction is determined based on the direction corresponding to the portrait information and the distance from the sound source corresponding to the audio clip, and the audio clip is associated with the portrait information to determine the target speaker; When it is determined to be a multi-person scene, omnidirectional sound reception is started. Based on the portrait information of the person whose lip shape changes, the time when the sound appears, the direction corresponding to each portrait information, and the sound source distance corresponding to each audio segment, each portrait information is associated with each audio segment, and the person whose lip shape changes is identified as the target speaker.
3. The target speech control method based on smart glasses according to claim 2, characterized in that: When it is determined that the scene is a multi-person scene, the method further includes: Respond to the preset scheme start request and match the audio clips collected in the current multi-person scene with the preset data in the database; If there is preset data matching the audio clip in the database, determining portrait information associated with the audio clip in the database to determine the target speaker based on the portrait information; If there is no preset data matching the audio clip in the database, the portrait information in the current scene and the audio information corresponding to each portrait information are collected, and a correspondence between the portrait information and the audio information is established, and the portrait information and the audio information are stored in the database based on the correspondence.
4. The target speech control method based on smart glasses according to claim 1, characterized in that: Before determining the sound preference information corresponding to the user in the current scene based on the mapping relationship between the noise level and the smart glasses usage preference data, the method further includes: Classifying historical usage data corresponding to the smart glasses based on different sound scenarios; wherein the historical usage data includes at least historical noise level data, historical volume adjustment data, and historical sound quality equalization parameter adjustment data; labeling the classified historical usage data with noise level labels; Based on different sound scenes, extracting sound adjustment features from the historical usage data corresponding to different noise level labels; Based on the sound adjustment features and the collaborative filtering algorithm, obtaining user sound preference data; Different sound scenes, different noise levels and the sound preference data are mapped to construct a smart glasses usage preference data table.
5. The target speech control method based on smart glasses according to claim 1, characterized in that: The analyzing the voice of the target speaker specifically includes: Determining, based on the audio segment corresponding to the target speaker, a sound wave frequency distribution feature and a waveform feature, and obtaining a timbre feature based on the sound wave frequency distribution feature and the waveform feature; Obtaining the output vocabulary and syllable variation of the target speaker within a preset time period, comparing the output vocabulary and syllable variation with preset variable thresholds, and obtaining a speech rate feature based on the comparison results; Based on the audio segment corresponding to the target speaker, the slope of the fundamental frequency curve between adjacent frames is determined, and the peak and trough data corresponding to the fundamental frequency curve are determined. The intonation feature is obtained based on the slope and the peak and trough data.
6. The target speech control method based on smart glasses according to claim 5, characterized in that: The dynamic sound adjustment strategy based on the sound feature response obtained after analysis to optimize the voice of the target speaker specifically includes: In the event that there is a deviation between the timbre characteristic and the target timbre characteristic, performing gain adjustment on different frequency bands of the audio segment according to the deviation; In the event that there is a deviation between the speech rate feature and the target speech rate feature, denoising the audio segment using a preset sound signal processing algorithm, and performing syllable lengthening on the denoised audio segment; In the case where there is a deviation between the intonation feature and the target intonation feature, the fundamental frequency change trajectory is extracted based on the audio segment, the fundamental frequency change trajectory is compared with the target intonation template through dynamic time warping, and the segments where the comparison error does not meet the preset conditions are adjusted.
7. The target speech control method based on smart glasses according to claim 6, characterized in that: Denoising the audio segment by using a preset sound signal processing algorithm, and extending the syllables of the denoised audio segment, specifically includes: Performing denoising and reconstruction processing on the audio segment by wavelet transform, and performing signal-to-noise ratio enhancement processing on the audio segment by spectral subtraction algorithm in speech enhancement; Performing semantic analysis and sentiment analysis on the speech recognition text corresponding to the processed audio clip, and determining key semantics in the speech recognition text based on the analysis results; A syllable extension coefficient is determined according to the relationship between the deviation and the preset ratio, and the syllable of the key semantics is extended based on the extension coefficient.
8. The target speech control method based on smart glasses according to claim 1, characterized in that: The method of adjusting the frequency band of the optimized sound based on the sound preference information by using a power amplifier provided on the smart glasses to achieve target voice control specifically includes: Taking the smart glasses as the center, the horizontal space is divided into a plurality of sector-shaped areas; wherein each sector-shaped area corresponds to a set of power amplifier parameters; Matching the sound source position corresponding to the target speaker with the plurality of sector-shaped areas to determine the sector-shaped area to be adjusted, and determining the initial power amplifier configuration corresponding to the sector-shaped area to be adjusted; Based on the voice preference information, the initial power amplifier configuration is adjusted to enhance the frequency band of the target speaker's voice and to attenuate the frequency band of the non-target speaker's voice.
9. A target voice control device based on smart glasses, characterized in that: The device comprises: A segmentation unit, which obtains audio information of a speaker in a current scene and divides the audio information into multiple audio segments according to audio features; an associating unit, which obtains portrait information of a person in the current scene and associates the portrait information with the corresponding audio clip to determine a target speaker; an analysis unit, which determines, when the noise level of the current scene reaches a preset noise level, the user's sound preference information corresponding to the current scene based on a mapping relationship between the noise level and the smart glasses usage preference data; an optimization unit that analyzes the target speaker's voice and responds to a dynamic voice adjustment strategy based on the voice characteristics obtained after the analysis to optimize the target speaker's voice; The adjustment unit adjusts the frequency band of the optimized sound through the power amplifier provided on the smart glasses based on the sound preference information to achieve target voice control.
10. A target voice control device based on smart glasses, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: execute the target voice control method based on smart glasses described in any one of claims 1-8 above.
Citation Information
Patent Citations
Voice and image composite interaction execution method and system for robot
CN105957521A
Audio signal processing equipment and method as well as electronic equipment
CN106653041A
Service robot noise reduction method based on video and audio localization technology
CN109147813A
Method for sound beauty and emotion modification
CN109599094A
Voice processing method and related equipment
CN110913073A
Cited By
Intelligent auxiliary agent response service method of customer service center in financial industry
CN121309727A
A method for intelligent auxiliary agent response service in customer service centers of the financial industry
CN121309727B
Audio intonation recognition method and device, computer equipment and readable storage medium
CN121963801A