Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

462 results about "Voice pitch" patented technology

Pitch is an integral part of the human voice. The pitch of the voice is defined as the "rate of vibration of the vocal folds" . The sound of the voice changes as the rate of vibrations varies. As the number of vibrations per second increases, so does the pitch, meaning the voice would sound higher.

Emotion prediction and disease derivation method and system based on multi-modal fusion

The invention discloses an emotion prediction and disease derivation method and system based on multi-modal fusion, and the system comprises a data collection and preprocessing module, an emotion fusion module, an abnormal condition detection and cloud uploading module, and a disease possibility derivation module. The data acquisition and preprocessing module comprises a video part, a text part and an audio part, and the video part comprises face emotion recognition and prediction and human motion recognition and prediction; the text part comprises text content emotion recognition and prediction; the audio part comprises voice-to-text and voice tone emotion recognition and prediction, the system comprehensively captures an emotion state by fusing multi-mode information such as video, text and voice, and the accuracy and prediction capability of emotion recognition are improved; and by predicting the future emotion trend, the abnormal condition is warned in advance, and the response timeliness is improved.
Owner:JIANGSU UNIV OF SCI & TECH IND TECH RES INST OF ZHANGJIAGANG

Speech cloning system and method fusing rhythm characteristics

The invention discloses a voice cloning system and method fusing rhythm characteristics, belongs to the technical field of voice synthesis and natural language, and is applied to the aspect of fine-grained rhythm control in zero-sample voice synthesis. The implementation method comprises the following steps of: 1, extracting rhythm features and audio features of an audio file, and further respectively acquiring pause features, speed features and tone features in the rhythm features by sequentially adopting transcriptional text inverse coding, syllable-level speed registration quantization and pitch sequence feature splicing modes; 2, fusing the features of the audio files in a manner of discarding feature screening without guidance of a classifier; 3, generating a target Mel spectrogram based on conditional flow matching; generating a target audio file from a to-be-cloned audio file through the trained voice cloning model controlled by the fusion rhythm; compared with the prior art, fine-grained rhythm control of tone and rhythm feature decoupling is realized in zero-sample speech synthesis, so that intonation accuracy based on a context scene is improved.
Owner:BEIJING INST OF TECH

Loudspeaker system and hearing correction system and method

PendingUS20250317697A1MicrophonesElectrical transducersHearing acuityHearing test
A loudspeaker system and method can measure a specific user's hearing and implement compensatory or corrective processing to address the user's hearing deficiencies. The user's hearing acuity may be determined by (i) headphone techniques in which frequency-based hearing thresholds are determined; (ii) a loudspeaker system set up as intended for use in an acoustic space to emit tonal stimuli; or (iii) inducing, acquiring, and interpreting otoacoustic emissions. Once the hearing profile is determined, appropriate signal processing parameters can be used to generate a corrected audio output signal adjustment. The administered hearing test results can include a plurality of hearing loudness threshold levels at respective test frequencies. The hearing correction system's corrected audio output can be generated from corrected loudness values at selected correction frequency sub-bands.
Owner:SOUND UNITED LLC

Voice interaction system and method for customer service based on artificial intelligence

The invention relates to the technical field of voice recognition, in particular to a voice interaction system and method for customer service based on artificial intelligence, and the system comprises a voice input processing module, an intention classification and routing module, a context dynamic adjustment module, a user behavior learning module, a multi-level intention fusion module and a final result module. According to the method, a multi-dimensional feature system is constructed by extracting tone intensity, speech speed frequency and emotional fluctuation amplitude, intention categories, priority weights and confidence scores are generated to realize accurate acquisition of appeals, and dialogue history, context and emotional change dynamic reconstruction path nodes, switching rules and response time sequences are tracked during interaction. Historical behavior mining preference features, habit fusion intention relevance, emergency calculation of an optimal strategy, construction of service steps, resource allocation schemes and execution timelines, adjustment of an interactive interface, a service process and a feedback mechanism according to multi-dimensional analysis, guarantee of differentiated service experience, and improvement of response accuracy and user satisfaction.
Owner:NANJING XIUGUO INTELLIGENT TECH CO LTD

Music generation method, music generation device, electronic equipment and storage medium

The invention provides a music generation method, a music generation device, electronic equipment and a storage medium, and the music generation method comprises the steps: after receiving a music generation instruction, generating an accompaniment audio and a singing audio according to the music generation instruction, carrying out the audio synthesis of the accompaniment audio and the singing audio, and generating and outputting target music. Therefore, when the method is applied to the vehicle intelligent cabin scene, the accompaniment audio and the singing audio are respectively generated by adopting an artificial intelligence technology, and the accompaniment audio and the singing audio are synthesized to obtain the target music, so that the adaptation degree of the accompaniment audio and the singing audio in music content creation can be effectively improved; the accompaniment audio and the singing audio are closely fused in multiple dimensions such as rhythm, tone and emotional expression, so that the user experience is greatly improved.
Owner:XIAOMI EV TECH CO LTD +2

Voice interaction method, server and computer readable storage medium

The invention discloses a voice interaction method, a server and a computer readable storage medium. The method comprises: determining an emotional state identifier according to acoustic features of a received voice request, the acoustic features including at least one of a tone feature, a speech speed feature and an energy feature; and according to the emotional state identifier and the voice request, generating an emotional adaptive pad call so as to complete the voice interaction. Therefore, by analyzing the acoustic features, identifying the emotional state identifier, accurately generating the emotional adaptive pad call corresponding to the voice request, and establishing emotional connection with the user, the user experience is enhanced. Moreover, by generating the emotion-adaptive pad call, a waiting blank period from the time after the user sends an instruction to the time before the system returns a formal result can be filled up, the experience of'instant response 'is created, the user is prevented from anxiety caused by too long waiting time, and the satisfaction degree and the credibility of the system are further improved.
Owner:GUANGZHOU XIAOPENG MOTORS TECH CO LTD

Multi-mode dysarthria speech reconstruction system and method

PendingCN120340510ASpeech analysisSpeech reconstructionDysarthria
The invention provides a multi-mode dysarthria speech reconstruction system and a multi-mode dysarthria speech reconstruction method. The system comprises an encoder and a decoder. The encoder is configured to: extract a multi-modal feature from a multi-modal input of dysarthria speech and generate phoneme embedding based on the multi-modal feature; predicting a sequence of phoneme durations based on the phoneme embedding and generating an extended phoneme embedding and a predicted tone sequence; and extracting a speaker identity characterization of the dysarthria speech speaker. The decoder is configured to reconstruct a normal speech waveform corresponding to dysarthria speech with the extended phoneme embedding, the predicted tone sequence, and the speaker identity representation as inputs.
Owner:CENT FOR PERCEPTUAL & INTERACTIVE INTELLIGENCE (CPII) LTD

Cross-modal music automatic generation system and method based on emotion recognition

The invention belongs to the technical field of emotion music generation, and particularly relates to a cross-modal music automatic generation system and method based on emotion recognition, and the method comprises the steps: synchronously collecting the facial expression, voice tone and ECG physiological signals of a user through a signal collection unit; processing the collected information through a multi-modal emotion recognition model to obtain a VAD three-dimensional continuous emotion vector, inputting the VAD three-dimensional continuous emotion vector into a music generation module, and constructing a shared cross-modal potential space through an emotion auto-encoder and a music auto-encoder in the music generation module; constraining the consistency of emotion-music in the potential space by adopting a comparative learning loss function; and generating a music file in an MIDI format based on the Mus-Decoder. The system can fully combine facial expressions, voice tones and ECG physiological signal multi-mode modes to generate music matched with the current emotion of the user, and emotion semantic consistency is achieved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

AI-based emotional text voice conversion method and device

The invention discloses an AI-based emotional text speech conversion method and device. The method comprises the following steps: acquiring speech segments and text records from historical data of a user; performing noise reduction processing and feature extraction according to the voice segments and the text records to obtain voice features; inputting the voice features into a pre-constructed emotional tendency model, and outputting emotional tendency and emotional intensity; according to the emotional tendency and the emotional intensity, adjusting a tone weight, a speech speed and a volume to obtain a speech parameter; extracting new voice features according to the voice parameters to perform scene emotion label matching, and determining voice adjustment parameters through a linear regression model; according to the emotional tendency and the new voice features, generating an emotional type through a pre-established emotional intention classification model, and calculating a voice parameter weight in combination with a pre-established emotional mapping table; and according to the voice adjustment parameter and the voice parameter weight, performing language synthesis to generate personalized voice. According to the method, personalized expression can be accurately generated according to the scene.
Owner:FUJIAN YUANZHI UNIVERSE CULTURE COMMUNICATION CO LTD

Wireless direction finder circuit and wireless direction finder

The invention discloses a wireless direction finder circuit and a wireless direction finder, and relates to the technical field of wireless direction finding. Wherein the wireless direction finder circuit comprises a local oscillation signal input end, a frequency mixing detection circuit, a rectifying circuit, a main control circuit and an inflexion output circuit. The frequency mixing detection circuit generates a frequency mixing signal according to the wireless signal and the local oscillator signal. The rectifying circuit converts the frequency mixing signal into a first direct current signal. And the main control circuit outputs an inflexion control signal when the voltage value of the first direct current signal is greater than or equal to a preset voltage value. When the inflexion output circuit does not receive the inflexion control signal, the inflexion output circuit outputs a corresponding first audio signal according to the frequency mixing signal; and when the inflexion control signal is received, the second audio signal of frequency conversion is output according to the inflexion control signal and the frequency mixing signal, so that a user can be prompted that the signal intensity is relatively high at the moment and the signal source has entered the preset distance range. According to the invention, positioning is assisted through a dual change mechanism of tone and volume, so that positioning efficiency and environmental adaptability are improved.
Owner:GUANGZHOU XINHAN TECHNOLOGY CO LTD

Systems and methods for automatic audio experience customization based on ridable device context

Devices, systems, and methods are disclosed for customizing audio experiences associated with operation of a ridable device. An electronic device having a processor detects contextual information such as speed, acceleration, ambient noise level, environment type, or location of the ridable device. Sensor data from accelerometers, gyroscopes, microphones, cameras, or positioning systems may be processed to identify these conditions. Based on the detected context, the electronic device determines audio configuration parameters, including selection of audio content, adjustment of volume, blending of tracks, or modification of tempo or pitch. The determined parameters are applied to output audio through one or more speakers of the electronic device, the ridable device, or associated accessories. In some implementations, playback is synchronized with other riders using a common time reference, geo-fence triggers are employed to provide location-specific content, or warning sounds are generated when nearby objects are detected, enhancing both rider enjoyment and environmental awareness.
Owner:SCOOTASONIX LTD

Short video copywriting tone automatic adjusting method driven by hierarchical rhythm mapping

The invention discloses a hierarchical rhythm mapping-driven short video copywriting mood automatic adjustment method, and relates to the technical field of video processing, and the method comprises the steps: 1, receiving a text character string and a language type identifier, and building an occupation column for bearing a tone mark, an accent mark and a duration mark at each level; 2, dividing each sentence into phrase segments based on the hierarchical index table, freezing boundaries by taking the phrase segments as units, presetting sentence end termination styles according to punctuations, determining kernel phrases according to semantic anchor points, initializing trends of the kernel phrases, and performing time sequence elastic alignment and hierarchical backfilling to obtain a sentence end termination pattern; and finally outputting a triple sequence which covers all syllables and is composed of tone marks, accent marks and duration marks as a target rhythm control sequence. And step 3, performing audio generation based on the target rhythm control sequence to obtain new dubbing. According to the method, the tone accuracy and expressive force of short video dubbing are improved, and the time and cost of manual adjustment are remarkably reduced.
Owner:CLOUD ATTACK NETWORK TECH HEBEI CO LTD

Sound equipment with tone-adjustable horse race lamp special effect

The invention discloses sound equipment for adjusting the special effect of a horse race lamp, and the sound equipment comprises a tone detection module which is used for monitoring the rotation angle and direction of a tone modification knob in real time, and outputting a corresponding tone grade signal; the light control module is used for responding to the tone level signal and generating a corresponding light control signal, and the light control signal comprises that relative to a reference tone level, when the tone level is increased by one level, the light irradiation range is expanded by a first preset proportion, the brightness is increased by a first set proportion, and when the tone level is reduced by one level, the light irradiation range is expanded by a second preset proportion; the lamplight irradiation range is shrunk by a second preset proportion, and the brightness is reduced by a second set proportion; and the horse race lamp assembly is used for receiving the light control signal and dynamically adjusting the light effect of the horse race lamp. According to the invention, the signal of operating the tone modification knob by the user can be detected, the tone grade signal can be determined, the mapping relation between the tone grade and the light parameter can be established, the correlation between tone change and the light effect can be ensured, accurate regulation and control of the light irradiation range and brightness can be realized, deep coupling of auditory adjustment and visual feedback can be realized, and the user interaction experience can be improved.
Owner:GUANGZHOU JIE LI ELECTRON CO LTD

Holistic and inclusive wayfinding

The technology employs a holistic approach to passenger pickups and other wayfinding situations. This includes identifying where passengers are relative to the vehicle and / or the pickup location. One aspect leverages camera imagery from a rider's client device when providing rider support. This enables an agent to receive the imagery to help guide the rider to the vehicle. Another aspect provides audio information to the rider to help them get to the vehicle. This can include selecting or curating various tones or melodies, giving the rider advanced notice of sounds to listen for, and modifying sounds as the rider approaches the vehicle or to address ambient noise in the environment. Different wayfinding tools may be selected for presentation to the rider based on their proximity state to a pickup location or to the vehicle. Gracefully transitioning between different tools can enhance their usefulness and elevate the rider's experience for a trip.
Owner:WAYMO LLC

Inaudible orthogonal signal communication

Methods are generally described for inaudible orthogonal signal communication. An example method includes determining, based on an input message, a series of symbols, where each symbol from the series of symbols represents a numerical value, wherein each symbol from the series of symbols corresponds to a band-limited white noise tone orthogonal to each other band-limited white noise tone corresponding to each other symbol. The example method also includes generating a sub-token comprising each symbol from the series of symbols, where each symbol overlaps in time with at least one other symbol. The example method also includes modulating a carrier wave with the sub-token to produce a signal audio waveform, where the modulation uses a spread spectrum technique, and providing the signal audio waveform to a television device, where the television emits the signal audio waveform as a low-amplitude audio signal.
Owner:AMAZON TECH INC

Yi language speech recognition method based on self-supervision and attention feature fusion

The invention relates to the technical field of natural language processing, and discloses a Yi language speech recognition method based on self-supervision and attention feature fusion. The method comprises a feature encoder module, a comparative learning module, a mask language modeling module, a joint optimization and feature fusion module and a decoder module, the feature encoder module adopts a convolutional neural network structure and converts continuous waveform signals into feature representation suitable for subsequent modeling, and the comparative learning module performs feature fusion on the continuous waveform signals through a Gumbel-Softmax technology. The method comprises the following steps that: a mask language modeling module and a feature fusion module are integrated, deviation caused by manual definition or clustering is avoided, the mask language modeling module obviously enhances semantic understanding and tone modeling capabilities of a model in a low-resource scene, a self-attention feature fusion mechanism is introduced into the feature fusion module, continuous features, discrete unit representation and semantic context representation from an acoustic level are spliced, and a self-attention feature fusion mechanism is introduced into the self-attention feature fusion mechanism. The decoder module adopts a decoder structure based on connection time sequence classification, and the corresponding relation between the voice and the text can be achieved without strict alignment labeling.
Owner:KUNMING UNIVERSITY

Privacy-preserving avatar voice transmission

Some implementations relate to methods, systems, and computer-readable media to reproduce a voice stream with removed personally identifiable information (PII). An original voice stream is received from a user by a hardware processor. The original voice stream is converted into text using a machine learning (ML) voice transcription model. Then, the text is used for generating an anonymized voice stream corresponding to the text having similar auditory properties to the original voice stream other than PII using a machine learning (ML) voice synthesis model. The similar properties are determined using metadata corresponding to the original voice stream. The metadata may be obtained by analyzing acoustic characteristics of the original voice stream. The metadata may include volume changes and pitch frequency changes stored as metadata annotations associated with tokens in the text as key-value pairs. The generation may use the metadata annotations to determine the similar audio properties.
Owner:ROBLOX CORP

Motor control method of oral care equipment and related device

The invention discloses a motor control method of oral care equipment and a related device. The oral care equipment comprises a motor, and the control method comprises the following steps: inputting a driving signal to the motor to drive the motor to generate cleaning vibration; the driving signal comprises a plurality of target driving waves, the target driving waves are used for driving the motor to generate sound of tones corresponding to the target driving waves, and most energy of the frequency spectrum of the sound is concentrated in a preset driving frequency range of the motor. In this way, the oral care equipment can provide an efficient cleaning effect for a user while playing music.
Owner:GUANGZHOU STARS PULSE CO LTD

System

An object of a system according to an embodiment is to accurately identify the pitch of a sound and provide visual and auditory feedback.SOLUTION: A system includes a sound recognition and analysis unit, a visual display unit, and an auditory feedback unit. The sound recognition and analysis unit identifies the pitch of the sound using the generated AI. The visual display unit visually displays the pitch of the sound identified by the sound recognition and analysis unit. The auditory feedback unit aurally feeds back the pitch of the sound specified by the sound recognition and analysis unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

AI call content optimization method and device based on voiceprint recognition

The invention relates to the technical field of intelligent voice processing and communication, and discloses an AI call content optimization method and device based on voiceprint recognition. The method comprises the following steps: acquiring an original voice data stream from a user call, and separating voice speed, tone and frequency characteristics to obtain a dynamic voice characteristic set; determining a microphone frequency response deviation and a speech speed change rate according to the microphone frequency response deviation and the speech speed change rate to form a speech feature parameter set; if the parameter exceeds the threshold value, redistributing a speech speed weight to generate an adjusted speech data stream; noise reduction is carried out to obtain a pure data stream, and features are fused to generate a personalized sound effect adjustment curve; adapting the equipment difference to obtain an adaptation curve; compressing the voice data according to the adaptive curve and optimizing the transmission priority to obtain an optimized transmission data stream; and combining the transmission data stream with the adaptive curve to generate final call voice output. According to the method, conversation content dynamic adaptation and whole-process optimization are realized, and conversation quality stability and cross-scene applicability are improved.
Owner:QUANZHOU YUANZHISHI ELECTRONIC COMMERCE CO LTD

Loudspeaker system and hearing correction system and method

PCT designated stageWO2025213113A1MicrophonesElectrical transducersHearing acuityHearing test
A loudspeaker system and method can measure a specific user's hearing and implement compensatory or corrective processing to address the user's hearing deficiencies. The user's hearing acuity may be determined by (i) headphone techniques in which frequency-based hearing thresholds are determined; (ii) a loudspeaker system set up as intended for use in an acoustic space to emit tonal stimuli; or (iii) inducing, acquiring, and interpreting otoacoustic emissions. Once the hearing profile is determined, appropriate signal processing parameters can be used to generate a corrected audio output signal adjustment. The administered hearing test results can include a plurality of hearing loudness threshold levels at respective test frequencies. The hearing correction system's corrected audio output can be generated from corrected loudness values at selected correction frequency sub-bands.
Owner:SOUND UNITED LLC

Pressure-touch sensor based on common electrode structure and application thereof

The invention relates to the field of sensors, in particular to a pressure-tactile sensor based on a common electrode structure and application thereof, and the pressure-tactile sensor comprises a first electrode layer, an air dielectric layer, a conductive layer, a mixed pressure sensitive layer and a second electrode layer from top to bottom, wherein the air dielectric layer is formed at the middle position by arranging insulating cushion blocks at the edge between the first electrode layer and the conductive layer, and an output electrode layer is arranged between the insulating cushion block at one side and the conductive layer. The sensor can be used in wearable health equipment to realize health monitoring, can be integrated to a signal acquisition and processing system, and is combined with a graphical programming platform to realize accurate control of electronic piano key tones and volumes. Good application prospects and important technical development potential are shown in the fields of intelligent interaction interfaces, intelligent musical instruments, human-computer interaction and the like.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Active noise cancellation of engine noise at sleeping locations on aircraft

PendingUS20260094597A1Sound producing devicesRest berthsNacelleNoise
In accordance with certain aspects, noise cancellation systems and methods are utilized to mitigate engine tonal noise at sleeping locations in the aircraft cabin space. In one embodiment, an onboard noise cancellation system is provided that includes a plurality of microphones, the plurality of microphones, and a plurality of speakers, and a noise cancellation circuit. In this embodiment the plurality of speakers are arranged proximate to a bed in the sleeping area of the aircraft cabin. Likewise, the plurality of speakers are arranged proximate to the proximate to the bed in the sleeping area. The noise cancellation circuit is further coupled to the plurality of speakers, and the noise cancellation circuit is configured to drive the plurality of speakers to generate noise cancellation audio to at least partially cancel engine noise in a region above the bed.
Owner:GULFSTREAM AEROSPACE CORP

Anonymization privacy protection method and system for voice information retention

The embodiment of the invention provides an anonymization privacy protection method and system for voice information retention. The method comprises the following steps: extracting speaker embedding of an original audio, eliminating the tone of the speaker, and keeping semantic and rhythm speaker irrelevant features; embedding and inputting a speaker into a speaker anonymous module matched with a three-stage stream based on a U-Net architecture to obtain anonymous embedding; and combining irrelevant features of the speaker with anonymous embedding by using a pre-trained voice reconstruction model to generate anonymized voice with tone privacy. According to the embodiment of the invention, voice anonymization facing content privacy and tone privacy reserved by voice information is realized, the effectiveness of the anonymized voice generated by using the method in a downstream task is superior to that of a baseline model, meanwhile, the privacy of a speaker is also guaranteed, and safe use of data is realized.
Owner:SHANGHAI JIAOTONG UNIV

Text-to-voice method, system and device and storage medium

The invention discloses a text-to-speech method, system and device and a storage medium, and the method comprises the steps: carrying out the chapter processing of a novel text, and enabling each chapter text to comprise a front text, a middle text and a rear text; extracting phonemes, tones, rhythm information and text content from a middle text in each chapter text, and splicing the extracted phonemes, tones, rhythm information and text content to obtain a text vector; obtaining a timbre vector of a target speaker, and splicing the timbre vector and the text vector to obtain a target vector; and inputting the target vector into a text-to-speech model for speech synthesis, and outputting an audio corresponding to the novel text. According to the invention, the end-to-end speech synthesis is realized, and there is no need to label the white, emotion and role of the dialogue in advance, so that the speech synthesis efficiency is improved.
Owner:GUANGZHOU QUYAN NETWORK TECH CO LTD

Method for determining the auditory threshold of a test subject, hearing aid system, method for setting hearing aid parameters and computer readable medium for performing the method

ActiveUS12350038B2Sets with translation techniquesAudiometeringAuditory thresholdsNoise
To determine an auditory threshold of a test subject, according to the method, multiple tone sets, which predominantly contain a plurality of noises, having properties remaining uniform within a tone set, are presented in succession to a test subject by use of an output transducer. The test subject is asked to indicate a perceived number of noises from the presented tone set after each presentation of one of these tone sets. As a function of the perceived number of noises, a property of at least one part of the noises is changed in relation to the preceding tone set for the presentation of a following tone set and the following tone set is presented to the test subject. As a function of the respective perceived number of noises in the presented tone sets, at least one value of the auditory threshold of the test subject is estimated.
Owner:SIVANTOS PTE LTD

Mandarin pronunciation evaluation system based on deep learning

The invention discloses a mandarin pronunciation evaluation system based on deep learning, and relates to the technical field of voice scoring, and the system comprises a data input port which is used for obtaining a copy, a standard pitch curve and a standard mouth shape video, shooting the face mouth shape motion and sound of a user during reading into a video, and extracting an audio from the video; the data processor is used for processing videos and audios; the tone evaluator is used for evaluating the audio to obtain a tone score; the mouth shape evaluator is used for obtaining a mouth shape score; the pronunciation evaluator is used for obtaining a pronunciation score; and the score output end is used for combining the tone score, the mouth shape score and the pronunciation score to generate a final score. The mandarin pronunciation scoring method and device have the effect of improving the accuracy and adaptability of mandarin pronunciation scoring.
Owner:AI TUER

Motor control method for oral care device, and related apparatus

The present application discloses a motor control method for an oral care device, and a related apparatus. The oral care device comprises a motor, and the control method comprises: inputting a driving signal into the motor so as to drive the motor to generate cleaning vibrations. The driving signal comprises a plurality of target driving waves, each target driving wave is used for driving the motor to generate sound at a pitch corresponding to the target driving wave, and most energy in the frequency spectrum of the sound is concentrated within a preset driving-frequency range of the motor.
Owner:GUANGZHOU STARS PULSE CO LTD

System and method for tone throne percussive guitar stool

A system and method for a stool with foot pedals that generate electronic drum sounds, whereby the stool includes a seat supported by a base and footstool, and two foot pedals attached to the footstool on either side of the seat, whereby foot pedals are capable of being pressed to generate electronic drum sounds, whereby the foot pedals are connected to an electronic control system that generates the electronic drum sounds such that electronic control system can be programmed to produce a wide range of drum sounds, from traditional acoustic drum sounds to more experimental electronic drum sounds.
Owner:MATTHEWS DAVID C

Voice dictation with audio large language model

A method comprising receiving audio data 102, generating a transcription 151 comprising a sequence of terms 152, such as “Buy some tomatoes and bananas. Change tomatoes to potatoes”, parallel processing the audio data and the transcription using a multimodal large language model (LLM) 150 to identify one or more revision terms 152R, for example “Change”, specifying a revision action to perform on at least one other term in the sequence, in this instance “tomatoes”, and modifying the transcription 151M accordingly – “Buy some potatoes and bananas”. Identifying the revision term(s) may be based on a corresponding user intent 154 determined for each respective term in the sequence, for example, the user 10 does not intend the final transcription to include “tomatoes”. For each term in the sequence, parallel processing may comprise correlating its speech characteristics 156 such as pitch, tone or prosody information determined from the audio data with its corresponding linguistic context 158 determined from the transcription. Transcription correction may be based on a revision token inserted into the sequence, the token indicating an N number of terms for replacement and their corresponding replacement terms. User context data 104 may be obtained to tailor the LLM to a particular user. [Figure 1A]
Owner:GOOGLE LLC