Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Speech acquisition" patented technology

Speech acquisition focuses on the development of spoken language by a child. Speech consists of an organized set of sounds or phonemes that are used to convey meaning while language is an arbitrary association of symbols used according to prescribed rules to convey meaning. While grammatical and syntactic learning can be seen as a part of language acquisition, speech acquisition focuses on the development of speech perception and speech production over the first years of a child's lifetime. There are several models to explain the norms of speech sound or phoneme acquisition in children.

Information processing device, information processing method, and information processing program

The present invention provides an information processing device that enables a more natural improvement in the accuracy of identifying past related utterances made by a user or identifying their relationship with other users. [Solution] The system includes: a speech acquisition unit 1 that acquires user utterances; a prompt generation unit 2 that generates prompts for a large-scale language model to generate tag information that includes at least one of the following based on the acquired utterances: the relationships between the constituent elements of the utterance, the attributes to which the utterance belongs, or a summary of the utterance; an interface 3 that transmits the generated prompts to the large-scale language model and receives response information that includes tag information generated by the large-scale language model in response to the transmitted prompts; and a recording unit 5 that records the acquired utterances and the received tag information in association.
Owner:PIONEER IP

Notification apparatus, notification method, and recording medium

A communication terminal functioning as a notification apparatus includes: an intellectual function evaluation unit that evaluates a plurality of intellectual functions relating to a physical condition of a driver being a user, based on biological information of the driver; a dialog unit that functions as a speech acquisition unit that acquires speech information of the driver; a speech evaluation unit that evaluate second physical condition information consisting of a speech analysis result relating to the physical condition of the driver, based on the speech information; and a notification control unit that notifies the driver of notification information corresponding to the physical condition of the driver, based on a combination of an evaluation result from the intellectual function evaluation unit and an evaluation result from the speech evaluation unit.
Owner:HONDA MOTOR CO LTD

Interactive voice acquisition method and system

The invention relates to the field of artificial intelligence, in particular to a voice interaction method and system, and the method comprises the steps: generating response audio data according to input audio data, and generating an output duration according to the response audio data; obtaining a timestamp of each audio unit in the response audio data according to the output duration; emotion words of audio units in the response audio data are recognized, target emotion expression information is generated for the audio units with the emotion words, and expression time of the target emotion expression information is determined according to timestamps of the audio units; and outputting interaction voice according to the response audio data, the target emotion expression information and the expression time. According to the method, efficient anthropomorphic response during voice interaction can be realized, and the use experience of the user and the dialogue realistic degree of voice interaction are improved.
Owner:GUANGZHOU HUYA INFORMATION TECH CO LTD

Order taking device and program

We provide an order taking device and program that can reduce the effort required of customers when placing an order. [Solution] The order receiving device comprises: a speech acquisition unit that acquires customer speech data (speech information) related to ordering from the menu; a text conversion unit that converts the speech data acquired by the speech acquisition unit into text data (text information); an order content estimation unit that analyzes the text data by referring to the features contained in the text data and the contents of the menu offered at the store to estimate the contents of the menu that the customer will order; and an order information presentation unit that presents the contents of the menu estimated by the order content estimation unit to the customer.
Owner:TOSHIBA TEC KK

Speech signal processing apparatus, method, electronic device, and sound amplification system

This invention provides a speech signal processing apparatus, method, electronic device, and loudspeaker system, relating to the field of speech processing technology. The apparatus includes: a first echo cancellation unit, used to input a reference signal and a current speech signal acquired by a speech acquisition device into an adaptive filter to obtain a target residual signal; and a second echo cancellation unit, used to input the target residual signal and a far-end speech signal into a preset speech processing model to obtain a near-end speech processing signal for the current frame. This invention performs echo cancellation processing on the target residual signal using a speech processing model. Since the speech processing model is trained based on residual signal samples and far-end speech signal samples, and both the residual signal samples and the target residual signal include nonlinear echoes and other signals that the adaptive filter cannot completely eliminate, the speech processing model can eliminate the echoes of nonlinear components in the target residual signal, thereby improving the accuracy of speech signal processing.
Owner:BEIJING ESWIN COMPUTING TECH CO LTD

A voice recognition-based nursing document automatic generation system

PendingCN122369461AClinical scenarioNursing documentation
This invention discloses an automatic nursing documentation generation system based on speech recognition. The system includes a speech acquisition module, a medical-specific speech recognition engine module, a structured nursing documentation parsing module, a system integration and adaptation module, a data security and compliance management module, and a terminal interaction module. This invention achieves automated, voice-driven generation of medical and nursing documentation throughout the entire process by adapting speech acquisition to all clinical scenarios, optimizing speech recognition specifically for nursing, and automatically generating structured nursing documentation. It can be seamlessly integrated into existing hospital medical document writing systems, bedside nursing cart systems, and PDA mobile nursing terminal systems as a standalone app or functional plugin, allowing for rapid deployment without modifying the original system architecture. This solves the core pain point of tedious medical document writing for frontline clinical medical staff, which consumes more than half of the consultation time, improving nursing documentation writing efficiency by over 80% and significantly reducing the time spent on non-medical documentation by medical staff.
Owner:JIANGSU PROVINCE INST OF TRADITIONAL CHINESE MEDICINE

A voice recognition system for the power industry

This invention provides a speech recognition system for the power industry, comprising: a speech acquisition model for collecting speech data from a designated area; a speech processing module for converting the speech data into general power industry text based on a pre-set speech dictionary in a database and outputting it to the speech recognition module; a speech recognition module for performing semantic recognition on the general power industry text based on input speech key information, recognizing general power industry semantics in the text, determining keywords, and outputting the recognition result; and a determination module for determining the recognition result based on input control specifications and determining the semantic recognition accuracy based on the determination result. This invention can perform speech recognition online or offline, can convert different power problems and solutions into text for storage and recording, and can simultaneously process the recorded text to extract key semantics and build models.
Owner:SHENZHEN POWER SUPPLY BUREAU

A photoacoustic signal enhancement method and device based on adaptive multi-band spectral subtraction and wiener filtering

PendingCN122454993AMoving averageSpectral subtraction
The application discloses a photoacoustic signal enhancement method and device based on adaptive multi-band spectral subtraction and Wiener filtering, and belongs to the technical field of speech signal enhancement. The application divides a signal into sub-bands according to Mel scale, determines a subtraction factor and a lower limit factor by a monotone decreasing function linkage according to real-time signal-to-noise ratio of each sub-band, and makes the two factors negatively correlated with the signal-to-noise ratio, so that the denoising strength and the spectral bottom filling depth are automatically adapted; the weighted moving average of the power spectrum of adjacent sub-bands is carried out to smooth isolated spectral peaks and spectral valleys left by spectral subtraction; a residual amplitude limiting is introduced before Wiener filtering, the maximum noise residual threshold is determined based on adjacent frame statistics and limiting, and Wiener filtering is carried out by taking the data after limiting as clean signal estimation; the noise spectrum is recursively updated by a forgetting factor which is dynamically adjusted according to the signal-to-noise ratio, and is only executed in the speech inactive section. The application has an output signal-to-noise ratio improvement of more than 50% when the input is 0dB, can stably enhance without damage under high signal-to-noise ratio, and is suitable for photoacoustic speech acquisition and enhancement scenes.
Owner:ANHUI ZHIBO PHOTOELECTRIC TECHNOLOGY CO LTD

Voice signal processing method and device, electronic equipment, medium and program product

PendingCN122135732ASpeech recognitionSpeech observationsIntelligent equipment
This disclosure relates to a speech signal processing method, apparatus, electronic device, medium, and program product. The speech signal processing method includes: separating and processing speech observation signals acquired by a speech acquisition array to obtain independent source speech signals; when human speech is detected, extracting a first audio mixing feature of each independent source speech signal; determining a target independent source speech signal based on the first audio mixing feature of each independent source speech signal and the target audio mixing feature of the target object; and performing speech recognition on the target independent source speech signal. Thus, even if the location of the target object moves during voice interaction, the target independent source speech signal containing the target object's speech can be accurately determined, thereby accurately distinguishing the target object's speech signal from the speech signal of interfering objects. This improves the accuracy of voice interaction between smart devices and target objects, enhancing the user experience.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Voice input noise reduction processing method based on deep learning

PendingCN121884844ASpeech analysisSpectral edgeFrequency spectrum
The invention discloses a voice input noise reduction processing method based on deep learning, relates to the technical field of voice noise reduction, is used for solving the problem that noise recognition errors are increased, and aims to solve the problem that the noise recognition errors are increased by monitoring voice acquisition microphones, acquiring pointing angles and sensitivity variations of the microphones and calculating pickup angle offset amplitudes according to the pointing angles and the sensitivity variations. The method comprises the following steps: comprehensively evaluating a microphone pickup state, screening a stable microphone and calling a voice signal segment covered by the microphone, detecting spectrum edge information of the signal segment by using a time-frequency domain convolution operator, carrying out frequency band division based on the edge information, and obtaining a noise power ratio and voice signal energy intensity of each frequency band. Evaluating a frequency band noise change trend according to the noise power ratio, generating a noise reduction weight index in combination with the voice signal energy intensity, and marking the frequency band according to the index; by accurately screening the stable microphone and effectively evaluating the frequency band noise characteristics, the noise reduction effect of voice signal processing is improved, and the voice recognition and communication quality is improved.
Owner:SHANGHAI MAIJUN TECHNOLOGY CO LTD

Display device and operation method thereof

A display device, according to one embodiment of the present disclosure, comprises: a display for displaying at least one input field; a speech acquisition unit for acquiring speech inputted in the input field; and a control unit for converting the speech to text, wherein the control unit may convert the speech to text by character unit or by phrase unit according to the type of the input field.
Owner:LG ELECTRONICS INC

Speech output device, speech output method, and program

The present invention provides a speech output device, method, and program that can output appropriate speech to a user using user behavior information. [Solution] The speech output device H1 comprises a receiving unit H12 that receives time-series information including one or more pieces of information about the user's actions in time series, a speech acquisition unit H131 that acquires speech information that identifies speech directed to the user and corresponds to the time-series information received by the receiving unit H12, and a speech output unit H141 that outputs the speech information acquired by the speech acquisition unit H131.
Owner:EXEVITA INC

Target character stroke-by-stroke display method, system and device based on voice semantic analysis

PendingCN122653739AData displayDisplay device
The application provides a kind of based on speech semantic analysis letter by letter display method, system and equipment, comprising the following steps: S1: receiving user voice instruction, speech activity detection is carried out to voice instruction, when detecting continuous silence reaches preset time length, end speech acquisition, the preset time length is 600~1200ms;S2: the semantic analysis of the collected speech text, based on the rule template, the target text is parsed from the speech text;S3: the voice segment, recognized text or target text is provided to writing guide data generation module, and the writing guide data corresponding to the target text is obtained;S4: on display device, based on guide layer data display background guide word and auxiliary grid line, and based on the stroke data stroke by stroke superimposed display current stroke.The application can reduce the threshold of use, provide coherent writing guide experience, and adapt to the display characteristics of display device.
Owner:叶则荣

Teaching speech recognition correction system based on voiceprint feature extraction

This invention relates to the field of speech signal processing and intelligent teaching assistance technology, specifically to a speech recognition and correction system for teaching based on voiceprint feature extraction. The system includes: a speech acquisition module, a feature extraction module, a speech feedback module, and a data processing center. The speech acquisition module is used to acquire input speech data in a multi-target interactive environment. The feature extraction module is used to extract current acoustic features and calculate the voiceprint feature contamination entropy. The speech feedback module is used to generate correction feedback data based on a standard pronunciation model. The data processing center is used to compare the voiceprint feature contamination entropy against a preset danger threshold and, in conjunction with the current acoustic features, generate voiceprint update instructions and correction strategy instructions to comprehensively manage the underlying voiceprint database and correction feedback strategy. This reduces the probability of incorrect attribution and misleading feedback. This invention effectively reduces the risk of contaminated speech samples being written into long-term voiceprint archives and improves the stability of target pronunciation recognition and correction control.
Owner:FUZHOU COLLEGE OF FOREIGN STUDIES & TRADE +1

A medical speech recognition-based gastroscopy report generation system and method

This invention belongs to the field of speech recognition technology. It discloses a system and method for generating gastrointestinal endoscopy reports based on medical speech recognition. The system includes a speech acquisition and processing module for acquiring real-time spoken speech from doctors during gastrointestinal endoscopy, performing adaptive noise reduction and mask occlusion compensation on the spoken speech, and outputting a speech segment sequence; a speech activity detection module for receiving the speech segment sequence, calculating the prior probability of speech activity based on the transition relationships of different operational states in the endoscopy procedure, and performing speech activity detection on the speech segment sequence based on the prior probability of speech activity to filter out speech event sequences consistent with the examination procedure; and a speech-image alignment module for transcribing the speech event sequences into medical speech, acquiring endoscopic images at corresponding times, and completing the ambiguous spoken content in the speech event sequence based on the endoscopic images to generate speech description data. This improves the automation level of examination recording.
Owner:SHANGHAI HAOKANGYUN MEDICAL TECHNOLOGY DEVELOPMENT CO LTD

Speech output device, speech output method, and program

Previously, it was not possible to output appropriate responses for a user based on their behavioral information. [Solution] A speech output device H1 comprises a receiving unit H12 that receives time-series information including one or more pieces of information on the user's actions in time, a speech acquisition unit H131 that acquires speech information that identifies speech directed to the user and corresponds to the time-series information received by the receiving unit H12, and a speech output unit H141 that outputs the speech information acquired by the speech acquisition unit H131. Using the user's action information, an appropriate speech can be output directed to the user by the speech output device H1.
Owner:EXEVITA INC

Deep learning-based adaptive speech recognition system

PendingCN122337201AFeature extractionSpeech code
This invention discloses a deep learning-based adaptive speech recognition system, relating to the field of speech recognition technology. The system processes raw speech signal data acquired by a speech acquisition module and dynamically adapts it to user identification information to obtain speech coding feature data. Furthermore, it extracts features from environmental metadata to obtain environmental embedding feature data. Based on the environmental embedding feature data and the speech coding feature data, feature recognition processing is performed to calculate recognition feature coefficients. The speech recognition module compares these coefficients with preset recognition feature thresholds and determines the speech quality based on the comparison results. This enables recognition under dynamically changing user and environmental conditions, improving the accuracy of speech recognition.
Owner:IANGSU COLLEGE OF ENG & TECH

Notification device, notification method, and computer program product

The invention provides a notification device, a notification method, and a computer program product, which can detect the physical condition of a user with high precision and perform notification corresponding to the physical condition of the user. A communication terminal (1) functioning as a notification device is provided with: an intellectual function evaluation unit (12) that evaluates a plurality of intellectual functions relating to the physical condition of a driver (D), which is a user, on the basis of biological information (Db) of the driver (D); a conversation unit (14) that functions as a voice acquisition unit that acquires voice information of the driver (D); a speech evaluation unit (15) that evaluates, on the basis of the speech information, second physical condition information comprising speech analysis results relating to the physical condition of the driver (D); and a notification control unit (16) that notifies notification information corresponding to the physical condition of the driver (D) on the basis of a combination of the evaluation result of the intellectual function evaluation unit (12) and the evaluation result of the voice evaluation unit (15).
Owner:HONDA MOTOR CO LTD

Role separation method, meeting summary recording method, role display method and apparatus, electronic device, and computer storage medium

A role separation method, a meeting summary recording method, a role display method and apparatus, an electronic device, and a computer storage medium, relating to the field of speech processing. The role separation method comprises: obtaining sound source angle data corresponding to a speech data frame, acquired by a speech acquisition device, of a role to be separated (S102); on the basis of the sound source angle data, performing identity recognition on the role to be separated to obtain a first identity recognition result of the role to be separated (S104); and separating the role on the basis of the first identity recognition result of the role to be separated (S106). The role is separated in real time, thus making user experience smooth.
Owner:ALIBABA GROUP HOLDING LTD

Carry-on partner based on artificial intelligence large model and voiceprint recognition and processing method

The invention discloses a carry-on partner based on an artificial intelligence large model and voiceprint recognition and a processing method, and the carry-on partner comprises a bracelet, a voice recognition model, a registration voiceprint model, the artificial intelligence large model, a database and a mobile control terminal, and the bracelet is internally provided with a microphone, a text-to-voice module and other core components. The database encrypts and stores interaction data, and the mobile control terminal supports privileged viewing. The processing method comprises the steps of training a voice recognition model, establishing a binding registration voiceprint model, collecting and preprocessing voice, converting the voice into a text and screening the voiceprint, generating a reply through a large model, feeding back the voice, optimizing corpora and storing and checking data. The method provides convenient, exclusive and interference-free voice interaction service for teenagers, assists mental healthy growth, optimizes parent-child communication, guarantees data safety, is suitable for multiple fields, and has a wide application prospect.
Owner:XINGMU TECH (HANGZHOU) CO LTD

system

We provide the system. [Solution] A voice acquisition means for receiving voice input, A speech recognition means that converts speech data acquired by the speech acquisition means into text data, A natural language processing means that analyzes the text data generated by the speech recognition means to extract transaction details, An authentication method that authenticates transactions based on extracted transaction details, A payment method that executes settlement processing when a transaction is authenticated, A notification method for notifying the transaction results, A system that includes this.
Owner:SOFTBANK GROUP CORP

Artificial throat sensor based on double-layer heterogeneous magnetic film magnetic repulsive force coupling and preparation method thereof

The invention discloses an artificial throat sensor based on double-layer heterogeneous magnetic film magnetic repulsive force coupling and a preparation method thereof. The sensor comprises an upper sensing unit and a lower sensing unit, each of the upper sensing unit and the lower sensing unit is provided with a first flexible magnetic layer magnetized directionally and a flexible hard magnetic film layer which forms a microscopic magnetic domain array through patterned magnetization, the same poles of the first flexible magnetic layer and the flexible hard magnetic film layer face each other to generate uniform magnetic repulsive force, and a dynamic adjusting gap is formed between friction electric layers. The gap can be passively and adaptively changed along with the vibration intensity of the throat, the repulsive force is sharply increased during strong vibration to prevent contact saturation, and the local tiny gap is maintained to keep high sensitivity during weak vibration. Through the full-film and patterned magnetizing technology, the problems that in a traditional magnetic repulsion scheme, magnetic force is not uniform, and wearing is uncomfortable are solved, unification of high signal fidelity, wide dynamic range and extreme wearing comfort is achieved, and the method is particularly suitable for a wearable artificial throat voice acquisition system.
Owner:FUZHOU UNIV

System

An object of a system according to an embodiment is to enable smooth and real-time communication between a person without hearing and a person with hearing.SOLUTION: A system includes a voice acquisition unit, a sign language acquisition unit, a conversion unit, and a notification unit. The voice acquisition unit acquires voice. The sign language acquisition unit acquires a sign language. The conversion unit converts the speech acquired by the speech acquisition unit into characters, and converts the sign language acquired by the sign language acquisition unit into speech and characters. The notification unit notifies the person without hearing of the characters converted by the conversion unit, and notifies the person with hearing of the voice and the characters converted by the conversion unit.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

A multi-disciplinary discussion automatic generation system based on evidence-based knowledge graph

This invention relates to the field of medical information management technology and discloses an automatic generation system for multidisciplinary discussions based on evidence-based knowledge graphs. The system includes: a speech acquisition and speaker separation module, which acquires and separates speech signals in real time during multi-person discussions, outputting segmented speech data; a data fusion and correction module, which converts the segmented speech data into initial text via speech recognition, dynamically corrects the recognition results, and generates a patient treatment event sequence; a causal-enhanced entity recognition module, which receives the patient treatment event sequence and identifies a set of medical entities with timestamps and preset causal role attributes; an evidence-based verification relation extraction module, which queries relevant evidence for potential entity pairs in the medical entity set, determines or marks conflicts in the semantic relationships between entities, and outputs entity relation triples; and a treatment knowledge graph generation module, which integrates the medical entity set and entity relation triples to construct a patient treatment knowledge graph centered on the patient.
Owner:CARDIOVASCULAR HOSPITAL AFFILIATED TO XIAMEN UNIV

System

An object of a system according to an embodiment is to allow people with limited hearing to smoothly perform communication in daily life.SOLUTION: A system according to an embodiment includes a voice acquisition unit, a text generation unit, a voice generation unit, and an answer generation unit. The voice acquisition unit acquires voice. The text generation unit converts the voice acquired by the voice acquisition unit into text. The speech generation unit converts the text generated by the text generation unit into speech. The response generation unit understands the content of the dialogue and generates an appropriate response.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Voice acquisition and preprocessing system and method based on software and hardware collaborative acceleration

The invention provides a voice acquisition and preprocessing system and method based on software and hardware collaborative acceleration, and relates to the technical field of voice signal processing, and the system comprises a subscriber line interface circuit module which is used for acquiring voice signals and carrying out analog-to-digital conversion on the voice signals to obtain digital voice data; the field programmable gate array module is used for preprocessing the digital voice data and packaging the preprocessed data into an Ethernet frame with a specific voice identifier; the network acceleration processing module is used for identifying an Ethernet frame of which an Ethernet type field is a specific voice identifier on a data link layer, and routing the Ethernet frame to a user mode application program; according to the system and the method, the Ethernet frame of a specific voice identifier is set, so that the voice data are processed more preferentially, kernel drive and protocol stacks are not needed, delay is reduced, and the problems that a traditional system is large in resource overhead and high in delay jitter are solved.
Owner:BEIJING HUAHUAN ELECTRONICS