Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

11 results about "Voice problem" patented technology

Voice Problems. Voice problems occur with a change in the voice, often described as hoarseness, roughness, or a raspy quality. People with voice problems often complain about or notice changes in pitch, loss of voice, loss of endurance, and sometimes a sharp or dull pain associated with voice use.

Exhibition hall intelligent guide voice question and answer optimization system based on multi-modal data

The invention discloses an exhibition hall intelligent guide voice question and answer optimization system based on multi-modal data, and relates to the technical field of intelligent voice question and answer data processing. The exhibition hall intelligent guide voice question-answer optimization system based on the multi-modal data comprises a voice signal noise interference monitoring module; a key semantic recognition monitoring module; and a voice question-answer response efficiency monitoring module. According to the method, voice signal noise interference evaluation is carried out to determine whether to take an adaptive noise reduction measure, then, based on a voice problem key semantic recognition analysis result, whether to take key semantic recognition correction is determined, and finally, based on a voice problem answer response judgment result, whether to take a triple retrieval strategy is determined. The effect of improving the response efficiency of complex voice question answering is achieved, and the problem that in the prior art, multi-modal data-based exhibition hall intelligent guide voice question answering optimization data is low in accuracy is solved.
Owner:GUANGZHOU FRONTOP DIGITAL ORIGINALITY TECH CO LTD

Voice quality optimization method and device, equipment, storage medium and product

The invention discloses a voice quality optimization method and device, equipment, a storage medium and a product, and relates to the technical field of communication, and the method comprises the steps: based on an RSRP sequence of each user equipment, recognizing each fast fading equipment with an RSRP fluctuation value exceeding a preset fluctuation threshold value from each user equipment through a preset recurrent neural network; determining a communication cell to which each fast fading device belongs according to the temporary identifier of each fast fading device, and determining the communication cell meeting a preset device scale condition and a fast fading device distribution condition as a fast fading cell; determining a voice problem cell from the fast fading cells according to the service quality identifier of each fast fading device in the fast fading cells; and voice quality optimization is carried out on voice problem equipment in the voice problem cell, and the voice problem equipment is fast fading equipment for executing voice services. Therefore, the influence of the fast fading phenomenon on the voice quality is reduced, and the voice quality is improved.
Owner:CHINA MOBILE GROUP JILIN BRANCH +1

Work order generation method and device based on large language model, equipment and medium

PendingCN122655711AEffectively identify and resolve logical conflictsIdentify and resolve logical conflictsData ingestionEngineering
The application relates to the technical field of artificial intelligence, and discloses a work order generation method and device based on a large language model, equipment and a medium, belongs to the technical field of artificial intelligence, is applied to financial and medical technology scenes, and comprises the following steps: extracting problem extraction data at least one of audio problem information and audio-video problem information from business problem feedback data, and the video problem information comprises video picture information and video voice problem information; if there is logical conflict information between the video voice problem and the video picture, correcting the logical conflict information, and generating a work order for problem feedback information at least one of the audio problem information and target conflict correction information. Through the audio-video information conflict detection and correction mechanism, logical conflicts between video voice problem information and video picture information are effectively identified and eliminated, and the accuracy of generated business problem work order content is significantly improved.
Owner:KANG JIAN INFORMATION TECH (SHENZHEN) CO LTD

Voice interaction method and device, equipment and storage medium

The invention discloses a voice interaction method and device, equipment and a storage medium, and relates to the technical field of voice processing, and the method comprises the steps: determining an optimal voice collection mode of a voice signal under the condition that the voice signal of a user is detected; acquiring a real-time voice signal of the user through the optimal voice acquisition mode; detecting whether abnormal voice signals exist in the real-time voice signals or not, wherein the abnormal voice signals are voice signals with pronunciation problems; if yes, repairing the real-time voice signal; and generating and executing a voice control instruction according to the restored voice signal. According to the invention, when the real-time voice signal of the user has the pronunciation problem, the real-time voice signal can be repaired, and the voice control instruction is generated and executed according to the repaired voice signal, so that the problem that the existing far-field voice recognition system cannot accurately recognize the voice instruction with the voice problem is solved, and the voice recognition efficiency is improved. And the voice interaction is limited.
Owner:SHENZHEN SKYWORTH DISPLAY TECH CO LTD

Speech interaction method and storage medium based on decision-making and large language model of hybrid training strategy

The present invention relates to a speech interaction method and storage medium based on a decision-making and large language model of a hybrid training strategy. The purpose of the present invention is to solve the problem that the existing large language model cannot provide accurate answers due to the lack of sufficient domain knowledge, and the use of two large language models will bring high computational costs and long response time. The process is: setting a specific answer format; constructing a decision data set; constructing a dialogue data set for dialogue questions and answers in a specific scenario; fine-tuning the large language model for the first time using full parameter fine-tuning based on the dialogue data set to obtain a large language model after the first fine-tuning; fine-tuning the large language model after the first fine-tuning for the second time using LoRA based on the decision data set to obtain a large language model after the second fine-tuning; connecting the speech recognition module and the speech synthesis module to the large language model after the second fine-tuning, processing the user's voice questions to be tested, and generating speech to interact with the user.
Owner:HARBIN INST OF TECH

Sample data construction method and device, equipment, medium and product

The invention provides a sample data construction method and device, equipment, a medium and a product, and relates to the technical field of natural language processing, and the method comprises the steps: obtaining seed voice data and a multi-dimensional attribute tag of the seed voice data; according to the prompt information, utilizing a large language model to generate a text question; the prompt information comprises at least one attribute tag in the multi-dimensional attribute tags; generating a voice question corresponding to the text question according to the at least one attribute tag of the text question and the multi-dimensional attribute tag of the seed voice data; according to the text question corresponding to the voice question and the multi-dimensional attribute tag, generating a voice response matched with the voice question by using a large language model; and according to the voice problem and the voice response, constructing a corresponding voice interaction sample pair. According to the method, the voice interaction sample pairs with diversified contents and specific attributes are automatically generated on a large scale, the cost and time of sample data construction are reduced, and the diversity of the sample data is improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Speech question answering method based on audio retrieval enhancement generation and double-layer rearrangement

The application discloses a voice question and answer method based on audio retrieval enhancement generation and double-layer rearrangement, which comprises the following steps: dividing a plurality of original documents contained in a knowledge base document according to a plurality of blocking strategies to obtain a plurality of types of document blocks divided based on different blocking strategies; converting a voice question input by a user end into a user question text and generating a user query vector based on the user question text; determining the relevance ranking results of the plurality of types of document blocks through a plurality of retrieval methods according to the user query vector; reordering the plurality of document blocks ranked higher than a preset ranking threshold in the relevance ranking results according to the semantic relevance between the user question text and the plurality of document blocks ranked higher than the preset ranking threshold, to obtain a reordering result; generating a question answer text according to the user question text and the corresponding document blocks in the reordering result, and generating a voice answer based on the question answer text and returning the voice answer to the user end.
Owner:BEIJING JIAOTONG UNIV

An embedded audio processing system and method supporting AI language enhancement

The application relates to the technical field of audio processing technology, in particular to an embedded audio processing system and method supporting AI language enhancement, which comprises a data acquisition module, an audio analysis module, an AI enhancement module, an audio processing module and an interactive output module. The modules of the system cooperate, the data acquisition module collects audio, the analysis module accurately classifies and calibrates, the enhancement module improves the quality of fuzzy audio, the processing module generates corresponding content, the output module outputs, accurate, efficient and high-quality voice interaction is realized, the speed of vehicle-mounted AI voice answering voice questions is improved, and the accuracy of vehicle-mounted AI voice answering voice questions is improved.
Owner:BEIJING HANGYU XINGZHOU TECHNOLOGY CO LTD

Fine-grained multi-dimensional speech evaluation method and system based on reinforcement learning and thinking chain

The invention discloses a fine-grained multi-dimensional speech evaluation method and system based on reinforcement learning and a thinking chain, and belongs to the technical field of artificial intelligence and speech signal processing. Batched structured evaluation problems are designed for a plurality of preset voice quality evaluation dimensions, a training set containing voice-problem pairs is obtained, and a pre-trained large language model is adopted to carry out thinking chain labeling; performing supervision fine tuning on a basic audio language model by adopting the training set and the annotation data thereof, and initializing an independent evaluation sub-model for each evaluation dimension based on the fine tuning model; the training set is divided into subsets according to evaluation dimensions, sub-models of all dimensions are trained under a reinforcement learning framework, and finally a unified multi-dimensional speech evaluation model is formed after weighted fusion and used for reasoning a speech-problem pair to be evaluated and outputting a thinking chain reasoning text and a binary evaluation result. According to the invention, a high-precision evaluation result containing logical reasoning can be output, and automatic and explainable synthetic speech quality evaluation is realized.
Owner:ZHEJIANG UNIV +1

Embedded audio processing system and method supporting AI language enhancement

The invention relates to the technical field of audio processing, in particular to an embedded audio processing system and method supporting AI language enhancement, and the system comprises a data collection module, an audio analysis module, an AI enhancement module, and an audio processing module. And an interactive output module. The modules of the system cooperate, the data acquisition module collects audios, the analysis module accurately classifies and calibrates the audios, the enhancement module improves the fuzzy audio quality, the processing module generates corresponding contents, and the output module outputs the contents, so that accurate, efficient and high-quality voice interaction is realized, and the speed of answering voice questions by vehicle-mounted AI voice is improved. And the accuracy of answering the voice question by the vehicle-mounted AI voice is improved.
Owner:BEIJING HANGYU XINGZHOU TECHNOLOGY CO LTD

system

The system according to the embodiment aims to extract text from an image, play it aloud, and provide appropriate answers to questions from the user. [Solution] A system according to an embodiment includes an image acquisition unit, an analysis unit, a voice synthesis unit, a playback unit, a question reception unit, an answer generation unit, and an answer provision unit. The image acquisition unit acquires an image. The analysis unit analyzes the image acquired by the image acquisition unit and converts it into text data. The voice synthesis unit converts the text data converted by the analysis unit into voice. The playback unit plays back the voice generated by the voice synthesis unit. The question reception unit accepts questions from users. The answer generation unit generates answers to questions accepted by the question reception unit. The answer provision unit provides the answers generated by the answer generation unit by voice.
Owner:SOFTBANK GROUP CORP