Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

415 results about "Speech output" patented technology

Like speech input, speech output is a familiar and natural form of communication, so it is also an appropriate complement in a character-based interface. However, speech output also has its liabilities. In some environments, speech output may not be preferred or audible.

Video generation method and interaction method based on digital human, and device, storage medium and program product

Provided in the embodiments of the present application are a video generation method and interaction method based on a digital human, and a device, a storage medium and a program product. In the embodiments of the present application, text-to-speech processing is performed on the basis of voice features of a user and an emotion label, speech-to-expression processing is performed on the basis of a mapping relationship between the voice features of the user and expression coefficients, and a digital human model is rendered on the basis of speech signals and the expression coefficients, so as to obtain video data of the digital human model. Thus, voice features of a user are accurately simulated, so as to ensure that a speech output of a digital human sounds natural and is also highly personalized, thereby realizing personalized driving of the digital human, and improving the realism of the digital human in terms of voice and dynamic images. Thus, the user experience is improved, and the interactivity of the digital human and the authenticity and immersion are enhanced.
Owner:TAOBAO CHINA SOFTWARE

Multi-mode-based AI digital human intelligent interaction method, system and equipment

The invention relates to the technical field of computer vision and human-computer interaction, and discloses an AI digital human intelligent interaction method, system and equipment based on multiple modalities, and the method comprises the steps: pre-awakening a digital human when a human face is detected, and further thoroughly awakening the digital human based on recognized preset voice information or preset gesture information; voice and video information of a user in the interaction process is obtained, a keyword extraction result, a gesture recognition result and an emotional state tag are generated, a pre-constructed knowledge base is utilized to retrieve related information, a big language generation model module is combined to generate an answer text, and the answer text is input into a preset voice synthesis model to generate emotional voice output. And based on the current emotional state label of the user, driving the digital human animation to be output in an emotional manner. According to the method and the system, the digital human for understanding the emotion of the user, generating personalized answers, providing voices with rich emotions and displaying natural expressions and actions can be created, better interaction with the user can be realized, and more humanized and effective services can be provided.
Owner:BEI JING WAN JIE SHU JU KE JI YOU XIAN ZE REN GONG SI WU HAN FEN GONG SI +1

Digital human interaction system and method based on multi-modal emotion recognition

ActiveCN121116129ASemantic analysisSpeech analysisInteractive modelingData stream
The embodiment of the invention provides a digital human interaction system and method based on multi-modal emotion recognition, and belongs to the technical field of digital human interaction. The system comprises a multi-modal sensing module used for collecting multi-modal data and preprocessing the multi-modal data to generate a standardized data stream; the cross-modal fusion and emotion recognition module is used for carrying out interactive modeling on the multi-modal features and outputting a current emotion label and emotion intensity; the reaction planning module is used for generating a composite reaction strategy; and the digital human rendering module is used for mapping the composite reaction strategy into control signals corresponding to the voice, the facial expression and the action respectively, and driving a digital human to execute corresponding voice output, facial expression change and limb action through the control signals so as to realize interaction. According to the method, multi-modal data are deeply fused through the cross-modal graph neural network and comparative learning, the weight is dynamically adjusted in combination with the modal confidence, and the emotion recognition accuracy and robustness are improved.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Voice text bidirectional conversion method and device, equipment and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice text bidirectional conversion method, device, equipment and medium, and the method comprises the steps: respectively executing voice recognition or voice synthesis operation according to the type of input information; for the voice information, noise suppression parameters are generated in combination with the lip movement video data, noise reduction processing is executed, and the recognition accuracy is improved; for text information, a pre-generated speaker style vector is obtained, the vector is cited in the speech synthesis process to generate natural personalized speech, and lip movement information and tactile feedback which are synchronous with speech output are generated. According to the method, complex noise is suppressed by fusing lip movement data, personalized voice is generated by using the style vector, and lip movement and touch information is output, so that bidirectional real-time conversion of voice and text in a complex environment is realized, and recognition accuracy, voice naturalness and interaction synchronism are effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Lightweight intelligent traditional Chinese medicine inquiry system and construction method thereof

The invention relates to the field of artificial intelligence medical application, and discloses a lightweight intelligent traditional Chinese medicine inquiry system and a construction method thereof, and the system comprises a multi-dialect adaptive speech recognition module, a traditional Chinese medicine intelligent dialogue large language model module, a natural speech synthesis module, and a continuous learning mechanism module. The multi-dialect adaptive speech recognition module is used for converting dialect speech input of a patient into a standard text; the traditional Chinese medicine intelligent dialogue big language model module is the core of the system and is used for carrying out natural language understanding, dialectical reasoning and inquiry dialogue generation, and the natural speech synthesis module is used for converting a text response generated by the system into speech output; and the continuous learning mechanism module realizes continuous optimization of the large language model through incremental learning architecture and clinical feedback integration. According to the method, while the professional traditional Chinese medicine diagnosis capability is maintained, the calculation complexity is remarkably reduced, and the universality and sustainable development capability of system application are improved.
Owner:SUZHOU ANGSHENG NETWORK TECHNOLOGY CO LTD

Privacy information desensitization method and system for voice generation type large model

The invention discloses a privacy information desensitization method and system for a voice generation type large model, and relates to the technical field of artificial intelligence. Input voice data is discretized, Gaussian noise disturbance sensitive features are injected, and a cross-modal voice generation model is constructed in combination with three-stage training; and meanwhile, a cross-modal privacy enhancement mechanism is applied to detect and fuzzify sensitive information in real time in an output stage. According to the invention, the adaptability of a large-scale voice generation model in a privacy protection scene is improved, and comprehensive protection of user privacy is realized. And moreover, the capability of extracting user privacy information by an adversarial attacker is effectively limited, and the risk of sensitive information leakage in the data transmission, storage and generation process of the voice generation model is reduced. On the premise that privacy is ensured, the voice generation model can still keep high-quality generation performance, generated voice output has high naturalness and accuracy, and actual application requirements are met.
Owner:ZHEJIANG UNIV

Voice interaction method and system of mobile digital human

The invention relates to the technical field of voice interaction, in particular to a voice interaction method and system for a mobile digital human, and the method comprises the steps: inputting data through a voice collection channel, reading a sound wave signal time sequence, extracting a waveform amplitude value, and recording a timestamp, and synchronously extracting and coding time domain and frequency domain changes implied in a conventional voice input signal. The method improves the capturing precision of tiny mutation characteristics in voice signals, introduces a generation and discrimination mechanism of an adversarial neural network through the generation and screening process of a mutation candidate characteristic group, enables a generator to dynamically generate diversified amplitude and frequency change modes, further improves the expressive force and naturalness of voice output, and improves the voice recognition accuracy. By performing weighting coefficient addition processing on the product of the amplitude and the frequency change, the voice output intensity can be reflected more accurately, the response performance of the voice is optimized, and the response efficiency and the stability of the system are further improved by setting the bidirectional mapping and the freezing state flag.
Owner:XIAN WUKONG INTELLIGENT TECH CO LTD

Intelligent health robot combining intelligent conversation and health intervention

The invention belongs to the field of intelligent voice, and particularly relates to an intelligent health robot combining intelligent dialogue and health intervention, which comprises a voice recognition module and a rehabilitation guidance module, and is characterized in that the voice recognition module comprises a voice acquisition unit, a voice compensation unit and a multi-mode recognition unit; the method comprises the following steps: acquiring user voice by a voice acquisition unit, compensating formant features and pause discrimination by a discrimination filter, and obtaining a to-be-recognized voice feature space; correcting voice missing words and tones through a voice compensation unit to obtain a corrected to-be-recognized voice feature space; the multi-modal recognition unit outputs the voice text and the demand intention space according to the voice text and the demand intention space, and obtains an intention emotion state space by combining a preset evaluation interval; and finally, the rehabilitation guidance module calls an enhanced reasoning strategy library to generate a rehabilitation strategy intervention scheme according to the demand intention space and the emotional state space, and performs real-time voice output to realize intelligent dialogue and health intervention.
Owner:NANJING MEDICAL UNIV

Information processing system, information processing method and program

An information processing system, method, and program for conducting interviews with applicants is provided. [Solution] The method includes starting a conference session in an interview preparation mode, and switching the conference session to the interview mode when an instruction to switch to the interview mode is received from a user, starting an interview with an avatar selected in the interview preparation mode as the interviewer, displaying an avatar in the interview preparation mode, outputting a first utterance for breaking the ice as the avatar's speech, outputting a second utterance for setting the avatar as the avatar's speech, and displaying selectable options for multiple setting items of the avatar, and accepting an avatar selection including a first response to the first utterance and a second response or selection of an option to the second utterance, and a switching instruction, a first display control step of displaying an avatar changed in accordance with the avatar selection, and a second display control step of displaying the changed avatar.
Owner:BIZREACH INC

Vehicle intelligent interaction method and system based on AI

The embodiment of the invention provides an AI-based vehicle intelligent interaction method and system, and belongs to the field of artificial intelligence interaction. The method comprises the following steps: dynamically judging a cognitive load level in a current driving scene through collaborative perception of a driver state and an environment state; determining a target man-machine interaction strategy of the vehicle based on the cognitive load level; wherein the target man-machine interaction strategy comprises a voice output mode, information density of screen display content and a feedback prompt mode; determining a corresponding trigger rule based on a switching relationship between the current man-machine interaction strategy and the target man-machine interaction strategy; and responding to a trigger signal of the determined trigger rule, executing man-machine interaction strategy switching, and executing vehicle interaction based on the switched target man-machine interaction strategy. According to the scheme of the invention, closed-loop adaptive control of a human-vehicle interaction strategy on a cognitive state is realized.
Owner:SHENZHEN YOUBIKANG TECH CO LTD

Real-time voice interaction and adaptive content generation system based on RTC and AIGC

The invention belongs to the technical field of artificial intelligence, and particularly relates to a real-time voice interaction and adaptive content generation system based on RTC and AI GC, which comprises the steps of capturing voice input of a user through a microphone of equipment, and converting the voice input into text information by using a voice recognition technology; the recognized text input is transmitted to a natural language understanding module, semantic analysis is carried out, and user intention and information are recognized; according to the intention and demand of the user, the AI GC technology is used for dynamically generating personalized content; and the generated content is optimized in real time according to the feedback in interaction, the demand change of the user is adapted, the generated text content is converted into voice to be output and provided for the user, and the effects of meeting the demands of different types of users and promoting real-time interactive propagation and automatic content generation are achieved.
Owner:SHENZHEN UASCENT TECH CO LTD

Interactive English teaching method and system based on AI vision

The invention relates to the technical field of intelligent education, in particular to an interactive English teaching method and system based on AI vision, and the method comprises the steps: obtaining the voice and mouth dynamic image frame segments of a student, extracting an offset frame segment to generate a semantic motion disjunction interval, analyzing the time sequence consistency of a behavior signal and a motion cause word triggering frame segment, and calculating the motion matching degree. Time synchronism of the speech output and the visual action is evaluated. According to the invention, through real-time analysis of the voice and mouth dynamic image of the student, the alignment offset of the pronunciation and the mouth action is identified, and in combination with behavior signals such as eye fixation and head rotation of the student, the interactivity between the language action and the behavior of the learner is comprehensively evaluated, so that the efficient alignment of the language expression and the standard template is ensured; by analyzing the distribution frequency and the matching degree of the voice behaviors, a personalized teaching path is optimized. In addition, delay frames and contour offset are deeply analyzed, the synchronism of voice and visual actions is enhanced, and the real-time feedback and interaction effect in the learning process is improved.
Owner:HUNAN INST OF INFORMATION TECH

Voice interaction method and system based on quantum heuristic algorithm, and computer equipment

The invention relates to an artificial intelligence technology, and discloses a voice interaction method and system based on a quantum heuristic algorithm, and computer equipment, and the method comprises the steps: carrying out the preprocessing of a voice signal after receiving voice input, and extracting voice features; taking different combinations of the voice features as superposition of quantum states, dynamically adjusting the selection probability of the voice features based on quantum revolving door operation, and iterating and screening feature subsets with the probability meeting a preset requirement; inputting the screened feature subset into a character recognition model for training, and generating a character recognition result; performing semantic understanding on a character recognition result based on a large language model; after the text information is obtained, extracting voice feature parameters based on the text information; the voice feature parameters are optimized and adjusted through quantum evolution operation; and based on the optimized voice feature parameters, performing voice output by adopting a voice synthesis algorithm. The invention further discloses a computer readable storage medium. The invention aims to improve the efficiency and accuracy of voice interaction.
Owner:深圳玄源科技有限公司

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Information acquisition system and method suitable for visually impaired people

The invention discloses an information acquisition system and method suitable for visually impaired people. The system comprises a hardware layer, a driving layer, an operating system layer, a middleware layer and an application layer which are sequentially in communication connection through interfaces. The application layer comprises a voice interaction module, a large model integration module and a braille conversion module; the hardware layer is used for collecting a voice input instruction and outputting a braille character sequence and a voice output result; the voice interaction module is used for converting a voice input instruction into a text input instruction and converting a text output result into a voice output result; the large model integration module is used for analyzing the text input instruction and generating a text output result; and the braille conversion module is used for converting the text output result into a braille character sequence, so that a hardware layer outputs the braille character sequence based on a timestamp alignment technology. Therefore, based on powerful semantic understanding and information retrieval of the large model integration module, accurate information acquisition is realized, and a five-layer decoupling architecture is adopted, so that the convenience of later system maintenance is improved.
Owner:ZHEJIANG LAB

Medical pre-inquiry real-time interaction method and system, terminal and medium

The invention belongs to the field of artificial intelligence, and particularly relates to a medical pre-inquiry real-time interaction method and system, a terminal and a medium. Screening medical knowledge contexts similar to the consultation text from a professional medical knowledge base, and recording the medical knowledge contexts as related texts; based on a pre-constructed prompt template, constructing a current consultation prompt through the consultation text and the related text; inputting the current consultation prompt into the large language model, and outputting consultation reply content; performing voice synthesis on the consultation reply content to generate consultation reply voice data, storing the consultation reply voice data in the buffer area, extracting reply voice data from the main buffer area for voice output, and transferring data in the auxiliary buffer area to the main buffer area when the data in the main buffer area is less than a threshold value; speech synthesis and speech output are asynchronously processed. According to the invention, the professionality and accuracy of medical pre-inquiry are improved, the playing lagging caused by waiting for synthesis is effectively avoided, and the interaction experience effect is improved.
Owner:山东浪潮智能生产技术有限公司

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

Questionnaire apparatus

To improve the accuracy of a questionnaire collected from a user who drove a test-drive vehicle.SOLUTION: A questionnaire apparatus which can be disposed in a test-drive vehicle includes an input unit to which voice can be input, an output unit which can output voice, and a control unit. The control unit outputs, while a user is driving the test-drive vehicle, voice for a questionnaire to a user via the output unit at a predetermined timing, and executes processing to receive answers of the user to the questionnaire via the input unit.SELECTED DRAWING: Figure 4
Owner:TOYOTA JIDOSHA KK

Automobile seat motor intelligent sliding rail adjusting system based on AI voice interaction

The invention discloses an automobile seat motor intelligent sliding rail adjusting system based on AI voice interaction, and the system comprises the following modules: a local voice wake-up module which is used for monitoring the voice input of a user, and activating a voice recognition module after detecting a preset wake-up keyword; the voice recognition module is used for generating a standardized text instruction; the semantic modeling module is used for executing word vector modeling processing on the standardized text instruction; the control parameter mapping module is used for inputting the semantic feature vector into a trained neural network model; the slide rail state detection module is used for collecting slide rail state information in real time; the voice synthesis module is used for converting the content in the prompt information field into a voice output content template; the vehicle-mounted audio output module is used for generating broadcast content; and the voice interaction log module is used for recording information involved in the voice interaction process. Semantic modeling and control mapping are fused, and the intelligent voice-driven seat adjusting system is constructed.
Owner:ZHEJIANG WANZHAO AUTO PARTS CO LTD

Vending machine intelligent shopping guide method and system based on voice interaction

The invention discloses a vending machine intelligent shopping guide method and system based on voice interaction. The objective of the invention is to realize accurate recognition of user demands through voice recognition and natural language processing technologies, and perform intelligent recommendation in combination with commodity big data. According to the method, through integration of a high-sensitivity microphone and voice recognition, a voice instruction of a user is converted into text information, and a purchase intention behind the user is analyzed. Meanwhile, detailed information of commodities sold by the vending machine is collected, and after standardization processing and feature extraction, the detailed information serves as training data of the deep learning model. Personalized commodity recommendation can be generated according to user requirements through a trained and optimized model, and the personalized commodity recommendation is fed back to a user through a high-resolution display screen or voice output. According to the invention, the shopping experience of the user is improved, and the purchase conversion rate is improved through intelligent recommendation.
Owner:SHANGHAI QUZHI NETWORK TECH CO LTD

Inner speech iterative learning loop

Methods and systems are disclosed for iteratively training a user and a ML model to produce accurate inner speech outputs. The methods and systems access a ML model and perform a first training iteration in which EMG data corresponding to inner speech is processed by the machine learning model to decode the EMG data into a set of predicted phonemes, phoneme sounds, words or phrases. The methods and systems present the set of predicted phonemes, phoneme sounds, words or phrases to the user and form a first set of training data comprising the set of predicted phonemes, phoneme sounds, words or phrases, the EMG data, and the set of specified phonemes, phoneme sounds, words or phrases as ground truth information. The methods and systems update parameters of the ML model based on the first set of training data prior to starting a second training iteration.
Owner:SNAP INC

Sign language recognition method based on double-arm electromyographic signals

The invention discloses a sign language recognition method based on double-arm electromyographic signals, and belongs to the technical field of human-computer interaction and biological signal processing. The method comprises the steps that multi-channel electromyographic signals generated when sign language gestures are executed are synchronously collected through electromyographic arm rings worn on the left forearm and the right forearm of a user; performing preprocessing and feature extraction on the signal to obtain a time sequence feature sequence of left and right arms; the time sequence feature sequence is input into a pre-trained two-arm collaborative recognition model, and the model outputs a sign language gesture recognition result by fusing the spatial-temporal features of the left arm and the right arm; and finally, the recognized text information is converted into voice to be output. According to the method, the cooperation and time sequence relation of the double-arm electromyographic signals is creatively utilized, the problems that a traditional visual recognition method is greatly interfered by the environment, privacy is invaded, and double-hand linkage complex gestures cannot be effectively analyzed through single-arm electromyographic recognition are solved, and natural, accurate and real-time recognition and translation of the double-hand sign language gestures are achieved.
Owner:宋飞 +1

Doctor inquiry training system based on generative AI technology

The invention discloses a doctor inquiry training system based on a generative AI technology, which comprises a medical document database and a processing unit in information interaction with the medical document database, the voice input module, the voice output module, the training module, the federation module, the vector database, the embedding module, the inspection module and the evaluation module are in information interaction with the processing unit. The system has the advantages that a real inquiry scene can be dynamically simulated to provide efficient and accurate inquiry ability training and evaluation support, and the problems that an existing inquiry training system is difficult to simulate real patient communication, the inquiry process cannot be deeply evaluated and continuously improved, a cross-hospital cooperative training mechanism is lacked, and the training effect is poor can be solved.
Owner:THE FIRST AFFILIATED HOSPITAL OF SUN YAT SEN UNIV

Voice interaction product test method and device based on artificial intelligence, and product

The invention discloses a voice interaction product test method and device based on artificial intelligence and a product, and the method comprises the steps: generating a corresponding question according to a selected question class through a first artificial intelligence large model; the questioning question is converted into questioning voice; the voice interaction product needing to be tested responds to the questioning voice and outputs corresponding answering voice; the answer voice is converted into a character answer; and evaluating the matching accuracy of the character answers and the corresponding questions by using the second artificial intelligence large model, and outputting the accuracy as a test result. The invention discloses a voice interaction product test method and device based on artificial intelligence and a product, and the method comprises the steps: generating a questioning question through a first artificial intelligence large model, evaluating the matching accuracy of a character answer and the corresponding questioning question through a second artificial intelligence large model, and outputting a test result. The system has the characteristics of high automation degree of voice interaction product testing, full testing and accurate testing effect.
Owner:BEIJING POLYTECHNIC

Controllable voice generation method and device based on multi-agent dynamic scheduling

The invention discloses a controllable voice generation method and device based on multi-agent dynamic scheduling, and belongs to the technical field of voice synthesis, and the method comprises the steps: constructing a multi-agent voice generation frame comprising a central scheduling module, an identity agent, an emotion agent and an environment agent; a central scheduling module analyzes a user instruction and outputs a structured task plan to drive each agent to generate primary voice output of identity, emotion and environment dimensions, and a cooperation cost matrix between the agents is constructed according to the primary voice output; the optimal execution path of the multiple agents is solved through an optimization algorithm based on the cost matrix, the agents are executed in a cascading mode based on the optimal execution path, and finally sound mixing is completed and high-quality voice is output. According to the method, the naturalness, the semantic consistency and the overall quality of the synthesized voice in a complex scene can be remarkably improved, and the method is suitable for various man-machine interaction application scenes such as intelligent voice assistants, virtual digital humans, immersive entertainment, barrier-free voice services, personalized content creation and the like.
Owner:ZHEJIANG UNIV OF TECH

Multi-modal interaction method and device for synchronously displaying voice and sign language

PendingCN120339476ASpeech analysisBiological modelsSpeech rhythmSemantic feature
The invention relates to a multi-modal interaction method and device for synchronous speech and sign language display, and belongs to the technical field of speech image data processing.The multi-modal interaction method for synchronous speech and sign language display comprises the steps that distance measurement is determined based on semantic difference loss between sign language semantic feature vectors and speech rhythm feature vectors, and the distance measurement is obtained; performing time synchronization on the sign language semantic feature vector and the voice rhythm feature vector based on distance measurement and a DTW algorithm; fusing the emotion feature vector with the sign language semantic feature vector and the voice rhythm feature vector after time synchronization to generate a multi-modal feature sequence; and generating sign language actions, facial expressions and lip shapes based on the multi-modal feature sequence, and controlling the digital human to display. According to the invention, the emotion information is expressed while the sign language action of the digital human is consistent with the voice output, and the user experience is improved.
Owner:WUHAN UNIV OF TECH

Noise isolation and target sound enhancement system based on specific voice pre-storage

The invention relates to the technical field of voice processing and enhancement, in particular to a noise isolation and target voice enhancement system based on specific voice pre-storage, and the system extracts pre-stored features from a pre-stored target voice database module through collecting environment audio signals in real time and carrying out frame segmentation and frequency domain transformation; adaptive modulation of amplitude and phase of each sub-band signal is realized by combining energy sensing sub-band modulation, then environmental noise energy is identified through dynamic noise isolation and multi-stage suppression is carried out, and finally continuous and clear target voice output is generated through enhanced fusion. The method can achieve the effective enhancement and environmental noise suppression of the target voice in a complex and changeable environment, improves the voice recognition precision and definition, supports the dynamic switching of multiple teachers and multiple classes, adapts to the target voiceprint change, and achieves the high-robustness voice enhancement and noise reduction in teaching, conference and multi-target scenes.
Owner:UNIV OF JINAN

Voice interaction control method and device, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice interaction control method, device, equipment and medium, and the method comprises the steps: obtaining multi-mode perception data, and constructing a dynamic scene model; session feature parameters are extracted and input into a decision model to generate a strategy parameter set; a dialogue process and expected voice output are generated according to the strategy parameter set and voice characteristics of the voice interaction device; controlling the voice interaction device to execute voice output and collect actual voice output; monitoring an error between the actual voice output and the expected voice output, and performing compensation adjustment; and collecting feedback data after the voice interaction session is ended, and optimizing the model and the control parameters according to the feedback data. According to the method, closed-loop control of perception, decision, output and feedback is established, so that the voice interaction equipment can dynamically adjust voice output according to a real-time environment and user characteristics and continuously perform self-learning optimization, and the accuracy and naturalness of voice interaction are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Work machine and operation support system

The present disclosure provides a work machine capable of setting a restricted area on the basis of an instruction in a natural language. This work machine (shovel 100) comprises an attachment (AT) for work, an information acquisition unit (301), a language conversion unit (303), an instruction acquisition unit (302), and an area setting unit (306). The information acquisition unit (301) acquires information pertaining to the posture of the attachment (AT) and the surrounding environment. The language conversion unit (303) converts the information acquired by the information acquisition unit (301) into a natural language. The instruction acquisition unit (302) acquires an instruction in the natural language from an operator. The area setting unit (306) sets restricted areas (RA1 to RA10) in which entry or operation speed is restricted at a work site. The area setting unit (306) sets the restricted areas (RA1 to RA10) on the basis of a result obtained by interpreting, using a language model (LM), the instruction acquired by the instruction acquisition unit (302) and the information converted into the language by the language conversion unit (303), or on the basis of a result obtained by interpreting the instruction acquired by the instruction acquisition unit (302) using the language model (LM), in combination with the information acquired by the information acquisition unit (301).
Owner:SUMITOMO HEAVY IND LTD

Voice processing system, voice processing method, and recording medium in which voice processing program is recorded

A voice processing system includes an acquisition processing unit that acquires voices uttered by users and input to respective microphones of a plurality of audio devices arranged in the same space, a determination processing unit that determines a degree of similarity among a respective plurality of voices acquired from the plurality of audio devices, and an output processing unit that outputs a specific first voice from among the plurality of voices to a voice recognition processing unit and a voice synthesis processing unit in a case where the degree of similarity among the plurality of voices is equal to or greater than a threshold value.
Owner:SHARP KK