Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Live voice" patented technology

Live voice synthetization

PendingUS20260097308A1Video gamesSpeech synthesisLive voiceTelecommunications
Systems, devices, methods, and machine-readable media configured to provide voice synthetization in a multiplayer video game are provided. A system can include a multiplayer video game including a character selection interface through which a player selects a character to represent them in playing the video game, and a voice model trained to convert audio from the player directly into audio in a voice of the character and provide an output that includes the audio in the voice of the character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Digital human live broadcast voice interaction system fused with emotion calculation

ActiveCN121393436ASpeech recognitionLive voiceData stream
The invention relates to the technical field of digital human voice interaction, and discloses a digital human live voice interaction system fused with emotion calculation. The system constructs an emotional response time window by acquiring a user voice input data stream, the starting point of the emotional response time window is the end moment of the user voice input data stream, and the end point is obtained by subtracting the necessary duration of voice response synthesis from the preset maximum response cut-off moment. The system obtains an emotional state vector of the user in real time and judges whether the emotional state vector reaches an emotional intensity threshold value or not; and if so, predicting the total generation duration of the digital human voice response. And when the residual duration of the emotional response time window is equal to the total generation duration, taking the window starting point as a voice response starting moment, and controlling the digital human to start voice response generation. According to the system, the voice response matched with the emotional state of the user is generated through refined time window management and emotional state perception, invalid response and interaction delay are reduced, and the naturalness of digital human live broadcast voice interaction and the user experience are improved.
Owner:BEIJING ZHONGSHENGSHENG DIGITAL TECHNOLOGY CO LTD

Agricultural product wholesale order automatic generation method and device, equipment and medium

The invention discloses an agricultural product wholesale order automatic generation method and device, equipment and a medium, and relates to the technical field of order management. The method comprises the following steps: firstly, receiving a plurality of field sound signals collected by a sound pickup array on an agricultural product selling communication field in real time, and carrying out voice signal extraction processing by applying an empirical mode decomposition technology, a sound source positioning technology and a voiceprint matching technology in real time so as to obtain speech signals of a seller and a buyer; then, the extraction result is integrated into an agricultural product selling communication dialogue text data stream in real time, the text data stream is imported into a large language model in real time to carry out semantic understanding and entity information extraction processing, then a matched agricultural product wholesale order template is retrieved according to the extraction result, and content filling and dynamic updating are carried out on the template; and generating a new agricultural product wholesale order, and finally pushing the order to the seller terminal equipment in real time and outputting and displaying the order in real time, so that the order generation efficiency can be improved, the order information is ensured to be generated correctly, and the seller experience is improved.
Owner:CHENGDU GAUSS ZHIDA INFORMATION TECHNOLOGY CO LTD

Live voice synthesis with voice mixing

PCT designated stageWO2026075705A1Video gamesSpeech synthesisLive voiceTheoretical computer science
Systems, devices, methods, and machine-readable media configured to provide voice inference in a video game are provided. A video game system can include an encoder configured to generate a first encoding representative of physical characteristics of a specified entity, a similarity operator configured to determine similarity values between (i) corresponding stored encodings of multiple characters, the stored encodings representative of physical characteristics of respective characters of the multiple characters and (ii) the first encoding, identify a selected character from the multiple characters based on the similarity values, and provide an identifier of the selected character, a voice database configured to provide audio or a spectrogram of the selected character, and a video game configured to provide the audio of a player-selected character in a voice of the selected character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Comprehensive diagnosis method and system for ultrahigh-pressure converter station valve cooling system

The invention relates to the technical field of ultrahigh-pressure converter station valve cooling system monitoring, in particular to an ultrahigh-pressure converter station valve cooling system comprehensive diagnosis method and system.The method comprises the steps that on-site sound signals are obtained and analyzed to generate a first analysis result, and a second analysis result is generated by combining meter image analysis; and the water leakage condition of the spraying pump is judged, a third analysis result is generated, and dynamic monitoring and deterioration trend analysis are conducted on abnormal points. According to the invention, the operation state and parameter abnormity of the valve cooling system can be rapidly identified through intelligent analysis of voiceprint features and meter images, a water leakage detection and abnormal point monitoring mechanism is introduced, the fault identification capability and the early warning level are improved, the monitoring efficiency and the intelligent degree are significantly improved, and the problem of low efficiency caused by dependence on manual operation in the prior art is solved.
Owner:GUANGZHOU BUREAU CSG EHV POWER TRANSMISSION

Audio and video stream processing method, device and equipment for cross-platform live voice chat and medium

The application relates to the technical field of network live broadcast, and discloses a cross-platform live broadcast microphone connection audio and video stream processing method and device, electronic equipment and a storage medium, the method saves an audio and video file in a second live broadcast platform server, extracts audio frame information and video frame information from the audio and video file according to real-time microphone connection audio and video stream playing abnormal event information, and the cause of the real-time microphone connection audio and video stream playing abnormal event can be automatically and quickly determined according to the audio frame information and the video frame information, without manual participation, so that the cause determination efficiency of the audio and video stream playing abnormality is improved.
Owner:GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD

Offline voice interaction monitoring device and system

The utility model discloses an off-line voice interaction monitoring device and system, and the device comprises a holder, a communication unit, an energy storage unit, an audio unit, an off-line voice recognition module, and a control unit. The holder is integrated with a camera and has a preset rotation angle; the communication unit is used for performing remote communication with a back-end platform; the energy storage unit is used for supplying power to equipment; the audio unit comprises a loudspeaker and a pickup microphone; the loudspeaker is used for voice broadcasting; the pickup microphone is used for acquiring field sound; and the off-line voice recognition module is used for responding to sound acquired on site according to a preset off-line voice instruction. The technical scheme of the utility model breaks through the dependence of the traditional scheme on network and wired operation, improves the operation convenience and reliability in a complex scene, avoids the potential safety hazard caused by data transmission, and is suitable for scenes with high requirements on real-time performance and autonomy, such as electric power inspection, emergency rescue, field operation and the like.
Owner:SHENZHEN TELICANG TECH CO LTD

Remote monitoring system and monitoring method for sound of key process actions in rolling processes

ActiveCN116274421BLive voiceRemote control
The present application relates to a kind of based on sound sensor's rolling process key process action remote monitoring system and monitoring method, belong to the remote monitoring technical field of sound signal.The system includes audio signal acquisition module, data communication transmission module, audio signal filtering processing module, audio signal storage module, audio signal display module, audio signal analysis module and comprehensive system management module.The system passes through the audio signal of important site in rolling process field acquisition, solves the problem that operator cannot hear the sound of field after remote control, realizes the monitoring and analysis to field situation, is convenient for operator to appear in production process possibly Problem is promptly found, and the intelligentization of operation control is promoted.
Owner:UNIV OF SCI & TECH BEIJING

News live voice anomaly real-time monitoring correction method and system thereof

ActiveCN122135742ATelevision system detailsSpeech analysisLive voiceSpeech error
This invention relates to the field of speech signal processing and live broadcast quality control technology, and discloses a method and system for real-time monitoring and correction of speech anomalies in news live broadcasts. The method includes: extracting multi-dimensional acoustic feature vectors from the announcer's speech signal in real time by frame segmentation and storing them in a circular feature buffer; performing parallel detection of abnormal pauses, speech errors, and sudden volume changes based on the feature vectors; pushing prompts through in-ear monitors, highlighting deviation positions on the teleprompter, or triggering a dynamic compressor to smooth the volume according to the type of abnormal event; recording abnormal events to generate a broadcast quality analysis report and using this report to correct the detection threshold. This invention achieves millisecond-level real-time perception and immediate correction assistance for broadcast anomalies, significantly shortening the response delay of traditional manual monitoring.
Owner:GUIZHOU NORMAL UNIVERSITY

Live voice synthesis with voice mixing

Systems, devices, methods, and machine-readable media configured to provide voice inference in a video game are provided. A video game system can include an encoder configured to generate a first encoding representative of physical characteristics of a specified entity, a similarity operator configured to determine similarity values between (i) corresponding stored encodings of multiple characters, the stored encodings representative of physical characteristics of respective characters of the multiple characters and (ii) the first encoding, identify a selected character from the multiple characters based on the similarity values, and provide an identifier of the selected character, a voice database configured to provide audio or a spectrogram of the selected character, and a video game configured to provide the audio of a player-selected character in a voice of the selected character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Vehicle-based remote interaction method, device, equipment and computer program product

The invention discloses a vehicle-based remote interaction method, device and equipment and a computer program product, and the method comprises the steps: obtaining a human body time sequence image and / or a vehicle time sequence image through extraction according to field image data, and obtaining voice information and / or road traffic noise through recognition according to field sound data; according to the sound source position, fusing the voice information and / or the road traffic noise with the human body time sequence image and / or the vehicle time sequence image to obtain multi-modal human body data and / or multi-modal vehicle data, inputting the multi-modal human body data and / or the multi-modal vehicle data into a vehicle moving demand recognition model, and judging whether a vehicle moving demand exists or not; when the vehicle moving demand exists, vehicle moving reminding information is sent to a user terminal corresponding to the target vehicle; and in response to a remote interaction confirmation operation of the user terminal, establishing an audio and video transmission channel between the user terminal and the external audio and video device of the target vehicle. The method can identify the vehicle moving demand and notify the vehicle owner to carry out remote interaction with the field personnel, improves the parking experience of the user, and can be applied to the technical field of vehicle monitoring.
Owner:GAC HONDA AUTOMOBILE CO LTD +1

An agricultural product wholesale order automatic generation method, device, equipment and medium

ActiveCN121544346BSound sourcesData stream
The application discloses a kind of agricultural product wholesale order automatic generation method, device, equipment and medium, it is related to order management technical field.The method is first received by pickup array to agricultural product selling communication field real-time acquisition multiple field sound signals, and real-time application empirical mode decomposition technique, sound source positioning technique and voiceprint matching technique are carried out speech signal extraction processing to obtain seller and buyer speech signal, then extraction result is integrated into agricultural product selling communication dialogue text data stream in real time, and real-time import large language model is carried out semantic understanding and entity information extraction processing, then according to extraction result, matched agricultural product wholesale order template is searched out, and the template is filled with content and dynamically updated, generates new agricultural product wholesale order, finally order is pushed to seller terminal equipment in real time and is outputed and shown in real time, so it can improve order generation efficiency, and guarantee order information generation correct, improve seller experience.
Owner:CHENGDU GAUSS ZHIDA INFORMATION TECHNOLOGY CO LTD

Digital human live voice interaction system fused with affective computing

ActiveCN121393436BSpeech recognitionLive voiceData stream
The application relates to the technical field of digital human voice interaction, and discloses a digital human live broadcast voice interaction system fusing emotional computing. The system acquires user voice input data streams, constructs an emotional response time window, the starting point of which is the end time of the user voice input data streams, and the ending point of which is the preset maximum response cutoff time minus the necessary duration of voice response synthesis. The system acquires the emotional state vector of the user in real time, and judges whether the emotional state vector reaches an emotional intensity threshold value; if the emotional state vector reaches the threshold value, the total generation duration of the digital human voice response is predicted. When the residual duration of the emotional response time window is equal to the total generation duration, the starting point of the window is taken as the starting time of the voice response, and the digital human is controlled to start voice response generation. Through fine time window management and emotional state perception, the system generates voice responses that are adapted to the emotional state of the user, reduces invalid responses and interaction delays, and improves the naturalness of digital human live broadcast voice interaction and user experience.
Owner:BEIJING ZHONGSHENGSHENG DIGITAL TECHNOLOGY CO LTD

Vehicle wire harness pickup (one)

ActiveCN309858232SLive voiceIn vehicle
1. Name of the designed product: Car wire harness sound pickup (one kind). 2. Use of the designed product: used to collect live sound. 3. Design points of the designed product: in shape. 4. Picture or photo that best shows the design points: perspective view.
Owner:JRMEMS TECH (WUXI) CO LTD

Live voice synthetization

PCT designated stageWO2026075704A1Video gamesSpeech synthesisLive voiceTelecommunications
Systems, devices, methods, and machine-readable media configured to provide voice synthetization in a multiplayer video game are provided. A system can include a multiplayer video game including a character selection interface through which a player selects a character to represent them in playing the video game, and a voice model trained to convert audio from the player directly into audio in a voice of the character and provide an output that includes the audio in the voice of the character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A live voice speech false propaganda identification method and system

PendingCN122347961ALive voiceIntent recognition
The present application relates to live audio and video content security audit and natural language processing technical field, especially in kind of live voice language technique false propaganda identification method and system, including: obtaining live stream real-time audio segment, extracting first acoustic representation and second acoustic representation according to real-time audio segment;Based on the joint acoustic representation sequence, the deep propaganda intention vector is preliminarily determined;With deep propaganda intention vector as condition, combined with joint acoustic representation sequence, output standard semantic text sequence;According to standard semantic text sequence and deep propaganda intention vector, the false propaganda intention recognition result of real-time audio segment is determined.The present application makes the system not dependent on the serial processing path of text transcription to audio segment first and then variant literal reduction, thereby reducing the information loss caused by the hard decision of the front-end reduction module, and overcoming the problem of insufficient use of acoustic paralinguistic information in pure text processing.
Owner:HANGZHOU NAT E-COMMERCE PROD QUALITY MONITORING & DISPOSAL CENT

Real-time echo cancellation method based on machine learning

PendingCN121171243ASpeech analysisLive voiceVoice communication
The invention relates to the field of echo elimination methods for conference site voice interaction, in particular to a real-time echo elimination method based on machine learning, which comprises the following steps: extracting multi-channel space-time frequency features, splicing the space features and the time frequency features into a fusion vector, and mapping the fusion vector to a preset interval as model input; tracking a dynamic echo path according to the path model, monitoring feature drift in real time, triggering model updating, and finely adjusting the path sub-model by using local data of each microphone; taking the output of the path model as an initial value, carrying out basic filtering cooperation, minimizing a multi-channel filtering residual error through KL divergence, and carrying out time-frequency masking suppression on filtered residual echoes; clustering the mixed voice based on voiceprint features, setting an independent echo suppression threshold value for each speaker voice, and focusing the current non-silent speaker voice through a time sequence attention mechanism; and carrying out model compression on the long reverberation path model, and carrying out parallel reasoning optimization. Echo can be efficiently suppressed in real time, and voice communication definition is improved.
Owner:AEROSPACE XINTONG TECH CO LTD

Communication method and device, electronic equipment and storage medium

The invention relates to a communication method and device, electronic equipment and a storage medium. The communication method comprises: obtaining target audio data, the target audio data comprising at least one of first audio data and second audio data, the first audio data being collected by a first device, the second audio data being collected by at least one second device, and the first device being in networking connection with the at least one second device; and taking the target audio data as uplink audio data of a preset application in the first equipment. According to the method, the target audio data collected by all networking devices are uniformly used as the uplink audio data of instant messaging through the first device, and all networking devices are used as a communication main body, so that the definition of the collected audio data can be ensured, meanwhile, the problem of onsite sound chaos is avoided, and the communication experience of a user is improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

News live voice anomaly real-time monitoring correction method and system thereof

ActiveCN122135742BImplement adaptive correctionLower Detection LatencyLive voiceSpeech error
The application relates to the technical field of speech signal processing and live broadcast quality control, and discloses a news live broadcast speech abnormality real-time monitoring and correction method and system, which comprises the following steps: extracting a multi-dimensional acoustic feature vector from a real-time frame of a broadcaster's speech signal and storing the feature vector in a ring feature buffer; synchronously performing parallel detection of abnormal pause detection, speech error deviation detection and volume mutation detection based on the feature vector; pushing a prompt sound through an ear return, highlighting a deviation position on a teleprompter or triggering a dynamic compressor to smooth the volume according to the type of an abnormal event; recording the abnormal event to generate a broadcast quality analysis report and correcting the detection threshold according to the report; and the application realizes millisecond-level real-time perception and instant correction assistance for broadcasting abnormalities, and significantly shortens the response delay of traditional manual monitoring.
Owner:GUIZHOU NORMAL UNIVERSITY