Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Live voice" patented technology

Live voice synthetization

PendingUS20260097308A1Video gamesSpeech synthesisLive voiceTelecommunications
Systems, devices, methods, and machine-readable media configured to provide voice synthetization in a multiplayer video game are provided. A system can include a multiplayer video game including a character selection interface through which a player selects a character to represent them in playing the video game, and a voice model trained to convert audio from the player directly into audio in a voice of the character and provide an output that includes the audio in the voice of the character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Digital human live broadcast voice interaction system fused with emotion calculation

ActiveCN121393436ASpeech recognitionLive voiceData stream
The invention relates to the technical field of digital human voice interaction, and discloses a digital human live voice interaction system fused with emotion calculation. The system constructs an emotional response time window by acquiring a user voice input data stream, the starting point of the emotional response time window is the end moment of the user voice input data stream, and the end point is obtained by subtracting the necessary duration of voice response synthesis from the preset maximum response cut-off moment. The system obtains an emotional state vector of the user in real time and judges whether the emotional state vector reaches an emotional intensity threshold value or not; and if so, predicting the total generation duration of the digital human voice response. And when the residual duration of the emotional response time window is equal to the total generation duration, taking the window starting point as a voice response starting moment, and controlling the digital human to start voice response generation. According to the system, the voice response matched with the emotional state of the user is generated through refined time window management and emotional state perception, invalid response and interaction delay are reduced, and the naturalness of digital human live broadcast voice interaction and the user experience are improved.
Owner:BEIJING ZHONGSHENGSHENG DIGITAL TECHNOLOGY CO LTD

Industrial field high-frequency voice recognition method and storage medium

ActiveCN120496577ASpeech analysisLive voiceAlgorithm
The invention relates to the field of voice recognition, in particular to an industrial field high-frequency voice recognition method and a storage medium. The method comprises the following steps: performing short-time Fourier transform on an industrial field sound signal by using a double-branch window to obtain a double-branch spectrogram; performing channel stacking on the double-branch spectrogram to obtain a three-dimensional tensor; after the extracted features are classified and calculated, a classification head outputs score vectors of probability scores of three dimensions of target sound, environment sound and strong noise; softening the probability score by using a temperature coefficient, and inputting the softened probability score into a classification function to obtain probability distribution; and calculating an energy score, when the energy score is lower than a preset energy threshold value, calculating an attenuation coefficient to carry out scaling suppression on the probability distribution to obtain final probability distribution, judging whether the highest value in the final probability distribution is lower than a rejection threshold value or not, and obtaining an identification result. According to the method, the fault sound detection rate is greatly improved in industrial actual measurement, and the false alarm rate is remarkably reduced.
Owner:SHANGHAI SANTONG AUTOMATION TECH CO LTD

Agricultural product wholesale order automatic generation method and device, equipment and medium

The invention discloses an agricultural product wholesale order automatic generation method and device, equipment and a medium, and relates to the technical field of order management. The method comprises the following steps: firstly, receiving a plurality of field sound signals collected by a sound pickup array on an agricultural product selling communication field in real time, and carrying out voice signal extraction processing by applying an empirical mode decomposition technology, a sound source positioning technology and a voiceprint matching technology in real time so as to obtain speech signals of a seller and a buyer; then, the extraction result is integrated into an agricultural product selling communication dialogue text data stream in real time, the text data stream is imported into a large language model in real time to carry out semantic understanding and entity information extraction processing, then a matched agricultural product wholesale order template is retrieved according to the extraction result, and content filling and dynamic updating are carried out on the template; and generating a new agricultural product wholesale order, and finally pushing the order to the seller terminal equipment in real time and outputting and displaying the order in real time, so that the order generation efficiency can be improved, the order information is ensured to be generated correctly, and the seller experience is improved.
Owner:CHENGDU GAUSS ZHIDA INFORMATION TECHNOLOGY CO LTD

Live voice synthesis with voice mixing

PCT designated stageWO2026075705A1Video gamesSpeech synthesisLive voiceTheoretical computer science
Systems, devices, methods, and machine-readable media configured to provide voice inference in a video game are provided. A video game system can include an encoder configured to generate a first encoding representative of physical characteristics of a specified entity, a similarity operator configured to determine similarity values between (i) corresponding stored encodings of multiple characters, the stored encodings representative of physical characteristics of respective characters of the multiple characters and (ii) the first encoding, identify a selected character from the multiple characters based on the similarity values, and provide an identifier of the selected character, a voice database configured to provide audio or a spectrogram of the selected character, and a video game configured to provide the audio of a player-selected character in a voice of the selected character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Comprehensive diagnosis method and system for ultrahigh-pressure converter station valve cooling system

The invention relates to the technical field of ultrahigh-pressure converter station valve cooling system monitoring, in particular to an ultrahigh-pressure converter station valve cooling system comprehensive diagnosis method and system.The method comprises the steps that on-site sound signals are obtained and analyzed to generate a first analysis result, and a second analysis result is generated by combining meter image analysis; and the water leakage condition of the spraying pump is judged, a third analysis result is generated, and dynamic monitoring and deterioration trend analysis are conducted on abnormal points. According to the invention, the operation state and parameter abnormity of the valve cooling system can be rapidly identified through intelligent analysis of voiceprint features and meter images, a water leakage detection and abnormal point monitoring mechanism is introduced, the fault identification capability and the early warning level are improved, the monitoring efficiency and the intelligent degree are significantly improved, and the problem of low efficiency caused by dependence on manual operation in the prior art is solved.
Owner:GUANGZHOU BUREAU CSG EHV POWER TRANSMISSION

Audio and video stream processing method, device and equipment for cross-platform live voice chat and medium

The application relates to the technical field of network live broadcast, and discloses a cross-platform live broadcast microphone connection audio and video stream processing method and device, electronic equipment and a storage medium, the method saves an audio and video file in a second live broadcast platform server, extracts audio frame information and video frame information from the audio and video file according to real-time microphone connection audio and video stream playing abnormal event information, and the cause of the real-time microphone connection audio and video stream playing abnormal event can be automatically and quickly determined according to the audio frame information and the video frame information, without manual participation, so that the cause determination efficiency of the audio and video stream playing abnormality is improved.
Owner:GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD

Offline voice interaction monitoring device and system

The utility model discloses an off-line voice interaction monitoring device and system, and the device comprises a holder, a communication unit, an energy storage unit, an audio unit, an off-line voice recognition module, and a control unit. The holder is integrated with a camera and has a preset rotation angle; the communication unit is used for performing remote communication with a back-end platform; the energy storage unit is used for supplying power to equipment; the audio unit comprises a loudspeaker and a pickup microphone; the loudspeaker is used for voice broadcasting; the pickup microphone is used for acquiring field sound; and the off-line voice recognition module is used for responding to sound acquired on site according to a preset off-line voice instruction. The technical scheme of the utility model breaks through the dependence of the traditional scheme on network and wired operation, improves the operation convenience and reliability in a complex scene, avoids the potential safety hazard caused by data transmission, and is suitable for scenes with high requirements on real-time performance and autonomy, such as electric power inspection, emergency rescue, field operation and the like.
Owner:SHENZHEN TELICANG TECH CO LTD

Remote monitoring system and monitoring method for sound of key process actions in rolling processes

ActiveCN116274421BLive voiceRemote control
The present application relates to a kind of based on sound sensor's rolling process key process action remote monitoring system and monitoring method, belong to the remote monitoring technical field of sound signal.The system includes audio signal acquisition module, data communication transmission module, audio signal filtering processing module, audio signal storage module, audio signal display module, audio signal analysis module and comprehensive system management module.The system passes through the audio signal of important site in rolling process field acquisition, solves the problem that operator cannot hear the sound of field after remote control, realizes the monitoring and analysis to field situation, is convenient for operator to appear in production process possibly Problem is promptly found, and the intelligentization of operation control is promoted.
Owner:UNIV OF SCI & TECH BEIJING

Remote sound acquisition and processing device

ActiveCN223274209UTransducer circuitsLive voiceNoise
The utility model discloses a remote sound acquisition and processing device, which comprises an on-site sound acquisition module and a remote sound receiving and processing module, on-site audio signals are directionally acquired through pickup equipment and are transmitted to the remote through an on-site radio interphone; the remote radio interphone receives and demodulates the signals, and divides the audio signals into two paths, one path is sent to the sound recording equipment for sound recording, the other path is sent to the audio signal level processing display unit, the normal audio signals on site are subjected to noise processing and converted into level signals, and the level signals are sent to the remote radio interphone; and on-site abnormal audios are amplified and are respectively displayed through the LED display equipment, and abnormal sound exceeding a preset amplitude is alarmed through the buzzer. The sound acquisition and processing device is simple to operate and easy to implement, and can acquire, store and analyze on-site audio signals of equipment in a steel workshop on the premise that normal operation of operators is not affected.
Owner:JIANGSU SHAGANG STEEL CO LTD +2

Live voice and media publishing and distribution platform

A network based media distribution platform includes a server connected to the network, the server including at least one data repository, and a non-transitory medium with a set of machine readable instructions executable therefrom coupled to the server, the instructions causing the server to (a) receive media content from a creator using a media capture device or system to create and upload content, (b) buffer the content for distribution over the network to one or more creator channels accessible to consumers operating media playback devices or systems connected to consumer platform, (c) authenticate the media content to the content creator using at least token data, (d) register the media content and creator authentication data on a connected blockchain storing information in distributed ledger form, and (e) distribute the content to creator channel(s) accessible to the consumers operating the network-connected playback devices or systems connected to the consumer platforms.
Owner:ATTNLIVE INC

Industrial site high-frequency sound recognition method and storage medium

ActiveCN120496577BSpeech analysisLive voiceAlgorithm
The present application relates to the field of sound recognition, and in particular to a method and storage medium for high-frequency sound recognition in industrial sites. The method comprises: performing short-time Fourier transform on the industrial site sound signal using a dual-branch window to obtain a dual-branch spectrogram; performing channel stacking on the dual-branch spectrogram to obtain a three-dimensional tensor; after classifying and calculating the extracted features, the classification head outputs a scoring vector containing probability scores of the three dimensions of target sound, ambient sound, and strong noise; softening the probability score using a temperature coefficient, and inputting the softened probability score into a classification function to obtain a probability distribution; calculating the energy score, and when the energy score is lower than a preset energy threshold, calculating the attenuation coefficient to scale and suppress the probability distribution to obtain a final probability distribution, determining whether the highest value in the final probability distribution is lower than the rejection threshold, and obtaining the recognition result. The method of the present application has greatly improved the detection rate of fault sounds and significantly reduced the false alarm rate in industrial measurements.
Owner:SHANGHAI SANTONG AUTOMATION TECH CO LTD

News live voice anomaly real-time monitoring correction method and system thereof

ActiveCN122135742ATelevision system detailsSpeech analysisLive voiceSpeech error
This invention relates to the field of speech signal processing and live broadcast quality control technology, and discloses a method and system for real-time monitoring and correction of speech anomalies in news live broadcasts. The method includes: extracting multi-dimensional acoustic feature vectors from the announcer's speech signal in real time by frame segmentation and storing them in a circular feature buffer; performing parallel detection of abnormal pauses, speech errors, and sudden volume changes based on the feature vectors; pushing prompts through in-ear monitors, highlighting deviation positions on the teleprompter, or triggering a dynamic compressor to smooth the volume according to the type of abnormal event; recording abnormal events to generate a broadcast quality analysis report and using this report to correct the detection threshold. This invention achieves millisecond-level real-time perception and immediate correction assistance for broadcast anomalies, significantly shortening the response delay of traditional manual monitoring.
Owner:GUIZHOU NORMAL UNIVERSITY

Live voice synthesis with voice mixing

Systems, devices, methods, and machine-readable media configured to provide voice inference in a video game are provided. A video game system can include an encoder configured to generate a first encoding representative of physical characteristics of a specified entity, a similarity operator configured to determine similarity values between (i) corresponding stored encodings of multiple characters, the stored encodings representative of physical characteristics of respective characters of the multiple characters and (ii) the first encoding, identify a selected character from the multiple characters based on the similarity values, and provide an identifier of the selected character, a voice database configured to provide audio or a spectrogram of the selected character, and a video game configured to provide the audio of a player-selected character in a voice of the selected character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Vehicle-based remote interaction method, device, equipment and computer program product

The invention discloses a vehicle-based remote interaction method, device and equipment and a computer program product, and the method comprises the steps: obtaining a human body time sequence image and / or a vehicle time sequence image through extraction according to field image data, and obtaining voice information and / or road traffic noise through recognition according to field sound data; according to the sound source position, fusing the voice information and / or the road traffic noise with the human body time sequence image and / or the vehicle time sequence image to obtain multi-modal human body data and / or multi-modal vehicle data, inputting the multi-modal human body data and / or the multi-modal vehicle data into a vehicle moving demand recognition model, and judging whether a vehicle moving demand exists or not; when the vehicle moving demand exists, vehicle moving reminding information is sent to a user terminal corresponding to the target vehicle; and in response to a remote interaction confirmation operation of the user terminal, establishing an audio and video transmission channel between the user terminal and the external audio and video device of the target vehicle. The method can identify the vehicle moving demand and notify the vehicle owner to carry out remote interaction with the field personnel, improves the parking experience of the user, and can be applied to the technical field of vehicle monitoring.
Owner:GAC HONDA AUTOMOBILE CO LTD +1

An agricultural product wholesale order automatic generation method, device, equipment and medium

ActiveCN121544346BSound sourcesData stream
The application discloses a kind of agricultural product wholesale order automatic generation method, device, equipment and medium, it is related to order management technical field.The method is first received by pickup array to agricultural product selling communication field real-time acquisition multiple field sound signals, and real-time application empirical mode decomposition technique, sound source positioning technique and voiceprint matching technique are carried out speech signal extraction processing to obtain seller and buyer speech signal, then extraction result is integrated into agricultural product selling communication dialogue text data stream in real time, and real-time import large language model is carried out semantic understanding and entity information extraction processing, then according to extraction result, matched agricultural product wholesale order template is searched out, and the template is filled with content and dynamically updated, generates new agricultural product wholesale order, finally order is pushed to seller terminal equipment in real time and is outputed and shown in real time, so it can improve order generation efficiency, and guarantee order information generation correct, improve seller experience.
Owner:CHENGDU GAUSS ZHIDA INFORMATION TECHNOLOGY CO LTD

Training sample selection method and device, electronic equipment and storage medium

The invention relates to the technical field of computers, and provides a training sample selection method and device, electronic equipment and a storage medium. The method comprises the following steps: firstly, converting live broadcast voice of an anchor into a live broadcast text, and filtering interference information in the live broadcast text to obtain a target text; then determining semantic features and style features of each sentence in the target text, and storing the semantic features and the style features of each sentence to an anchor corpus of an anchor; and finally, according to a preset hierarchical clustering strategy, performing clustering based on the semantic features and style features of each sentence in the anchor corpus to obtain a plurality of target sentences, and taking each target sentence as a training sample of a role playing model corresponding to the anchor. Wherein the role playing model is used for simulating the language style of the anchor. The text capable of accurately representing the anchor language style is selected to train the role playing model, so that the training effect and the imitation ability of the model are improved.
Owner:GUANGZHOU HUYA INFORMATION TECH CO LTD

Digital human live voice interaction system fused with affective computing

ActiveCN121393436BSpeech recognitionLive voiceData stream
The application relates to the technical field of digital human voice interaction, and discloses a digital human live broadcast voice interaction system fusing emotional computing. The system acquires user voice input data streams, constructs an emotional response time window, the starting point of which is the end time of the user voice input data streams, and the ending point of which is the preset maximum response cutoff time minus the necessary duration of voice response synthesis. The system acquires the emotional state vector of the user in real time, and judges whether the emotional state vector reaches an emotional intensity threshold value; if the emotional state vector reaches the threshold value, the total generation duration of the digital human voice response is predicted. When the residual duration of the emotional response time window is equal to the total generation duration, the starting point of the window is taken as the starting time of the voice response, and the digital human is controlled to start voice response generation. Through fine time window management and emotional state perception, the system generates voice responses that are adapted to the emotional state of the user, reduces invalid responses and interaction delays, and improves the naturalness of digital human live broadcast voice interaction and user experience.
Owner:BEIJING ZHONGSHENGSHENG DIGITAL TECHNOLOGY CO LTD

Fault self-checking device for security sound pick-up

The utility model discloses a fault self-checking device for a security sound pickup, which is characterized in that an MCU (Microprogrammed Control Unit) controls an ultrasonic transmitter to emit an ultrasonic self-checking signal, and a silicon microphone simultaneously picks up field sound and ultrasonic waves. The silicon microphone converts field sound and ultrasonic waves into electroacoustic signals, the electroacoustic signals are amplified by the preamplifier, and the amplified signals are output in two paths: one path passes through the 20kHz low-pass filter to output sound signals in an auditory frequency range; and one path passes through a 40kHz high-pass filter and is output to an MCU (Microprogrammed Control Unit) connected with the 40kHz high-pass filter. And the MCU compares the received ultrasonic signal with the ultrasonic signal transmitted by the ultrasonic transmitter, and judges whether the sound pickup function fails or not. According to the utility model, the fault self-detection is not interfered by field environment sound, the field voice is not interfered, the sound pickup is not influenced to output audio signals, the versatility is strong, and the cost is low.
Owner:HANGZHOU DIANZI UNIV

Vehicle wire harness pickup (one)

ActiveCN309858232SLive voiceIn vehicle
1. Name of the designed product: Car wire harness sound pickup (one kind). 2. Use of the designed product: used to collect live sound. 3. Design points of the designed product: in shape. 4. Picture or photo that best shows the design points: perspective view.
Owner:JRMEMS TECH (WUXI) CO LTD

Live voice synthetization

PCT designated stageWO2026075704A1Video gamesSpeech synthesisLive voiceTelecommunications
Systems, devices, methods, and machine-readable media configured to provide voice synthetization in a multiplayer video game are provided. A system can include a multiplayer video game including a character selection interface through which a player selects a character to represent them in playing the video game, and a voice model trained to convert audio from the player directly into audio in a voice of the character and provide an output that includes the audio in the voice of the character.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A live voice speech false propaganda identification method and system

PendingCN122347961ALive voiceIntent recognition
The present application relates to live audio and video content security audit and natural language processing technical field, especially in kind of live voice language technique false propaganda identification method and system, including: obtaining live stream real-time audio segment, extracting first acoustic representation and second acoustic representation according to real-time audio segment;Based on the joint acoustic representation sequence, the deep propaganda intention vector is preliminarily determined;With deep propaganda intention vector as condition, combined with joint acoustic representation sequence, output standard semantic text sequence;According to standard semantic text sequence and deep propaganda intention vector, the false propaganda intention recognition result of real-time audio segment is determined.The present application makes the system not dependent on the serial processing path of text transcription to audio segment first and then variant literal reduction, thereby reducing the information loss caused by the hard decision of the front-end reduction module, and overcoming the problem of insufficient use of acoustic paralinguistic information in pure text processing.
Owner:HANGZHOU NAT E-COMMERCE PROD QUALITY MONITORING & DISPOSAL CENT

Live picture brightness adjustment method and device for live voice chat, computer device and medium

ActiveCN115802064BSelective content distributionLive voiceData control
The application relates to the network live broadcast technical field and proposes a live broadcast microphone connection picture brightness adjustment method and device, computer equipment and a storage medium, the method comprises the following steps: in response to a microphone connection opening request, obtaining a plurality of anchor identifiers, establishing microphone connection session connections between a plurality of anchor identifiers corresponding to anchor terminals; according to the exposure compensation range of the video picture of each anchor terminal, determining the best exposure compensation value of the video picture of each anchor terminal, and downloading the best exposure compensation value to the corresponding anchor terminal; wherein the exposure compensation range is determined according to the brightness of the video picture collected by the anchor terminal; receiving target video picture data collected by each anchor terminal according to the best exposure compensation value; according to the target video picture data sent by each anchor terminal, the microphone connection live broadcast picture is displayed in the respective live broadcast room interface of the client in the live broadcast room, so that the brightness deviation of the microphone connection picture is reduced, and the retention rate of the audience in the live broadcast room is improved.
Owner:GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD

Real-time echo cancellation method based on machine learning

PendingCN121171243ASpeech analysisLive voiceVoice communication
The invention relates to the field of echo elimination methods for conference site voice interaction, in particular to a real-time echo elimination method based on machine learning, which comprises the following steps: extracting multi-channel space-time frequency features, splicing the space features and the time frequency features into a fusion vector, and mapping the fusion vector to a preset interval as model input; tracking a dynamic echo path according to the path model, monitoring feature drift in real time, triggering model updating, and finely adjusting the path sub-model by using local data of each microphone; taking the output of the path model as an initial value, carrying out basic filtering cooperation, minimizing a multi-channel filtering residual error through KL divergence, and carrying out time-frequency masking suppression on filtered residual echoes; clustering the mixed voice based on voiceprint features, setting an independent echo suppression threshold value for each speaker voice, and focusing the current non-silent speaker voice through a time sequence attention mechanism; and carrying out model compression on the long reverberation path model, and carrying out parallel reasoning optimization. Echo can be efficiently suppressed in real time, and voice communication definition is improved.
Owner:AEROSPACE XINTONG TECH CO LTD

Communication method and device, electronic equipment and storage medium

The invention relates to a communication method and device, electronic equipment and a storage medium. The communication method comprises: obtaining target audio data, the target audio data comprising at least one of first audio data and second audio data, the first audio data being collected by a first device, the second audio data being collected by at least one second device, and the first device being in networking connection with the at least one second device; and taking the target audio data as uplink audio data of a preset application in the first equipment. According to the method, the target audio data collected by all networking devices are uniformly used as the uplink audio data of instant messaging through the first device, and all networking devices are used as a communication main body, so that the definition of the collected audio data can be ensured, meanwhile, the problem of onsite sound chaos is avoided, and the communication experience of a user is improved.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

News live voice anomaly real-time monitoring correction method and system thereof

ActiveCN122135742BImplement adaptive correctionLower Detection LatencyLive voiceSpeech error
The application relates to the technical field of speech signal processing and live broadcast quality control, and discloses a news live broadcast speech abnormality real-time monitoring and correction method and system, which comprises the following steps: extracting a multi-dimensional acoustic feature vector from a real-time frame of a broadcaster's speech signal and storing the feature vector in a ring feature buffer; synchronously performing parallel detection of abnormal pause detection, speech error deviation detection and volume mutation detection based on the feature vector; pushing a prompt sound through an ear return, highlighting a deviation position on a teleprompter or triggering a dynamic compressor to smooth the volume according to the type of an abnormal event; recording the abnormal event to generate a broadcast quality analysis report and correcting the detection threshold according to the report; and the application realizes millisecond-level real-time perception and instant correction assistance for broadcasting abnormalities, and significantly shortens the response delay of traditional manual monitoring.
Owner:GUIZHOU NORMAL UNIVERSITY