Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

189 results about "Audio recognition" patented technology

Modular Audio Recognition Framework (MARF) is an open-source research platform and a collection of voice, sound, speech, text and natural language processing (NLP) algorithms written in Java and arranged into a modular and extensible framework that attempts to facilitate addition of new algorithms.

End-to-end dialect audio recognition method and system based on multi-layer information fusion

PCT designated stageWO2025255947A1Speech recognitionSpeech soundAudio frequency
An end-to-end dialect audio recognition method (10) based on multi-layer information fusion. The method (10) comprises: performing audio preprocessing on dialect audio to generate an acoustic feature (S102); inputting the acoustic feature into a coder, and then performing, by the coder, a progressive downsampling operation on the acoustic feature to generate multi-layer fine-grained acoustic features (S104); performing, by means of a layer adaptation module, multi-layer information fusion on the multi-layer fine-grained acoustic features to generate a fused acoustic feature (S106); performing, by means of a cross-attention mechanism, cross-fusion on the fused acoustic feature to generate a corrected acoustic feature (S108); and inputting the corrected acoustic feature into an end-to-end dialect recognition model to generate a dialect audio recognition result (S110). In the method, complex speech signals and multi-accent features can be efficiently captured and processed, and dialect audio can also be classified and decoded online in real time, thereby improving the accuracy and robustness of speech recognition.
Owner:SHANGHAI QIYUE INFORMATION TECH CO LTD

Audio recognition method and apparatus, device, storage medium and computer program product

Provided is an audio recognition method. The method includes that audio data is encoded to an audio encoding feature; the audio encoding feature is decoded to a first decoded feature; a preset word text and the first decoded feature are encoded to a first text encoding feature, the first text encoding feature including a feature representing semantics of the preset word text and a feature representing semantics of the audio data; and the first text encoding feature is decoded to a predicted audio text, where the predicted audio text represents the semantics of the preset word text and the semantics of the audio data. An audio recognition apparatus and a storage medium are also provided.
Owner:MASHANG CONSUMER FINANCE CO LTD

Audio processing model quantification method and system

The invention relates to the technical field of AI audio chips, and discloses an audio processing model quantification method and system, and the method comprises the steps: applying an incremental voltage pulse to a multi-layer memristor cross array, measuring a conductivity value, and generating a memristor quantification configuration table; programming the neuron units in the multi-layer memristor cross array to obtain quantized neuron units; calculating a synaptic weight matrix; the audio signal to be processed is input into the quantization neuron unit to execute pulse calculation, the audio features are obtained, and the audio recognition result is generated on the end-side processor, so that the problems of nonlinearity and element difference of the memristor are effectively solved, the robustness of the system is improved, and the optimal balance of precision and power consumption can be realized in different scenes.
Owner:SHENZHEN ULTRA EASY TECH CO LTD

Audio identification method and device based on marking and backtracking correction

The invention relates to an audio recognition method and device based on marking and backtracking correction, and the method comprises the steps: segmenting a target audio, and enabling adjacent segments after segmentation to have a partial overlapping region; if a slice point exists in a non-mute segment of the target audio and the acoustic feature similarity of a preset number of frames before and after the slice point is smaller than a set similarity threshold value, the slice point is marked as a cut-off risk point, and the cut-off risk point is used for indicating that continuous semantics before and after the slice point has a cut-off risk; after a text corresponding to the audio of each slice is recognized through the speech recognition model, the language confidence degree of an overlapping area where the truncation risk point is located is determined, and the language confidence degree is used for indicating context language logic of the overlapping area; and if it is determined that the language confidence is lower than a set confidence threshold, correcting the recognition texts of the slices before and after the truncation risk point. According to the invention, the accuracy of audio recognition is improved.
Owner:FIBOCOM WIRELESS

Method and device for detecting and judging bullying behavior based on audio and behavior feature recognition

The invention provides a bullying behavior detection and judgment method and device based on audio and behavior feature recognition, and the method comprises the steps: extracting and analyzing the emotion, keyword and abnormal sound features through an audio recognition module, and judging whether a bullying-related language or audio behavior exists in real time; the behavior recognition module is used for behavior data processing and time sequence feature extraction, behavior classification is performed in combination with a GRU network, and detection and recognition of bullying related behaviors are completed; in combination with the recognition results of the audio recognition module and the behavior recognition module, whether the bullying behavior occurs in the current scene is judged, and if the judgment result is yes, an alarm is given and corresponding personnel are notified to go to the current scene; otherwise, recording the identification result to the cloud. The beneficial effects of the invention are that through multi-dimensional technical means and process design, the bullying behavior can be efficiently, accurately and sustainably monitored, pre-warned and recorded, and powerful technical guarantee is provided for building a safe and friendly environment.
Owner:TIANJIN JINHU DATA CO LTD

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Intelligent sound signal sensing method and system based on time-frequency feature fusion

The invention discloses an intelligent sound signal sensing method and system based on time-frequency feature fusion, and relates to the technical field of intelligent sound signal processing, and the method comprises the steps: obtaining an original audio signal; performing short-time Fourier transform on the original audio signal to obtain a plurality of frequency spectrum features; inputting the plurality of frequency spectrum features into a plurality of neural networks for processing, and extracting frequency domain depth features; processing the original audio signal by using a one-dimensional convolutional neural network, and extracting time domain depth features; splicing and fusing the frequency domain depth feature and the time domain depth feature in a channel dimension to obtain a time-frequency fusion feature; and inputting the time-frequency fusion feature into a full connection layer classifier to obtain a final audio scene recognition result. According to the invention, by designing a double-branch network architecture of time domain and frequency domain parallel processing, time domain waveform information and frequency domain amplitude and phase information are effectively fused, and in an end-to-end training scene, the accuracy of audio recognition is remarkably improved.
Owner:ZHEJIANG UNIV +1

Audio data processing method and device, equipment and storage medium

The invention relates to the field of audio data analysis, in particular to an audio data processing method and device, equipment and a storage medium. The method comprises the following steps: carrying out multi-angle audio continuous acquisition and scene noise suppression on a conference room through a surrounding microphone array, and generating a filtering optimization audio signal; performing three-dimensional time difference positioning calculation according to the filtered and optimized audio signal to obtain accurate sound source positioning information; personalized voiceprint analysis is carried out on the filtered and optimized audio signals, audio distribution modeling is carried out based on accurate sound source positioning information, and a multi-person conference audio field is constructed; performing parallel audio stream separation according to the multi-person conference audio field to generate an intelligent spliced audio segment; and performing deep semantic analysis and semantic logic correction on the intelligent spliced audio segment to generate an audio analysis result. According to the invention, rapid and accurate speaker audio recognition of a multi-person parallel conference is improved.
Owner:SHENZHEN ULTRA EASY TECH CO LTD

Harmful voice detection method in combination with side language information

The invention provides a harmful voice detection method in combination with side language information, and the method comprises the steps: collecting a multi-source voice sample, carrying out the preliminary screening and automatic marking of the multi-source voice data through a multi-modal model, and obtaining a preliminary marking sample; performing manual annotation and data reorganization based on the preliminary annotation sample to obtain a high-quality annotation data set, and performing audio unified preprocessing on the high-quality annotation data set to obtain standardized input; based on standardized input, self-supervised high-dimensional features are extracted, and a dual-task model topology is constructed to perform audio recognition; and for the dual-task model topology, carrying out source, category and joint capability training to obtain a joint optimization model so as to output a source label and a type label of harmful information of the to-be-detected audio. According to the method, the optimal performance of a text harmful scene can be considered through high-quality multi-dimensional data support and source and category dual judgment, and the paranguage harmful detection capability is remarkably improved.
Owner:ZHEJIANG UNIV +1

Wind-noise-resistant self-adaptive volume adjusting method and system for riding earphone

The invention relates to the technical field of audio signal processing, and discloses a wind-noise-resistant self-adaptive volume adjustment method and system for a riding earphone, and the method comprises the following steps: S1, collecting environment sound and a source audio signal in parallel; s2, analyzing environment sound, and determining macroscopic adjustment intensity; s3, analyzing the source audio in parallel, and identifying a content type and a transient signal of the source audio; s4, based on the psychological acoustic model, generating a frequency-related target method tone quality compensation gain; s5, performing dynamic smoothing and frequency weighting on the target gain according to the content and the transient characteristics to generate an application gain; s6, fusing the macroscopic adjustment intensity and the application gain, and synthesizing a final gain parameter; and S7, applying the final gain to the source audio, and outputting the compensated audio. According to the method, the psychological acoustic model and audio content parallel analysis are fused, frequency-level accurate compensation is realized, the problems of traditional degradation and sudden hearing sense change are solved, and the audio definition and comfort are improved.
Owner:SHENZHEN ASMAX INFINITE TECH CO LTD

Classroom teaching comprehensive quality evaluation method and system based on teacher and student behavior multi-mode intelligent analysis

The invention discloses a classroom teaching comprehensive quality evaluation method and system based on teacher and student behavior multi-mode intelligent analysis, and the method comprises the steps: firstly obtaining a multi-view classroom teaching video, and constructing a multi-view teacher behavior data set, a teacher behavior audio data set and a student behavior data set; secondly, constructing a multi-view behavior recognition network based on a Mama structure, and respectively obtaining a teacher behavior visual recognition model, a teacher behavior audio recognition model, a teacher behavior multi-modal fusion model and a student learning state detection model through the three data sets; and then acquiring a multi-view classroom teaching video in real time, and based on the model, obtaining a comprehensive teacher behavior identification result and the concentration degree of the students. And finally, according to the teacher behavior comprehensive identification result and the concentration degree of the students, performing quantitative evaluation on the classroom teaching video, and outputting a classroom comprehensive quality evaluation result. According to the method, the problem of one-sided single-modal analysis is solved, and the balance between the multi-view behavior recognition precision and the running speed is realized.
Owner:HANGZHOU DIANZI UNIV

Audio recognition method and device and electronic equipment

The invention provides an audio recognition method and device and electronic equipment. The method comprises the steps of obtaining to-be-recognized audio; determining at least one first audio clip and at least one second audio clip from the to-be-recognized audio through a pre-trained first recognition model; the ratio of the song duration in the first audio clip meets a first specified condition, and the ratio of the song duration in the second audio clip meets a second specified condition; determining a first probability that the first audio clip belongs to the AI flipping audio and a second probability that the second audio clip belongs to the AI generated audio through a second identification model; determining an audio recognition result of the to-be-recognized audio according to the first probability and the second probability; the audio recognition result is AI audio turning and singing, AI generated audio or normal audio. According to the method, audio recognition is carried out on the singing sound segments and the non-singing sound segments, misleading of the non-singing sound segments to the recognition result is avoided, the recognition accuracy of the singing sound segments is improved, and then the recognition accuracy of the whole audio is improved.
Owner:HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD

Automatic string music tremolo detection method

The invention belongs to the technical field of audio recognition, and particularly relates to a string music tremolo automatic detection method which comprises the following steps: S1, preprocessing and framing an input audio, and extracting time sequence features; s2, obtaining an emotion intensity score through an emotion intensity evaluation function; s3, mapping the emotion intensity score into an expected tremor speed, and obtaining a stable tremor speed through exponential smoothing and amplitude limiting; s4, generating a dynamic threshold value changing along with time; and S5, the main tremolo detector calculates the criterion in parallel, compares the criterion with a dynamic threshold value on a synchronization time axis to determine frame-level tremolo, connects adjacent frames to obtain a tremolo segment, and drives threshold value self-adaption by emotional intensity, so that the detection criterion is synchronized with a playing situation, wrong division and over-division are reduced, and the boundary consistency and the stability and accuracy of tremolo estimation are improved.
Owner:SHANGQIU NORMAL UNIVERSITY

A method and system for remotely recognizing lip language by a drone

PendingCN122369454AData streamNoise
This invention provides a method and system for long-range lip-reading recognition using unmanned aerial vehicles (UAVs), belonging to the fields of artificial intelligence and UAV technology. The system utilizes a camera, IMU (Integrated Mutor Unit), and microphone mounted on the UAV to acquire video streams, IMU data streams, and audio streams. Motion compensation is applied to the video stream using IMU data to stabilize image frames. Then, image sequences of the target speaker's lips and nasal / jaw region are extracted and input into a lip-reading recognition model to obtain initial recognition text and visual confidence scores. Simultaneously, the signal-to-noise ratio (SNR) of the audio stream is calculated. When the visual confidence score is low and the SNR is high, the audio recognition process is triggered, generating audio recognition text and an audio confidence score. Finally, the visual and audio confidence scores, along with the lip movement-audio temporal matching degree, are weighted and fused to generate the final recognition text. This invention effectively solves the problems of video instability and background interference in long-range lip-reading recognition, improving recognition accuracy in complex environments.
Owner:FUZHOU PLANNING DESIGN & RES INST

Image capturing apparatus having audio recognition, control method thereof, and storage medium

An image capturing apparatus obtains audio of an utterance that occurs in a vicinity of the image capturing apparatus, captures an image, and controls image transmission such that in response to a determination that an expression indicating a particular person is included in the audio of the utterance, transmits a first image related to the obtainment of the audio of the utterance among captured images to an external apparatus associated with the expression indicating the particular person. In addition, a second image captured in the external apparatus and related to playing of the first image is received from the external apparatus.
Owner:CANON KK

Machine learning based video recognition method, device, server and storage medium

Embodiments of the present application disclose a video recognition method and device based on machine learning, a server and a storage medium. Embodiments of the present application can obtain a target video; obtain a source video corresponding to the target video, the target video being created by processing the source video; compare the target video and the source video in content to obtain a content type of the target video; when the content type of the target video is a funny content type, perform audio recognition on the target video and the source video to determine an audio type of the target video; when the audio type of the target video is a funny dubbing type, determine the target video as a funny dubbing video, so as to push the funny dubbing video to a user. Embodiments of the present application compare the target video with the source video in dimensions such as content and audio to identify whether the target video is a funny dubbing video created by processing the source video. Thus, the present application can accurately identify a funny dubbing video from a plurality of videos, and improves the efficiency of video recognition.
Owner:TENCENT TECH (BEIJING) CO LTD

Audio recognition method and device, electronic equipment and storage medium

The invention provides an audio recognition method and device, electronic equipment and a storage medium, and belongs to the technical field of audio processing, and the method comprises the steps: carrying out the Fourier transform of a non-voice signal, obtaining a noise frequency domain signal, and calculating the first power spectrum density of a background noise signal in the non-voice signal; performing Fourier transform on the voice signal to obtain a voice frequency domain signal, and calculating a second power spectral density corresponding to the initial voice signal based on the voice frequency domain signal and the first power spectral density; calculating a power ratio of an initial voice signal to a background noise signal in the voice signals; performing signal enhancement on the initial voice signal based on the power ratio, the voice frequency domain signal, the first power spectral density and the second power spectral density to obtain a target voice signal; and identifying the target voice signal based on the target model. According to the audio recognition method and device, the electronic equipment and the storage medium provided by the invention, the audio recognition precision can be improved.
Owner:BEIJING SUPERHEXA CENTURY TECH CO LTD

Audio processing method, device and system and storage medium

The invention relates to the technical field of audio processing, and provides an audio processing method, device and system and a storage medium, and the method comprises the steps: receiving a mixed audio signal which at least comprises two kinds of audio signals; identifying the mixed audio signal through at least one audio identification module to identify a target audio signal in the mixed audio signal; and carrying out audio processing on the target audio signal by adopting corresponding preset processing rules through the at least two audio processing modules. By adopting an intelligent identification and differential processing mechanism, on the premise of keeping the integrity of the original audio, different target audio signals are subjected to respective corresponding better audio processing, for example, in a game scene, environment sound effects such as shocking explosive sound and fine footstep sound are enhanced in an immersive low-frequency mode, and the effect of improving the sound quality is achieved. And meanwhile, the teammate voice instruction is kept clear and transparent without interference, so that the audio processing effect is greatly improved.
Owner:GUANGDONG DINGCHUANG SMART MANUFACTURING CO LTD

Speech determination model training method, object speech extraction method, device, electronic equipment and storage medium

The present disclosure relates to a speech determination model training method, an object speech extraction method, a device, an electronic device and a storage medium. The method comprises: obtaining a second speech set, the second speech set comprising a plurality of second sample speeches and a second object identifier corresponding to each second sample speech, each second sample speech comprising a fifth speech and a sixth speech, the fifth speech comprising at least speech corresponding to a fifth speech non-associated object identifier, and the sixth speech being silent audio; performing speech extraction on the fifth speech in each second sample speech based on a speech extraction model to obtain a seventh speech corresponding to each second sample speech; and training a candidate speech determination model based on the seventh speech and the sixth speech to obtain a target speech determination model. The present application increases the training samples by adding the speech corresponding to the fifth speech non-associated object identifier and the silence, trains the candidate speech determination model, obtains a model suitable for richer scenarios, and reduces the error rate of audio recognition.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

A breathing training system using mantras

The present application relates to a kind of breathing training systems using six-word formula, including user end hardware and cloud server, using the VR head-mounted integrated machine of user end hardware and elastic fabric waistband, guide and detect the breathing training condition of user, realize the accurate classification of abdominal breathing, chest breathing, mixed breathing by means of EMG, displacement, triple fusion calculation mode of flow, utilize mouth shape identification auxiliary phoneme confidence correction, accurately identify different exhale mode of six-word formula, introduce mouth shape identification, phoneme model, physiological sensing fusion multi-modal scoring mechanism, compared with traditional single chest and abdominal displacement monitoring scheme or single audio identification scheme, accurately evaluate the training quality of user from many aspects and comprehensive score, it is convenient for user to understand the improvement direction and progress space of self breathing training.
Owner:THE SECOND HOSPITAL AFFILIATED TO WENZHOU MEDICAL COLLEGE

Efficient extension to recognize new languages

The present disclosure describes techniques for efficiently extending to recognize new languages. A first data flow pipeline of a machine learning model can be maintained. The first data flow pipeline comprises pre-trained parameters and is pre-trained to recognize existing languages based on input audio. A second data flow pipeline of the machine learning model is configured. The second data flow pipeline is configured to utilize the pre-trained parameters of the first data flow pipeline and leverage additional trainable parameters. The machine learning model is fine-tuned by exclusively updating the additional trainable parameters of the second data flow pipeline using data from the new languages. The machine learning model is fine-tuned to recognize the new languages based on input audio while preserving performance in recognizing the existing languages.
Owner:LEMON INC(GB) +1

Audio recognition method, apparatus, device, and storage medium

ActiveCN119905108BNoiseEngineering
This application relates to an audio recognition method, apparatus, device, and storage medium, and pertains to the field of vehicle technology. The method includes: in response to detecting an abnormal noise audio, determining an abnormal noise region based on the abnormal noise audio; if a fault code is detected, determining the faulty component corresponding to the fault code based on the fault code; if the region where the faulty component is located coincides with the abnormal noise region, determining that the abnormal noise audio is emitted by the faulty component. This is used to improve the accuracy of vehicle operation audio recognition.
Owner:CHONGQING CHANGAN TECH CO LTD

Video content recognition methods, devices, equipment, storage media, and software products

PendingCN122313344ANoise (video)Noise
This application provides a video content recognition method, apparatus, device, storage medium, and program product, belonging to the field of data processing technology. The method includes: acquiring a video to be recognized, comment data of the video to be recognized, and descriptive text, wherein the video to be recognized consists of video data and audio data; inputting the comment data into a semantic analysis model to obtain a judgment result of the background music volume; if the volume judgment result indicates a large volume, inputting the video data and audio data into an audio-video fusion extraction model to obtain fusion information text output by the audio-video fusion extraction model; using the fusion information text and the descriptive text to remove background noise from the audio data to obtain a noise-reduced frequency; and determining video content information based on the noise-reduced frequency and the video data. This method solves the problem of low accuracy in audio recognition and information extraction when background music is present.
Owner:CHENGDU TD TECH LTD

Emotion analysis and feedback device based on visual and audio recognition

PendingCN122296895Areduce distractionsAvoid collection blind spotsHuman bodyBlind zone
This invention discloses an emotion analysis and feedback device based on visual and audio recognition in the field of psychological instruments and equipment technology. The device includes: a mobile base; a functional panel, vertically fixed to the top of the mobile base; a positioning ring, vertically suspended to the side of the functional panel, with an internal opening larger than the frontal contour of the human body, used for centering the user and restricting their sitting posture; a connecting arm, one end connected to the lower end of the functional panel and the other end connected to the positioning ring, used to support the positioning ring in a suspended state; and a flexible covering assembly located above the connecting arm, its surface forming a sloped structure facing the functional panel, used to guide the user's gaze to focus on the functional panel. This device features a hollow, large-sized positioning ring, which can center and align the user's torso and correct their sitting posture, avoiding blind spots caused by posture deviation. Combined with multi-directionally deployed visual and audio acquisition components, it significantly improves the accuracy of emotion analysis.
Owner:ANHUI SUNSHINE HEART HEALTH TECH DEV CO LTD

Multilingual speech and semantic intelligent translation method and system applied to exhibition scene

The invention discloses a multilingual speech semantic intelligent translation method and system applied to an exhibition scene, and belongs to the technical field of machine translation, and the method comprises the following steps: S1, obtaining a multi-person question judgment result; s2, if the multi-person questioning judgment result is multi-person questioning, audio identification information is obtained through analysis, and otherwise, the audio identification information is directly obtained through analysis; s3, obtaining each storage question keyword, each contrast question keyword and a key matching weighting factor corresponding to each question keyword; s4, obtaining a same-group evaluation result, if the same-group evaluation result is the same group, analyzing to obtain the comprehensive matching similarity of the storage groups, otherwise, analyzing the comprehensive matching similarity of each parallel storage group; s5, obtaining a comprehensive matching judgment result, if the comprehensive matching judgment result is unqualified, performing secondary refining processing to obtain a question and answer, and otherwise, directly obtaining the question and answer; and S6, voice broadcasting is carried out, and accurate separation and language recognition of voice sources of different questioning users are achieved.
Owner:ZHEJIANG HUIZHAN ELF TECHNOLOGY CO LTD

Method and device for determining audio emotion and computer equipment

The application discloses a kind of audio emotion determination method, device and computer equipment, belong to artificial intelligence technical field.The method comprises: obtaining the audio representation and first text representation of audio, first text representation is the text representation of audio text, and audio text is obtained by audio recognition to audio;Audio recognition error detection is carried out based on audio representation and first text representation, to obtain predicted error probability, and predicted error probability indicates the identification error probability of audio recognition to audio;First text representation is weighted based on predicted error probability, to obtain weighted text representation, and weighted processing is used to set the confidence of first text representation in audio emotion classification process;Audio emotion classification is carried out based on weighted text representation and audio representation, to obtain the predicted audio emotion of audio.The method can improve the robustness of audio emotion classification.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio recognition method and device, computer equipment, storage medium and program product

The invention relates to an audio recognition method and device, computer equipment, a storage medium and a program product, and relates to the technical field of audio recognition. The method is executed by computer equipment and comprises the steps that audio identification information of an audio file is obtained, the audio identification information is information obtained after a target frequency band in the audio file is converted, and the target frequency band comprises at least one of an ultrasonic frequency band and an infrasonic wave frequency band; based on the audio identification information, executing an identification operation on the audio file to obtain an identification result; specification processing is performed based on the recognition result. According to the audio recognition scheme, the audio recognition accuracy can be improved on the basis of ensuring the concealment and security of the watermark information of the audio file.
Owner:GUANGZHOU KUGOU COMP TECH CO LTD

X-ray imaging system

PCT designated stageWO2026018862A1Radiation diagnosticsControl cellAcoustics
This X-ray imaging system (100) comprises: a voice input unit (5) that receives voice input; a voice output unit (15); and a control unit (8) that, on the basis of the voice received by the voice input unit (5), performs voice recognition of a keyword (86), thereby performing control based on the keyword (86). The control unit (8) is configured to reproduce and output a sample voice of the keyword (86), by using the voice output unit (15), on the basis of an operation performed by an operator.
Owner:SHIMADZU CORP

Data processing method and apparatus

This specification provides a data processing method and apparatus, wherein the data processing method includes: preprocessing initial audio data to obtain audio data, and inputting the audio data into an audio recognition model to obtain initial text containing prosody identifiers; determining at least one text unit corresponding to the initial text, and determining time information corresponding to each of the at least one text unit based on the audio data; updating the initial text in the prosody identifier dimension based on the time information corresponding to each of the at least one text unit to obtain target text corresponding to the audio data, wherein the target text and the audio data are used to train an audio generation model.
Owner:BEIJING YUANLI WEILAI SCI & TECH CO LTD

Audio recognition method and device, computer device and computer readable storage medium

The application discloses an audio recognition method and device, computer equipment and a computer readable storage medium, and applies to the technical field of computers. The method comprises the following steps: inputting to-be-recognized audio data into an audio recognition model to obtain an audio fingerprint of the to-be-recognized audio data output by the audio recognition model; wherein the audio recognition model is obtained based on contrast learning of first training audio data and second training audio data and the first training audio data and third training audio data; the audio identifiers of the first training audio data and the second training audio data are the same; the audio identifiers of the first training audio data and the third training audio data are different; target audio fingerprints satisfying a preset condition are determined from an audio fingerprint library according to the audio fingerprint of the to-be-recognized audio data; and an identification result is determined according to the target audio fingerprints, wherein the identification result comprises an audio identifier corresponding to the to-be-recognized audio data. Through the method, the accuracy of audio recognition can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD