Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

556 results about "Audio segment" patented technology

Video editing method and device, electronic equipment and nonvolatile storage medium

The invention discloses a video editing method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: acquiring an audio data stream of a target video, and segmenting the audio data stream into a plurality of audio clips; determining a classification result corresponding to the audio clip, and determining a time period corresponding to the audio clip as a candidate ball hitting time period when the classification result is that the audio clip contains the ball hitting sound; acquiring a video frame corresponding to the candidate ball hitting time period in the target video, and judging that the candidate ball hitting time period is a real ball hitting time period under the condition that the visual feature of the video frame is matched with a preset ball hitting rule; and determining an editing time point according to the real ball hitting time period, and editing the target video according to the editing time point to obtain a ball hitting round video clip. The technical problem that the accuracy of ball hitting detection segment detection is low due to the fact that ball hitting detection is conducted only through sound and is easily interfered by environmental noise in the prior art is solved.
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Conference summary method based on AI large model

The invention relates to a conference summary summarization method based on an AI large model, and belongs to the technical field of voice processing, and the technical scheme specifically comprises the following steps: obtaining audio information, recognizing all spokesmen involved in the audio information by using a preset spokesman separation algorithm, and splitting the audio information into a plurality of audio segments according to the spokesmen; analyzing and processing each audio segment by using a preset voice transcription algorithm to generate a transcription text; and based on a preset large language model, extracting and summarizing the transcriptional text, and generating and outputting summarized contents for describing the audio information, so that a user can know the summarized contents. The method has the effect of optimizing the information definition and the text readability of the transcribed text.
Owner:SUZHOU CHUANGLUTIANXIA INFORMATION TECH CO LTD

Target voice regulation and control method, device and equipment based on intelligent glasses

The embodiment of the invention discloses a target voice regulation and control method, device and equipment based on intelligent glasses, belongs to the technical field of intelligent wearing, and solves the problem that in the prior art, a user is difficult to obtain a voice signal meeting the own requirement due to the fact that standardization processing is generally performed on a voice signal. Comprising the following steps: acquiring audio information of a speaker in a current scene, and dividing the audio information into a plurality of audio clips; acquiring portrait information in the current scene, and associating the portrait information with the corresponding audio clip to determine a target speaker; determining corresponding sound preference information of the user in the current scene based on a mapping relation between the noise level and intelligent glasses use preference data; analyzing the sound of the target speaker, responding to a dynamic sound adjustment strategy based on the obtained sound characteristics, and optimizing the sound of the target speaker; and based on the sound preference information, performing frequency band adjustment on the optimized sound through a power amplifier arranged on the intelligent glasses to realize target voice regulation and control.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Text to synchronized joint video and audio generation

Methods and apparatus for generating a synchronized video-audio pair. According to an example embodiment, a method of generating a synchronized video-audio pair includes: applying one or more text inputs to a generative model, the one or more text inputs including a speech text, the generative model including a neural network; with the generative model, converting the speech text into an audio segment for the synchronized video-audio pair; and with the generative model, generating a video segment for the synchronized video-audio pair, the video segment including a talking head having lip movements corresponding to the speech text and in synchronization with the audio segment.
Owner:DOLBY LABORATORIES LICENSING CORP

Voice interaction large model family health assistant dialogue method, device and equipment and medium

The invention relates to a voice interaction large model family health assistant dialogue method and device, equipment and a medium. The method comprises the following steps: carrying out fragmentation processing according to an original voice stream of a user to generate an audio fragment with a medical mark, and carrying out voice recognition and entity extraction on the audio fragment to generate a dynamic entity map; generating an evidence-based decision prompt based on the map, inputting the prompt into a preset medical big model for processing, and outputting a result containing an essential symptom list; and according to the symptom matching degree of the necessary symptom list and the dynamic entity map, generating a diagnosis report or a question-asking list, if the diagnosis report is output, performing medical rule chain verification operation on the diagnosis report to generate a quality control report, and based on the question-asking list or the quality control report, generating a synthetic voice stream. According to the method, through medical intention directional screening, map entity analysis, large model diagnosis, voice synthesis and the like, the voice recognition accuracy, the diagnosis suggestion reliability and the inquiry interaction efficiency of the family health assistant in the medical scene are improved.
Owner:SHANGHAI LOHAS YUAN MEDICAL TECHNOLOGY CO LTD

Speaker role separation method, related device and computer program product

The invention discloses a speaker role separation method, related equipment and a computer program product, and the method comprises the steps: determining a transcription result of target audio data, the transcription result comprising a transcription text and speaker turning point information; and carrying out time alignment on the transliteration text and the target audio data, and segmenting to obtain a single-person audio clip. The voiceprint information is extracted by sliding the window in each single-person audio clip, and the single-person audio clip only contains the content of a single speaker, so that the window length can be increased without worrying about the risk that multiple persons speak in the window, and the voiceprint accuracy can be improved through the larger window length. All sliding window audios are clustered according to voiceprint information in a global range, text fragments corresponding to the sliding window audios of the same cluster are endowed with the same speaker identity, text fragments corresponding to the sliding window audios of different clusters are endowed with different speaker identities, the accuracy of speaker role separation on a transliteration text level can be improved, and the user experience is improved. The method is especially suitable for long audio scenes.
Owner:HKUST IFLYTEK (SHANGHAI) TECH CO LTD

Audio processing method and device, storage medium and electronic device

The invention discloses an audio processing method and device, a storage medium and an electronic device, and relates to the technical field of data processing, and the audio processing method comprises the steps: obtaining voice feature data and semantic feature data of original audio data; the original audio data is audio data extracted from the target audio and video data; generating segmentation position data according to the voice feature data and the semantic feature data; and generating audio paragraph data and paragraph visualization data according to the segmentation position data, and providing the paragraph visualization data to a user interaction interface. According to the technical scheme of the embodiment of the invention, the utilization rate of the audio information can be improved, the data processing cost in an application scene is remarkably reduced, and the data application efficiency and scene adaptability are improved.
Owner:QINGDAO HAIER TECH +2

Virtual digital human lip synchronization optimization method, device, equipment and storage medium

The present invention relates to the field of computer vision technology, and discloses a virtual digital human lip synchronization optimization method, device, equipment and storage medium. The virtual digital human lip synchronization optimization method includes: obtaining the target audio segment to be output by the virtual digital human at the next moment; judging whether the target audio segment belongs to the audio type to be processed; if the target audio segment belongs to the audio type to be processed, then based on a preset lip synchronization optimization strategy, generating a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment; based on the 3D human face mouth shape parameter frame sequence, generating a corresponding 3D human face mouth shape image frame sequence and rendering it into the virtual digital human. The present invention can adapt to various audio types and improve the fluency and naturalness of the virtual digital human's mouth shape under different audio types.
Owner:GUANGZHOU HUYA TECH CO LTD

Audio processing method and device, electronic equipment, computer readable storage medium and computer program product

The invention provides an audio processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring a structured list pre-generated for an audio to be edited; analyzing the structured list to obtain metadata of each audio clip; in response to an editing operation for the first audio clip, updating the metadata of the first audio clip to obtain updated metadata; obtaining the first audio clip based on the index of the first audio clip, and decoding the first audio clip to obtain corresponding audio data; creating an audio buffer source node, configuring playing parameters of the audio buffer source node according to the updated metadata, and calling the configured audio buffer source node to play audio data; and in response to receiving the file export request, constructing an offline audio context based on the updated metadata, and generating an audio file through the offline audio context. According to the invention, the audio editing efficiency can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-modal emotion recognition method and device, electronic equipment, storage medium and product

The embodiment of the invention provides a multi-mode emotion recognition method and device, electronic equipment, a storage medium and a product, and relates to the technical field of emotion recognition. The method comprises the following steps: acquiring a to-be-recognized audio / video which comprises an audio stream and a video stream, segmenting the audio stream to obtain at least one audio segment, inputting each audio segment into an audio recognition model to obtain an audio recognition result, determining a corresponding video segment in the video stream according to a target audio segment of which the audio recognition result is an emotion result, and inputting the video segments into a video recognition model to obtain a video recognition result, and determining a target emotion result of the to-be-recognized audio and video based on the audio recognition result and the video recognition result. According to the embodiment of the invention, the video emotion recognition is used for assisting the audio emotion recognition to complete the emotion recognition of the audio and the video, errors possibly caused by single audio recognition are avoided, and the recognition accuracy can be improved.
Owner:BEIJING VISION WORLD TECH CO LTD

Audio identification method and device based on marking and backtracking correction

The invention relates to an audio recognition method and device based on marking and backtracking correction, and the method comprises the steps: segmenting a target audio, and enabling adjacent segments after segmentation to have a partial overlapping region; if a slice point exists in a non-mute segment of the target audio and the acoustic feature similarity of a preset number of frames before and after the slice point is smaller than a set similarity threshold value, the slice point is marked as a cut-off risk point, and the cut-off risk point is used for indicating that continuous semantics before and after the slice point has a cut-off risk; after a text corresponding to the audio of each slice is recognized through the speech recognition model, the language confidence degree of an overlapping area where the truncation risk point is located is determined, and the language confidence degree is used for indicating context language logic of the overlapping area; and if it is determined that the language confidence is lower than a set confidence threshold, correcting the recognition texts of the slices before and after the truncation risk point. According to the invention, the accuracy of audio recognition is improved.
Owner:FIBOCOM WIRELESS

Multi-application audio splicing

Systems and methods for multi-application audio splicing may include (1) receiving, from a third-party application installed on a user's device, an audio segment of an audio file selected by the user via an audio-segment-selection interface presented by the third-party application in association with the audio file, (2) loading the audio segment into a post-creation interface of a social media application installed on the user's device and prompting the user to add visual content via the post-creation interface, (3) creating a social media post that includes (i) the audio segment and (ii) visual content added by the user via the post-creation interface, and (4) posting the social media post to a social media consumption channel. Various other methods, systems, and computer-readable media are also disclosed.
Owner:META PLATFORMS INC

Chinese acoustic text depression detection system based on channel attention mechanism fusion

The invention provides a Chinese acoustic text depression detection system based on channel attention mechanism fusion, and the system obtains an original audio signal of a target object through each unit in the system, carries out the preprocessing, obtains audio data, carries out the text transcription, obtains the corresponding text data, carries out the fragment division of the audio data, and carries out the segmentation of the audio data. Corresponding audio clips are obtained, and audio-text data are formed; performing feature extraction on the audio clips and the text data to obtain low-level descriptor features, high-level time-frequency features, audio waveform features and text features; fusing the low-level descriptor features, the high-level time-frequency features and the audio waveform features step by step to determine acoustic features, performing cross-modal fusion on the acoustic features and the text features to obtain fusion features, performing depression detection according to the fusion features, and determining a depression detection result of the target object. According to the scheme, more comprehensive acoustic expression and text features are used, voice text modes are effectively fused, depression detection is achieved, and clinical doctors are assisted in diagnosis.
Owner:SHENZHEN NANSHAN DISTRICT CHRONIC DISEASE CONTROL CENT (SHENZHEN NANSHAN DISTRICT MENTAL HEALTH CENT)

Automatic De-identification of Sensitive Conversational Audio Data

Techniques for automatically de-identifying sensitive information in audio conversations by combining un-transcribed voice activity detection (VAD) with large language model (LLM) analysis are disclosed. An audio de-identification system processes speech-to-text transcriptions while identifying segments where automatic speech recognition (ASR) failed to transcribe spoken content. These un-transcribed segments are represented as placeholders in prompts sent to an LLM, which analyzes the surrounding textual context to determine if sensitive information (such as PII or PHI) was likely spoken during these gaps. When sensitive content is identified, the system modifies the corresponding audio segments through an audio identification tactic. This approach addresses the technical challenge of incomplete de-identification in automated audio processing by leveraging LLMs' contextual understanding to detect sensitive information in segments that traditional ASR systems miss, particularly in scenarios involving poor audio quality or diverse accents. The result is a more comprehensive and reliable audio de-identification system.
Owner:ORACLE INT CORP

Dynamic systems and methods for media-aware transport of fragment of content in low-latency, over-the-top, and adaptive bitrate streaming

Low latency, over-the-top (OTT), and / or adaptive bitrate (ABR) content streaming is provided. Content delivery is enhanced by determining if a fragment of a content segment at a content delivery network (CDN) edge node meets a threshold for preferential encapsulation and transport. If met, preferential encapsulation and transport to the client device is provided; otherwise, it defaults to non-preferential encapsulation. The size of the fragment is quantified at a parser of the CDN edge node or an ABR segment encryption system. The ABR system may be connected between a content source and a CDN origin and may include an encryptor that sends CMAF video and audio segment's fragment byte offsets metadata. Also, the CDN edge node may include the ABR system and an encryptor that sends an encrypted CMAF segment's fragment size to a threshold calculator of an HTTP server. Related apparatuses, devices, techniques, and articles are also described.
Owner:ADEIA GUIDES INC

Display control method of word mentioning content, terminal equipment and storage medium

The embodiment of the invention provides a prompt content display control method, terminal equipment and a storage medium. The method comprises the following steps: acquiring a target audio clip; obtaining K text blocks most similar to the semantics of the target audio clip from a plurality of pre-stored text blocks as K candidate text blocks; determining context text blocks of the candidate text blocks, and determining a context consistency coefficient according to the similarity between the target audio clip and each context text block and the similarity between the candidate text block and each context text block; according to the similarity between the target audio clip and the candidate text blocks and the context consistency coefficient, determining the target similarity between the target audio clip and the candidate text blocks, and determining the candidate text block corresponding to the highest target similarity in the K candidate text blocks as a target text block; and switching the displayed prompt content into the manuscript content related to the target text block in the electronic manuscript. And the accuracy of the displayed word mentioning content is improved.
Owner:ZHUHAI MOJIE TECH CO LTD

AI-based short video creative material cutting method and device, equipment and medium

PendingCN121284336ASelective content distributionArtistic renderingAlgorithm
The invention relates to an AI-based short video creative material cutting method and device, equipment and a medium. The method comprises the following steps: performing space-time segmentation on an original video material through a pre-training model to generate a video clip set carrying semantic tags and emotion vectors, and performing beat segmentation on an audio material to generate an audio clip set containing semantic tags, emotion vectors and rhythm features; calling knowledge graph analysis based on the target artistic effect type to generate an inconsistency rule set; calculating a conflict value of video and audio clip combinations through a creative tension scoring function, and screening high-score combinations to construct a candidate set; utilizing a sequence optimization algorithm to generate a fragment sequence for maximizing artistic tension and a cutting time point and a transition instruction of the fragment sequence; and synthesis is realized through an editing engine. Therefore, control over artistic audio-visual conflicts is achieved on the premise that basic continuity is guaranteed, the limitation that an existing AI editing system can only be mechanically matched with sounds and pictures is broken through, and the ability of building deep emotion expression is given to a machine for the first time.
Owner:CHENGDU YISHI INNOVATION INFORMATION TECHNOLOGY CO LTD

OSAHS screening method based on sleep breathing sounds and related equipment

PendingCN120690234ASpeech analysisBiological modelsNoiseAudio segmentation
The invention discloses an OSAHS screening method and related equipment based on sleep breath sounds, and the method comprises the steps: inputting an obtained audio signal into an OSAHS screening model, outputting an AHI value, and then obtaining an OSAHS symptom level; the training steps of the model include: performing audio segmentation and automatic labeling on sleep recording signals, and dividing tags into reliable tags and unreliable tags according to the reliability of the tags; training a noise label classifier by using a label correction algorithm, and iteratively identifying and correcting potential error labels in the unreliable labels in the training process of the noise label classifier; after the tag correction is completed, training an audio clip classifier by adopting the corrected tag; and fitting a linear regression model to predict an AHI value based on a classification result of the whole sleep recording signal. The label correction link is introduced in the model training process, potential wrong labels in unreliable labels are corrected, the accuracy of the model is effectively improved, and the method can be widely applied to the technical field of disease diagnosis equipment.
Owner:SOUTH CHINA UNIV OF TECH

Systems and methods for transforming digital audio content

A system for platform-independent visualization of audio content, in particular audio tracks utilizing a central computer system in communication with user devices via a computer network. The central system utilizes various algorithms to identify spoken content from audio tracks and identifies “great moments” and / or selects visual assets associated with the identified content. Audio tracks, for example Podcasts, may be segmented into topical audio segments based upon themes or topics, with segments from disparate podcasts combined into a single listening experience, based upon certain criteria, e.g., topics, themes, keywords, and the like.
Owner:TREE GOAT MEDIA LLC

Double-recording quality inspection method, system and device and electronic equipment

The invention provides a double-recording quality inspection method, system and device and electronic equipment. The method comprises the steps that a double-recording information stream is acquired, and the double-recording information stream represents audios and videos acquired by real-time recording of conversations between salesmen and clients; performing preprocessing operation on the double-recording information flow to obtain audio segments and picture frames, performing visual navigation classification processing on the picture frames to obtain classified pictures, and performing preliminary detection processing on the audio segments and the picture frames through local end quality inspection service to obtain a first detection result; the preliminary detection processing comprises blank screen detection and / or leaving detection of picture frames and audio detection of audio segments; and performing deep detection processing on the audio segments and the classified pictures through the cloud quality inspection service to obtain a second detection result, and merging the first detection result and the second detection result to obtain a double-recording quality inspection report corresponding to the double-recording information stream. The problem of low double-recording quality inspection efficiency of a double-recording quality inspection scheme in the prior art is solved.
Owner:中国邮政储蓄银行股份有限公司

Voice analyzer for interactive care system

A support interaction is guided in real time by generating from audio content featurized audio data that includes audio segments and audio features; generating in real time classification scores associated with certain audio segments; and displaying in real time the classifications scores and information associated with the corresponding audio segments.
Owner:LIVE CIRCLE INC

Text segmentation voice streaming processing method, system and device, medium and program product

The invention discloses a text segmentation voice streaming processing method, system and device, a medium and a program product, and the method comprises the steps: firstly, carrying out segmentation processing according to a semantic structure of a long text, and obtaining ordered text segments with complete semantics; speech synthesis is carried out on the first text segment to obtain a first audio segment, and SSE connection with the client is established while the first audio segment is synthesized; after the synthesis of the first audio clip is completed, the first audio clip is immediately transmitted to a segmentation buffer queue of the client through SSE connection, and meanwhile, speech synthesis is sequentially performed on subsequent text clips; and after the segmented buffering queue receives the first audio clip, the first audio clip is played immediately, and subsequent audio clips are continuously received and buffered while the first audio clip is played. According to the invention, the coordinated control of the AI speech synthesis progress and the streaming transmission progress is realized, the response speed of AI speech synthesis is obviously improved, and the switching fluency of segmented audios is effectively improved.
Owner:BEIJING DIANFU TECHNOLOGY CO LTD

Service quality detection method and device

The embodiment of the invention discloses a service quality detection method and device. The method comprises the following steps: acquiring a to-be-detected service voice audio of a target service scene to which a target object belongs; determining an overlapped voice audio segment and a single-person voice audio segment based on the voice overlapping attribute and the voice sounding change attribute of each audio frame in the to-be-tested service voice audio; based on each single-person voice audio segment and preset voiceprint information of the target object, determining a target single-person audio segment corresponding to the target object; for each overlapped voice audio segment, based on the current overlapped voice audio segment, the corresponding video stream data and a voice separation model, separating a target separation audio segment corresponding to the target object; and the service quality attribute corresponding to the target object is determined based on the target single-person audio segment and the target separation audio segment, so that the problem of disordered voice recognition results is solved, the service quality of the server is detected and evaluated based on the voice recognition results, and the accuracy and reliability of the service quality evaluation result are improved.
Owner:BEIJING YIBAIYISHIYI MEDICINE SCI & TECH CO LTD

Audio data processing method and device, equipment and storage medium

The invention relates to the field of audio data analysis, in particular to an audio data processing method and device, equipment and a storage medium. The method comprises the following steps: carrying out multi-angle audio continuous acquisition and scene noise suppression on a conference room through a surrounding microphone array, and generating a filtering optimization audio signal; performing three-dimensional time difference positioning calculation according to the filtered and optimized audio signal to obtain accurate sound source positioning information; personalized voiceprint analysis is carried out on the filtered and optimized audio signals, audio distribution modeling is carried out based on accurate sound source positioning information, and a multi-person conference audio field is constructed; performing parallel audio stream separation according to the multi-person conference audio field to generate an intelligent spliced audio segment; and performing deep semantic analysis and semantic logic correction on the intelligent spliced audio segment to generate an audio analysis result. According to the invention, rapid and accurate speaker audio recognition of a multi-person parallel conference is improved.
Owner:SHENZHEN ULTRA EASY TECH CO LTD

Automated segmentation and transcription of unlabeled audio speech corpus

ActiveUS12512100B2Speech recognitionTimestampAudio segmentation
A method includes obtaining initial transcription for input natural speech; performing segmentation of initial transcription into text portions, based on punctuation marks in initial transcription; determining segment-level timestamps for text portions based on the input natural speech; performing audio segmentation on input natural speech, by cutting input natural speech based on segment-level timestamps, to obtain audio chunks; generating transcription portions for each of the audio chunks; merging transcription portions to form re-transcription; determining word-level timestamps for re-transcription, by aligning input natural speech against re-transcription; calculating silence time periods, each corresponding to silence between each two adjacent words of input natural speech, based on word-level timestamps; performing a final segmentation on input natural speech and re-transcription, based on silence time periods, to generate final audio segments and corresponding final transcription portions. The final audio segments and corresponding final transcription portions may be included in training dataset for training a model.
Owner:ORACLE INT CORP

Voice wake-up interaction method and system based on microphone array

The invention discloses a voice wake-up interaction method and system based on a microphone array, and the method comprises the steps: carrying out VAD processing, so as to judge whether a target audio segment has voice or not; voice and noise source directions are obtained; selecting beam parameters of voice and noise in combination with a pre-designed fixed beam; whether GSC module processing is carried out or not is selected according to the difference between the beam parameters of the voice and the noise, so that an enhanced audio signal is obtained, a wake-up task is carried out, and a final voice wake-up result is obtained; and according to whether the wake-up is successful, determining whether to lock the voice beam direction in the current interaction stage for enhancement, thereby preventing interference of voice in other directions on subsequent interaction tasks. According to the voice enhancement mode based on the microphone array, sound source positioning can be realized, interference in a non-target direction can be suppressed, the voice quality in the target direction can be improved, the wake-up success rate can be effectively improved when the voice enhancement mode is applied to voice wake-up, and then the experience of back-end voice interaction is improved.
Owner:PANOVASIC TECHNOLOGY CO LTD

Methods and systems to provide a playlist for simultaneous presentation of a plurality of media assets

Systems and methods are described herein for generating a playlist for a simultaneous presentation of a plurality of media assets. The system retrieves a user preference associated with a user profile and receives a selection of a first media asset and a second media asset from the plurality of media assets for presentation on a user device. The system parses the respective audio streams of the first media asset and the second media asset to identify one or more preferred audio segments based on the user preference and generates the playlist of the identified one or more preferred audio segments. Based on a generated audio playlist, the system generates, for presentation on the user device, the video stream for each of the first media asset and the second media asset and the playlist of the identified one or more preferred audio segments.
Owner:ADEIA GUIDES INC

Audio processing method and device, equipment and storage medium

The embodiment of the invention relates to an audio processing method and device, equipment and a storage medium. The method provided herein includes: extracting background audio content and text audio content corresponding to text content from first media content; based on a request for replacing a first text in the text content with a second text, generating a second audio segment corresponding to the second text by using timbre information associated with a first audio segment of the text audio content, the first audio segment corresponding to the first text; adjusting a third audio segment corresponding to the first text in the background audio content based on the second audio segment; and generating second media content based on the first media content, the second audio segment, and the adjusted third audio segment. In this way, the embodiment of the invention can support editing of the media content by modifying the text corresponding to the media content, and can make the modified audio content more real, thereby improving the quality of the edited media content.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Multi-language voice content recognition method and system

The invention relates to a multi-language voice content recognition method and system, and belongs to the technical field of voice signal processing, and the recognition method comprises the steps: collecting original audio stream data, executing the noise reduction filtering processing, and segmenting the original audio stream data into a plurality of audio segments; extracting acoustic feature vectors of the audio clips; inputting the audio clip into the speech recognition model group, and obtaining a text clip and a confidence score; fusing all the text fragments to generate a preliminary recognition text, segmenting the preliminary recognition text into text blocks, and translating the text blocks into a target language text block by block through a streaming translation model in combination with a semantic feedback tag received in real time; executing bidirectional translation verification on the text blocks of which the confidence scores are lower than a set threshold value; inputting the translated text stream and the acoustic feature vector into a semantic analysis model, and fusing an output result to obtain a semantic feedback tag; and constructing a cross-modal association graph through the graph neural network, outputting a key information abstract and triggering an alarm signal. The speech recognition accuracy can be improved while the recognition efficiency is guaranteed, and risk response is carried out.
Owner:BEIJING HIZHI TECH CO LTD

Personalized voice interaction method and related equipment

The embodiment of the invention provides a personalized voice interaction method and related equipment, and belongs to the technical field of intelligent voice interaction. The method comprises the following steps: receiving an object question audio stream by adopting a streaming processing technology in response to a voice interaction service request; performing segmentation operation and preprocessing operation on the object question audio stream to obtain object question audio clips; performing object feature analysis on the object question audio clip to obtain a service object tag; carrying out object group classification on the service object labels, and selecting a corresponding voice recognition model according to a classification result of an object group; sequentially performing voice recognition on the object question audio clips by adopting a voice recognition model to obtain an object question text; and performing text analysis on the object question text to obtain an answer response. According to the scheme, the questioning audio can be accurately recognized, and the user experience of intelligent voice interaction service is improved.
Owner:E-SURFING DIGITAL LIFE TECH CO LTD