Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1863 results about "Voice data" patented technology

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Digital large screen interaction method and device based on multi-agent cooperation and medium

The invention discloses a digital large-screen interaction method and device based on multi-agent cooperation and a medium, and relates to the technical field of digital large-screen interaction.The method comprises the steps that postures, expressions and voice data of a user are collected in real time, the attention weight is calculated in combination with spatial position information, the emotional state change rate is monitored, and the user experience is improved in combination with personal historical browsing preferences; acquiring a main interaction user, an emotional state feature and an emotional change trend; according to the main interaction user, the emotional state features and the emotional change trend, obtaining an initial confidence coefficient of the suspected interested field through a semantic understanding model; and based on the candidate guide images, identifying a user selection intention through a multi-modal fusion processing mechanism, calculating a selection probability, triggering content activation of the main interaction area and content dynamic generation of the auxiliary information area, synchronizing state information, and starting an immersive collaborative interaction process. Through three-level collaborative service configuration, the beneficial effects of improving large screen content organization efficiency and enhancing immersive interaction experience are achieved.
Owner:YLZ INFORMATION TECHNOLOGY CO LTD

Knowledge graph construction method and device based on multi-modal clinical data, equipment, medium and product

PendingCN121235062AMedical data miningSemantic analysisMedical knowledgePatient communication
The invention discloses a knowledge graph construction method and device based on multi-modal clinical data, equipment, a medium and a product, and relates to the field of knowledge graph construction, and the method comprises the steps: constructing a medical knowledge type knowledge graph; based on the dynamic event data of the patient and the doctor-patient dialogue data, constructing a medical event type knowledge graph; fusing the medical knowledge type knowledge graph and the medical event type knowledge graph to obtain a fused medical knowledge graph; collecting multi-modal data in the diagnosis process of the patient, and fusing the multi-modal data with the fused medical knowledge graph to obtain a finally constructed knowledge graph; the multi-modal data comprises text data, image data, voice data and physiological signal data. The problems that clinical knowledge is fragmented, deep mining is difficult, data cannot be associated in a unified mode, and doctor-patient communication knowledge is fragmented can be solved.
Owner:ZHEJIANG UNIV

Blind guiding robot interactive navigation system combining vision, inertial navigation and voice

The invention relates to the technical field of blind guiding robot interactive navigation, and particularly discloses a blind guiding robot interactive navigation system combining vision, inertial navigation and voice, which is characterized in that environmental vision, movement and voice data are synchronously acquired through a multi-modal perception data acquisition module, and are analyzed through a situation-user joint state feature extraction module; a user potential intention deduction module is combined to predict a user intention and generate a self-adaptive guide strategy; a self-adaptive guiding path optimization module adjusts a navigation path according to a strategy to ensure safety and high efficiency; finally, the situational multi-mode interaction instruction output module provides voice and tactile feedback, and interaction experience and navigation autonomy are enhanced. According to the method, the travel problem of the visually impaired person in a complex environment is solved, the individuation, context perception capability and safety of blind guiding are improved, and more natural and efficient navigation assistance is brought to the visually impaired person.
Owner:SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD

Method and system for generating chronic disease intervention scheme based on reinforcement learning and multi-modal data

The invention relates to the technical field of reinforcement learning, and discloses a chronic disease intervention scheme generation method and system based on reinforcement learning and multi-modal data, and the method comprises the steps: obtaining the multi-modal data of a patient, the multi-modal data at least comprising electronic medical record data, voice data, image data, text data and physiological time sequence data; performing feature extraction on the multi-modal data to obtain a multi-modal feature vector; on the basis of the multi-modal feature vectors, health state vectors are constructed, and the health state vectors at least comprise a physiological risk score, a treatment compliance score and a lifestyle health degree score; designing a reward function according to the dynamic change of the health state vector; optimizing the strategy network in a predefined intervention action space according to the reward function by utilizing a reinforcement learning algorithm so as to output an optimal intervention action; and converting the optimal intervention action into personalized natural language interaction content through a generative AI model. According to the invention, the efficiency, precision and patient compliance of chronic disease management can be significantly improved.
Owner:GUANGDONG URBAN & RURAL PLANNING & DESIGN INST

Voice data recognition method and system based on AI voice algorithm

The invention discloses a voice data recognition method and system based on an AI voice algorithm, relates to the technical field of AI voice recognition, and solves the problem that the voice data recognition capability is low. The method comprises the following steps: S1, multi-mode cooperative triggering collection: synchronously collecting lip electromyographic signals and voiceprint features through a multi-mode sensor, an activation instruction is generated through feature fusion, and voice acquisition starting is triggered; s2, AI adaptive noise reduction processing: carrying out noise separation on the original audio signal by adopting a generative adversarial network, separating environmental noise features to generate a dynamic noise reduction mask, and keeping the integrity of human voice features; s3, beam dynamic optimization adjustment: analyzing real-time audio quality based on a reinforcement learning algorithm, dynamically adjusting beam pointing and gain parameters of a microphone array, and focusing a target sound source; and S4, semantic association cache enhancement: carrying out real-time semantic analysis on the collected voice data. According to the invention, the voice data recognition capability of an AI voice algorithm is greatly improved.
Owner:HUAQIAO UNIVERSITY

Voice data compression and index storage method based on voiceprint template

The invention discloses a voice data compression and index storage method based on a voiceprint template, and relates to the technical field of voice compression processing. The voice data compression and index storage method based on the voiceprint template comprises the following steps: S1, collecting and preprocessing acoustic state data, coded data and edge synchronization data, and constructing a standardized voice input data set; s2, evaluating the individual sound channel difference degree, and stripping non-universal components highly related to individuals in the voice; s3, analyzing the audio compression content, and constructing a structured total compression amount; s4, performing joint evaluation on real-time transmission regulation and control requirements of the compressed data, and dynamically adjusting an uploading priority and a continuous transmission strategy; and S5, constructing a multi-dimensional reverse index system at the cloud. The problems that for a high-noise far-field pickup scene, an existing compression method does not consider that individual sound channel characteristics are stripped from original signals, so that compression redundancy is high, and retrieval interference is large are solved.
Owner:HUNAN UNIV +1

Mild cognitive impairment screening and diagnosing system based on deep learning

The invention discloses a mild cognitive impairment screening and diagnosis system based on deep learning, and relates to the technical field of cognitive impairment screening and diagnosis, and the system comprises a data collection and extraction module which is used for collecting voice data and a digital drawing track of a subject; the low-frequency window determination module is used for determining a low-frequency window according to the voice envelope and the handwriting speed; the lag index calculation module is used for obtaining a speech writing low-frequency lag index based on the phase relation and the correlation intensity; the disturbance quantity extraction module is used for extracting a listening disturbance quantity and a motion disturbance quantity; the depolarization residual calculation module is used for calculating the depolarization residual of the speech writing low-frequency lag index; and the risk calculation and grading module is used for calculating the risk probability and determining a grading result. According to the method, the interference of non-cognitive factors on mild cognitive impairment screening can be eliminated, and the screening accuracy is improved.
Owner:SHENZHEN SECOND PEOPLES HOSPITAL (SHENZHEN INST OF TRANSLATIONAL MEDICINE)

Language barrier execution type intervention effect evaluation method based on deep learning

The invention relates to the technical field of language barrier evaluation, and discloses a deep learning-based language barrier executive intervention effect evaluation method. The method comprises the following steps: acquiring real-time voice data and facial expression data of a language barrier patient in intervention training through a multi-modal data acquisition device to form an original behavior feature set; performing acoustic feature hierarchical analysis on the voice data by adopting a time sequence feature extraction network to generate a voice time sequence feature vector; performing micro-expression dynamic capture on the expression data through a three-dimensional convolutional neural network to generate an expression state feature vector; inputting the two types of vectors into a multi-modal feature fusion layer to carry out cross-modal correlation analysis, and generating a comprehensive behavior evaluation matrix; on the basis of the matrix, an intervention effect analysis model driven by an attention mechanism is adopted, a behavior improvement degree index of the current intervention stage is calculated, accurate evaluation of the intervention effect is achieved, and support is provided for dynamic adjustment of language barrier rehabilitation intervention.
Owner:SHANDONG VOCATIONAL COLLEGE OF SPECIAL EDUCATION

Lip shape synchronization model training method, digital human video generation method and device

The invention provides a training method of a lip shape synchronization model and a generation method and device of a digital human video. The invention relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, virtual reality, large models, digital human broadcast, digital human live broadcast and the like. According to the specific scheme, the method comprises the steps of determining a two-dimensional face feature truth value and a three-dimensional face feature truth value based on a video of a target object; extracting a three-dimensional face feature prediction value from the sample voice data through the first network sub-model; generating a two-dimensional face feature predicted value from the three-dimensional face feature predicted value through a second network sub-model; based on the two-dimensional face feature truth value, the three-dimensional face feature truth value, the two-dimensional face feature predicted value and the three-dimensional face feature predicted value, a first network sub-model and a second network sub-model are trained, a lip shape synchronization model is obtained, and the lip shape synchronization model comprises the first network sub-model and the second network sub-model.
Owner:BEIJING DUSHANG SOFTWARE TECH CO LTD

Model compression and data enhancement fused lightweight deep counterfeit voice detection method

The invention provides a model compression and data enhancement fused lightweight deep forged voice detection method. The method comprises the steps of obtaining and processing a public voice data set and a large-scale self-supervision pre-training voice model; performing structured pruning and knowledge distillation to obtain a lightweight voice model; performing audio preprocessing and diversified data enhancement on the true and false voice samples to obtain an enhanced true and false voice data set; performing faking task joint fine tuning on the lightweight voice model to obtain a lightweight deep faking voice detection model; and locally deploying the model to obtain a localized counterfeit voice detection system, and carrying out real-time authenticity judgment. According to the method, calculation complexity and reasoning time delay are remarkably reduced, local deployment is carried out on a resource-limited end side platform, detection generalization and robustness are improved, and low-delay, low-power-consumption and high-robustness detection performance is achieved.
Owner:ZHEJIANG UNIV

High-performance voice processing method

The invention relates to a high-performance voice processing method, and the method comprises the steps: obtaining voice data transmitted through a network, generating corresponding scene fingerprint information, determining a frame parameter of a lightweight multi-scene adaptation frame, and a corresponding processing parameter, and if the duration of a lost segment is smaller than a preset threshold value, determining that the segment is lost. If yes, generating a corresponding first compensation result according to the acoustic compensation parameter; if the duration is larger than or equal to a preset threshold value, whether the acoustic compensation process and the semantic processing process are performed in parallel or not is judged according to the frame parameters, if not, a corresponding second compensation result is generated, and if yes, a corresponding second compensation result is generated according to the acoustic compensation process and the semantic processing process; and outputting the corresponding target voice transmission data. The definition and the stability of the voice signal in a complex environment can be effectively improved, different compensation modes are flexibly switched or executed in parallel according to the packet loss duration and the resource state, and the effectiveness and the stability of a recovery mechanism are improved.
Owner:SHENZHEN SOUNDFIT TECH CO LTD

Intelligent call auxiliary system for old people

The invention relates to the technical field of data processing, in particular to an intelligent call auxiliary system for old people, which comprises an acquisition module used for acquiring voice data flow of a user in a call process in real time; the processing module is used for analyzing the voice data flow to obtain analyzed data; performing risk keyword identification processing on the analysis data to obtain a risk keyword identification result; performing emotion feature analysis processing on the analysis data to obtain an emotion feature analysis result; constructing a three-dimensional feature distribution model based on the risk keyword recognition result and the emotion feature analysis result; calculating volume parameters of the three-dimensional feature distribution model in a preset feature space; and the risk keyword recognition result, the emotion feature analysis result and the volume parameters jointly form multi-modal feature data. According to the invention, the perception, evaluation and control capabilities of the auxiliary system in a complex call scene are improved, and the call experience and communication efficiency of the elderly user are guaranteed.
Owner:EAST CHINA NORMAL UNIV

Voice QoS optimizing and playing method, device and equipment facing end-to-end information source encryption and medium

The invention discloses a voice QoS optimizing and playing method, device and equipment for end-to-end information source encryption and a medium, and relates to the field of instant messaging, and the method comprises the steps: storing a received encrypted voice data packet into a jitter buffer area in a ciphertext form; determining a current target buffer water level of the jitter buffer based on the current network jitter degree and a delay factor introduced by the current decryption processing; generating a play control strategy and a decryption scheduling strategy according to the current actual buffer water level and the current target buffer water level of the jitter buffer; the decryption scheduling strategy comprises a decryption operation frequency, a triggering opportunity and an execution priority, and the decryption scheduling strategy is adjusted based on a real-time playing control strategy and a jitter buffer area state; and scheduling the corresponding encrypted voice data packet from the jitter buffer area by using a decryption scheduling strategy for decryption, and decoding and playing the decrypted voice data packet according to a playing control strategy. According to the invention, the overall fluency of encrypted real-time voice communication is improved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Time-frequency domain speech enhancement method based on KAN channel attention

The invention relates to the technical field of speech enhancement, in particular to a time-frequency domain speech enhancement method based on KAN channel attention, which comprises the following steps: processing a speech data set to obtain frequency domain representation; extracting local features in the frequency domain representation through an encoder to obtain output features; the output features are input into a TF-Transform block, noise components are recognized and suppressed, and output components are obtained; inputting the output component into a KAN-based channel attention module, sequentially obtaining a channel feature map and a spatial feature map, and successively multiplying the output component by the channel feature map and the spatial feature map and introducing jump connection to obtain refined features; and obtaining a time domain voice signal based on refined feature recovery, and processing the recovered time domain voice signal through an amplitude decoder and a phase decoder to correspondingly obtain an amplitude spectrum and a phase spectrum. And through a KAN-based channel attention module, the denoising and reconstruction performance of sparse and sensitive high-frequency components in the voice is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

Method and device for demonstration-based robot programming supplemented by video

A method of programming an industrial robot, comprising: recording movements of a robot manipulator during a demonstration-based programming session, to obtain a robot trajectory (301); using a natural-language model (171), parsing speech data describing the programming session into at least one robot parameter (305) value and annotating the robot trajectory with the robot parameter value; and, on the basis of the annotated robot trajectory (306), generating a robot program (307). The method further comprises: capturing a video (302) of the programming session; using a vision-enabled language model, VLM (172), generating descriptions (303) of movements of the robot manipulator which are visible in the video; and generating said speech data (304) on the basis of the descriptions of the robot manipulator movements.
Owner:ABB (SCHWEIZ) AG

Cloud edge collaborative early warning method for early recognition of cerebral apoplexy

The invention discloses a cloud edge collaborative early warning method for early recognition of cerebral apoplexy, and the method comprises the steps: collecting the posture data and voice data of an edge end, and carrying out the preprocessing of the data; on the edge end, lightweight model reasoning is performed on the attitude data and the voice data based on a sub-modal processing and feature fusion strategy, and finally the cerebral apoplexy risk probability is output; and cooperative and multi-terminal linkage early warning is carried out based on the cloud and the edge terminal. According to the method, active capture of the cerebral apoplexy risk is realized through a framework of edge-end multi-modal real-time processing, cloud collaborative optimization and multi-end linkage early warning, the mode is converted from passive help calling to active early warning-rapid linkage, and the early warning precision and the response speed are improved.
Owner:NANTONG UNIV

Marketing verbal skill generation method and device, equipment, storage medium and program product

The embodiment of the invention provides a marketing verbal skill generation method and device, equipment, a storage medium and a program product, and relates to the field of financial science and technology and artificial intelligence. The method comprises the following steps: acquiring multi-modal data; performing cross-modal feature fusion on the voice data and the text data to generate a joint feature vector; based on the joint feature vector, through an attention mechanism, a multi-dimensional emotion evaluation result is generated, and the multi-dimensional emotion evaluation result comprises an evaluation result of at least one preset emotion dimension; and according to the multi-dimensional emotion evaluation result, generating and pushing a marketing verbal skill corresponding to the multi-dimensional emotion evaluation result. According to the method provided by the invention, through joint modeling of voice and text features and in combination with an attention mechanism, a key emotion region is focused, so that the accuracy of emotion scoring is remarkably improved, and the attitude of a customer can be reflected more truly; according to the emotion evaluation result, the optimization suggestion is generated, the generated verbal skill accurately matches the demand of the customer, the response efficiency of the marketing verbal skill is improved, and the customer emotion improvement rate is greatly improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Sign language-voice conversion system

The invention discloses a sign language-voice conversion system, and belongs to the technical field of auxiliary communication and wearable computing. The system comprises a wearable myoelectricity acquisition module used for acquiring double-arm myoelectricity signals when a user executes sign language; the mobile terminal module is wirelessly connected with the acquisition module and is used for receiving and preprocessing the signal and uploading the signal; the cloud processing module is used for receiving the signal, converting the signal into text information through a sign language recognition model, and further calling a voice synthesis service to convert the text into voice data; and the wearable audio output module is used for receiving and playing the voice data. Through an innovative end-to-end hardware system architecture, natural, accurate and real-time translation and voice output of sign language gestures are realized, communication barriers between hearing-impaired people and healthy hearing people are effectively solved, and the system has the advantages of flexible deployment, user friendliness and privacy protection.
Owner:宋飞 +1

Intelligent quality inspection method and device based on RAG and dual-enhancement mechanism

The invention provides an intelligent quality inspection method and device based on RAG and a dual-enhancement mechanism, and the method comprises the steps: determining a retrieval text based on a transcriptional text of voice data to be subjected to quality inspection and an initial quality inspection result; based on the semantics of the retrieval text and the scene label of the voice data, retrieving from a preset knowledge base to obtain matched entries and the relevancy of the matched entries; whether an IDK enhancement mechanism is triggered or not is judged based on the matching items, if not, a direct preference optimization DPO model is called to conduct preference matching degree scoring on the initial quality inspection result, and the DPO matching degree of the initial quality inspection result is obtained; and determining the accuracy and confidence of the initial quality inspection result based on the correlation degree of the matched items, the confidence of the large model, the DPO matching degree and the accuracy of the historical similar cases, and determining a final quality inspection result of the voice data based on the accuracy and confidence. According to the method and the device provided by the invention, the result accuracy, the illusion resistance and the service adaptation efficiency of intelligent quality inspection are remarkably improved.
Owner:湖北消费金融股份有限公司

Method and device for evaluating speech expression ability of neurological and mental diseases

ActiveCN121483566AHealth-index calculationSemantic analysisMicrolinguisticsNeuropsychiatric disease
The invention discloses a method and device for evaluating speech expression ability of neurological and mental diseases. According to the method, speech is induced through a visual stimulation material, speech data are collected, and after speech recognition, semantic calibration and text error correction, evaluation indexes are calculated from multiple dimensions of fluency, microcosmic linguistics and macroscopic linguistics in combination with a natural language processing technology, a large language model and a vector embedding technology, so that the evaluation accuracy is improved. And the result credibility is guaranteed through a robustness verification mechanism of iteration reflection of a double-big language model. According to the method, the problems of large subjective deviation, low efficiency and shallow analysis dimension of traditional evaluation are solved, and objective, accurate, efficient and comprehensive evaluation is realized.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

A Human-Computer Voice Interaction Control Method and System Based on Smart TV

This application relates to a human-computer voice interaction control method and system based on a smart TV. The method includes acquiring voice data within a preset range, processing the voice data to obtain voice feature data carrying control commands; performing feature analysis on the voice feature data using a preset voice analysis model to extract wake-up keywords and compare them with a preset command library to obtain control command comparison results; sending a secondary confirmation request to the user based on the control command comparison results, and combining the confirmation voice information from the user feedback to perform command recognition evaluation and correct command deviation processing to obtain the corrected control command; performing function switching processing on the TV according to the correct control command, and optimizing the display effect by linking and adjusting related devices based on the program display requirements after the switch to obtain human-computer voice interaction control data. This application has the effect of improving the intelligence of voice interaction control of smart TVs.
Owner:GUANGZHOU XIANYOU INTELLIGENT TECH CO LTD

Method for waking up application, and electronic device

The method includes: A breath wake-up processing apparatus sends voice data in first data when detecting that the obtained first data is used for indicating to wake up a first application through breath. A breath wake-up software module stores the voice data, starts the first application, and controls the breath wake-up processing apparatus to stop detecting breath wake-up of the first application. The first application sends a first notification when successfully calling the breath wake-up software module. The breath wake-up software module sends the voice data to the first application. The first application performs voice recognition on the voice data. The first application sends a second notification when determining, based on the voice data, that the voice recognition ends. The breath wake-up software module controls, in response to the second notification, the breath wake-up processing apparatus to start detecting next breath wake-up.
Owner:HONOR DEVICE CO LTD

System

A system is provided.SOLUTION: A system, comprising: means for collecting audio; means for pre-processing the collected audio; means for converting the pre-processed audio to text using a AI recognition model; means for translating the converted text using a multilingual generative AI model; means for displaying and playing the translated text; means for expert review and completion of the translated text; means for data encryption; and means for training the generative speech model to account for cultural differences.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Public network interphone with quantum chip

The invention relates to the technical field of wireless communication, in particular to a public network interphone with a quantum chip. According to the interphone, a master control module and a quantum security chip work cooperatively, a quantum random number generator is used for dynamically deriving a session key, and offline security distribution of a group master key is realized by encrypting a two-dimensional code; in a communication process, a system adaptively switches a voice coding mode according to a network condition, end-to-end authentication encryption is carried out on voice data by using a hardware encryption engine, and low-delay transmission is guaranteed through a network priority mark; the device establishes a complete key life cycle management mechanism, supports regular update and forward security of session keys and instant update and backward security of a group master key when members change, and ensures that all key operations are completed in a quantum chip. According to the invention, various eavesdropping and tampering attacks are effectively resisted, and safe, real-time and clear high-confidentiality voice communication in a public network environment is realized.
Owner:ZHEJIANG HAIGAOSI COMM TECH CO LTD

Training method and device of voice large model, equipment and medium

The invention provides a large voice model training method and device, equipment and a medium, and the method comprises the steps: inputting a training voice data subset into a large voice model, and obtaining the prediction probability distribution of the training voice data subset outputted by a large language model module in the large voice model; according to the prediction probability distribution and the real probability distribution, an entropy weighted cross entropy loss function is determined, and the entropy weight of the entropy weighted cross entropy loss function is determined based on the distribution entropy of the prediction probability distribution at the current moment; and with the purpose of minimizing an entropy weighted cross entropy loss function, updating parameters of the large language model module and training a voice recognition tag of the voice data subset. In the training method of the large voice model, entropy weight is added in an entropy weighted cross entropy loss function, the problem of uncertainty of prediction probability distribution is solved to a certain extent, iteration is performed by using a self-feedback signal of distribution entropy based on prediction probability distribution, and the voice recognition effect of the large voice model is continuously improved.
Owner:NEW ORIENTAL EDUCATION & TECH GRP CO LTD

Spatial document system and method

A computing system captures image data using a camera and captures spatial information using one or more sensors. The computing system receives voice data using a microphone. The computing system analyzes the voice data to identify a keyword. The computing system analyzes the image data and the spatial information to identify an object corresponding to the keyword. The computing system generates text based on the voice data and the keyword. The computing system stores the text in association with the object. The computing system generates and provides output comprising the text linked to the object or a derivative thereof.
Owner:ADOBE INC

Noise-fused voice data set construction method and device, and storage medium

The invention discloses a construction method and device of a noise-fused voice data set and a storage medium. The method comprises the steps of collecting noise data in a target environment; the noise segments of the target duration in the noise data are fused in the overlapping area of the voice segments, a noise-fused voice data set is obtained, the voice segments are segments in an initial voice set, the initial voice set does not contain the noise data, and the overlapping area is an area meeting a preset condition in the voice segments; and carrying out labeling processing on the voice data set fused with the noise to obtain a target voice data set. Through the method and the device, the problem of high cost of collecting the voice data set with noise in a real environment in the prior art is solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Data Integration and Governance System and Method for Multimodal Large Models

ActiveCN121144855BData setData graph
This invention provides a data integration and governance system and method for multimodal large models, relating to the field of data processing technology. The method includes: acquiring a data input set of a multimodal large model and preprocessing it to form a preliminary multimodal data set; performing timestamp mapping processing to align the collection timestamps of text data, image data, and speech data to a unified time index table; performing semantic consistency detection processing to identify semantic misalignments across modalities by comparing symptom descriptions in text data, structured annotations in image data, and spoken content in speech data, forming a consistent multimodal data set; performing dynamic label alignment processing to generate correction labels for detected misalignments and update the correction labels to the corresponding modal annotation content; and inputting the data into the large model training pipeline to perform unified governance processing of cross-modal features, forming a training data set. This invention improves the accuracy of data integration and governance.
Owner:江苏数兑科技有限公司