Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3039 results about "Voice data" patented technology

Text prediction-based large-model real-time voice text intention recognition method and system

The invention discloses a large-model real-time voice text intention recognition method and system based on text prediction, and the method comprises the steps: obtaining the real-time voice data of a user, carrying out the real-time voice recognition processing through a streaming voice recognition interface, and obtaining a part of transcriptional text; inputting the partial transcription text into a mask language model for text prediction, and generating a plurality of high-credibility complete sentence candidates; based on the complete sentence candidates, the complete sentence candidates are input into a large language model in parallel for intention recognition, a corresponding intention result is obtained, and a mapping relation between the candidate sentences and the intention recognition result is established; and obtaining a sentence completely expressed by the user, calculating the similarity between the complete actual sentence and a plurality of high-credibility complete sentence candidates through a multi-level text similarity algorithm, selecting the candidate sentence with the highest similarity score, and directly obtaining a corresponding final intention recognition result based on the mapping relationship. The objective of the invention is to solve the technical problem of high response delay of an existing voice intention recognition system.
Owner:BEIJING YULORE INNOVATION TECH

Speech recognition authentication method and system based on multi-modal features and dynamic evaluation

The invention discloses a speech recognition and authentication method and system based on multi-modal features and dynamic evaluation in the technical field of speech recognition and authentication, and the method comprises the steps: collecting an original speech signal of a user through a microphone, and carrying out the preprocessing of the original speech signal, and obtaining the preprocessing speech data; and extracting feature data of the preprocessed voice data by adopting a multi-dimensional feature hierarchical extraction technology, and injecting a multi-source noise sample into an acoustic feature space of a voiceprint feature model based on an initial training stage established by the voiceprint feature model to construct an anti-noise mixed voiceprint map. Through integrating voiceprint, semantics, behavior characteristics and an environment adaptation mechanism, an authentication threshold is adjusted in real time according to dynamic risk assessment, meanwhile, a risk scoring model is utilized to calculate a comprehensive risk value, authentication modes of different levels are started according to risk scenes of different degrees, and two-factor authentication is forcibly implemented for high-risk scenes. And the authentication security and reliability can be obviously enhanced.
Owner:JIANGSU VARIABLE SUPERCOMP TECH

Multi-modal language learning auxiliary system and method based on artificial intelligence

The invention relates to the technical field of artificial intelligence and language learning, in particular to a multi-modal language learning auxiliary system and method based on artificial intelligence, and the system comprises a multi-modal input module, a cross-modal feature fusion module, a dynamic adaptive learning module and an interactive feedback generation module. Wherein the multi-mode input module is used for receiving original text data, original image / video data and original voice data; the cross-modal feature fusion module is used for extracting a visual feature vector and a semantic coding vector and generating a cross-modal joint feature vector; the dynamic adaptive learning module is used for generating dynamic scene parameters; and the interactive feedback generation module is used for outputting a multi-mode feedback data packet through a prompt learning engine. According to the method, by fusing multi-modal feature alignment and dynamic context parameter modeling, linkage generation of grammar, culture and pronunciation feedback is achieved, and the context understanding ability and interaction feedback precision in the language learning process are improved.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Speech recognition method and related device

ActiveCN114360510AImprove fault tolerancePrecise Syllable Probability DistributionSpeech recognitionSyllableAcoustic model
The embodiment of the invention discloses a speech recognition method and a related device, and at least relates to a speech recognition technology in artificial intelligence, speech data to be recognized are used as input data of a time delay neural network in an acoustic model, and an output layer of the time delay neural network comprises acoustic modeling units corresponding to a plurality of syllables respectively, so that the speech recognition efficiency is improved. And the syllable probability distribution corresponding to the voice frames included in the voice data can be obtained by taking the syllables as the recognition granularity through the time delay neural network. When syllable recognition is carried out through the output layer, auxiliary judgment can be carried out on the syllables to which the voice frames belong on the basis of pronunciation rules in combination with front and back syllable information of the voice frames, so that more accurate syllable probability distribution is output. Moreover, since the syllables are generally composed of one or more phonemes, the method has higher fault-tolerant capability, not only can more accurately determine the speech recognition result based on the probability distribution of the syllables, but also has low requirements for the quality of the speech data to be recognized, and effectively expands the application scenarios of the speech recognition technology.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Model-based interaction method and system, wearable device and storage medium

The invention provides a model-based interaction method and system, wearable equipment and a storage medium, and belongs to the technical field of intelligent interaction.The method comprises the steps that in response to a received interaction instruction, voice data, a gesture image, eye movement data and an environment image are obtained based on the interaction instruction; extracting user intention features based on the voice data, the gesture image and the eye movement data, and determining scene type features based on the environment image; determining an interaction theme based on the interaction instruction, obtaining user historical interaction information associated with the interaction theme from a context memory database, and generating a context feature vector based on the user historical interaction information; and inputting the user intention feature, the scene type feature and the context feature vector into an intention recognition model based on quantum enhancement to obtain a user intention, and generating interaction response data based on the user intention. According to the invention, the accuracy of user intention recognition can be improved, and the intelligence of interaction is improved.
Owner:BEIJING SUPERHEXA CENTURY TECH CO LTD

Environmental acoustic simulation voice generation method, apparatus and device, and medium

ActiveCN120612917ASpeech recognitionSpeech synthesisPhonetic environmentData set
The invention relates to the technical field of voice processing, can be applied to business scenes such as financial science and technology and medical health, and discloses an environmental acoustics simulation voice generation method, device and equipment and a medium, and the method comprises the steps: obtaining and separating mixed voice data, and generating original voice content and original environmental acoustics information; converting the original voice content into first text information; determining a target environment acoustic tag in combination with the original environment acoustic tag, the first text information and the target geographical location information; acquiring target environmental acoustic information from a preset sound data set based on the tag, and adjusting the amplitude characteristic of the target environmental acoustic information to match the original environmental acoustic information; and synthesizing the adjusted target environment acoustic information and the original voice content into simulated voice data. According to the method, the target geographic position information is introduced to participate in acoustic feature determination and amplitude adjustment, so that the generated simulated voice data is more consistent in geographic semantics and acoustic performance, and the authenticity and concealment of voice environment disguise are effectively improved.
Owner:PING AN TECH (BEIJING) CO LTD

Education robot-based intelligent system and method for fusing mental health evaluation in interest scene

PendingCN120636769AEnsemble learningDigital data protectionNumber generatorPsychometric testing
The invention discloses an intelligent system and method for fusing mental health evaluation in an interest scene based on an education robot. The system realizes non-intrusive mental health assessment by constructing a technical architecture of'multi-modal biological feature anonymization-quantum hybrid encryption-federated learning privacy protection-dynamic risk assessment '. The method is characterized by comprising the following steps of: (1) generating irreversible biological characteristic codes by adopting a chaos random number generator, and realizing anonymized association by combining a double hash chain technology; (2) designing a quantum enhanced hybrid encryption system, fusing an SM4 algorithm, NTRU post quantum cryptography and quantum key distribution; (3) developing a dynamic scene generator, and naturally fusing psychological test questions into interest topics by using a generative adversarial network (GAN) and a curiosity engine; (4) constructing a multi-modal emotion analysis system, and combining an LSTM + CNN hybrid model to realize physiology-behavior-voice data fusion analysis; and (5) deploying a federated learning framework, and protecting model training data through differential privacy noise injection (epsilon = 0.5).
Owner:张景飞

Speech signal detection device

Methods and systems are disclosed for collecting EMG speech signals using a speech signal detection device. The methods and systems collect a combination of signals comprising electromyograph (EMG) data signals and one or more non-EMG data signals. The methods and systems process the combination of signals by a machine learning (ML) model to detect inner speech of the user, the ML model trained to establish a relationship between training signals comprising training EMG data signals and training non-EMG data signals and ground-truth inner speech data and, in response, performing one or more operations associated with the speech signal detection device.
Owner:SNAP INC

Pilot dynamic evaluation method, system and equipment based on TEM model and storage medium

The invention relates to the technical field of TEM models, provides a pilot dynamic evaluation method, system and device based on a TEM model and a storage medium, and solves the problems of low accuracy of flight training evaluation and poor adaptability of a training scheme. The method comprises the steps that flight control, physiological monitoring and cockpit voice data are acquired, and a multi-modal data stream is formed through time synchronization; performing threat perception, error management and non-technical skill three-level evaluation on the feature vector by using a TEM model, generating corresponding indexes, and fusing the indexes into a comprehensive feature vector; respectively processing the time sequence and cognitive features by means of a dual-channel long-short-term memory network, and extracting deep features; performing decision tracing by adopting an SHAP value so as to locate a capability defect node; finally, defect types are matched through a personalized improvement scheme recommendation mechanism, and self-adaptive training content is generated. According to the invention, the accuracy of flight training evaluation and the adaptability of the training scheme are improved.
Owner:CHINA SOUTHERN TECHNOLOGY (GUANGDONG HENGQIN) CO LTD +2

Multi-modal decision-making auxiliary method and system for legal compliance examination

The invention provides a multi-modal decision-making auxiliary method and system for law compliance examination. The method comprises the steps of S1, obtaining multi-source data of a to-be-examined object; s2, preprocessing the multi-source data; s3, extracting modal feature vectors of the text data, the voice data, the image data and the video data; s4, inputting each modal feature vector into a trained first convolutional neural network model, and outputting a unified feature representation matrix; s5, calling a law and regulation knowledge graph retrieval interface, and generating a law and regulation constraint vector corresponding to the associated law and regulation terms; and S6, inputting the unified feature representation matrix and the regulation constraint vector into a trained second convolutional neural network model, and outputting an illegal clause matching result and interpretable information. According to the method, the latest regulation terms are automatically matched, the visual risk evidence chain is output, and the review coverage, accuracy and timeliness are greatly improved.
Owner:XINXIANG UNIV +1

AI digital human interactive response method based on large language model

The invention discloses an AI digital human interactive response method based on a large language model, and relates to the technical field of digital human interaction, and the method comprises the steps: analyzing collected user voice data and visual data through a natural language processing method, generating a cross-modal feature vector, carrying out the cross-modal association analysis of the cross-modal feature vector, and carrying out the cross-modal association analysis of the cross-modal feature vector. Generating a semantic association topological graph; calculating a vertex coordinate and a joint activity threshold value of the semantic association topological graph through high-digital human correlation, inputting the vertex coordinate and the joint activity threshold value into a constructed coordinate index database to execute attention weight calibration, and outputting a multi-dimensional association graph; and performing information density analysis based on the multi-dimensional association map, generating an information density gradient vector field, and dividing a high-density core region and a low-density edge region, the high-density core region generating a semantic core coding tensor, and the low-density edge region generating an edge feature package. According to the method, the cross-modal fusion vector is converted into the cross-modal feature vector, so that the modeling of the cross-modal association relationship is realized.
Owner:BEI JING XIN ZHI YUAN LANG WANG LUO KE JI YOU XIAN GONG SI

Dietary structure evaluation method and system based on digital intelligent analysis

The invention discloses a dietary pattern evaluation method and system based on digital intelligent analysis, and relates to the technical field of dietary pattern evaluation.The method comprises the steps that a mobile terminal is used for synchronously collecting multi-modal data, and the multi-modal data are preprocessed; food pictures are retrieved based on keywords to form a data set, an image recognition model is constructed, the volume and area of food materials are calculated in combination with a 3D reconstruction technology, a food material density database is matched, and automatic estimation of the quality of the food materials is achieved; building a voice interaction model fusing characters and regional voice features, and recognizing food material types and food material quality by using voice data; a dynamic allocation weight is calculated through multi-modal confidence, a graph attention mechanism is utilized to detect and correct data conflicts, and a space-time constraint rule base is combined to align dispersed intake records; and constructing a data visualization chart library. According to the method, the rise of dietary structure analysis from extensive statistics to refined intelligent evaluation is realized, and the problem of error accumulation caused by dependence on a single data source in a traditional method is effectively solved.
Owner:BEIJING BONNIE YINGCE TECH CO LTD

Digital large screen interaction method and device based on multi-agent cooperation and medium

The invention discloses a digital large-screen interaction method and device based on multi-agent cooperation and a medium, and relates to the technical field of digital large-screen interaction.The method comprises the steps that postures, expressions and voice data of a user are collected in real time, the attention weight is calculated in combination with spatial position information, the emotional state change rate is monitored, and the user experience is improved in combination with personal historical browsing preferences; acquiring a main interaction user, an emotional state feature and an emotional change trend; according to the main interaction user, the emotional state features and the emotional change trend, obtaining an initial confidence coefficient of the suspected interested field through a semantic understanding model; and based on the candidate guide images, identifying a user selection intention through a multi-modal fusion processing mechanism, calculating a selection probability, triggering content activation of the main interaction area and content dynamic generation of the auxiliary information area, synchronizing state information, and starting an immersive collaborative interaction process. Through three-level collaborative service configuration, the beneficial effects of improving large screen content organization efficiency and enhancing immersive interaction experience are achieved.
Owner:YLZ INFORMATION TECHNOLOGY CO LTD

Voice quality inspection method and device, computer equipment and storage medium

The invention discloses a voice quality inspection method and device, computer equipment and a storage medium, belongs to the technical field of artificial intelligence, and is applied to voice quality inspection scenes in the fields of finance, health medical care, old-age care and the like. The multi-modal feature fusion technology is introduced, the context semantic features of the text and the acoustic features of the voice are extracted, the emotion features are obtained by combining the pre-trained emotion recognition model, comprehensive understanding of the voice data from the three dimensions of semantics, acoustics and emotions is achieved, the deep fusion of the three feature vectors is carried out, and the voice recognition efficiency is improved. Compared with a traditional method which only depends on text or acoustic features, the method has the advantages that multi-modal features are realized by combining emotional features on the basis of the text or acoustic features, and information contained in voice content can be reflected more comprehensively and meticulously, so that the accuracy and practicability of voice quality inspection are improved, and the voice quality inspection efficiency is improved. And the requirements of application scenes such as intelligent customer service and voice auditing on high-quality automatic quality inspection are met.
Owner:PING AN TECH (BEIJING) CO LTD

Smart home control method, readable storage medium and smart home system

The invention provides a smart home control method, a readable storage medium and a smart home system. The method comprises the steps that character characteristics of target home equipment are set at least according to equipment functions of the target home equipment, the character characteristics comprise character dimensions, and the character dimensions comprise at least one of gentle, cold, humor and fine; according to at least part of the voice data, the behavior data and the indoor environment data of the user, the emotional state of the user is determined, and the behavior data comprises use condition data of the user for the multiple intelligent devices; analyzing the emotional state and the character traits by using a large model, and at least determining an interaction strategy of the target home device, the interaction strategy comprising a response style and response content; and under the condition that the emotional interaction function of the target home equipment is started, at least controlling the target home equipment to interact with the user according to the interaction strategy. According to the invention, the personification and intelligence level of the smart home equipment is improved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1

Server, display device and digital human processing method

The embodiment of the invention provides a server, display equipment and a digital human processing method. The method comprises the following steps: receiving voice data input by a user and sent by the display equipment; broadcast voice is determined based on the voice data; extracting voice features of the broadcast voice; determining mouth shape parameters based on the voice features; determining emotion parameters and acquiring user image data; generating digital human image data based on the user image data, the emotion parameters and the mouth shape parameters; and sending the broadcast voice and the digital human image data to the display device, so that the display device plays the broadcast voice and displays a digital human image based on the digital human image data. According to the embodiment of the invention, the expression parameters and the mouth shape parameters are determined according to the voice data input by the user, the expression parameters and the mouth shape parameters are combined to generate the digital human image with better facial expression expression, and emotion customization and control are realized.
Owner:HISENSE VISUAL TECH CO LTD

Knowledge graph construction method and device based on multi-modal clinical data, equipment, medium and product

The invention discloses a knowledge graph construction method and device based on multi-modal clinical data, equipment, a medium and a product, and relates to the field of knowledge graph construction, and the method comprises the steps: constructing a medical knowledge type knowledge graph; based on the dynamic event data of the patient and the doctor-patient dialogue data, constructing a medical event type knowledge graph; fusing the medical knowledge type knowledge graph and the medical event type knowledge graph to obtain a fused medical knowledge graph; collecting multi-modal data in the diagnosis process of the patient, and fusing the multi-modal data with the fused medical knowledge graph to obtain a finally constructed knowledge graph; the multi-modal data comprises text data, image data, voice data and physiological signal data. The problems that clinical knowledge is fragmented, deep mining is difficult, data cannot be associated in a unified mode, and doctor-patient communication knowledge is fragmented can be solved.
Owner:ZHEJIANG UNIV

Blind guiding robot interactive navigation system combining vision, inertial navigation and voice

The invention relates to the technical field of blind guiding robot interactive navigation, and particularly discloses a blind guiding robot interactive navigation system combining vision, inertial navigation and voice, which is characterized in that environmental vision, movement and voice data are synchronously acquired through a multi-modal perception data acquisition module, and are analyzed through a situation-user joint state feature extraction module; a user potential intention deduction module is combined to predict a user intention and generate a self-adaptive guide strategy; a self-adaptive guiding path optimization module adjusts a navigation path according to a strategy to ensure safety and high efficiency; finally, the situational multi-mode interaction instruction output module provides voice and tactile feedback, and interaction experience and navigation autonomy are enhanced. According to the method, the travel problem of the visually impaired person in a complex environment is solved, the individuation, context perception capability and safety of blind guiding are improved, and more natural and efficient navigation assistance is brought to the visually impaired person.
Owner:SHANDONG SAIFEITE SAFETY ENG TECH DEV CO LTD

Real-time sound duplicating method and system based on end-cloud fusion

The invention provides a real-time sound copying method and system based on end-cloud fusion. The method comprises the following steps: a cloud end carries out real-time tone copying and voice synthesis on a small amount of voice data of a user based on an AI large model; timbre samples are collected when a user registers voice audio data, and user timbre voice data of a preset text are synchronously generated by a large model and are used as fine tuning training data of an end-side voice synthesis model; user tone voice data of a preset text and voice audio data registered by a user are utilized to carry out migration fine tuning training on an end-side voice synthesis model to adapt to the personalized tone of the user, so that high-quality output of the end-side voice synthesis model is ensured, and personalized voice replication is realized; and issuing the trained end-side speech synthesis model to the user equipment, and independently completing speech replication in a network-free or weak network environment. According to the invention, through automatic generation of the user tone data and adaptive fine tuning of the model, the tone of the user is deployed to the end side after fine tuning, and high-quality and high-adaptability sound replication of end-cloud collaboration is realized.
Owner:PACHIRA TIMES (ZHUHAI HENGQIN) INFORMATION TECH CO LTD

LED display screen interaction method and system supporting voice interaction

The invention provides an LED display screen interaction method and system supporting voice interaction, and the method comprises the steps: carrying out the voice feature extraction processing through obtaining a continuous voice data stream sent by a user, and obtaining an acoustic feature sequence and a semantic association feature set of a voice segment; and calling a pre-trained voice semantic understanding model to perform joint semantic analysis processing on the acoustic feature sequence and the semantic association feature set, generating a semantic understanding result of the voice segment, and determining a user interaction intention type corresponding to the continuous voice data stream and key information positioning features of the user interaction intention in the voice content based on the semantic understanding result. And generating an LED display control instruction containing an interaction content identifier based on the user interaction intention type and the key information positioning feature, and sending the LED display control instruction to the target LED display screen to execute an interaction display operation. The real intention and key information in the continuous voice of the user can be accurately understood, the accuracy and naturalness of voice interaction between the LED display screen and the user are remarkably improved, and the interaction experience of the user is effectively improved.
Owner:SHANXI LAMPSON TECHNOLOGY CO LTD

Multi-modal emotion recognition method and system for service-oriented robot

The invention belongs to the technical field of artificial intelligence, and particularly relates to a service-oriented robot-oriented multi-modal emotion recognition method and system, and the method comprises the steps: collecting audio and video stream data of emotion changes of a user, and separating visual and voice data; extracting visual and voice emotion features through a pre-training model, and calculating prediction probability distribution of each mode; constructing a bimodal confidence quantitative model based on the distribution to obtain each modal confidence; and fusing the features by adopting a sectional type dynamic weight distribution strategy so as to identify the emotional state of the user. Visual and voice modes are fused, feature alignment is realized in combination with dynamic time warping, spatial optimization performance is shared and expressed through a confidence model, a dynamic weight strategy and a cross-modal time sequence cooperation module, and the method has high recognition accuracy, high robustness and real-time processing capacity in a complex environment and is suitable for various service scenes.
Owner:SUZHOU CITY UNIV

Multi-modal robot interaction system

The invention discloses a multi-modal robot interaction system, and relates to the technical field of robots, the system comprises a robot subsystem and a cloud processing subsystem which are in communication connection with each other, the robot subsystem comprises a multi-modal sensing module and an interaction execution module, and the cloud processing subsystem comprises a cloud processing module. The multi-mode sensing module is used for collecting target data of at least one target user and sending the target data to the cloud processing module, and the target data comprises at least one of touch data, distance data, posture data, position data, action data and voice data; the cloud processing module is used for analyzing the target data, generating an interaction instruction and sending the interaction instruction to the interaction execution module; and the interaction execution module is used for receiving and executing the interaction instruction, so that the higher expectation of the user in the aspects of emotion accompanying and personalized service is effectively met.
Owner:SHENZHEN XINBAN TECHNOLOGY CO LTD

Abnormal behavior recognition method and device, electronic equipment and nonvolatile storage medium

The invention discloses an abnormal behavior recognition method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: acquiring voice data in a call process, and acquiring user portrait data corresponding to the call; extracting a first voiceprint feature and a synthetic voice detection feature of the voice data, and determining a user behavior feature according to the user portrait data; determining a call content text corresponding to the voice data, and extracting a text semantic feature corresponding to the call content text; and adopting an abnormal behavior recognition model to analyze the first voiceprint feature, the synthetic voice detection feature, the user behavior feature and the text semantic feature to obtain a recognition result corresponding to the call, the recognition result being used for representing whether the call relates to an abnormal behavior. According to the method and the device, the technical problem that the abnormal behavior recognition accuracy is poor due to the fact that the abnormal behavior in the call process is detected only by depending on a text or voiceprint single mode in the prior art is solved.
Owner:CHINA TELECOM CORP LTD

Non-autoregressive end-to-end dialect identification method for power supply service telephone system

The invention provides a non-autoregressive end-to-end dialect recognition method for a power supply service telephone system. The method comprises the following steps of dual-track telephone recording corpus preprocessing, automatic pre-labeling and context construction, manual labeling and dialect word library construction, dialect sub-area division and classification modeling, and dialect recognition model training and optimization. Based on the ffmpeg audio processing tool, the standardization and automation level of voice data processing in a telephone system scene is improved, and the manual annotation efficiency is greatly improved. And meanwhile, aiming at the problems that dialects are various in type, significant in difference and the like, a unified dialect labeling rule system is constructed, a model training process is optimized in combination with a transfer learning strategy, and the robustness and generalization ability of the model in a complex telephone voice environment are enhanced while high accuracy is ensured.
Owner:STATE GRID HUBEI ELECTRIC POWER RES INST +1

Teaching speech emotion recognition method based on dynamic time sequence modeling and multi-scale fusion

The invention discloses a teaching speech emotion recognition method based on dynamic time sequence modeling and multi-scale fusion. The method comprises the steps of speech data set preprocessing, speech feature extraction, speech emotion classification network construction, speech emotion classification network training, inputting a test set into the trained speech emotion classification network, and outputting the probability of each type of emotion. According to the method, the direction of the information flow is dynamically adjusted through the adaptive time displacement module, the features of different time scales are extracted by using the multi-scale convolution branch, the modeling capability of the model for a complex time sequence structure and variation data is improved, the feature expression is enhanced, the time sequence features are extracted by using the Wav2Vec2.0 pre-training model, and the time sequence features are extracted by using the Wav2Vec2.0 pre-training model. A speech emotion classification network comprising an AdaShiftFormer learning module and a multi-scale time sequence fusion module is constructed, training is carried out in combination with classification loss and comparison loss, and classification performance is optimized. The method is superior to the prior art in emotion classification accuracy and feature expression ability, and can assist in teacher speech behavior analysis and classroom interaction optimization in an intelligent education scene.
Owner:SHAANXI NORMAL UNIV

Speech processing method and apparatus, device, and medium

PendingUS20250329334A1Speech analysisSpeech segmentationAcoustics
A speech processing method includes: obtaining overlapping speech data; obtaining reference speech data of a specified object; extracting a voiceprint representation vector of the specified object from the reference speech data, the voiceprint representation vector representing a voiceprint characteristic of the specified object, and inputting the overlapping speech data and the voiceprint representation vector into a preset speech segmentation model, and segmenting, by the speech segmentation model based on an attention mechanism, the overlapping speech data to obtain a target speech signal matching the voiceprint characteristic; and generating a speech file of the specified object based on the speech signal.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Multi-modal interactive virtual teaching method, system, equipment and medium

The invention discloses a multi-mode interactive virtual teaching method and system. The method comprises the following steps: acquiring multi-source data; extracting voice data acoustic features, and inputting the voice data acoustic features into a learning model to obtain a text triple; key terms are extracted from the text data, a traceable operation chain is generated, semantic analysis is carried out, and an operation scheme is output in combination with a knowledge graph; performing abnormal state recognition on the image data through a target detection model, positioning abnormal equipment in combination with a character recognition model, performing action mapping through gesture recognition and a spatial constraint rule, and outputting a corresponding instruction; fusing the three types of outputs to generate a scheduling event chain; and performing semantic analysis, evaluation and optimization on the scheduling event chain, interacting with personnel, updating operation suggestions and providing an operation analysis result. According to the method, manual rechecking requirements are reduced through multi-modal data collaboration, a new man-machine interaction database and a new man-machine interaction standard in the power industry can be formed through multi-modal interaction rules and event chain construction, and data intelligent driving is achieved while the training efficiency is improved.
Owner:GUANGXI POWER GRID CORP

Intelligent chart dynamic generation method based on voice recognition and multi-modal interaction

The invention belongs to the technical field of speech recognition, and discloses an intelligent chart dynamic generation method based on speech recognition and multi-modal interaction. Comprising the following steps: collecting user voice data and denoising to obtain accurate user voice data; performing semantic structure segmentation and combination, and outputting complete user semantic data; executing semantic analysis and outputting user query intention information; extracting a query target data subset; constructing a field structure portrait based on the query target data subset, generating a target chart format, and outputting target chart parameters; performing structure correction on the target chart parameters based on the multi-modal interaction information collected in real time, and outputting a dynamically corrected target chart; performing trend evolution display on the dynamically corrected target chart to obtain an optimized target chart, and sending the dynamically optimized target chart to a preset data large screen; the voice recognition accuracy and robustness are improved, the interaction function of the chart is optimized, the interaction requirement of the user is met, and the interaction experience of the user is improved.
Owner:DONGQU INTELLIGENT TRANSPORTATION INFRASTRUCTURE TECH (JIANGSU) CO LTD +1

Internet of Things multi-source heterogeneous data management method based on artificial intelligence

InactiveCN121189325ASemantic analysisBiological modelsTransliterationEngineering
The invention discloses an Internet of Things multi-source heterogeneous data management method based on artificial intelligence, and the method comprises the following steps: collecting voice data and text data in an Internet of Things system, and constructing a voice and text pair; inputting the voice data into an improved Whisper model, and outputting an enhanced transliteration text; inputting the enhanced transliteration text and the text data into a semantic encoder based on dynamic weighting to generate a fused semantic embedding vector; calculating semantic similarity between the fused semantic embedding vectors, and screening semantic consistent sample pairs; performing event aggregation on the semantically consistent sample pairs; performing semantic level coding and confidence adaptive fusion processing to generate a unified semantic fusion vector; and constructing a backmarking training set based on expert feedback, and executing periodic increment optimization on the updatable parameters. The voice and text multi-modal data fusion quality and semantic consistency recognition precision are improved, and the method has good expandability and is suitable for intelligent data governance tasks in a complex Internet of Things environment.
Owner:CHONGQING PAILING INFORMATION TECHNOLOGY CO LTD

Ambient sound feature-based authenticity analysis method, apparatus and device, and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses an authenticity analysis method, device and equipment based on environmental sound characteristics and a medium. Performing voice separation processing on the original voice data to generate environment voice data and pure voice data, analyzing the environment voice data and retrieving knowledge information associated with the environment voice data, and identifying the content of the pure voice data to generate dialogue text data, and inputting the environment sound features, the knowledge information, the dialogue text data and the user declaration information into an analysis model, and outputting a authenticity analysis result. According to the method, the environment sound data and the voice content are separated, the available features of the environment sound data and the voice content are extracted respectively, and the background knowledge and the user declaration information are combined to perform fusion reasoning in the unified analysis model, so that the accuracy of authenticity judgment and the adaptability to complex scenes can be improved.
Owner:PING AN TECH (BEIJING) CO LTD