Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

463 results about "Recognition speech" patented technology

Speech recognition. Speech recognition is the inter-disciplinary sub-field of computational linguistics that develops methodologies and technologies that enables the recognition and translation of spoken language into text by computers. It is also known as automatic speech recognition (ASR), computer speech recognition or speech to text (STT).

Intelligent oral English training system based on Prompt engineering

The invention discloses an intelligent oral English training system and method based on Prompt engineering, and relates to the technical field of oral English training. According to the method, a user is allowed to freely describe training requirements by using a natural language, the system automatically constructs a training Prompt, and personalized customization is realized; content deviating from a training target is accurately recognized through a deviation question detection module, a user is guided to return to a theme through real-time voice or text, and effectiveness and pertinence of training are guaranteed; a clear template definition ensures systematicness and reproducibility of a training task, and clear control and evaluation of a teaching target are facilitated; in combination with voice recognition, voice synthesis and text generation, natural interaction in a real context is realized, and the training experience is improved; on the basis of multi-dimensional automatic evaluation of a keyword hit rate, pronunciation quality and the like, personalized learning suggestions are supported; teachers are supported to generate training tasks in batches, students are supported to customize scenes, and different education environment requirements are met.
Owner:CHUXIONG NORMAL UNIV

Multi-mode emotion recognition method, system, electronic device and storage medium

Disclosed are a multi-mode emotion recognition method, a system, an electronic device, and a storage medium. The method includes obtaining a spectrogram of a voice to be recognized and a corresponding text and inputting the spectrogram and the text into a multi-mode emotion recognition model to obtain an emotion recognition result output by the multi-mode emotion recognition model. The multi-mode emotion recognition model is trained based on a sample spectrogram, and a corresponding sample text, and a sample emotion recognition result, and is configured to extract a feature from the spectrogram and the text by a self-attention mechanism to obtain the voice features and the text feature, fuse the text feature and voice feature to obtain a multi-mode fusion feature, and make an emotion classification decision to obtain an emotion recognition result based on the text feature, the voice feature, and the multi-mode fusion feature.
Owner:HUAZHONG NORMAL UNIV

Enterprise intelligent finance and tax system fusing supervised learning and block chain

The invention provides an enterprise intelligent finance and taxation system fusing supervised learning and a block chain, and solves the core technical problems of insufficient data credibility, transparent contradiction between privacy protection and supervision, AI model training data island and the like of a traditional finance and taxation system. According to the system, an innovative technical fusion scheme is adopted, bank API, OCR recognition, voice input and other multi-source financial data are integrated through a multi-modal data fusion unit, and high-quality data fusion is achieved through confidence evaluation and an intelligent conflict resolution algorithm; an intelligent classification unit based on BERT deep learning integrates a pre-training model with a business rule engine and a gradient boosting decision tree, and accurate and automatic classification of financial transactions is achieved. A complete technical solution is provided for enterprise finance and taxation digital transformation, intelligent, automatic and credible processing of finance and taxation businesses is achieved on the premise that data safety and privacy protection are guaranteed, and the method has wide market application prospects.
Owner:张宏

Speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis

The invention relates to a speech recognition and transcription method and system based on multi-modal fusion and sentiment analysis, and relates to the field of speech recognizing.The speech recognition and transcription method comprises the steps that a target speech signal and auxiliary modal information of synchronous visual information and text context information are obtained firstly, and the speech signal is segmented and recognized to obtain speech feature vectors; the method comprises the following steps: extracting text context information to obtain a text auxiliary feature vector, carrying out multi-modal fusion on the text auxiliary feature vector and the text auxiliary feature vector to generate fusion feature representation so as to carry out voice transcription to obtain an initial transcription text, and carrying out sentiment analysis according to visual information and the initial transcription text to generate a sentiment feature tag; and finally, optimizing and correcting the initial transliteration text based on the label to obtain a target transliteration text, thereby solving the technical problems that the speech recognition transliteration is difficult to adapt to dialect diversity and the recognition accuracy and robustness are insufficient due to neglect of emotion information, and improving the recognition accuracy and robustness through fusion of multi-modal information and emotion analysis. The voice content can be recognized more accurately, the transcription text can be optimized, and the accuracy and quality of voice recognition transcription are improved.
Owner:山西益通电网保护自动化有限责任公司

Voice segmentation intelligent editing system based on deep learning

PendingCN121260170ASpeech recognitionSpeech segmentationInformation density
The invention relates to the technical field of voice signal processing, and discloses a voice segmentation intelligent editing system based on deep learning. The system comprises a voice feature extraction module, a segmentation boundary detection module, a semantic content analysis module, an editing strategy generation module and a real-time quality evaluation module. The voice feature extraction module collects multi-dimensional voice features and timestamp information, and verifies feature integrity and timeliness; the segmentation boundary detection module identifies voice pause intervals and semantic turning nodes and divides segmentation units and boundary types; a semantic content analysis module extracts text content and emotion features of each segment, and analyzes semantic topic relevance and information density; an editing strategy generation module formulates a segmentation retention rule and a sequence adjustment scheme, and matches user preferences and scene demands; the real-time quality evaluation module monitors voice fluency and information integrity in the editing process and analyzes splicing errors and user feedback. According to the system, intelligent processing of the whole voice editing process is realized.
Owner:SHENZHEN JYEOO NETWORK TECH CO LTD

Artificial intelligence-based speech emotion recognition method, device, equipment and medium

The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment, and medium for speech emotion recognition based on artificial intelligence. The method performs frame segmentation and windowing processing on speech information to be recognized to obtain a speech frame sequence, extracts a speech feature tensor and a text feature tensor of the speech information to be recognized, aligns the speech feature tensor and the text feature tensor and performs feature extraction to obtain multimodal features, performs average pooling processing and global maximum pooling processing on the speech frame sequence using a local window to obtain enhanced speech features, performs feature fusion on the enhanced speech features and the multimodal features, determines a fusion result, and obtains an emotion recognition result based on the fusion result. The method obtains low-level enhanced speech features through pooling processing, and performs feature fusion on the enhanced speech features and the multimodal features, thereby avoiding the degradation problem of deep networks, effectively improving the accuracy of speech emotion recognition while improving the generalization ability of the model.
Owner:PING AN TECH (SHENZHEN) CO LTD

Interactive question and answer task processing method based on AI large model

The invention discloses an interactive question and answer task processing method based on an AI large model, and relates to the technical field of AI questions and answers, and the method comprises the following steps: performing correlation screening and function label labeling on high-confidence sub-queries in a retrieval result set, performing weight reduction on low-confidence sub-queries, constructing a cross-source consistency constraint vector, and generating an input data set; based on the input data set, a multi-source evidence consistency verification channel is constructed, entity-by-entity alignment and conflict detection are executed, the credibility interval of answers is calculated, meanwhile, voice emotions are recognized, and an answer data set is generated; and converting the answer data set into multi-modal feedback, monitoring user behaviors in real time, calculating behavior response strength indexes, dynamically adjusting a feedback form, and generating an interactive question and answer data set. The semantic consistency constraint of the multi-modal evidence is realized, and the robustness of answer credibility evaluation in the natural language processing task is improved.
Owner:张婧

Systems and methods for orchestrating interaction with an artificial intelligence application

Systems and methods for orchestrating interaction with an artificial intelligence (AI) application in a contact center environment receive, via an AI agent, a voice message from a user; convert the message from voice to text; generate an initial computational inference process based on the text message; determine whether or not all information required to execute the initial computational inference process is available to the processor; when a determination is made that all information required is available: execute the initial computational inference process; generate a text reply based on the initial computational inference process; convert the text reply to a voice reply; and send the voice reply to the user via the AI agent; when a determination is made that information is unavailable: generate a text query requesting the information; convert the text query to a voice query; and send the voice query to the user via the AI agent.
Owner:THE BANK OF NEW YORK MELLON

Speech recognition method and device, equipment, storage medium and program product

The invention provides a voice recognition method and device, equipment, a storage medium and a program product, and relates to the technical field of image processing. The speech recognition method comprises the following steps: acquiring speech to be recognized, and extracting speech features based on the speech to be recognized; obtaining video content corresponding to the to-be-recognized voice, and extracting video features based on the video content; acquiring a historical voice recognition text of the to-be-recognized voice, and extracting historical text features based on the historical voice recognition text; obtaining a first multi-modal fusion feature based on the voice feature and the historical text feature; obtaining a second multi-modal fusion feature based on the video feature and the historical text feature; and generating a speech recognition text corresponding to the speech to be recognized based on the first multi-modal fusion feature and the second multi-modal fusion feature.
Owner:SHANGHAI HODE INFORMATION TECH CO LTD

Classroom teacher teaching performance description method, model and system based on multi-modal data fusion and storage medium

The invention relates to the technical field of data processing, in particular to a classroom teacher teaching performance description method, model and system based on multi-modal data fusion and a storage medium. By introducing a multi-modal data fusion strategy, cooperative processing of classroom teacher visual information, audio information and text information is realized, and the accuracy and time sequence continuity of teacher target perception are significantly improved. Compared with a traditional single-mode method, the method not only can extract the posture, expression and action characteristics of the teacher from the visual mode, but also can recognize the voice emotion and the side language signal from the audio mode, and achieves the full-dimensional description of the teaching behavior of the teacher in combination with the text semantics. The objective of the invention is to solve the problem of how to perform multi-dimensional evaluation on teaching performance of a classroom teacher based on data of multiple modalities.
Owner:YUNNAN NORMAL UNIV

Voice recognition method based on acoustic model, computer equipment and storage medium

The invention belongs to the field of voice recognition, and discloses a voice recognition method based on an acoustic model, computer equipment and a storage medium. The method comprises the following steps: acquiring voice features of to-be-recognized voice; inputting the voice features into an acoustic model, and outputting a recognition result by the model; wherein the time sequence processing network layer firstly determines the ratio of current input future frames needing to be pre-watched to context information through a pre-trained gating fusion unit, then calculates the number of the future frames needing to be pre-watched based on the ratio, obtains the corresponding future frames, calculates long-time context representation in combination with the future frames, processes the long-time context representation and outputs the long-time context representation to the next layer of network. According to the method and the device, the problem of static binding of delay and accuracy in the prior art is solved by dynamically adjusting the number of the future frames to be pre-watched, low-delay response to simple command words is realized, the recognition accuracy is improved through multiple future frames to be pre-watched for easily-confused instructions, the balance of the delay and the accuracy is realized, and the performance of a voice recognition system and the user experience are improved.
Owner:深圳市友杰智新科技有限公司

Voice conversion authentic identification method and system based on knowledge distillation alignment

The invention discloses a voice conversion authentic identification method and system based on knowledge distillation alignment, and is applied to the technical field of voice authentic identification. The method comprises the following steps that a double-branch model used for voice authentic identification is constructed, pure voice serves as input of a teacher branch, noisy voice serves as input of a student branch, and the teacher branch and the student branch share a feature extraction network of the same structure; applying a speech enhancement technology at the front end of a student branch to generate an enhanced speech signal; deep features are extracted, and alignment of the deep features in a hidden space is restrained through a knowledge distillation loss function; carrying out dynamic weight fusion on the aligned deep features through a fusion weight; and training a classifier, performing joint optimization in combination with classification loss and knowledge distillation loss, and outputting a voice authentic identification result. According to the method, distribution alignment of pure and noise features is realized through knowledge distillation, forged traces in voice conversion are effectively recognized, and high detection precision is still kept in complex noise and unknown attack scenes.
Owner:ZHEJIANG UNIV

System for identifying companion animal and method therefor

ActiveUS12424016B2Speech analysisNeural architecturesAnimal scienceCCTV - Closed circuit television
The present invention relates to a technology capable of identifying a companion animal by analyzing video, image, and voice data collected through CCTV or a camera, for management of companion animals and tracking in case of loss thereof, wherein, by simultaneously or sequentially using at least one identification method among facial recognition, nose print recognition, voice recognition, and motion recognition by analyzing a video, image, or voice, an effect of greatly improving the reliability of object identification for companion animals can be provided.
Owner:AJIRANG RANGIRANG INC

Voice recognition method and device and electronic equipment

The invention discloses a voice recognition method and device and electronic equipment. The method comprises the following steps: acquiring a voice to be recognized, the voice to be recognized being a voice in a dialect form; performing voice recognition on the to-be-recognized voice through the voice recognition model to obtain an initial recognition result of the to-be-recognized voice, the initial recognition result including a word sequence of the to-be-recognized voice and a confidence coefficient corresponding to the word sequence; under the condition that the confidence coefficient is smaller than a confidence coefficient threshold value, a plurality of candidate vectors with the similarity with the word sequence in a preset number are determined in the professional domain knowledge base and the dialect knowledge base respectively; and determining a target prompt word at least according to the to-be-recognized voice, the confidence coefficient and the plurality of candidate vectors, and correcting the word sequence according to the target prompt word through a voice correction model to obtain a target recognition result. According to the invention, a technical problem of low speech recognition accuracy in a scene of combining a professional field and a dialect in a speech recognition method in the related technology is solved.
Owner:CHINA TELECOM CORP LTD

Voiceprint recognition method and system based on wavelet transform convolution, terminal and medium

The invention provides a voiceprint recognition method and system based on wavelet transform convolution, a terminal and a medium. The method comprises the steps of obtaining a to-be-recognized voice of a current speaker; extracting Mel-frequency cepstral coefficient features based on the voiceprint features of the voice to be recognized; inputting the Mel-frequency cepstral coefficient features into a trained voiceprint recognition model so as to perform feature extraction on the Mel-frequency cepstral coefficient features, and obtaining voiceprint feature embedding vectors; and identifying the voiceprint feature embedded vector to judge whether the current speaker is a registered user or not. According to the method, the multi-dimensional hybrid wavelet transform convolution, the attention mechanism and the Res2Net network structure are organically combined, the capturing capability of different frequency features is remarkably improved, a multi-segment voice similarity calculation strategy is further adopted, the adaptability to complex environments and noise is improved, and the recognition precision of voiceprint features is effectively enhanced.
Owner:VERISILICON MICROELECTRONICS (NANJING) CO LTD +2

Multi-language voiceprint recognition method based on pre-training voice model

The invention discloses a multi-language voiceprint recognition method based on a pre-training voice model. The pre-training voice model WavLM and a traditional voiceprint recognition model ECAPA-TDNN are fused. According to the method, a multi-layer perceptron (MLP) module is introduced for further refining and converting the features extracted by the WavLM, so that the features are more suitable for the input requirement of an ECAPA-TDNN model, and the abstraction and expression ability of the model to the features is enhanced. In the aspect of multilingual voiceprint recognition, the model is finely adjusted by using a small amount of voice data sets, and the method comprises the following basic steps of: firstly, freezing parameters of a pre-trained voice model WavLM, so that the pre-trained voice model WavLM keeps learned knowledge; and then, parameters of the MLP module and the ECAPA-TDNN model are continuously adjusted in training, so that the multilingual voiceprint recognition capability is learned. In the application, a to-be-recognized voice passes through the fusion model to obtain a feature vector, and after the vector is subjected to judgment and decision making, a voiceprint recognition result is obtained.
Owner:BEIJING UNIV OF TECH

Contextually boosted aviation speech recognition

A variety of applications can include a system having a speech recognition system responsive to the speech input, where the speech recognition system can be configured to recognize the speech input using an aviation vocabulary including words extracted using state information of an aircraft or intent information of the aircraft associated with the received speech input. A control system can be implemented to automatically perform an action in the system in response to analysis of the recognized speech input, where the action is associated with flight of the aircraft.
Owner:CIRRUS DESIGN CORP D B A CIRRUS AIRCRAFT

Business process processing method and device, electronic equipment and storage medium

The invention discloses a business process processing method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: converting target voice data of a to-be-handled service of a user into a voice text; identifying a target business intention corresponding to the voice text; determining a business flow chart corresponding to the target business intention; and in the process of handling the to-be-handled business by the user, generating operation guidance information based on the business flow chart, and displaying the operation guidance information. According to the technical scheme provided by the invention, the user can be guided to complete the business process operation, the use threshold of the digital service is reduced, and the business handling efficiency and the use experience feeling of various users are improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Vehicle control method and device and vehicle

The invention discloses a vehicle control method and device and a vehicle, and relates to the technical field of vehicle control and voice interaction. The vehicle comprises a vehicle machine screen. The method comprises the following steps: determining an operation object on the vehicle machine screen according to a voice operation instruction input by a user; and when the number of the operation objects is greater than 1, identifying an intention operation object of the voice operation instruction in all the determined operation objects according to the historical operation of the user on the vehicle machine screen and / or the pointing information contained in the voice operation instruction and / or the vehicle machine application where the operation objects are located. According to the method, accurate matching of the voice operation instruction can be achieved, especially under the condition that multiple operation objects exist on the vehicle-mounted terminal screen, the operation object intended by the user can be accurately recognized, and therefore the accuracy and reliability of voice control or voice interaction are improved, and the user has good voice interaction experience.
Owner:GREAT WALL MOTOR CO LTD

Instant messaging-based proper noun speech recognition processing method and computer device

The invention discloses a proper noun speech recognition processing method based on instant messaging and a computer device. The method comprises the following steps: firstly, constructing annotation data sets of different scenes and a user-specific custom word list, and dynamically obtaining related data and a hot word list according to a to-be-recognized voice scene; afterwards, a training voice recognition model is subjected to fine tuning by using the annotation data set to obtain a first model, and after a recognition instruction is received, a hot word list is loaded for recognition to obtain a voice initial recognition text; and after the initial recognition text is obtained, dynamic optimization is carried out by utilizing a user-specific custom word list, and recognition errors are corrected. According to the method, matching data can be dynamically obtained, the model is combined with a scene to understand proper nouns, and the recognition difficulty caused by pronunciation and meaning complexity is reduced; in combination with targeted data and hot word information training, the proper noun recognition accuracy is improved, the defects of an existing model error correction mechanism are overcome, the accuracy of a final recognition result is remarkably improved, and powerful support is provided for speech recognition and subsequent application.
Owner:BEIJING VRV SOFTWARE CO LTD

System and Method for Flowsheet Population

A method, computer program product, and computing system for flowsheet population by voice. Dictation or conversation is transcribed from voice to text. The keys and values relevant to the documentation are extracted from the transcript and used to populate the flowsheet by selecting a subset of rows and examples relevant to the transcription of the information from a plurality of rows in the flowsheet and based on the transcript, each row corresponding to a key and a value for the flowsheet and extracting an instance of information from the transcript of information associated with a key from the subset of rows by processing a prompt including at least a portion of the transcript of information and the subset of rows with a generative artificial intelligence (AI) model using retrieval augmented generation (RAG).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Voice interaction method and device of refrigeration equipment, refrigeration equipment and electronic equipment

The invention discloses a voice interaction method and device of refrigeration equipment, the refrigeration equipment and electronic equipment, and belongs to the technical field of refrigeration equipment. The method comprises the following steps: acquiring voice interaction data of a target user to refrigeration equipment; target identity information corresponding to the target user is determined in a user library based on the voice interaction data, and the user library comprises identity information of multiple users; processing the voice interaction data based on the target identity information to obtain a first processing result; inputting the first processing result into a multi-task processing model to obtain a task embedding vector output by the multi-task processing model; and inputting the task embedding vector into a language model to obtain an interaction result corresponding to the voice interaction data. According to the method, voice identity recognition and multi-task recognition are combined, the accuracy of voice recognition and interaction is high, and the interaction effect on complex tasks is good.
Owner:QINDAO HAIER REFRIGERATOR CO LTD +1

Speech recognition optimization method and system applied to online AI legal consultation

The embodiment of the invention relates to the technical field of information, in particular to a speech recognition optimization method and system applied to online AI legal consultation. The method comprises the following steps: constructing a legal case library; voice content is recognized in real time, case plots are extracted, and the anxiety degree of a consultant is recognized; obtaining a plurality of prediction similar cases according to the extracted case plots and the identified anxiety degree; reading related legal terms, related subject descriptions and plot descriptions corresponding to all the prediction similar cases, and extracting keywords to obtain a keyword set; according to the plot influence singularity and the judgment trend uncertainty corresponding to all the prediction similar cases, obtaining a fusion singularity; and when the fusion degree exceeds a preset threshold value, matching the recognized words without definite semantics with the keywords in the keyword set again, and taking the keyword with the highest matching degree as a voice recognition result.
Owner:ZHEJIANG NAGU TECHNOLOGY CO LTD

Voice interaction method, device and equipment and computer storage medium

The invention relates to the technical field of artificial intelligence, and provides a voice interaction method and device, equipment and a computer storage medium. The method comprises the following steps: receiving a voice instruction input by a user; according to the voice instruction, obtaining an instruction text sequence corresponding to the voice instruction and historical interaction data corresponding to the user; the instruction text sequence and the historical interaction data are input into a semantic analysis model, enhanced semantic representation features are generated, and the enhanced semantic representation features are obtained based on fusion of text semantic features of the instruction text sequence, tag semantic features of the historical interaction data and a device topological relation; the equipment topological relation is obtained based on a pre-constructed knowledge graph of the household equipment; identifying intention information corresponding to the voice instruction and target equipment in the associated household equipment based on the enhanced semantic representation feature; and sending the equipment control instruction to the target equipment, wherein the equipment control instruction at least comprises the intention information and the equipment identification information of the target equipment.
Owner:CHINA MOBILEHANGZHOUINFORMATION TECH CO LTD +1

Electric wheelchair voice instruction recognition method and system

The invention relates to a voice instruction recognition technology of an electric wheelchair, in particular to a voice instruction recognition method and system of the electric wheelchair. The method comprises the following steps: receiving a voice instruction of a user, and identifying a fuzzy instruction in the voice instruction; on the basis of the fuzzy instruction, sensing data related to the current environment of the electric wheelchair are obtained; identifying a physical target in the sensing data and a relative distance between the physical target and the electric wheelchair; according to a preset attribution strategy, generating a customized clarification inquiry according to the physical target and the relative distance thereof; receiving a voice response, generating an instruction parameter according to the voice response, and controlling the electric wheelchair to execute a corresponding action according to the instruction parameter; and recording an interaction log, and optimizing the attribution strategy based on the interaction log. The problems of information loss, inaccurate instruction execution and system error attribution caused by improper clarification of fuzzy instructions can be effectively solved, and the accuracy of voice instruction recognition of the electric wheelchair is remarkably improved.
Owner:深圳复成医疗科技有限公司

Voice data processing method and device, equipment, storage medium and program product

The embodiment of the invention provides a voice data processing method and device, equipment, a storage medium and a program product, and relates to the field of artificial intelligence. The method comprises the following steps: converting to-be-recognized voice into to-be-recognized text data, and recognizing the to-be-recognized text data by using a verbal skill recognition model to obtain first text data of which the recognition result of the verbal skill recognition model is not passed; identifying whether the first text data contains a preset risk tag or not through the text data large model to obtain a first identification result; matching the first text data with a pre-established violation rule base to obtain a second recognition result; and based on the first recognition result and the second recognition result, determining whether the to-be-recognized voice is illegal. According to the method provided by the invention, the risk content suspected to be privately sold by the flight bill is effectively identified, the method has an identification capability for unknown flight bill privately sold verbal skills, the false alarm rate of the risk content is reduced, and the efficiency and accuracy of risk content identification are improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA +1

Language recognition method and device, electronic equipment and product

The invention provides a language recognition method and device, electronic equipment, a storage medium and a product, and the method comprises the steps: carrying out the sliding extraction of a voice segment from a to-be-recognized voice according to a sliding window of a preset size, and detecting a voiceprint turning point from the to-be-recognized voice; under the condition that the voiceprint turning point is detected from the voice segment in the first sliding window, carrying out displacement adjustment on the first sliding window according to the voiceprint turning point so as to enable the voiceprint turning point to be located at the end point position of the first sliding window; and performing language recognition on the shifted voice segment in the first sliding window to obtain a language recognition result. According to the scheme, the voiceprint turning point in the sliding window can be detected, the jumping moment of the speaker in the sliding window can be recognized, the sliding window is moved according to the jumping moment to divide the voice segments, language recognition of the voice segments of multiple languages and a single language in the sliding window can be avoided, and the user experience is improved. Therefore, the accuracy of language recognition of the to-be-recognized speech is higher, and the accuracy of speech translation is further improved.
Owner:IFLYTEK CO LTD

Sliding hybrid model construction method

The invention discloses a sliding hybrid model construction method, and particularly relates to the technical field of natural language processing and artificial intelligence, and the method specifically comprises the following steps: S1, voice-to-text and small model prediction; s2, performing Flag preliminary judgment and output; s3, threshold table judgment and model selection; and S4, building and predicting a large model prompt. The invention relates to a sliding hybrid model construction method, aims to solve the problems of low efficiency, poor accuracy, large resource consumption and the like in related business applications, and provides a sliding hybrid model construction method through construction of a long-tail intention threshold table, information extraction and classification based on prompt, 'prior information + PR curve 'threshold analysis and the like. The long-tail intention is quickly processed, the multi-label classification accuracy is improved, and the model performance is stabilized. The application effect is good in scenes such as automobile sales and after-sales, an efficient and intelligent solution is provided for multi-label classification and related services, and user experience and enterprise benefits are effectively improved.
Owner:深圳溥泉科技有限公司

Voice control method and device, electronic equipment and storage medium

The invention provides a voice control method and device, electronic equipment and a storage medium, and relates to the technical field of smart home. The method comprises the steps of obtaining a to-be-recognized voice signal of a user; inputting a to-be-recognized voice signal into the instruction simplification model to obtain a target control instruction output by the instruction simplification model; the instruction simplification model comprises a voice recognition model and a language simplification model, and the voice recognition model is used for converting a to-be-recognized voice signal into a text instruction; the language simplification model is used for simplifying the text instruction according to a preset semantic rule and a vocabulary mode to obtain a target control instruction; and executing the target control instruction. According to the method, the target control instruction is generated through voice recognition and language simplification, when a user sends out a tedious to-be-recognized voice signal with low accuracy, the language simplification model can filter redundant vocabularies and reduce semantic ambiguity, and then the problem that execution of equipment is affected due to the fact that the accuracy of the to-be-recognized voice signal sent by the old is low is solved.
Owner:GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1

Artificial intelligence-based message generation device and method

A message generation device method based on artificial intelligence are disclosed. The message generation device includes a memory and a processor electrically connected to the memory, wherein the processor is configured to receive selection of a user for a chat room from a user terminal, receive a voice file of the user recorded on the user terminal, recognize a voice of the voice file and generate a script converted into text and a summary message, and display the summary message as a conversation message in the chat room associated with the selection of the user.
Owner:CHOI JAE HO +2