Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

46 results about "Expressive communication ability" patented technology

Voice synthesis method and system based on VITS improvement

The invention provides a voice synthesis method and system based on VITS improvement, and the method comprises the steps: optimizing a text encoder of a VITS model, introducing a large language model, and enabling the emotion, intention and speaking style of an input text to be captured when the text is encoded; a random disturbance item is introduced when the Q value is dynamically planned and solved, the alignment flexibility in the initial training stage is improved, meanwhile, monotonicity constraint is strictly kept, and it is avoided that a suboptimal solution is obtained through convergence too early; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is utilized to efficiently reconstruct a time domain signal, so that waveform up-sampling is realized, redundant calculation of traditional transpose convolution is avoided, and reasoning is accelerated. According to the method, the reasoning speed, the emotion expression ability and the style control flexibility of speech synthesis can be effectively improved, a new solution is provided for cross-language diversified speech synthesis, and a reference is provided for the more efficient and more intelligent development of the speech synthesis technology.
Owner:豫章师范学院

Sound authenticity identification method and system based on multi-modal feature deep interactive fusion

The invention discloses a sound authenticity identification method based on multi-modal feature deep interactive fusion. According to the method, a double-flow architecture is adopted, and a pre-trained BEATs model and a CNN14 network are respectively utilized to extract Transform sequence features and convolution time-frequency embedding features of an audio; an adaptive MobileFormer fusion device is innovatively proposed, an original MobileFormer structure used in the image field is transformed into a bidirectional cross-modal interaction module suitable for one-dimensional time sequence audio features, dynamic complementary modeling of local details and global semantic features is achieved through a cross attention mechanism, and dynamic nonlinearity is introduced to activate and enhance the expression ability; and the fused enhanced features are input into a hierarchical graph attention network ASSIST, and high-order semantic modeling and authenticity classification are completed by combining spectrogram and time sequence double-flow reasoning. Experimental results show that the performance of the method on an ASVspoof2021LA data set is superior to that of an existing baseline model. The method can effectively detect AI generation voice, replay attack and other forged voice, and is suitable for a voice authentication system.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Multi-mode knowledge representation learning method fusing multi-attention mechanism and semantic enhancement

The invention relates to a multi-modal knowledge representation learning method fusing a multi-attention mechanism and semantic enhancement, and aims to solve the problems that the multi-modal knowledge representation method is insufficient in multi-modal feature fusion and difficult to effectively model complex semantic relationships such as symmetric and anti-symmetric. The method comprises the following steps: firstly, extracting entity image and text features from a multi-modal knowledge graph by using a CLIP model, and extracting entity audio features by using a VGGish model; then image features are enhanced through spatial attention, and image-text feature fusion is realized through cross attention; designing a cross-modal fusion module to dynamically fuse the image-text features and the audio features to obtain multi-modal fusion features; and finally, extracting structured features based on a ComplEx model, integrating the structured features with multi-modal fusion features through an adaptive dual-channel scoring function, and optimizing a training process by adopting a contrast learning loss function. The multi-modal features can be fully fused, the semantic expression ability is enhanced, and the performance of downstream tasks such as intelligent question and answer is improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Method and device for evaluating speech expression ability of neurological and mental diseases

ActiveCN121483566AHealth-index calculationSemantic analysisMicrolinguisticsNeuropsychiatric disease
The invention discloses a method and device for evaluating speech expression ability of neurological and mental diseases. According to the method, speech is induced through a visual stimulation material, speech data are collected, and after speech recognition, semantic calibration and text error correction, evaluation indexes are calculated from multiple dimensions of fluency, microcosmic linguistics and macroscopic linguistics in combination with a natural language processing technology, a large language model and a vector embedding technology, so that the evaluation accuracy is improved. And the result credibility is guaranteed through a robustness verification mechanism of iteration reflection of a double-big language model. According to the method, the problems of large subjective deviation, low efficiency and shallow analysis dimension of traditional evaluation are solved, and objective, accurate, efficient and comprehensive evaluation is realized.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

Estrus-sharing dialogue generation method and system fusing psychological information

The invention discloses an emotional dialogue generation method and system fused with psychological information, and the emotional expression capability of generating response can be improved through the obtained global emotional latent variable fused with a speaker and a listener. By determining a multi-hop knowledge reasoning path set, obtaining context semantic representation of each knowledge reasoning path, and fusing the context semantic representation with dialogue history global context representation to obtain knowledge-enhanced context representation, the knowledge reasoning ability and semantic relevance in a dialogue context are improved; by constructing a knowledge representation matrix, taking a global emotion latent variable as a bias item, and dynamically adjusting the knowledge representation matrix and a value vector generated by knowledge-enhanced context representation through semantic features, multi-scale attention fusion representation is obtained, and the context adaptive capacity and emotion representation naturalness of generated response are effectively improved. According to the method, the psychological state of the user can be identified more accurately, and the dialogue response with humanization, estrus-sharing expression and interaction adaptability is generated.
Owner:ZHEJIANG NORMAL UNIV

Multi-modal dialogue emotion recognition method based on uncertainty adaptive weighting

The invention belongs to the technical field of multi-modal emotion recognition, and particularly relates to a multi-modal dialogue emotion recognition method based on uncertainty adaptive weighting. The method comprises the following steps: processing original modal information of each modal to obtain an alignment feature vector of each modal; calculating a cross-modal fusion feature of each cross-modal path based on the alignment feature vector of each modal; obtaining an uncertainty score of each cross-modal path according to the cross-modal fusion features of each path; obtaining the optimized weight of each path based on the initial weight and the uncertainty score of each path; and obtaining a predicted emotion category based on the optimized weights of all the paths and the cross-modal fusion features. According to the method, the expression ability of the non-text mode in the key emotion scene is remarkably enhanced while the semantic advantages of the text are kept.
Owner:ZHEJIANG UNIV OF TECH

Natural emotional singing method and emotional singing audio processing method

PendingCN122347963AAcoustic areaMusic class
The application discloses a natural emotional singing method and an emotional singing audio processing method, and belongs to the singing information field, and aims to solve the problems of fragmented training elements and single feedback system of a traditional vocal music teaching method. The technical scheme comprises six training modules combined by action guidance, emotion awakening and sound area training, and a training system for improving singing emotional expression capability based on a deep learning model for real-time extraction of user singing audio emotional indexes and generation of dynamic visual feedback.
Owner:BEIJING FABIAN EDUCATION TECH CO LTD

Voice keyword recognition method based on CNN-LSTM and online knowledge distillation

The invention discloses a voice keyword recognition method based on CNN-LSTM and online knowledge distillation, and relates to the technical field of voice processing. According to the method, an improved contraction residual attention module is introduced into a CNN-LSTM-based neural network model and is used for discovering and inhibiting redundant information and noise in features, and the expression ability of the features is enhanced. According to the method, a training method based on online knowledge distillation is introduced, the current training iteration is supervised by using information in the previous two training iterations, and an extra regularization effect is provided for the training process by using the difference in prediction probabilities output by a neural network model. The overfitting risk of the neural network model is reduced, and the generalization ability of the neural network model is enhanced.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Voice generation method and device based on preference alignment, equipment and medium

The invention relates to the technical field of speech synthesis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a speech generation method and device based on preference alignment, equipment and a medium, and the method comprises the steps: obtaining a pre-training speech generation model and a preference training sample pair, constructing a preference alignment model and a non-preference alignment model, and taking the pre-training model as a reference; determining a loss function based on the preference samples and the non-preference samples, and updating model parameters to obtain a trained model; and receiving a target condition item and generating an agent prompt in a reasoning stage, fusing speed prediction results of the two types of models, and generating a target voice based on a fusion result in a stream matching process. According to the method, human preference signals are fused through a preference alignment mechanism, so that the naturalness and semantic consistency are considered in the speech generation process, and the speech quality and the personalized expression ability are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Augmented reality content recommendation and expressive presentation method based on user behavior perception

ActiveCN121685902BImprove interactive experienceImprove the effect of information transmissionInformation transmissionThree-dimensional space
The application provides a kind of based on user behavior perception's augmented reality content recommendation and expressive presentation method, including constructing the state representation of multiple virtual roles in augmented reality scene;Joint objective function for evaluating virtual role layout rationality is constructed;According to the attribute information of virtual role, the adaptive weight is generated by the weight prediction model constructed in advance, and the adaptive weight is used to adjust the cost item related to the importance of role in the space cost function;Based on the joint objective function, the position and orientation of all virtual roles are iteratively optimized to obtain an optimized layout;Expressive presentation is carried out in the augmented reality scene The optimized layout is realized in three-dimensional space Automatic optimization of role position and orientation, avoid visual conflict, improve the naturalness, coordination and semantic expression ability of layout, thereby significantly enhance the interactive experience and information transmission effect of AR system.
Owner:BEIJING TECH & BUSINESS UNIV

Empathy-Based Wearable Systems for Perceptual and Communicative Assistance

PendingUS20260179644A1Speech recognitionHeadphonesEmpathy
The invention provides a unified assistive technology system, the ADA Empathy Wearable Bundle, enhancing perception, communication, and self-expression for individuals with visual, auditory, or speech impairments.It integrates Empathy Glasses, Empathy Headphones, Word Articulation Guidance (WAGS), Word Articulation Response Monitoring (W.A.R.M.), and Conversational MIDI (C-MIDI), using AI-driven narrative interpretation, augmented reality, and blockchain-based security.The system delivers vivid narrative audio for blind users, multimodal AR overlays for hearing-impaired users, and real-time visual and auditory articulation guidance for speech-impaired or non-verbal users. A tone-based trust protocol ensures operational integrity, while blockchain storage preserves privacy and consent.Grounded in ethical AI, empathy, and human dignity, the system enables users to perceive, understand, and engage with the world fully, offering a holistic, compassionate, and transformative solution in assistive technology.
Owner:BOWES DEAN

A sign language sentiment recognition and teaching feedback method based on machine learning

PendingCN122454639AComputer aided instructionComputer-aided
The application relates to the technical field of artificial intelligence, affective computing and computer-assisted teaching, in particular to a sign language emotion recognition and teaching feedback method based on machine learning, which collects learner sign language videos and extracts face, hand and posture holographic key points; action semantic features and emotion expression features are obtained in parallel by using a double-branch time sequence coding network; the sign language content and the emotion state containing continuous values of valence-arousal-dominance and discrete categories are obtained by a semantic recognition subnetwork and an emotion recognition subnetwork respectively; the recognition result is compared with standard semantics and emotion labels, and emotion intensity, naturalness, semantic matching degree scores and comprehensive quality scores are calculated; emotion correction instructions, reinforcement learning training sequence planning and key point trajectory visualization feedback are generated based on the scores; and teaching strategies are iteratively improved through closed-loop optimization and emotion resonance models. The application realizes synchronous recognition and quantitative evaluation of sign language semantics and emotions, and provides an adaptive feedback means for sign language emotion expression ability training.
Owner:杨莉红

An emotion dialogue generation method based on an improved generative adversarial network

The emotion dialogue generation method based on improved generative adversarial network belongs to the dialogue system field under natural language processing, realizes the interaction between memory matrices through a self-attention mechanism, enhances the long-distance transmission energy between information, thereby improves the expression ability and feature extraction ability of the model, and solves the problem that the existing model has weak expression ability and generates short sentences. Meanwhile, a multi-class discriminator is set, which respectively discriminates the true and false of the text and the emotion category, calculates the category relative loss and the emotion information loss two parts to feed back and update the generator, so as to improve the consistency of the emotion information, and make the emotion expression of the reply generation sentence more obvious and clear. Finally, the iterative evolution algorithm is used in the generator update, the temperature parameter and the quality parameter are controlled respectively, and the child generator in the optimal direction of the evolution temperature is selected to complete the generation task. The dialogue generation method of the present application realizes the emotion embedding which considers the quality and diversity of the reply.
Owner:BEIJING UNIV OF TECH

A script semantic-driven multi-digital human collaborative generation method

This invention relates to the field of artificial intelligence technology, specifically to a script semantic-driven multi-digital human collaborative generation method. The method includes script input, script semantic feature encoding, script semantic dual-graph construction, dynamic-static graph joint optimization based on a constraint-driven mechanism, script temporal semantic hierarchical modeling, multi-digital human role behavior strategy generation, retrieval-enhanced digital human speech generation processing, triple consistency constraint fusion, and multimodal scene generation. By constructing a script semantic dual-graph structure containing role nodes, scene nodes, and event nodes, and introducing a joint modeling and optimization mechanism of a global static relationship graph and a local dynamic interaction graph, the accuracy and expressive power of complex script semantic modeling are improved. Furthermore, by introducing a multi-digital human behavior generation mechanism based on role relationship constraints, and combining retrieval-enhanced speech generation with semantic, emotional, and temporal triple consistency constraints, the consistency and naturalness of multi-digital human collaborative expression are enhanced.
Owner:BEIJING XILIANLIAN TECHNOLOGY CO LTD

Machine-learned attention models featuring echo-attention layers

The present disclosure provides echo-attention layers, a new efficient method for increasing the expressiveness of self-attention layers without incurring significant parameter or training time costs. One intuition behind the proposed method is to learn to echo, i.e., attend once and then get N echo-ed attentions for free (or at a relatively cheap cost). As compared to stacking new layers, the proposed echoed attentions are targeted at providing similar representation power at a better cost efficiency.
Owner:GOOGLE LLC

An adaptive piano automatic performance control method and system

PendingCN122347934APianoEngineering
The application discloses a kind of self-adapting piano automatic performance control method and system, method includes calibration stage and performance stage.Calibration stage obtains the multidimensional response data reflecting the mechanical response characteristic of each key and pedal of specific piano, constructs the individualized response model representing independent mechanical response characteristic of each key, and determines performance compensation parameter set.Performance stage obtains target music score data and emotion indication information, generates ideal performance control sequence through emotion mapping model, then carries out note-by-note correction based on performance compensation parameter set, generates adapted performance control sequence and outputs to piano automatic performance device execution.The application realizes the closed-loop architecture of "key-by-key response modeling-key-by-key compensation determination-emotion mapping generation-note-by-note fusion execution", makes the automatic performance effect adapt to the individual differences of keys of different pianos, and gives human-like emotion expression ability, significantly improves performance consistency and realism.
Owner:刘晋恺

Speech synthesis method and device, electronic equipment and storage medium

The invention provides a speech synthesis method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a to-be-synthesized text and a mood description text for describing the non-semantic information of a to-be-synthesized target speech signal, carrying out the joint coding of the to-be-synthesized text and the mood description text, and obtaining a mixed lexical element sequence, and inputting the mixed lexical element sequence into a mood control synthesis model to obtain an audio lexical element sequence output by the mood control synthesis model, and decoding the audio lexical element sequence to obtain a target voice signal. According to the speech synthesis method and device, the electronic equipment and the storage medium provided by the invention, the mood description text in a natural language form is used as an additional input parameter, so that the model can directly understand and precisely regulate and control the non-semantic attribute of the speech; the technical problems of dependence on fixed labels, coarse control granularity and single expression ability in the prior art are solved, and the controllability, diversity and anthropomorphic expressive force of synthetic speech are remarkably improved.
Owner:IFLYTEK CO LTD

Multi-round interaction emotion speech synthesis method and system based on context self-adaption

The invention provides a multi-round interactive emotion speech synthesis method and system based on context self-adaption, and belongs to the technical field of speech synthesis, and the method comprises the steps: obtaining an instant multi-round dialogue with an adjacent single-round historical speech as a context window and a current to-be-synthesized text; directly deconstructing and predicting emotional and multi-dimensional side language acoustic features from historical voice signals through a two-stage trained context adaptive feature predictor; through a feature parameter mapping mechanism of a preset standardized template, the prediction features are converted into standardized control parameters available for synthesis; based on the standardized parameters and the to-be-synthesized text, driving a speech synthesis model to generate natural and coherent target speech; the system comprises a data acquisition module, a context adaptive feature prediction module, a feature parameter mapping module and a speech synthesis module. According to the method, dual optimization of emotion accuracy and voice naturalness in a multi-round interaction scene is realized, and the comprehensive performance of acoustic quality and emotion expression ability of the synthesized voice is remarkably improved.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Speaker voice segmentation method based on non-local space U-Net and mixed features

PendingCN121983030AImprove discrimination abilityFully explore the spatial and temporal correlationsSpeech recognitionSpeech segmentationSemantic feature
The invention belongs to the technical field of voice signal processing, and particularly provides a speaker voice segmentation method based on non-local space U-Net and mixed features. According to the invention, the non-local space U-Net network is constructed, and a non-local space attention module is introduced, so that the long-range dependency relationship in voice signals is effectively captured, and the expression ability of spatial features is improved; and meanwhile, a mixed feature fusion mechanism is adopted, and time-frequency domain features and deep semantic features are combined, so that the discrimination of voice features is enhanced. In addition, a dynamic weight loss function is designed, the fusion efficiency of the features under different scales is optimized, and the contribution degrees of various voice segments are balanced. According to the method, the space-time relevance in the voice signals can be fully mined, the dependence of a traditional method on local features and single features is overcome, and the accuracy and robustness of speaker voice segmentation are remarkably improved.
Owner:CHINA CRIMINAL POLICE UNIV

Multi-scale time domain feature learning and audio-visual emotion fusion method for depression recognition

PendingCN122348057ATime domainBi modal
The application discloses a multi-scale time domain feature learning and audio-visual emotion fusion method for depression degree recognition, and specifically comprises the following steps: step 1, collecting a sample set with depression degree labels, and preprocessing each sample to obtain a preprocessed sample set; step 2, building an MTFL-DTAF model, inputting each sample in the preprocessed sample set into the MTFL-DTAF model one by one, predicting the depression score of each sample, and recognizing the depression degree of the sample. In the depression degree recognition method, the multi-scale time domain feature learning module enhances the time expression ability of the depression features; and the dual-modal time attention fusion module can further select more discriminative depression features, thereby improving the performance of the visual and audio dual-modal depression degree recognition network.
Owner:XIAN UNIV OF TECH

Mobile English learning system based on interaction

The invention discloses an interaction-based mobile English learning system, which belongs to the technical field of English learning and comprises a user management module, an English level test module, a level analysis and formulation module, a multi-aspect learning module, an interactive practice module, a feedback module and a data management and storage module. The user management module is used for being responsible for basic information management and authority control of users; and the English level test module is used for evaluating the initial English level of the user and providing a basis for making a subsequent learning plan. According to the invention, a special interactive practice module is arranged, a real language application scene is simulated by using a voice recognition technology, real-time interactive dialogue practice between a user and AI is supported, the oral expression ability and the actual application level are significantly improved, and meanwhile, a gamification breakthrough mechanism is introduced, so that the learning interestingness is improved, and the interestingness of learning is improved. And the participation sense and continuous learning motivation of the user are enhanced.
Owner:HEBEI VOCATIONAL & TECH UNIV OF SCI & TECH

Robot emotion expression system under specific cultural background

The invention discloses a robot emotion expression system under a specific culture background. The system comprises a culture adaptation module, an emotion perception module, a scene construction module, a trigger logic module, a dialogue strategy library, a behavior expression module, a memory optimization module and an emotion adjustment module. The output ends of the culture adaptation module, the emotion perception module and the scene construction module are connected to the input end of the logic trigger module, the logic trigger module is in bidirectional connection with the dialogue strategy library, the output end of the logic trigger module is connected to the input end of the behavior expression module, and the output end of the behavior expression module is connected to the input end of the memory optimization module. The output end of the memory optimization module is connected to the scene construction module and the logic trigger module, and the output end of the emotion adjustment module is connected to the logic trigger module and the behavior expression module. According to the invention, a life scene under a specific culture background can be simulated, natural expression is triggered, and the emotional expression ability of the robot is significantly improved.
Owner:GUANGZHOU UNIV OF CHINESE MEDICINE SHENZHEN HOSPITAL (FUTIAN)

Personalized synthesis and recognition enhancement of dysarthric speech

ActiveCN120412540BSpeech synthesisDysarthric speechSpeech disorder
This invention discloses a personalized synthesis and recognition enhancement method for speech disorders. The speech disorder synthesis model includes a long-range dependent feature encoding module, a non-stationary feature encoding module, and a decoding module. The input of the speech disorder synthesis model includes samples, and the output includes synthesized speech disorder speech. The samples are speech disorder text sequences. The input of the long-range dependent feature encoding module includes samples, and the output is an alignment vector z. The input of the non-stationary feature encoding module includes the alignment vector z, and the output is the final embedding representation. The input of the decoding module is the final embedding representation, and the output is the synthesized speech disorder speech. The speech disorder synthesis model of this invention improves the ability to extract personalized features of speech disorder speech, enhances speech synthesis performance, and improves the fine-grained expression of speech disorder speech features.
Owner:TIANJIN UNIV

TMD detection model based on voice time-frequency feature fusion, construction method and system

The invention relates to a TMD detection model based on voice time-frequency feature fusion, and a construction method and system thereof, and the model comprises a time sequence feature extraction module which is used for extracting the time sequence features of a voice signal according to MFCC features; the frequency domain global feature embedding module is used for embedding the acoustic features related to the TMD in the frequency domain of the voice signal into the time sequence features; the classification module is used for judging whether the patient corresponding to the voice signal is a TMD patient or not; speech signals are extracted, MFCC features and acoustic features of the speech signals of a TMD patient and a non-TMD patient are input into a time sequence feature extraction module and a frequency domain global feature embedding module respectively, an initial TMD detection model is trained, and a TMD detection model used for detection is obtained; according to the method, the overall feature expression capability is enhanced, and the accuracy and robustness of the TMD detection model are remarkably improved, so that the recognition effect of the TMD detection model on TMD-related voice anomalies is improved, and pathological information hidden in voice signals can be more comprehensively mined.
Owner:ZHONGNAN HOSPITAL OF WUHAN UNIV

Speech synthesis method and system based on improved vits

The application provides a speech synthesis method and system based on an improved VITS, wherein a large language model is introduced by optimizing a text encoder of a VITS model, so that the input text can capture the emotion, intention and speaking style when the text is encoded; a random disturbance term is introduced when a Q value is solved by dynamic programming, so that the alignment flexibility in the early training stage is improved, and the monotonicity constraint is strictly maintained to avoid early convergence to a suboptimal solution; a ConvNeXt module is used as a basic backbone network of a decoder, and ISTFT is used to efficiently reconstruct a time domain signal, so that the upsampling of the waveform is realized, the redundant calculation of a traditional transpose convolution is avoided, and the inference is accelerated. The application can effectively improve the inference speed, emotion expression ability and flexibility of style control of speech synthesis, provides a new solution for cross-language diversified speech synthesis, and provides a reference for the development of speech synthesis technology in a more efficient and intelligent direction.
Owner:豫章师范学院

A method and apparatus for assessing speech expression ability in neuropsychiatric diseases

ActiveCN121483566BHealth-index calculationSemantic analysisMicrolinguisticsNeuropsychiatric disease
The application discloses a method and device for evaluating speech expression ability of neuropsychiatric diseases. The method induces speech through visual stimulation materials and collects voice data. After voice recognition, semantic calibration and text correction, the method combines natural language processing technology, large language model and vector embedding technology to calculate evaluation indexes from multiple dimensions of fluency, micro linguistics and macro linguistics, and guarantees the result reliability through the robustness verification mechanism of double large language model iteration reflection. The application solves the problems of large subjective bias, low efficiency and shallow analysis dimension of traditional evaluation, and realizes objective, accurate, efficient and comprehensive evaluation.
Owner:THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV

Speech synthesis method and device, electronic equipment and storage medium

The application provides a speech synthesis method and device, electronic equipment and storage medium, belonging to the technical field of artificial intelligence, comprising: obtaining a text to be synthesized and a tone description text for describing non-semantic information of a target speech signal to be synthesized, jointly encoding the text to be synthesized and the tone description text to obtain a mixed word sequence, inputting the mixed word sequence into a tone control synthesis model to obtain an audio word sequence output by the tone control synthesis model, and decoding the audio word sequence to obtain the target speech signal. The speech synthesis method and device, electronic equipment and storage medium provided by the application use the tone description text in natural language form as an additional input parameter, so that the model can directly understand and accurately control the non-semantic attributes of the speech, solve the technical problems of the prior art, such as dependence on fixed labels, coarse control granularity and single expression capability, and significantly improve the controllability, diversity and humanization of the synthesized speech.
Owner:IFLYTEK CO LTD

Doll with replaceable accessories

The utility model relates to the technical field of toy accessories, in particular to a toy with replaceable accessories. The toy comprises a doll body, a plurality of eye assemblies and connecting structures, the eye assemblies are movably connected with the doll body through the connecting structures, mounting holes are formed in the doll body, each eye assembly is provided with the connecting structure matched with the doll body, and the eye assemblies comprise a plurality of eye patterns with different expressions. By replacing the eye assembly, the doll can display various expressions, increase the interestingness of interaction with a user, stimulate the imagination and creativity of the user, replace appropriate eyes for the doll according to different situations, simulate different emotions, and contribute to cultivating the emotional expression ability and the like of the user.
Owner:SHANDONG SHANGNUO INVESTMENT HLDG CO LTD