Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

360 results about "Speech generation" patented technology

Speech generation. Speech generation and recognition are used to communicate between humans and machines. Rather than using your hands and eyes, you use your mouth and ears. This is very convenient when your hands and eyes should be doing something else, such as: driving a car, performing surgery, or (unfortunately) firing your weapons at the enemy.

Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Three-dimensional digital human generation method and system capable of voice interaction

The invention belongs to the technical field of three-dimensional reconstruction, and discloses a three-dimensional digital human generation method and system capable of voice interaction. According to the invention, brand new speaking audios in different languages are automatically generated according to different languages of the input target text and the sampled human voice audios; the sequential stability and detail reduction capability of three-dimensional human motion are guaranteed by using multi-model joint estimation and a sequential loss function, and facial expression details and hand postures in the image can be accurately estimated. After the high-precision three-dimensional human body model is obtained through estimation, human body action and expression generation is carried out based on voice driving, accurate synchronization of actions and expressions generated through voice is achieved, and facial expression movement and body posture movement, namely a whole-body three-dimensional human body model, conforming to brand-new speaking audio are accurately generated; and finally, rendering the whole-body three-dimensional human body model into a real digital human capable of voice interaction by using a three-dimensional neural rendering model. According to the invention, the realization of single person picture input, high-precision three-dimensional digital person generation and voice interaction is facilitated.
Owner:NANJING UNIV OF SCI & TECH

Robot control method and device based on pulse neural network, equipment and medium

The invention relates to the technical field of robot control, and discloses a spiking neural network-based robot control method, which comprises the following steps of: preprocessing collected multi-modal data to obtain an emotion pulse signal; inputting the emotion pulse signal into a pre-constructed emotion pulse neural network, and outputting a comprehensive emotion pulse; inputting the comprehensive emotion pulse into a central pattern generator, and outputting a behavior rhythm; acquiring an environment feedback signal generated by executing the behavior rhythm, and adjusting a connection weight according to the environment feedback signal; optimizing the comprehensive emotion pulse according to the connection weight, and converting the optimized comprehensive emotion pulse into emotion interaction voice information; and adjusting the emotion intensity according to the optimized comprehensive emotion pulse and the multi-modal data. According to the method, data are collected through the multi-mode sensor to generate emotion pulses, the emotion pulses are input into the central mode generator to generate behavior rhythms after SNN processing, weight optimization, emotional speech generation and emotional steady-state control setting are combined with the STDP algorithm, and the efficiency of a robot service scene is improved.
Owner:SHENZHEN ZHONGSHEN ZHIHUI TECHNOLOGY CO LTD

Voice generation method and device based on pseudo-autoregression modeling, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a voice generation method, device and equipment based on pseudo-autoregression modeling and a medium, and the method comprises the steps: obtaining a training sample containing a text sequence, a prompt voice segment and a target semantic token sequence; performing continuous fragment mask training on the text-to-semantic model to obtain a pseudo-autoregression trained text-to-semantic model; generating candidate speech output by using the text-to-semantic model and the initial semantic-to-acoustic model which are subjected to pseudo-autoregression training, and constructing a preference data pair; updating the semantics-to-acoustics model based on the preference data pair to obtain a preference optimized semantics-to-acoustics model; and generating target voice output based on the target text and the target prompt voice. According to the method, the time sequence modeling capability of the model is enhanced through pseudo-autoregression training, and the voice generation quality is directly optimized through the preference data pair, so that the voice alignment precision and the subjective listening feeling performance are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

High-fidelity end-to-end acoustic modeling method combining speech intelligibility reconstruction and autoregressive feedback optimization

PendingCN121565148ASpeech recognitionFrequency spectrumWeighting filter
The invention belongs to the technical field of speech recognition and acoustic modeling, and particularly relates to a high-fidelity end-to-end acoustic modeling method combining speech intelligibility reconstruction and autoregressive feedback optimization, which comprises the following steps of: firstly, establishing an intelligibility loss function based on a human speech intelligibility model; subjective definition features of the voice are reconstructed through a frequency spectrum reconstruction network, intelligibility related features are extracted and embedded in combination with a perceptual weighted filter bank, and the intelligibility related features and Mel frequency spectrum features are input into an acoustic encoder in parallel; based on this, by introducing a voice intelligibility reconstruction and autoregressive feedback mechanism, while the conciseness of an end-to-end voice modeling framework is maintained, systematic improvement of voice sharpness, naturalness and stability is realized, and the method can be widely applied to scenes such as intelligent voice assistants, voice transcription, virtual anchors, voice restoration, cross-language voice generation and the like. And the method has extremely high practical application value and popularization prospect.
Owner:GUANGDONG UNIV OF TECH

Multi-speech synthesis model bearing method and device based on virtual GPU

The invention provides a multi-speech synthesis model bearing method and device based on a virtual GPU, and relates to the technical field of graphics processing units, and the method comprises the steps: carrying out the virtualization processing of a physical graphics processing unit, dividing the physical graphics processing unit into a plurality of virtual processing units with independent video memories and calculation quotas, and combining a resource scheduling mechanism, and deploying the speech synthesis language model instances in a plurality of service containers, and constructing a plurality of speech synthesis model bearing units. After a voice synthesis request is accessed, the scheduling module carries out load balancing according to the request connection number of each bearing unit, the request is distributed to a target bearing unit with the minimum connection number, and a voice generation task is completed by a virtual processing unit bound with the target bearing unit. According to the invention, resource division can be carried out on the physical graphic processing unit, and efficient operation of the multi-speech synthesis model is realized.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Marketing verbal skill generation method and device, equipment, storage medium and program product

The embodiment of the invention provides a marketing verbal skill generation method and device, equipment, a storage medium and a program product, and relates to the field of financial science and technology and artificial intelligence. The method comprises the following steps: acquiring multi-modal data; performing cross-modal feature fusion on the voice data and the text data to generate a joint feature vector; based on the joint feature vector, through an attention mechanism, a multi-dimensional emotion evaluation result is generated, and the multi-dimensional emotion evaluation result comprises an evaluation result of at least one preset emotion dimension; and according to the multi-dimensional emotion evaluation result, generating and pushing a marketing verbal skill corresponding to the multi-dimensional emotion evaluation result. According to the method provided by the invention, through joint modeling of voice and text features and in combination with an attention mechanism, a key emotion region is focused, so that the accuracy of emotion scoring is remarkably improved, and the attitude of a customer can be reflected more truly; according to the emotion evaluation result, the optimization suggestion is generated, the generated verbal skill accurately matches the demand of the customer, the response efficiency of the marketing verbal skill is improved, and the customer emotion improvement rate is greatly improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Real-time intention recognition method and system based on streaming incremental reasoning

The invention relates to the technical field of voice processing, in particular to a real-time intention recognition method and system based on streaming incremental reasoning, and the method comprises the following steps: receiving the voice input of a user through a voice collection module, slicing the voice into a plurality of audio frames, and carrying out the recognition of the intention of the user through an incremental large language model module by adopting an Early-Exit reasoning mechanism; side outlets are arranged at multiple levels of the model, incremental reasoning is performed on token streams based on a QLoRA4-bit quantization technology, prediction results of multiple tokens are smoothed by using a stream ASR decoding module and an accumulative fusion module, and a stable final label is generated; the method has the beneficial effects that by combining streaming ASR decoding with incremental large language model reasoning, intention recognition and risk assessment can be immediately performed on each speech token after the speech token is generated. Through an Early-Exit reasoning mechanism, under the condition of high confidence, the system can output a fraud intention in advance in a middle layer in the reasoning process and stop subsequent calculation, and unnecessary calculation overhead is reduced.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Video data processing method and device and electronic equipment

The invention provides a video data processing method and apparatus, and an electronic device. The method comprises the steps of obtaining initial video data; the initial video data comprises initial video stream data and initial audio stream data corresponding to the initial video stream data; determining a role feature of at least one audio generation role in the initial video data; determining scene features of the initial video data; based on the role features, the scene features and the content translation text corresponding to the initial audio stream data, determining input data; inputting the input data into a preset voice generation model, and generating target audio stream data through the voice generation model; and generating target video data based on the initial video stream data and the target audio stream data. According to the mode, in the process of generating the translated audio corresponding to the video, tone cloning and emotion restoration of the translated audio are realized through the role features of the speaker of the original video and the scene features embodied by the video, the translation effect of the video is improved, and the video translation cost is reduced.
Owner:WANGYIYOUDAO INFORMATION TECH BEIJING CO LTD

Multi-modal emotion fusion robot voice style conversion method and device

The invention discloses a multi-mode emotion fusion robot voice style conversion method and equipment. The method comprises the following steps: extracting emotion feature vectors of text, audio and image input through a multi-mode emotion feature fusion extraction network; constructing a probability path taking Gaussian noise as a starting point and a target Mel spectrum as a terminal, taking the emotional characteristics as a condition guide vector, and predicting a vector field from the noise to target distribution by using a time condition U-Net; wasserstein-2 regularization constraint based on an optimal transmission theory is introduced to learn a smooth and efficient probability flow path; and finally, reconstructing a high-fidelity voice waveform with the target emotion through an acoustic decoder. According to the method, the problems of insufficient emotion expressive force and low generation efficiency in a traditional speech synthesis technology are solved, end-to-end optimization of emotion feature cross-modal mapping and speech generation is realized, and the emotion naturalness and the generation stability of synthesized speech are remarkably improved.
Owner:WUHAN UNIV

Healthcare provider assistant system and computer-implemented method

A healthcare provider assistant system receives voice data from a microphone; recognizes a first voice as a healthcare provider voice; assigns a second voice as a patient voice; generates a transcription of the voice data that identifies words spoken by the healthcare provider voice and by the patient voice; provides the transcription and / or the voice data to a trained language model; and provides a set of prompts to the trained language model including a first subset of prompts associated with the healthcare provider voice, each prompt including one or more tasks to complete using the transcription and / or voice data. At least one prompt of the first subset relates to obtaining healthcare data based on words spoken by the healthcare provider voice. The trained language model processes the transcription and / or the voice data and the set of prompts to provide responses to one or more prompts from the set of prompts.
Owner:VODAFONE GROUP SERVICES LTD

Multi-modal driving system based on streaming big language model output

The invention discloses a multi-modal driving system based on streaming big language model output, which belongs to the technical field of multi-modal interaction and comprises the following steps: a streaming big language model module receives input information and then generates a text unit sequence in a streaming manner; when each text unit is generated, a multi-dimensional emotion label sequence containing emotion types, emotion intensity values and semantic association degrees is synchronously output; the speech generation module maps the text unit sequence and the corresponding multi-dimensional emotion label sequence into speech synthesis parameters, and drives a streaming text-to-speech model to generate an emotional speech stream; the label generation module generates an expression label sequence and an action label sequence of the virtual image in real time; the emotion synchronous control module performs synchronous interpolation and coordination control by taking the time step of the text unit as a reference to generate a synchronous multi-mode driving signal; and the rendering and driving module drives the virtual image based on the signal to realize collaborative output of voice, mouth shape, expression and limb movement, and realizes synchronous presentation of multi-modal emotion expression.
Owner:SHANGHAI JIDOU TECH CO LTD

Speech synthesis method and related device

The invention provides a speech synthesis method and a related device, and relates to the technical field of speech synthesis. The speech synthesis method comprises the following steps: acquiring a first emotion feature of a target historical interaction speech; predicting a second emotion feature of a target voice to be generated according to the first emotion feature, a historical interaction text and a target text; wherein the historical interaction text is a text corresponding to the target historical interaction voice, the target text is a reply text generated based on an input voice in the latest round of voice interaction, and the target voice is a reply voice to be generated in the latest round of voice interaction; and generating the target voice according to the first emotion feature, the historical interaction text, the second emotion feature and the target text. According to the technical scheme provided by the invention, the problem that the emotional rhythm of the reply voice does not accord with the current context when the reply voice is generated based on the reply text in the prior art can be solved.
Owner:IFLYTEK CO LTD

Improving speech recognition by a machine learning model

The present disclosure describes techniques for improving speech recognition using a machine learning model. The machine learning model comprises a speech encoder configured to generate acoustic representations based on input speech, an adapter configured to generate adapted representations based on the acoustic representations, and a decoder configured to generate text corresponding to the input speech. A matching loss is applied during training the machine learning model. The matching loss is configured to explicitly force acoustic representations generated by the adapter to align with text embeddings. The machine learning model is fine-tuned by employing parameter-efficient low-rank adaptation. The machine learning model is trained to perform automatic speech recognition with performance improvement and parameter efficiency.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD +1

Controllable voice generation method and device based on multi-agent dynamic scheduling

The invention discloses a controllable voice generation method and device based on multi-agent dynamic scheduling, and belongs to the technical field of voice synthesis, and the method comprises the steps: constructing a multi-agent voice generation frame comprising a central scheduling module, an identity agent, an emotion agent and an environment agent; a central scheduling module analyzes a user instruction and outputs a structured task plan to drive each agent to generate primary voice output of identity, emotion and environment dimensions, and a cooperation cost matrix between the agents is constructed according to the primary voice output; the optimal execution path of the multiple agents is solved through an optimization algorithm based on the cost matrix, the agents are executed in a cascading mode based on the optimal execution path, and finally sound mixing is completed and high-quality voice is output. According to the method, the naturalness, the semantic consistency and the overall quality of the synthesized voice in a complex scene can be remarkably improved, and the method is suitable for various man-machine interaction application scenes such as intelligent voice assistants, virtual digital humans, immersive entertainment, barrier-free voice services, personalized content creation and the like.
Owner:ZHEJIANG UNIV OF TECH

Closed-loop medical voice interaction system and method

The invention provides a closed-loop medical voice interaction system and method. The closed-loop medical voice interaction system comprises a voice generation unit, a voice acquisition unit, a semantic understanding unit, a dialogue management unit and a data flow controller. According to the system, a structured inquiry summary can be automatically generated and converted into a standard medical document through the seq2seq model, and the document workload of a doctor is greatly relieved. For special groups such as patients with mobility difficulties, residents in remote areas and old people, the system provides a convenient remote inquiry way, and the problem of non-uniform distribution of medical resources is effectively relieved.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Experimental commentary real-time generation method

The invention discloses an experimental explanation real-time generation method, which comprises the following steps that: a training database is constructed, a knowledge enhancement language model is trained, and the training database comprises an experimental title, a video, an audio, voice transcription text information and sensor time sequence data; the training process specifically comprises the following steps: splicing and fusing data in a database to obtain a structured input sequence, inputting the structured input sequence into a knowledge enhancement language model, and generating a corresponding basic operation instruction; the knowledge enhancement language model judges whether a knowledge module needs to be called for content supplementation or not, the content comprises scientific principles or operation specifications, and an experiment explanation text is output; and generating an experimental explanation speech corresponding to the experimental explanation text by adopting an end-to-end real-time speech generation mechanism based on semantic driving. The system has a real-time response capability, and can output voice content which is moderate in information amount, clear in structure and rich in guiding significance in real time according to sudden changes, dynamic events and the like in the experiment process.
Owner:SOUTH CHINA UNIV OF TECH

System

An object of a system according to an embodiment is to reduce an adoption cost and prevent a mismatch.SOLUTION: According to an embodiment, a system includes a AI interviewer. The AI interviewer refers to points to be checked in interviews according to the company and past employment information using the RAG. The AI interviewer can use the voice generation AI to conduct a natural conversation. The AI interviewer analyzes the candidate's responses in real-time.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Implementation method of personalized artificial intelligence assistant system based on end-cloud collaborative architecture

The invention discloses an implementation method of a personalized artificial intelligence assistant system based on an end-cloud collaborative architecture, the personalized artificial intelligence assistant system based on the end-cloud collaborative architecture is constructed, face recognition and voice recognition tasks are deployed at a local terminal, user privacy is guaranteed, response delay is reduced, and the user experience is improved. Meanwhile, high-computing tasks such as large language model reasoning and personalized speech synthesis are deployed on the cloud, the semantic comprehension and speech generation capacity is improved, the system achieves bidirectional collaboration of the terminal and the cloud through an encrypted communication channel, a user is supported to upload knowledge documents and tone samples on a webpage side, and the user experience is improved. An exclusive question and answer knowledge base and a personalized voice model are constructed, and natural language interaction experience of thousands of people and thousands of faces is achieved. The system constructed by the method has the advantages of quick response, safety, reliability, high customizability and the like, and is suitable for multi-scene intelligent voice interaction requirements of education, accompanying, inquiry and the like.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Artificial Intelligence Based Character-Specific Speech Generation

A system includes a hardware processor and a memory storing software code, a character database, a language model and an artificial intelligence (AI) model trained to emulate speech by a character. The software code is executed to receive interaction data including a description of speech by a human to a performer impersonating the character and a description of a facial expression by the performer in response, obtain, from the character database, one or more communication trait(s) of the character, and generate, by the language model using the description of the speech and the communication trait(s) as inputs, a character-specific response to the speech. The software code is further executed to synthesize, by the AI model using the character-specific response and the description of the facial expression as inputs, audio data of the character-specific response in a voice of the character, and output the audio data for use by the performer.
Owner:DISNEY ENTERPRISES INC

Dialect speech enhancement method and device based on cross-modal and confrontation verification

The invention discloses a dialect speech enhancement method and device based on cross-modal and confrontation verification, and belongs to the technical field of speech recognition and enhancement. According to the method, the dialect voice and the lip movement video are jointly modeled, so that the accuracy and the naturalness of a dialect generation task are remarkably improved; a set of generation-adversarial-feedback closed-loop enhancement framework is constructed, and a multi-dimensional adversarial verification mechanism is introduced, so that the model is self-evolved, and enhancement data is ensured to be diversified and accord with real use habits of dialects. The dialect knowledge graph and the voice generation model are creatively fused, double enhancement of semantic and cultural levels is achieved, the generated dialect voice content better conforms to the real language environment and cultural background of the dialect voice content, and the method is particularly suitable for protection and inheritance of endangered or small dialects. The method can be directly applied to subsequent dialect recognition, synthesis or protection and the like.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Voice data generation method based on large model and method for training large model

The invention provides a voice data generation method based on a large model and a method for training the large model, and relates to the technical field of artificial intelligence, in particular to the technical fields of voice generation, intelligent customer service, video production and the like. The large model-based voice data generation method comprises the following steps of: receiving a rhythm description text and a voice text, wherein the rhythm description text describes a pronunciation rhythm intention of a plurality of text words in the voice text; semantic fusion is conducted on the rhythm description text and the voice text through a large model, rhythm fusion features are obtained, sub-features in the rhythm fusion features represent the pronunciation rhythm of the voice segments for the text words, and the to-be-generated target voice data comprise the voice segments; and according to the rhythm fusion feature and a specified pronunciation attribute related to the specified object, generating target voice data which represents that the specified object pronunciates according to the pronunciation rhythm intention and corresponds to the voice text.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Method, device and equipment for generating action of three-dimensional virtual object and storage medium

The application discloses a three-dimensional virtual object action generation method, device and equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: obtaining target voice data; performing audio feature coding on the target voice data to obtain first audio common features; wherein the audio common features refer to features corresponding to actions in audio features; obtaining sampling action features, which are obtained by randomly sampling an action-specific feature set; and performing feature decoding on the first audio common features and the sampling action features to obtain actions of a three-dimensional virtual object. The application can generate a variety of actions, such as different actions based on the same voice, which greatly improves the richness of actions.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Dynamic audio content generation

ActiveUS12572323B2Cryptography processingAutomatic exchangesSocial graphQuestion selection
There is provided a computer implemented method of dynamic generation of audio content of a panel including questions asked by a moderator and responses by responders, comprising: accessing user interest(s) of a target user, accessing a social network graph that includes the target user, selecting questions correlated with the user interest(s) of the target user, selecting responses to questions by responders, wherein the responders are linked to the target user in the social network graph, wherein the responders are associated with user interests correlated with the user interest(s) of the target user, dynamically assembling audio content(s) from voice records and / or from audio recordings generating by converting text to voice, each audio content including a sub-set of the selected questions and a sub-set of the responses, and providing dynamically assembled audio content(s) for selection thereof for playing on a speaker.
Owner:WEBTALK LTD

Controllable diffusion-based speech generation model

Systems and techniques described herein relate to a diffusion-based model for generating converted speech from source speech based on target speech. For example, a device may extract first rhythm data from input data, and may generate content embedding based on the input data. The device may extract second rhythm data from the target speech, generate a speaker insert from the target speech, and generate a rhythm insert from the second rhythm data. The device may generate converted rhythm data based on the first rhythm data and the rhythm embedding. The device may then generate a converted spectrogram based on the converted rhythm data, the speaker embedding, and the content embedding.
Owner:QUALCOMM INC

Intelligent dialogue method and system for accompanying old people based on user portraits

The invention discloses an elder accompanying intelligent dialogue method and system based on a user portrait, and relates to the technical field of intelligent interaction, and the method comprises the steps: calculating user features based on multi-modal data, splicing the user features into a user portrait vector, generating a frame sequence from voice, calculating an interaction feature vector, constructing an interaction sequence, generating a prefix tree, and extracting all paths. Calculating a support degree, and generating a candidate mode set; extracting an intention label corresponding to each mode in the candidate mode set, generating a frequent mode set, calculating similarity, generating an associated intention label and confidence, extracting frequency features in user portrait vectors, splicing the frequency features into interest feature sub-vectors, calculating joint utility, and selecting a maximum value as an optimal intention. According to the method, through combination of the prefix tree and frequent pattern mining and interaction sequence clustering, the capability of capturing long-term interaction behavior rules of old people is enhanced, and through combination of interest utility and matching utility, the accuracy of intention inference is improved.
Owner:KUAISHANGYUN (SHANGHAI) NETWORK TECHNOLOGY CO LTD

Food fraud detection device using terahertz spectroscopy

This invention provides a food fraud detection device using terahertz spectroscopy that achieves a compact and stable structural configuration. [Solution] The food fraud detection device 100 using terahertz spectroscopy according to the present invention comprises a base housing 20 and an upright housing 21 attached to the base housing 20. A display unit 1 is provided in the upright housing 21, and a user input unit 2 is provided in the base housing 20. The main body of the device 100 is equipped with a wireless connection antenna 7, a voice generation unit 8, a voice input port 81, first to fourth connection ports 31, 32, 33, 34, an external sensor connection port 4, a server connection port 5, and a port 6 for opening the device. An optical sensor 10, an ultrasonic sensor 11, and a chemical sensor 12 are arranged on the base housing 20 together with corresponding control switches 13, 14, 15.
Owner:シャイマ アブデルラオフ モハメド アブデルモフセン +2

Marketing verbal skill generation method and device, equipment, storage medium and product

The invention discloses a marketing verbal skill generation method and device, equipment, a storage medium and a product. The method comprises the following steps: determining a historical rejection reason set, a selling point abstract of each target product and a strong correlation label of each target product; based on the historical rejection reason set, the product information of each target product and the strong correlation tag, driving a first large language model through a scenarized cue word, and generating a cut-in verbal skill and a save verbal skill of each target product; and based on the strong correlation label, the cut-in verbal skill, the save verbal skill and the selling point abstract of each target product, determining the marketing verbal skill of each target product by using a preset template. The problems that an existing marketing verbal skill generation method is low in generation efficiency and insufficient in content matching degree are effectively solved.
Owner:CHINA MOBILE GROUP SHANDONG +1

Information processing device, information processing method, and information processing program

The present invention provides an information processing device that enables a more natural improvement in the accuracy of identifying past related utterances made by a user or identifying their relationship with other users. [Solution] The system includes: a speech acquisition unit 1 that acquires user utterances; a prompt generation unit 2 that generates prompts for a large-scale language model to generate tag information that includes at least one of the following based on the acquired utterances: the relationships between the constituent elements of the utterance, the attributes to which the utterance belongs, or a summary of the utterance; an interface 3 that transmits the generated prompts to the large-scale language model and receives response information that includes tag information generated by the large-scale language model in response to the transmitted prompts; and a recording unit 5 that records the acquired utterances and the received tag information in association.
Owner:PIONEER IP

Method and system for providing video conference service including artificial intelligence-based speech interpretation function

PCT designated stageWO2026146907A1Speech translationSpeech sound
Disclosed are a method and system for providing a speech conference service including an artificial intelligence-based speech interpretation function. According to one embodiment, the method for providing a video conference service may comprise the steps of: setting two or more interpretation bots participating as virtual participants in a video conference service; interpreting a speech of a first language input by participants of the video conference service into a speech of at least one other language different from the first language through processes of speech recognition, translation, and synthetic speech generation between the two or more interpretation bots and artificial intelligence; and providing the interpreted speech.
Owner:LINE PLUS