Method and system for adaptively acquiring image data to carry out personalized voice conversation

By acquiring user voice and image data and combining natural language processing and deep learning technologies, personalized voice dialogues are generated, solving the problem that existing chatbots cannot reflect user status and environment in real time, and realizing personalized voice dialogue services.

CN121641010APending Publication Date: 2026-03-10PEIXI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing natural language chatbots such as ChatGPT cannot provide real-time, personalized voice dialogue services, lack correlation with the user's state and environment, and cannot engage in dialogue based on individual needs and preferences.

Method used

By acquiring voice and image data through user devices, and utilizing natural language processing models and deep learning techniques, combined with voice features, image features, and environmental information, personalized voice dialogues are generated.

Benefits of technology

It enables personalized voice dialogue services that reflect users' emotions and environment in real time, and can provide personalized dialogue content based on users' real-time status and needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641010A_ABST
    Figure CN121641010A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for providing personalized voice conversation service and a transaction system for providing personalized model making. The cloud server is connected with a user device, provides user personalized voice conversation service through a conversation interface, and executes a method for providing the user personalized voice conversation service by a processor. Wherein the cloud server receives voice data generated by real-time voice of a user, can comprise image data, obtains current voice features and image features so as to learn current emotion of the user, recognizes semantics in the voice features through a natural language processing model, and sends the voice data to the user according to the semantics and current environment information. A voice dialogue based on the current emotion of the user is correspondingly generated, and then the voice dialogue is issued through a simulation object displayed on the user device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to personalized intelligent voice dialogue technology, and more particularly to a method and system for personalized voice dialogue using adaptive image data and personalized information obtained from a language model. Background Technology

[0002] Among the rapidly developing fields of artificial intelligence (AI), one type is the natural language chatbot, which can process natural language and automatically generate content. Examples include ChatGPT (Chat Generative Pre-trained Transformer), developed by OpenAI. These chatbots utilize generative AI technology, allowing them to be trained on large amounts of data and then generate new data related to the original data. This new data is then used for deep learning (such as generative adversarial networks, GANs) to build an intelligent model.

[0003] Take ChatGPT as an example. ChatGPT is trained by learning from a large amount of network information and can respond to users in natural language. However, the content that usually responds to users is a standard answer derived from learning and cannot adapt in real time to provide answers that are relevant to the user's current state. Although it is a natural language chatbot, it lacks content that is relevant to the user and consistent with the real-time situation. In other words, the dialogue service provided by current chatbots is general and does not provide personalized voice dialogue services (such as personal needs, preferences, background, etc.). Summary of the Invention

[0004] To provide a voice dialogue model that can be tailored to individual needs and learning preferences, the open book proposes a method and system for adaptively acquiring image data for personalized voice dialogue.

[0005] According to a system embodiment of a method for providing personalized voice dialogue services to users, the system includes a cloud server, a user device connected to the cloud server, the user device acquiring the user's voice through a human-machine interface, and acquiring environmental or user images through a camera lens, and after learning personalized voice and image features, providing personalized voice dialogue services to users through a dialogue interface, and the method of adaptively acquiring image data for personalized voice dialogue is executed by one or more processors.

[0006] In a method for adaptively acquiring image data for personalized voice dialogue, the user device receives real-time voice data generated by the user and obtains the user's current voice features. A natural language processing model is then used to identify the semantics within these voice features. Next, after determining that an image needs to be acquired, the user device's image capture function is activated to obtain image data. Voice dialogue is then generated based on the correspondence between the acquired voice features and image features. By repeatedly receiving voice data generated by the user through the user device and obtaining voice features, and receiving image data generated by the user device and obtaining image features, voice dialogue based on voice and image features is continuously generated.

[0007] Furthermore, in addition to acquiring the semantic features from the speech features, the system also learns the user's current emotions based on the speech features, and generates speech dialogues based on semantics and emotions.

[0008] Furthermore, the voice features include pitch, rhythm, and intensity variations in the voice transmitted by the user through the user device.

[0009] Furthermore, by querying the emotion database, the tone, rhythm, and volume changes that should correspond to the user's current emotion can be obtained, thus generating personalized voice dialogues.

[0010] Furthermore, if the user's semantics indicate a topic involving environmental information, the image capture function on the user's device is activated to capture an image of the user's environment.

[0011] Among them, image features can be obtained from image data of the user's environment, environmental features can be judged, and further voice dialogue based on environmental features can be generated accordingly.

[0012] Alternatively, if the user's words indicate a topic involving the user, the selfie function on the user's device can be activated to capture an image of the user. By obtaining image features from the user's selfie image data, the user's current emotion or needs can be determined based on these features, and further voice dialogue can be generated accordingly.

[0013] Furthermore, deep learning is performed on a cloud server to learn the user's preferences and personality based on the voice and image data continuously received from the user's device in the past and present, and to build a user profile.

[0014] Furthermore, by obtaining the user's current voice features, semantics, and image features, the user feature file can be obtained based on the user's recognition information. Then, voice dialogue can be generated according to the semantics, the user feature file, and / or the current environmental information.

[0015] Furthermore, the cloud server can retrieve realistic objects generated by an artificial intelligence simulation model from a database, transmit image signals and model data to the user device, display the realistic object on the user device's display, and simulate the voice dialogue generated at that moment.

[0016] Furthermore, the simulated object can be a lifelike chatbot, and the cloud server can generate an interactive interface through a web server program to provide multiple options for lifelike chatbots.

[0017] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description

[0018] Figure 1 This is a schematic diagram illustrating a scenario where adaptive image data is acquired for personalized voice dialogue.

[0019] Figure 2 A schematic diagram illustrating a system architecture embodiment of a method for adaptively acquiring image data to perform personalized voice dialogue;

[0020] Figure 3 An embodiment of a method for adaptively acquiring image data to conduct personalized voice dialogue is shown in the figure;

[0021] Figure 4 A flowchart of a first embodiment of a method for adaptively acquiring image data to perform personalized voice dialogue is shown;

[0022] Figure 5 A flowchart of a second embodiment of a method for adaptively acquiring image data to perform personalized voice dialogue is shown;

[0023] Figure 6 One of the schematic diagrams showing a scenario where adaptive image data is acquired for personalized voice dialogue; and

[0024] Figure 7 This is the second illustration of a scenario where adaptive image data is acquired for personalized voice dialogue. Detailed Implementation

[0025] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated beforehand. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.

[0026] It should be understood that while terms such as "first," "second," and "third" may be used in this document to describe various components or signals, these components or signals should not be limited by these terms. These terms are primarily used to distinguish one component from another, or one signal from another. Furthermore, the term "or" as used herein should, as appropriate, include any combination of one or more of the related listed items.

[0027] In order to provide users with personalized voice dialogue services and to enable natural language dialogue using a lifelike chatbot, this publication proposes a method and system for adaptively acquiring image data for personalized voice dialogue.

[0028] For examples of using image data to provide personalized voice dialogue services, please refer to Figure 1 The illustrated diagram shows that the system providing personalized voice dialogue services includes a cloud server 20. Through the cloud server 20, user 10 can engage in natural language dialogue with a lifelike chatbot 101 generated by generative artificial intelligence (AI) through learning knowledge, human speech, behavior, and facial expressions. Specifically, the intelligent model running behind the lifelike chatbot 101 learns the pitch and speed of user 10's current conversation, adding image features. Based on these voice and image features, it judges the user's meaning and emotion, or considers the current environment, enabling the lifelike chatbot 101 to engage in dialogue with corresponding audio and speech speed, reflecting the user 10's current emotions and environmental information.

[0029] For example, when the intelligent model analyzes changes in pitch, rhythm, and volume in voice data and determines that user 10 is currently sad, the lifelike chatbot 101 will adjust its audio and speaking speed accordingly to reflect user 10's current sadness. Furthermore, the intelligent model also determines the user's state or the current environment based on acquired image features, further providing voice dialogue that reflects the current context.

[0030] The diagram illustrates user 10 operating an application on user device 100. For example, by opening a chat interface through the application and connecting to cloud server 20, user 10 can select to load a lifelike chatbot 101 from cloud server 20, including relevant model data and image signals of the lifelike chatbot 101, just like two people having a video call in the real world. The diagram shows a chat interface starting on the display of user device 100, and the camera 105 on user device 100 is activated at the same time to display images 103 of two users having a conversation and the lifelike chatbot 101.

[0031] When providing personalized dialogue services, user device 100 connects to cloud server 20, activates the audio (microphone) and video (camera) functions on user device 100, and performs real-time audio recording and image capture. Voice and image data can be uploaded to cloud server 20 in real time. The language model running on cloud server 20 acquires voice features, and the image model acquires image features. Current environmental information can also be added. After using an intelligent model to determine the user's current emotion, a voice dialogue is generated. The basis for determining emotion includes analyzing the pitch, rhythm, and intensity changes of each segment of voice data, as well as semantics, through an intelligent model. Simultaneously, the user's facial image can be acquired to determine facial expressions.

[0032] It is worth mentioning that, in one embodiment, when generating voice dialogue using a language model, the relevant audio data is transmitted to the user device 100 and can be transcribed and displayed on the dialogue interface. Furthermore, in addition to the voice dialogue being emitted directly from the speaker of the user device 100, the voice dialogue can also be simulated by a lifelike chatbot 101 displayed on the dialogue interface. The running artificial intelligence can generate facial expressions and emoticons to match the current emotion. For example, the generated lifelike character's mouth movements can simulate human mouth movements, and the facial muscle movements can also match the emotion expressed in the voice.

[0033] The cloud server 20 implements a method for adaptively acquiring image data for personalized voice dialogue, which mainly involves the collaborative operation of various artificial intelligence algorithms and models by software and hardware. (See reference...) Figure 2 The diagram shows a system architecture embodiment.

[0034] The cloud server 20 utilizes the processing circuitry, memory, and software of a computer system to implement various functional modules, such as a natural language processing module 201 for processing content from conversations with the user; an instruction processing module 203 for providing corresponding services based on user requests generated through the dialogue interface; and a user interface module 205 for generating an interface for conversations with the user, which interfaces with the user interface generated by a specific application running on the user device 100. The cloud server 20 also includes a machine learning module 207 for training a natural language model that meets user needs, a database module 209 for providing content from built-in or external databases, and an image processing module 211 for processing acquired image data.

[0035] The cloud server 20 has a built-in or external database 22, which includes user data 221 and multiple realistic objects generated by an artificial intelligence simulation model 223. It can also include an emotion database 225 derived from big data analysis. The emotion database 225 stores data such as audio, speech rate, and word choice for speech under various emotions, allowing the realistic objects to query and generate dialogue corresponding to specific emotions. The cloud server 20 provides personalized voice dialogue services via network 200. Users operate the user device 100 on their terminals and use the software programs running on it to access the services of the cloud server 20 via network 200.

[0036] According to the embodiment, the software program provided by the system for the user device 100 to run related services includes, according to its functions, an audio-visual dialogue processing module 111 for processing audio-visual dialogue interaction with simulated objects, a voice acquisition module 113 for receiving and acquiring user voice signals, an image acquisition module 115 for capturing images of the user or environment, and a communication module 117 for establishing a connection with an external system (such as a cloud server 20).

[0037] The database module 209 in the cloud server 20 is used to access and manage the built-in or external database 22, including maintaining user data 221 generated by users using personalized voice dialogue services, and also provides simulated objects generated by the artificial intelligence simulation model 223. The simulated objects can be humanoid, which may include a variety of simulated humanoid chatbots from different fields provided by multiple users, as well as a variety of simulated humanoid chatbots provided by the system.

[0038] According to the embodiment, the database 22 of the cloud server 20 provides realistic objects generated by the artificial intelligence simulation model 223, which can be used to generate multiple realistic humanoid chatbots (excluding realistic humanoids, animals, or various objects). In addition to the realistic humanoids having different appearances and accents, they can also be designed as various realistic humanoid chatbots with different professional knowledge, so as to provide users with personalized voice dialogue services according to their needs. It is worth mentioning that the method of adaptively acquiring image data for personalized voice dialogue proposed in the publication can use the realistic objects generated by the artificial intelligence simulation model 223 to perform natural language dialogue. It uses a generative artificial intelligence, which uses machine learning algorithms to learn from a large amount of human data to build a realistic human character generation model, and then uses generative artificial intelligence based on 3D drawing technology to generate images, which can generate specific realistic objects according to user needs (providing prompts).

[0039] Furthermore, to give the lifelike chatbot a professional background, this part is achieved through machine learning algorithms. By learning from data in specific domains, a domain-specific intelligent model is built. For example, a Retrieval Augmented Generation (RAG) technique can be used to implement a natural language model for a specific domain based on a large language model (LLM).

[0040] When a user operates the user device 100 to execute a software program, the user generates a dialogue request to the cloud server 20 through the software program. After receiving the user's request, the cloud server 20 processes various requests transmitted from the user device 100 through its instruction processing module 203, such as requests to execute a dialogue, requests to select a lifelike chatbot that performs personalized voice dialogue services, and various requests generated during the voice dialogue.

[0041] The cloud server 20 transmits various information, such as audio and video content, text content and image files, to the user device 100 through the user interface module 205 (including necessary encoding, decoding, compression and decompression programs). It also processes the dialogue content displayed by two or more parties through the dialog interface and uses it to generate various functions and graphics in the dialog interface.

[0042] In the cloud server 20, the machine learning module 207 runs machine learning algorithms, uses neural network deep learning technology to learn the speech features spoken by the user, implements a speech model by training a large amount of speech data, and establishes a natural language processing (NLP) model to obtain speech features in the speech data, derive the semantics, and generate response dialogue content.

[0043] Furthermore, the image processing module 211 processes image data transmitted from the user device 100, which can be an image of the user or an image of the environment where the user is located. This data becomes machine learning data in the cloud server 20 and can be used to calculate and learn the user's current situation. Therefore, the machine learning module 207 can run machine learning algorithms to learn image features. For example, by training on a large number of user facial image features, it can build an image model that can learn the user's expressions and emotions. Thus, in addition to judging the user's emotions based on voice features, the cloud server 20 can also use image models to process image features to assist in judging the user's emotions.

[0044] The cloud server 20 uses the natural language processing module 201 to process the voice data obtained from the user device 100. In addition to textualizing the voice content, it also runs the natural language processing model to obtain the voice features in the voice data and identify the semantics therein, as well as obtain the image features in the image data and identify the information involving the user or environment therein. After generating dialogue content, it responds to the user in natural language.

[0045] Furthermore, in the process of obtaining semantics and judging user emotions, the speech features used include changes in pitch, rhythm, and intensity. Speech models can identify a user's current emotion based on these speech features (such as sound wave waveforms). They can also learn about individual users' personalities and preferences over a long period, creating a user profile. Similarly, image models can learn about the relationship between user behavior and emotions based on pixel changes in image data, the relationships between objects in an image, and changes in the identification of specific objects (such as the user's facial features). This allows for accurate judgment of user emotions, combined with speech feature analysis, to create a feature file describing the user's characteristics.

[0046] By using a computer system to run machine learning algorithms to learn human language and classify it, and by deep learning language structure and the relationships between sentences, a natural language processing model is generated. In the method proposed in the published document, the natural language processing model can be trained using speech data generated from conversations between users to form a personalized language model. Furthermore, after training the natural language processing model to generate a personalized language model, deep learning algorithms can be used to learn features such as pitch, rhythm, and intensity changes in speech, thereby learning emotions in the language.

[0047] It is worth mentioning that the system proposed in the open book can adopt large language models (LLM) and can be a multimodal model, which can be used to process data of different data types (such as text, images and sound), and can integrate data of various types to build intelligent models that can predict user behavior.

[0048] Furthermore, the natural language processing module 201 does not simply utilize existing natural language search / dialogue services such as ChatGPT, which employ large language models. This is because existing ChatGPT can only respond to user dialogue in a one-way manner and fails to learn the user's emotions or reference environmental information. Instead, the natural language processing module 201 further utilizes user feature files, current semantics, and real-time environmental information (such as current geographical location, weather, and time) to enable the lifelike chatbot to engage in dialogue based on the user's current emotions, generating dialogues that meet the user's real-time needs. According to an embodiment, the lifelike chatbot can query the emotion database 225 to retrieve data such as audio, speech rate, and vocabulary corresponding to various emotions, thereby eliciting dialogues corresponding to specific emotions.

[0049] For example, when it is noon, the system obtains the user's geographical location (such as the user's device location information obtained through the application) and weather (cold, hot, sunny or rainy) through environmental information. The lifelike chatbot can proactively suggest suitable restaurants based on the user's preferences. The user can respond whether they want to eat, what their acceptable price range is, or other needs, so that the lifelike chatbot can adjust its suggestions in real time or conduct dialogue to confirm the needs.

[0050] According to an embodiment, in the machine learning algorithm that determines the current emotion based on the image features of the user's face, a supervised machine learning method can be used to learn the changes in image features labeled by humans under different emotions, such as changes in the distance ratio between facial features or specific parts, line changes, shape changes, and color changes, thereby training an image model for judging emotions based on image features. Thus, in the cloud server 20, the voice model and image model can be used simultaneously to determine the user's emotion. Then, the natural language processing module 201, in conjunction with the image processing module 211, obtains semantic and environmental information to generate dialogue content. The audio, speech rate, and word choice of the dialogue can be obtained by querying the emotion database 225.

[0051] Based on the above system architecture and implementation method for providing image-based personalized voice dialogue services, Figure 3 Next, an embodiment diagram of a method for adaptively acquiring image data to perform personalized voice dialogue is shown.

[0052] This example demonstrates that the user's terminal uses the human-machine interface 303 to convert the user's voice / image 301 into text content displayed on the dialogue interface 305 or to issue a voice message, while the cloud server receives the text content or voice data. The user also operates an application running on the user's device to request a personalized voice dialogue service from the cloud server, which then provides the user with a personalized voice dialogue service through the dialogue interface 305.

[0053] A cloud server executes a method for adaptively acquiring image data for personalized voice dialogue through one or more processors of its computer system, acquiring user voice / image 301 generated from the user device through a dialogue interface 305. According to an embodiment, the system utilizes various sensors in the user device to acquire the user's voice and image data. For example, after the user device initiates the dialogue interface 305, it simultaneously activates a microphone to receive the user's voice, activates a camera to capture an image of the user's face, and generates speech (or inputs text) through the human-machine interface 303. The user can interact with realistic objects displayed in the dialogue interface 305 and engage in dialogue using natural language.

[0054] Voice and image data are transmitted from the user device to the cloud server. The cloud server runs a voice / image model 307, which can be divided into a voice model and an image model. For voice data, the voice model transcribes the voice data to obtain its semantics and determine the user's current emotion. For image data, the image model analyzes the image data to obtain the user's facial features. For example, it can analyze the color and lighting changes of pixels corresponding to the user's facial organs in the image to obtain changes in the images of facial features or specific parts such as the mouth and eyes. If the pixels of each part reflect the geometric relationship between facial features or specific parts, combined with semantics and voice features, the user's current emotion can be determined.

[0055] Furthermore, the voice / image model 307 can also obtain current environmental information 313 from an external database (or server). According to the embodiment, after the cloud server connects to the user device, it can obtain the location information in the user device and determine the user's geographical location. Therefore, it can obtain environmental information 313 based on the user's location and time from an external database, such as local weather, business information of surrounding stores, geographical environment, and holidays, to provide effective content to the user. When the cloud server runs the voice / image model 307, it also queries the user database 315 to obtain user feature files, from which basic user data, learned preferences, historical records, etc. can be obtained. Furthermore, the method proposed in the publication can adaptively obtain image data, such as obtaining an environmental image 317 of the user's location through the user device's camera. After the voice / image model 307 processes the environmental image 317, it determines the user's real-time state, such as the user's clothing, people or objects around them, etc., which become reference information when personalized voice dialogue is generated. Among them, the cloud server also stores and updates each user's data according to the time dimension through the user database 315, recording each user's historical dialogue records.

[0056] In this way, the cloud server can effectively provide personalized voice dialogue services to users through the voice / image model 307. Furthermore, when users engage in personalized voice dialogue, the cloud server continuously acquires the user's voice / image features 311 and continuously learns the voice and image data generated in the user's dialogue through the feature learning algorithm 309 to update or optimize the voice / image model 307.

[0057] Based on the above system architecture and system operation implementation examples, Figure 4 Next, a flowchart of an embodiment of a method for adaptively acquiring image data to perform personalized voice dialogue is shown.

[0058] The method runs on a cloud server. The user operates an application running on their device to establish a connection with the cloud server and initiates a dialogue interface through the application. The user device can acquire real-time voice data generated by the user's speech through sensing components such as a microphone (or with the addition of a camera), or add image data (step S401). Then, it learns the voice features, or adds image features. The voice features may include changes in tone, rhythm, and volume of the voice transmitted by the user through the user device, while the image features may include the user themselves or an image of the surrounding environment (step S403). Thus, the semantics can be recognized based on the voice features. When receiving the real-time voice data generated by the user, according to the embodiment, the user's real-time image data is also received, and the user's current image features are obtained. Therefore, the user's current emotion, or the current environmental status, can be learned simultaneously based on the voice features (including semantics) and image features (step S405).

[0059] According to the embodiment, while obtaining voice features, image features can also be obtained based on image data. Then, the formed voice and / or image features 41 are transmitted to the image generation model 43 of the cloud server to generate realistic images, such as the facial expressions, mouth and specific parts of the image of the realistic humanoid chatbot. They are also transmitted to the personalized natural language model 45 to generate voice dialogue with emotions derived from features 41.

[0060] In step S405, which determines the user's current emotion, the cloud server uses a natural language processing model to identify the semantics in the speech features and can add the user behavior obtained from the image features to determine the current environmental information. Based on the semantics and the current environmental information (which may include environmental information determined from the image of the user's location and environmental information provided by the external system), a voice dialogue based on the user's current emotion is generated, which is also a voice dialogue that responds to the user's real-time voice.

[0061] Furthermore, the cloud server uses image generation model 43 to generate images of realistic objects based on voice features or by adding image features. This can also include realistic objects available for user selection on the cloud server (step S407). One of these realistic objects is displayed on the user's device's screen in a dialogue interface (step S409). A personalized natural language model 45 then generates dialogue content based on the user's semantics, current emotion, or environmental information to conduct a voice dialogue with the user (step S411). For example, a realistic object could be a lifelike chatbot, whose images primarily reflect current emotions, such as facial expressions, mouth, and changes in specific body parts.

[0062] During the dialogue, the natural language model 45 in the cloud server continues to run and determines whether to obtain the image based on the semantics formed in the dialogue (step S413). If there is no intention to obtain the image (no), the process can return to step S401 and repeat the above steps.

[0063] Conversely, if the semantics determine that an image is needed (yes), the cloud server uses its user interface modules (such as...) Figure 2 205) The image capture function in the user device is activated through the human-machine interface. For example, the main camera or the self-portrait camera can be activated as needed (step S415), and image data is captured (step S417).

[0064] During this process, the cloud server continuously acquires voice and image features, processes them through natural language, and generates corresponding voice dialogues through a lifelike chatbot. This process is repeated, such as receiving voice data generated by the user through their device, adding image data to acquire voice features, or adding image features to learn the user's current emotions, to continuously generate personalized voice dialogues.

[0065] According to an embodiment, in the above-described process step S415, taking a user holding a mobile device as an example, when the system determines based on the user's semantics that the topic involves environmental information, it can activate the main camera in the user's mobile device through a software program executed within the user's device to capture an image of the user's environment. Therefore, the system can obtain image data of the user's environment, and through image processing, obtain image features. The natural language model determines environmental features based on the image features to generate further voice dialogue based on the environmental features, and can display the corresponding images.

[0066] According to Embodiment 2, in the above-mentioned process step S415, when the user's semantics determine that the topic involves the user, the selfie camera in the user's device can be activated to capture an image of the user. Similarly, the system obtains image features from the image data of the user's selfie image, determines the user's current emotion or need based on the image features, generates further voice dialogue based on the emotion or need, and can display the corresponding image.

[0067] Related embodiments can be referred to. Figure 5 A flowchart illustrating another embodiment of the method for adaptively acquiring image data for personalized voice dialogue.

[0068] When the system receives voice data transmitted by the user device through the dialogue interface (step S501), it similarly analyzes the voice features using a natural language model to determine the meaning and emotion (step S503) and conducts a corresponding voice dialogue (step S505).

[0069] Furthermore, the system determines the topic based on the user's semantics and whether it is necessary to obtain environmental information (step S507). If not, the process steps repeatedly perform voice dialogue service. Conversely, if it is determined that environmental information is needed, the system can activate the image capture function on the user device by executing the software program on the user device and start the camera (either the main lens or the front lens can capture environmental images) (step S509) to capture images of the user's environment.

[0070] Simultaneously, the system continuously acquires image data (step S511), obtains image features from the image data of the user's environment, and determines environmental features (step S513). According to the embodiment, through an image model, or including a natural language model (multimodal), a display image can be generated based on the acquired image features and displayed on the dialogue interface (step S515), while further voice dialogue based on environmental features can still be generated accordingly (step S517).

[0071] An embodiment of using the adaptive acquisition of image data for personalized voice dialogue can be found in [reference needed]. Figure 6 The scene diagram shown.

[0072] The relevant user terminal device is shown in the attached figure as a scene 1 6, which is a realistic character display billboard 60. The realistic character display billboard 60 is implemented by a computer system and can have a built-in natural language model or connect to a cloud server through the network to obtain adaptive image data for personalized voice dialogue services.

[0073] The scene shows a user 600 entering scene 1 6 (such as the restaurant entrance) and having a conversation with a realistic character 603 displayed on the realistic character display billboard 60. The conversation can involve introducing the restaurant, dishes, ordering food and making reservations. During the conversation, the built-in (or external) camera 601 of the realistic character display billboard 60 is activated to capture the user 600 and include the user's background image.

[0074] User 600 can converse with the virtual character 603. The virtual character displays the billboard 60 or the cloud server continuously obtains the user 600's voice and image data, analyzes it, and generates dialogue content and displays images according to the user's needs.

[0075] According to the attached diagram, the simulated character display billboard 60 is implemented as an information service kiosk providing customer service. When the natural language model running within it receives the user's voice data, it can deduce the meaning and reply to the user in natural language. During the dialogue, if it is determined that the user's meaning is about obtaining environmental information, the application running on the user's device will actively activate the camera to capture an image of the environment; if it is determined that the user's meaning is about a topic concerning the user themselves, the application can activate the camera to capture an image of the user.

[0076] The attached example shows a lifelike character display billboard 60 set up in a restaurant as an information service machine. It is equipped with a camera and a monitor. The system running in the machine provides an intelligent customer service robot that meets the needs of the restaurant. The lifelike character displayed on the monitor converses with the user (customer). When the camera needs to be activated based on the semantics, it will obtain the user's image and provide corresponding content, such as ordering food or meal suggestions.

[0077] Figure 7 This is an illustration of another scenario where adaptive image data is acquired for personalized voice dialogue.

[0078] According to the attached diagram, user 700 enters scene 2 7 and has a conversation with the realistic character 703 displayed on the realistic character display billboard 70. During the conversation, the camera 701 on the realistic character display billboard 70 is activated to capture the user 700, obtain the user 700's image data, and analyze it to obtain image features. Based on the voice and image features, dialogue content and images can be generated and displayed.

[0079] For example, in scenario 7, such as a clothing store, when user 700 is trying on clothes, they can ask the model on the virtual character display billboard 70 for clothing suggestions through natural language dialogue. While trying on clothes, the application can actively activate the camera 701 of the virtual character display billboard 70 to capture images of the user's full body or parts, and provide clothing suggestions based on the analyzed user needs. The virtual character 703 can also introduce the user and actually display the clothing outfit, or generate a simulated image of user 700's outfit by obtaining image data of clothing items and the user's image and synthesizing them, allowing the user to easily take pictures of themselves trying on clothes.

[0080] In summary, the publication proposes a method and system for adaptively acquiring image data to conduct personalized voice dialogue. The system establishes an intelligent robot that interacts with the user based on the above embodiments, especially a realistic character drawn from a realistic character model. It can determine whether to activate the shooting function on the user's device according to needs, and can analyze user needs based on real-time acquired voice and image features to provide personalized voice dialogue services.

[0081] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the claims of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of the claims of the present invention.

Claims

1. A method for adapting image data to conduct a personalized voice conversation, running in a cloud server, characterized in that The method comprises: receiving voice data generated by a user in real time through a user device, obtaining voice features of the user at the moment, and identifying the semantic meaning in the voice features through a natural language processing model; after determining that image data needs to be obtained, starting the image shooting function of the user device to obtain image data, and generating a voice dialogue based on the voice features and the image features; repeatedly receiving voice data generated by the user through the user device, obtaining voice features, and receiving image data generated by the user device, obtaining image features, and continuously generating the voice dialogue based on the voice features and the image features.

2. The method of claim 1, wherein the image data is adaptively acquired for the personalized voice conversation, and In addition to obtaining the semantic meaning in the voice features, the user's emotion at the moment is learned based on the voice features, and the voice dialogue is generated based on the semantic meaning and the emotion.

3. The method of claim 2, wherein the image data is adaptively acquired for personalizing the voice dialogue. The voice features include changes in tone, rhythm, and sound intensity in the voice transmitted by the user through the user device.

4. The method of claim 3, wherein the image data is adaptively acquired for the personalized voice dialogue. The tone, rhythm, and sound intensity changes corresponding to the user's emotion at the moment are obtained by querying an emotion database, and the personalized voice dialogue is generated.

5. The method of claim 1, wherein the image data is adaptively acquired for the personalized voice conversation, and If the user's semantic meaning is determined to be related to environmental information, the image shooting function of the user device is started to shoot images of the environment in which the user is located.

6. The method of claim 5, wherein the image data is adaptively acquired for the personalized voice conversation. Image features are obtained from image data of the environment in which the user is located, and environmental features are determined to correspondingly generate further voice dialogue based on the environmental features.

7. The method of claim 1, wherein the image data is adaptively acquired for the personalized voice conversation. If the user's semantic meaning is determined to be related to the user himself / herself, the selfie image shooting function of the user device is started to shoot images of the user himself / herself.

8. The method of claim 7, wherein the image data is adaptively acquired for the personalized voice dialogue. Image features are obtained from image data of the user's selfie images, and the user's emotion or demand at the moment is determined based on the image features to correspondingly generate further voice dialogue based on the emotion or the demand.

9. The method of adaptively taking image data for a personalized voice conversation of any one of claims 1 to 8, wherein, In the cloud server, deep learning is performed to learn the user's preferences and personality based on voice data and image data continuously received from the user device in the past and at the moment, and a user feature file is established.

10. The method of claim 9, wherein the image data is adaptively acquired for the personalized voice conversation. In the cloud server, the user feature file is obtained based on the user's identification information after obtaining the voice features, semantic meaning, and image features of the user at the moment, and the voice dialogue is correspondingly generated based on the semantic meaning, the user feature file, and / or the environmental information at the moment.

11. The method of adaptively taking image data for a personalized voice conversation of any one of claims 1 to 8, wherein The cloud server obtains the artificial object generated by an artificial intelligence simulation model from a database, transmits image signals and model data to the user device, displays the artificial object on a display of the user device, and simulates the voice dialogue.

12. The method of claim 11, wherein the image data is adaptively acquired for personalizing the voice dialogue. The artificial object is an artificial human-shaped chat robot, and the cloud server generates an interactive interface through a web server program to provide options for multiple artificial human-shaped chat robots.

13. A system for running a method of adapting to take image data for a personalized voice dialogue, characterized by The system comprises: a cloud server connected to a user device, providing personalized voice dialogue services for users through a dialogue interface, and executing the method of adaptively obtaining image data for personalized voice dialogue by one or more processors, comprising: receiving voice data generated by a user in real time through a user device, obtaining voice features of the user at the moment, and identifying the meaning of the voice features through a natural language processing model; judging that image data needs to be obtained, starting the image shooting function of the user device to obtain image data, and generating a voice dialogue based on the voice features and image features according to the voice features and the image features; and repeatedly receiving voice data generated by the user through the user device, obtaining voice features, and receiving image data generated by the user device, obtaining image features, and continuously generating the voice dialogue based on the voice features and the image features.

14. The system of claim 13, wherein, The voice features corresponding to the user's current mood are obtained by querying an emotion database, the tone, rhythm change and sound intensity change of the user's speech are generated, and the personalized voice dialogue is generated; in addition to obtaining the meaning of the voice features, the user's current mood is learned according to the voice features, and the voice dialogue is generated according to the meaning and the mood.

15. The system of claim 13, wherein, In the method of adaptively obtaining image data to generate personalized voice dialogue, if the user's meaning is judged to be related to environmental information, the image shooting function of the user device is started to shoot the image of the user's environment, and image features are obtained from the image data of the user's environment to judge the environmental features, so as to generate further voice dialogue based on the environmental features.

16. The system of claim 13, wherein, In the method of adaptively obtaining image data to generate personalized voice dialogue, if the user's meaning is judged to be related to the user himself, the self-shooting function of the user device is started to shoot the image of the user himself.

17. The system of claim 16, wherein, Image features are obtained from the image data of the user's self-shooting image, the user's current mood or demand is judged according to the image features, and further voice dialogue based on the mood or the demand is generated.

18. The system of any one of claims 13 to 17, wherein, In the cloud server, deep learning is performed to learn the user's preferences and personality according to the voice data and image data continuously received from the user device in the past and at the moment, and a user feature file is established.

19. The system of claim 18, wherein, In the cloud server, the user feature file is obtained according to the user's identification information after obtaining the voice features, meaning and image features of the user at the moment, and the voice dialogue is generated according to the meaning, the user feature file, and / or the current environmental information.

20. The system of any one of claims 13 to 17, wherein The cloud server obtains the artificial object generated by an artificial intelligence simulation model from a database, the artificial object is an artificial human-shaped chat robot, the cloud server generates an interactive interface through a web server program, and provides options of multiple artificial human-shaped chat robots; image signals and model data are transmitted to the user device, the artificial object is displayed on the display of the user device, and the voice dialogue is simulated to be sent out.