A method and system for adaptively acquiring video data and performing personalized voice dialogue.
Patent Information
- Application Number
- JP2025266592
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-05
- Filing Date
- 2025-12-19
- Publication Date
- 2026-09-17
Smart Images

Figure 2026148438000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to personalized smart voice dialogue technology, and in particular to a method and system for adaptively acquiring video data and implementing personalized voice dialogue by utilizing real-time video data and personalized data obtained by a language model.
Background Art
[0002] Currently, among artificial intelligence (AI), which is developing rapidly in various fields, one type thereof is a natural language chatbot capable of processing natural language and automatically generating content. For example, there is Chat Generative Pre-trained Transformer developed by OpenAI, which is referred to as ChatGPT in English. Such natural language chatbots use generative artificial intelligence technology to learn a large amount of data, then generate new data related to the original data, and build a smart model through deep learning (e.g., generative adversarial networks, GAN).
[0003] Taking ChatGPT as an example, ChatGPT can respond to users in a natural language manner by learning and training from a large amount of network information, but the general content it provides to users in responses is standard answers based on learned content, and it cannot provide responses related to the user's current state in real time. Although it is called a natural language chatbot, it lacks content related to the user and adapted to real-time situations. That is, the dialogue service provided by conventional chatbots is general, and not suitable for personalized voice dialogue services (e.g., personal needs, preferences, backgrounds, etc.).
Summary of the Invention
Problem to be Solved by the Invention
[0004] In order to provide a voice dialogue model that can learn individual preferences while being based on individual needs, the present invention provides a method and system for adaptively acquiring video data and performing personalized voice dialogue. [Means for solving the problem]
[0005] According to an embodiment of a system that performs a method for providing personalized voice dialogue services to a user, the system includes a cloud server, the cloud server is connected to a user device, and the user device acquires the voice emitted by the user via a human-machine interface, and at the same time acquires images of the environment or the user by a camera lens, and after learning the characteristics of the personalized voice and images, provides personalized voice dialogue services to the user via an interactive interface, and performs a method of performing personalized voice dialogue by adaptively acquiring the image data with one or more processors.
[0006] In a method for adaptively acquiring video data and conducting personalized voice dialogue, the user device receives voice data generated by the user's real-time voice, acquires the user's current voice characteristics, and identifies the meaning of words in the voice characteristics using a natural language processing model. Subsequently, when it is determined that video acquisition is necessary, the video recording function on the user device is activated to acquire video data, and then voice dialogue is generated correspondingly based on the acquired voice characteristics and video characteristics. The process of receiving voice data generated by the user through the user device and acquiring voice characteristics, and receiving video data generated by the user device and acquiring video characteristics is repeated to continuously generate voice dialogue based on the voice characteristics and video characteristics.
[0007] Furthermore, in addition to acquiring the meaning of words based on their speech characteristics, the system also learns the user's current emotions based on those characteristics, and simultaneously generates speech dialogues based on both the meaning and emotion.
[0008] Furthermore, the characteristics of the voice include the tone, rhythm, and intensity of the voice transmitted by the user via the user device.
[0009] Furthermore, by searching the emotion database, it is possible to generate personalized voice dialogues that correspond to changes in tone, rhythm, and voice intensity that should reflect the user's current emotions.
[0010] Furthermore, based on the user's meaning, if it is determined that the topic relates to environmental information, the video recording function on the user's device is activated and used to record video of the environment in which the user is located.
[0011] Here, features of the video are obtained from video data of the user's environment, the characteristics of the environment are determined, and further voice dialogue is generated in response to those characteristics.
[0012] Alternatively, if it is determined, based on the user's meaning, that the topic concerns the user, the selfie function on the user's device can be activated to record the user's own video. Features of the video can be obtained from the video data of the user's selfie video, and based on these features, the user's current emotions or needs can be determined, and further voice dialogue can be generated accordingly based on those emotions or needs.
[0013] Furthermore, by performing deep learning on a cloud server, a user profile is constructed based on the user's preferences and personality learned from audio and video data that the user's device has continuously received from the past to the present.
[0014] Furthermore, by acquiring the user's current voice characteristics, semantics, and features with accompanying video, and by acquiring the user profile based on the user's identification information, voice dialogue can be generated in a corresponding manner based on the semantics, user profile, and / or current environmental information.
[0015] Furthermore, the cloud server can retrieve a simulated object generated by an artificial intelligence pseudo-model from the database and send the video signal and model data to the user device, allowing this simulated object to be displayed on the user device's screen and enabling simulated voice interaction.
[0016] Furthermore, the pseudo-object may be a pseudo-humanoid chatbot, and a cloud server can generate an interactive interface via a web server program, after which multiple pseudo-humanoid chatbot options can be provided. [Brief explanation of the drawing]
[0017] [Figure 1] This is a schematic diagram illustrating a situation in which video data is acquired adaptively and personalized voice dialogue is performed. [Figure 2] This is a schematic diagram showing an embodiment of a system configuration that performs a method of adaptively acquiring video data and conducting personalized voice dialogue. [Figure 3] This figure shows an example of a method that adaptively acquires video data and performs personalized voice dialogue. [Figure 4] This is a first flowchart illustrating an embodiment of a method for adaptively acquiring video data and performing personalized voice dialogue. [Figure 5] This is a second flowchart illustrating an embodiment of a method for adaptively acquiring video data and performing personalized voice dialogue. [Figure 6] This is the first schematic diagram illustrating a scenario in which video data is adaptively acquired and personalized voice dialogue is performed. [Figure 7] This is a second schematic diagram illustrating a scenario in which video data is adaptively acquired and personalized voice dialogue is performed. [Modes for carrying out the invention]
[0018] In order to enable the provision of individual voice dialogue services to users and conduct natural language dialogue by utilizing a pseudo-human chatbot, the present invention provides a method and system for adaptively acquiring video data and conducting personalized voice dialogue.
[0019] For the scenario of providing personalized voice dialogue services to users by utilizing video data, reference can be made to the schematic diagram shown in FIG. 1. A system for providing personalized voice dialogue services is provided with a cloud server 20, and through the cloud server 20, a service that allows the user 10 to conduct natural language dialogue with the pseudo-human chatbot 101 generated by generative artificial intelligence (Artificial Intelligence, AI) through learning features such as knowledge, human conversation, behavior, and facial expressions can be provided. In particular, the smart model operating behind the pseudo-human chatbot 101 learns the level of the voice frequency and the speaking speed of the current dialogue of the user 10, further adds video features, judges the user's semantic meaning and emotion, or additionally judges the current environment based on the voice features and video features, enables the pseudo-human chatbot 101 to conduct dialogue with a corresponding voice frequency and speaking speed, and allows it to reflect the current user 10's emotion and environment information in a timely manner.
[0020] For example, when the smart model analyzes changes in tone, rhythm and voice intensity in voice data and determines that the user 10 is currently sad, the pseudo-human chatbot 101 adjusts the voice frequency and speaking speed in the conversation to correspond to this sad emotion, and reflects the user 10's current sad emotion in a timely manner. Furthermore, the smart model can determine the user's state or the current environmental state based on the acquired video features, and further provide a voice dialogue that reflects the current situation.
[0021] The diagram schematically shows a scenario in which user 10 operates an application program on user device 100. For example, after launching an interactive interface via the application program and connecting to cloud server 20, the user can selectively load the pseudo-humanoid chatbot 101, including related model data and the video signal of the pseudo-humanoid chatbot 101, from the cloud server 20, making it appear as if two people are having a video call in the real world. The diagram also shows the launch of an interactive interface on the display of user device 100, simultaneously activating the camera lens 105 on user device 100, showing the video 103 of the two users currently having a private conversation with the pseudo-humanoid chatbot 101.
[0022] When providing personalized conversational services, the user device 100 connects to the cloud server 20, activates its audio (microphone) and video (camera) functions, and performs real-time sound collection and video recording. The audio and video data are then uploaded to the cloud server 20 in real time. The language model running on the cloud server 20 captures the characteristics of the voice, the video model captures the characteristics of the video, and in addition, captures current environmental information. The smart model is then used to determine the user's current emotions before generating a voice conversation. The basis for determining emotions includes changes in voice tone, rhythm, and intensity for each segment of the audio data, as well as word meaning, all analyzed by the smart model. Simultaneously, video of the user's face is acquired, and facial expressions can be determined.
[0023] It should be particularly noted here that, when generating voice dialogue by means of a language model, according to an embodiment, the relevant audio data can be sent to the user device 100, transcribed into text, and then displayed on the interactive interface. Furthermore, the voice dialogue may be not only the voice directly output from the speaker of the user device 100, but also a voice dialogue pseudo-generated by the pseudo-humanoid chatbot 101 displayed on the interactive interface. The artificial intelligence operating herein can generate vocal expressions and facial expressions based on current emotions. For example, the generated movement of the mouth of the pseudo-human character can mimic the mouth movement of human conversation, and the movement of the muscles of the five facial features can further be adapted to the emotion indicated by the current voice.
[0024] The cloud server 20 implements a method for adaptively acquiring video data and conducting personalized voice dialogue, wherein various artificial intelligence algorithms and models are operated mainly through the collaboration of software and hardware (refer to the schematic diagram of the embodiment related to the system configuration shown in Figure 2).
[0025] The cloud server 20 implements various functional modules through the collaboration of processing circuits, memory and software in a computer system, for example: a natural language processing module 201 configured to process dialogue content with a user; an instruction processing module 203 configured to provide corresponding services based on needs generated by the user through the interactive interface; and a user interface module 205 configured to generate a dialogue interface with the user and connect to the user interface generated by a specific application program executed on the user device 100. The cloud server 20 may further include: a machine learning module 207 that can train a natural language model meeting the user's needs according to the user's needs; a database module 209 configured to provide content from a built-in or external database; and a video processing module 211 configured to process acquired video data.
[0026] The database 22, either built into or external to the cloud server 20, includes user data 221 and multiple pseudo-objects generated by the artificial intelligence pseudo-model 223. It may also include an emotion database 225 obtained through big data analysis. The emotion database 225 stores data such as voice frequencies, speaking speed, and terminology that should appear in conversations for each emotion, and the pseudo-objects are made to search for this data to perform dialogues corresponding to specific emotions. The cloud server 20 provides a personalized voice dialogue service via the network 200, and users access the cloud server 20's services via the network 200 by operating a user device 100 at a terminal and utilizing software programs running within it.
[0027] According to the embodiment, the software that causes the user device 100 to perform related services includes, depending on the function, a video call processing module 111 that handles video call interactions with a simulated object, an audio acquisition module 113 that collects and acquires the user's voice signal, an audio capture module 115 that captures and acquires images of the user or the environment, and a communication module 117 that establishes a connection with an external system (e.g., a cloud server 20).
[0028] The database module 209 in the cloud server 20 is for accessing and managing an internal or external database 22, and includes maintenance of user data 221 generated by the user using a personalized voice dialogue service, and further provides pseudo-objects generated by an artificial intelligence pseudo-model 223, the pseudo-objects may be humanoid and may include a variety of pseudo-humanoid chatbots from different fields provided by multiple users, and multiple types of pseudo-humanoid chatbots provided by the system.
[0029] According to the embodiment, the database 22 of the cloud server 20 provides pseudo-objects generated by the artificial intelligence pseudo-model 223, which can be used to generate multiple pseudo-humanoid chatbots (without excluding pseudo-people, animals, or various objects). The pseudo-people can have different appearances and accents, and by designing them to be multiple types of pseudo-humanoid chatbots with different expertise, personalized voice dialogue services can be provided to users as needed. It is worth noting that the adaptive method of acquiring video data and performing personalized voice dialogue provided by the present invention can perform natural language dialogue by utilizing pseudo-objects generated by the artificial intelligence pseudo-model 223. Here, generative artificial intelligence is utilized, and machine learning algorithms are used to learn a large amount of human data and build a pseudo-person generation model. At the same time, generative artificial intelligence based on 3D rendering technology is applied to generate video, and specific pseudo-objects can be generated according to the user's needs (by providing prompts).
[0030] Furthermore, in order to give the pseudo-humanoid chatbot a specialized background, this part uses machine learning algorithms to learn from data in specific fields, thereby building smart models for those fields. For example, Retrieval Augmented Generation (RAG) technology is employed to realize a natural language model that provides support for specific fields, based on a Large-Scale Language Model (LLM).
[0031] After the user operates the user device 100 and executes a software program, the software program generates a dialogue request to the cloud server 20. After the cloud server 20 receives the user's request, its instruction processing module 203 processes the various requests sent from the user device 100 (for example, a request to perform a dialogue, a request to select a pseudo-human chatbot that performs a personalized voice dialogue service, and various requests generated during the voice dialogue).
[0032] The cloud server 20 can transmit various messages (e.g., video content, text content, and video files, etc.) (including necessary encoding and decoding, compression, and decompression processes) via the user interface module 205 and the user device 100, and can also process two or more interactive contents through the interactive interface to generate various functions and patterns in the interactive interface.
[0033] On the cloud server 20, the machine learning module 207 executes machine learning algorithms, utilizes neural network deep learning techniques to learn the characteristics of the user's speech, and trains on a large amount of speech data to realize a speech model. This is used to build a natural language processing (NLP) model, acquire speech characteristics from the speech data, and generate dialogue content that responds after confirming the meaning of the words.
[0034] Furthermore, the video processing module 211 can process video data transmitted from the user device 100 (which may be video of the user themselves or video of the environment where the user is located, both of which become machine learning data on the cloud server 20) and use computational learning to acquire the user's current situation. Accordingly, the machine learning module 207 can learn video features by utilizing machine learning algorithms, and for example, by training with the features of a large amount of video of the user's face, it is possible to build a video model that can learn the user's facial expressions and emotions. In this way, not only can the user's emotions be judged by the features of their voice, but the cloud server 20 can also use the video model to process the features of the video and use that as a means to judge the user's emotions.
[0035] The cloud server 20 utilizes the natural language processing module 201 to process audio data acquired from the user device 100 and transcribe the audio content into text. In addition, it further utilizes the natural language processing model to acquire audio features in the audio data, identify the meaning of the words within it, acquire video features in the video data, identify information about the user or environment within it, generate dialogue content, and then respond to the user in natural language.
[0036] Furthermore, in the process of acquiring word meaning and determining the user's emotions, the speech features utilized include changes in tone, rhythm, and intensity. Based on these speech features (e.g., sound wave waveform features), the speech model can identify the user's current emotions. Moreover, through long-term learning, it can learn and acquire the individuality and preferences of each user to construct a user profile. Similarly, a corresponding video model can learn the relationship between the user's actions and emotions over a long period of time in response to changes in pixels in video data, the relationships between objects in the video, and changes in specific objects identified in the video (e.g., the user's five senses). This allows for a more accurate determination of the user's emotions and, combined with the judgment of speech features, can construct a profile that describes the user's characteristics.
[0037] In the method provided by the present invention, a natural language processing model can be generated by executing machine learning algorithms on a computer system to learn human language and classification, and by performing deep learning on the relationships between language structures and words. In this method, the natural language processing model can be trained with speech data generated by each user's dialogue to form a personalized language model. Furthermore, after the trained natural language processing model generates a personalized language model, the deep learning algorithm can be further utilized to learn features such as changes in tone, rhythm, and intensity in speech, thereby learning emotions in language.
[0038] It is worth noting that the system provided by the present invention can employ large language models (LLMs) and may also be multimodal models, and can be used to process data in different data formats (e.g., text, video, and audio), enabling the construction of a smart model that can integrate various data formats to predict user behavior.
[0039] Furthermore, the natural language processing module 201 is not simply a natural language search / dialogue service such as ChatGPT that easily utilizes conventional large-scale language models. Currently, ChatGPT can only respond unilaterally to user dialogue and cannot learn user emotions or refer to environmental information. Thus, the natural language processing module 201 further utilizes user profiles, current word meanings, and real-time environmental information (e.g., current geographical location, weather, time) to enable the pseudo-humanoid chatbot to engage in dialogue based on the user's current emotions, thereby generating dialogue that matches the user's real-time needs. According to the embodiment, the pseudo-humanoid chatbot can search the emotion database 225 for data such as speech frequencies, speaking speed, and terminology that should appear in conversations corresponding to various emotions, and engage in dialogue corresponding to a specific emotion.
[0040] For example, when it is noon, the system can obtain the user's geographical location (e.g., location information from the user's device through an application program) and weather (cold, hot, sunny, or rainy) through environmental information, and the virtual human-like chatbot can proactively recommend a suitable restaurant based on the user's preferences. By responding to whether the user wants to eat, what price range they are willing to pay, or other needs, the virtual human-like chatbot can adjust its recommendations in real time or continue the conversation to confirm the user's needs.
[0041] According to the embodiment, in a machine learning algorithm that determines the current emotion based on the features of the user's face in the video, supervised machine learning is employed to train a video model that determines emotions based on video features by learning changes in video features for different emotions of artificial marks (for example, changes in the ratio of the distance between the five senses or specific parts of the face, changes in lines, changes in shape, and changes in color). Thus, on the cloud server 20, the voice model and video model are used simultaneously to determine the user's emotion, and the natural language processing module 201 generates dialogue content by combining the semantic information and environmental information acquired by the video processing module 211. In this case, the voice frequency, speaking speed, and word choice in the dialogue are obtained by searching the emotion database 225.
[0042] Based on the configuration and implementation method of the system that provides personalized voice dialogue services to users based on the above video, Figure 3 shows an example of a method that adaptively acquires video data and performs personalized voice dialogue.
[0043] In this example, the user utilizes the human-machine interface 303 to convert their voice / video 301 to form text content displayed on the interactive interface 305, or they emit voice, and the cloud server receives the text content or voice data. The user further operates an application program running on their device to request a personalized voice interaction service from the cloud server, and the cloud server provides the personalized voice interaction service to the user via the interactive interface 305.
[0044] The cloud server performs a method of adaptively acquiring video data through one or more processors in its computer system to perform personalized voice dialogue, and acquires the user's voice / video 301 generated from the user device via the interactive interface 305. According to the embodiment, the system utilizes various sensors in the user device to acquire the user's voice data and video data. For example, after the user device activates the interactive interface 305, the microphone is activated to receive the user's voice, the camera is activated to capture a video of the user's face, and the user's voice is used by the human-machine interface 303 to generate speech (or input text), allowing interaction with a simulated object displayed in the interactive interface 305 or dialogue in natural language.
[0045] The user device sends audio and video data to a cloud server, and the cloud server operates the audio / video model 307, which can then be separated into an audio model and a video model. For audio data, the audio model transcribes the audio data, obtains its meaning, and determines the user's current emotion. For video data, the video model analyzes the video data to obtain features of the user's face. For example, it can analyze changes in the color and brightness of pixels corresponding to the organs of the user's face in the video to obtain changes in the video of the five senses or specific parts, such as the mouth and eyes. For example, the pixels of each part reflect the geometric relationships between the five senses or specific parts, and together with the meaning and audio features, it can determine the user's current emotion.
[0046] Furthermore, the voice / video model 307 can also acquire current environmental information 313 from an external database (or server). According to the embodiment, after the cloud server connects to the user device, it can acquire location information on the user device and determine the geographical location of the user. This helps in providing the user with useful content by acquiring environmental information 313 (e.g., local weather, business hours of nearby shops, geographical environment, seasonal activities, etc.) based on the user's location and time from an external database. When the cloud server operates the voice / video model 307, it further searches the user database 315 to acquire the user profile and obtains the user's basic data, preferences acquired through learning, past records, etc. Furthermore, the method provided by the present invention adaptively acquires video data (e.g., environmental video 317 of the user's location acquired by the user device's camera), processes the environmental video 317 with the voice / video model 307, and then determines the user's real-time state (e.g., the user's clothing, people or objects around them, etc.) to use as reference data when generating personalized voice dialogues. Here, the cloud server stores and updates each user's data via the user database 315, further based on chronological order, and records each user's past conversation history.
[0047] In this way, the cloud server can effectively provide personalized voice interaction services to users through the voice / video model 307. Furthermore, when a user engages in personalized voice interaction, the cloud server continuously acquires the user's voice / video features 311, and the feature learning algorithm 309 continuously learns from the voice and video data generated in the user's interaction, updating or optimizing the voice / video model 307.
[0048] Based on the system configuration and the embodiment of the system described above, Figure 4 shows the first flowchart of an embodiment relating to a method for adaptively acquiring video data and performing personalized voice dialogue.
[0049] This method is executed on a cloud server, and the user establishes a connection with the cloud server by operating an application program executed on the user device. Through the application program, an interactive interface is activated, and the user device acquires audio data generated by the user's real-time voice, or audio data and video data, using sensor elements such as a microphone (or microphone and camera) (step S401). Subsequently, the system learns the characteristics of the audio, or the characteristics of the audio and video. Here, the audio characteristics may include changes in tone, rhythm, and intensity of the voice transmitted by the user through the user device, and the video characteristics may include video of the user themselves or the environment at their location (step S403). Thus, based on the audio characteristics, the meaning of words can be identified. When receiving audio data generated by the user's real-time voice, according to the first embodiment, the system also receives the user's real-time video data to acquire the characteristics of the user's current video. Therefore, based on the audio characteristics (including the meaning of words) and video characteristics, the system can simultaneously learn the user's current emotions, or the user's current emotions and the current state of the environment (step S405).
[0050] According to the embodiment, while acquiring audio features, it is also possible to acquire video features based on video data. The resulting audio and / or video features 41 are then transmitted to a video generation model 43 on a cloud server and used to generate pseudo-video (for example, changes in the facial expressions, mouth, and specific parts of a pseudo-humanoid chatbot's face). They are also transmitted to a personalized natural language model 45 to generate voice dialogue based on the emotions obtained from the features 41.
[0051] In step S405, which involves acquiring the user's current emotions, the cloud server uses a natural language processing model to identify the meaning of words in the speech features and, combined with the user's actions acquired in the video features, can determine the current environmental information. Based on the meaning of words and the current environmental information (which may include environmental information determined from the video of the user's location and environmental information provided by an external system), it generates a voice dialogue based on the user's current emotions (which is also a voice dialogue that responds to the user's real-time voice).
[0052] Furthermore, the cloud server may also utilize the video generation model 43 to generate images of pseudo-objects based on audio characteristics or a combination of audio and video characteristics, and allow the user to select a pseudo-object on the cloud server (step S407). Additionally, one of the pseudo-objects may be displayed in an interactive interface on the user device's display (step S409), and the personalized natural language model 45 may generate dialogue content based on the user's meaning, current emotions, or these combined with environmental information, to engage in voice dialogue with the user (step S411). For example, with respect to a pseudo-object (e.g., a pseudo-humanoid chatbot), the image of the pseudo-object may primarily reflect changes in the facial expression, mouth, and specific body parts of the face that reflect the current emotions.
[0053] During the dialogue process, the personalized natural language model 45 on the cloud server continues to operate and, based on the meanings formed in the dialogue content, determines whether or not to capture video (step S413). If there is no point in acquiring any video (no), the process returns to step S401 and the above steps can be repeated.
[0054] On the other hand, when it is determined that it is necessary to acquire video from the meaning of the words (yes), the cloud server activates the function to capture video on the user device via a human-machine interface through its user interface module (for example, 205 in Figure 2). For example, it can activate the main camera or selfie camera as needed (step S415) to capture and acquire video data (step S417).
[0055] In this process, the cloud server continuously acquires audio and video features, and after natural language processing, a simulated human-like chatbot is displayed to generate corresponding voice dialogues. By repeating the above process (for example, receiving audio data generated by the user via the user device, or audio and video data, acquiring audio features, or audio and video features, and learning the user's current emotions), personalized voice dialogues are continuously generated.
[0056] According to the first embodiment, in step S415 of the above flow, if the system determines from the meaning of the user's words that the topic relates to environmental information, using a handheld mobile device as an example, a software program running on the user's device can activate the main camera on the user's mobile device and use it to capture video of the environment in which the user is located. Thus, the system can acquire video data of the environment in which the user is located, acquire video features through video processing, and the natural language model can determine the characteristics of the environment based on the video features, generate corresponding voice dialogue based on the characteristics of the environment, and display the corresponding video.
[0057] According to the second embodiment, in step S415 of the above flow, if it is determined from the user's vocabulary that the topic concerns the user themselves, the selfie camera on the user device can be activated and used to capture a video of the user. Similarly, the system can acquire video features from the video data of the user's selfie video, determine the user's current emotions or needs from the video features, and display a corresponding video while generating a corresponding further voice dialogue based on those emotions or needs.
[0058] For related embodiments, see the flowchart of another embodiment relating to a method for adaptively acquiring video data and performing personalized voice dialogue, shown in Figure 5.
[0059] The system receives voice data transmitted by the user device via an interactive interface (step S501), and similarly analyzes the characteristics of the voice using a natural language model to determine the meaning and emotion of the words (step S503), while conducting a corresponding voice dialogue (step S505).
[0060] Furthermore, the system determines from the user's utterance that the topic relates to environmental information and decides whether or not to acquire environmental information (step S507). If it does not need to acquire the information, the system repeatedly performs the voice dialogue service in the flow steps. On the other hand, if it determines that it needs to acquire environmental information, it activates the video capture function on the user device via a software program executed on the user device, activating the camera (either the main lens or the front lens can acquire environmental images) (step S509), which can then be used to capture images of the environment in which the user is located.
[0061] Simultaneously, the system continues to acquire video data (step S511), obtains video features from the video data of the user's environment, and determines the characteristics of the environment (step S513). According to the embodiment, the system can determine and generate a display image from the captured video features via a video model or a natural language model (multimodal) and display it on an interactive interface (step S515), and at the same time generate corresponding voice dialogue based on the characteristics of the environment (step S517).
[0062] An embodiment of the adaptive acquisition of video data and the subsequent personalized voice dialogue can be seen by referring to the schematic diagram of the scene shown in Figure 6.
[0063] The related user terminal device is a pseudo-person display sign 60 in the first scene 6 as shown in the drawing, which is implemented by a computer system and has a built-in natural language model or connects to a cloud server via a network to acquire video data adaptively and acquire a service that provides personalized voice dialogue.
[0064] This section describes how user 600 enters the first scene 6 (for example, the entrance to a restaurant) and interacts with a simulated person 603 displayed on the simulated person sign 60. The interaction may involve an introduction to the restaurant, dining, ordering food, and reserving a table. During this process, the camera lens 601 built into (or attached to) the simulated person sign 60 is activated to photograph user 600 while covering the user's background.
[0065] User 600 can interact with the simulated person 603, and the simulated person display sign 60 or a cloud server continuously acquires and analyzes the user 600's voice and video data to generate dialogue content and display video according to the user's needs.
[0066] According to the diagram, the pseudo-person display sign 60 is equipped with something like an information service machine (Kiosk) that provides services to customers. A natural language model operating within it receives the user's voice data, obtains the meaning of the words, and then responds to the user in natural language. In the course of the conversation, if the user's meaning is determined to be a topic that requires information about the environment, the application program running on the user's device actively activates the camera lens to capture and acquire images of the environment. If the user's meaning is determined to be a topic related to the user themselves, the application program activates the camera lens to photograph the user and acquire images of the user.
[0067] The example in the drawing shows a simulated human display sign 60 installed in a restaurant as an information service device. It is equipped with a camera lens and a display, and the system operating here provides a smart customer service robot that meets the restaurant's needs. The user (customer) interacts with the simulated human displayed on the screen, and when it determines from the meaning of the speech that it is necessary to activate the camera, it acquires the user's image and then provides appropriate content (for example, food orders or meal suggestions).
[0068] Figure 7 shows a schematic diagram of another scenario in which video data is adaptively acquired and personalized voice dialogue is performed.
[0069] According to the diagram, user 700 enters the second scene 7 and interacts with a simulated person 703 displayed on the simulated person display sign 70. During this process, the camera lens 701 on the simulated person display sign 70 is activated to photograph user 700, acquire video data of user 700, perform analysis, acquire video features, and generate dialogue content and display video based on the audio and video features.
[0070] For example, in the second scenario 7, for instance, an apparel shop, when a user 700 tries on clothes, they can use a virtual person display sign 70 to communicate with a model on the sign using natural language to request suggestions for how to wear the clothes. When trying on clothes, the application program can activate the camera lens 701 of the virtual person display sign 70 to capture full-body or partial images of the user, and based on the user's needs obtained through analysis, suggestions for how to wear the clothes can be made. Furthermore, a virtual person 703 can provide introductions, displaying actual images of the user trying on clothes, or generating a model (by acquiring and combining video data of the items being worn and the user's image) to generate a simulated video of the user trying on clothes, making it easier for the user to film themselves trying on clothes.
[0071] In short, the present invention provides a method and system for adaptively acquiring video data and performing personalized voice dialogue. Here, based on the above embodiments, the system constructs an AI robot (particularly a pseudo-person depicted by a pseudo-person model) that interacts with the user, and, if necessary, determines whether or not to activate the shooting function on the user's device. It then acquires voice and video features in real time, analyzes the user's needs, and provides a personalized voice dialogue service. [Explanation of Symbols]
[0072] 10...Users 100...User devices 101...Pseudo-humanoid chatbot 103...User video 105...Photography lens 20...Cloud Server 201...Natural Language Processing Module 203...Instruction Processing Module 205...User Interface Module 207...Machine Learning Module 209...Database Module 211...Video processing module 200... Network 111...Video call processing module 113...Voice acquisition module 115...Video capture module 117...Communication module 22...Database 221...User Data 223... Artificial Intelligence Pseudo-Model 225... Emotion Database 301...User's audio / video 303...Man-machine interface 305...Interactive Interface 307...Audio / Video Model 309...Feature Learning Algorithms 311...User's voice / video characteristics 313...Environmental information 315...User Database 317...Environmental video 41...Features 43...Image generation model 45... Personalized natural language models 6...Scene 1 60... Pseudo-human display sign 601...Photography lens 603... Pseudo-person 600...users 7...Scene 2 70... Pseudo-human display sign 701...Photography lens 703... Pseudo-person 700...users S401~S417...Process S501~S517...Process
Claims
1. It runs on a cloud server, The system receives audio data generated by the user's real-time voice via the user's device, acquires the characteristics of the user's current voice, and identifies the meaning of words in the audio characteristics using a natural language processing model. After determining that it is necessary to acquire video, the function to capture video on the user device is activated to acquire video data, and then, based on the characteristics of the audio and video, a corresponding voice dialogue is generated. The process includes repeatedly receiving the audio data generated by the user via the user device and acquiring the characteristics of the audio, and receiving the video data generated by the user device and acquiring the characteristics of the video, thereby continuously generating the audio dialogue based on the characteristics of the audio and the video. A method for adaptively acquiring video data and conducting personalized voice dialogue.
2. In addition to obtaining the meaning of words in the aforementioned voice features, the system learns the user's current emotions based on the voice features, and simultaneously generates the voice dialogue based on the meaning and emotions, wherein the voice features include the tone, rhythm changes, and intensity changes of the voice transmitted by the user through the user device. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 1.
3. The system searches an emotion database and generates personalized voice dialogues that correspond to changes in tone, rhythm, and voice intensity that should reflect the user's current emotions. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 2.
4. If it is determined from the user's words that the topic concerns environmental information, the function to capture video on the user device is activated and used to capture video of the environment in which the user is located. Features of the video are obtained from the video data of the environment in which the user is located, the characteristics of the environment are determined, and further voice dialogue is generated in accordance with the characteristics of the environment. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 1.
5. If it is determined from the meaning of the user's words that the topic concerns the user themselves, the function to take a selfie video on the user's device is activated and used to record a video of the user themselves. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 1.
6. The system obtains video features from the video data of the user's selfie video, and determines the user's current emotions or needs based on the video features, thereby generating corresponding further voice dialogue based on those emotions or needs. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 5.
7. On the aforementioned cloud server, deep learning is performed to construct a user profile based on the user's preferences and personality learned from the audio and video data that the user device has continuously received from the past to the present. A method for adaptively acquiring video data and performing personalized voice dialogue according to any one of claims 1 to 6.
8. The cloud server acquires the user's current voice characteristics, semantics, and video characteristics, and further acquires the user profile based on the user's identification information, and generates the corresponding voice dialogue based on the semantics, user profile, and / or current environment information. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 7.
9. The cloud server retrieves a pseudo-object generated by an artificial intelligence pseudo-model from a database, transmits video signals and model data to the user device, displays the pseudo-object on the user device's display, and simulates the voice interaction. A method for adaptively acquiring video data and performing personalized voice dialogue according to any one of claims 1 to 6.
10. The aforementioned virtual object is a virtual humanoid chatbot, and the cloud server generates an interactive interface via a web server program, providing a selection of multiple virtual humanoid chatbots. A method for adaptively acquiring video data and performing personalized voice dialogue as described in claim 9.
11. Includes a cloud server that connects to a user device and provides personalized voice interaction services to the user via an interactive interface, and performs a method of adaptively acquiring video data by one or more processors to perform personalized voice interaction. The method described above, which involves adaptively acquiring video data and performing personalized voice dialogue, The system receives audio data generated by the user's real-time voice via the user device, acquires the characteristics of the user's current voice, and identifies the meaning of words in the audio characteristics using a natural language processing model. After determining that it is necessary to acquire video, the function to capture video on the user device is activated to acquire video data, and based on the characteristics of the audio and video, a corresponding voice dialogue is generated. The process includes repeatedly receiving the audio data generated by the user via the user device and acquiring the characteristics of the audio, and receiving the video data generated by the user device and acquiring the characteristics of the video, thereby continuously generating the audio dialogue based on the characteristics of the audio and the video. A system that adaptively acquires video data and performs personalized voice dialogue.
12. The system searches an emotion database to obtain the voice features corresponding to the user's current emotion, the voice features include changes in tone, rhythm, and intensity of speech that should reflect the user's current emotion, generates a personalized voice dialogue corresponding to the voice features, obtains the meaning of words in the voice features, and further learns the user's current emotion based on the voice features while simultaneously generating the voice dialogue based on the meaning and emotion. The system according to claim 11.
13. In the method of adaptively acquiring video data and performing personalized voice dialogue, when it is determined from the user's utterance that the topic concerns environmental information, the video recording function on the user device is activated and used to record video of the environment in which the user is located. Features of the video are acquired from the video data of the environment in which the user is located, the characteristics of the environment are determined, and further voice dialogue corresponding to the characteristics of the environment is generated. The system according to claim 11.
14. In the method for adaptively acquiring video data and conducting personalized voice dialogue, when it is determined from the user's vocabulary that the topic concerns the user themselves, the video selfie function on the user device is activated and used to capture video of the user themselves. Features of the video are acquired from the video data of the user's selfie video, and based on the features of the video, the user's current emotions or needs are determined, thereby generating corresponding further voice dialogue based on those emotions or needs. The system according to claim 11.
15. On the cloud server, deep learning is performed to construct a user profile based on the user's preferences and personality learned from the audio and video data that the user device has continuously received from the past to the present, to obtain the user's current audio features, semantics, and video features, to obtain the user profile based on the user's identification information, and to generate the corresponding audio dialogue based on the semantics, user profile, and / or current environmental information. The system according to any one of claims 11 to 14.