AI Interaction Method and Apparatus, Wearable Device, Storage Medium
By building user portraits in wearable devices and combining multi-source data and voice correlation information to generate personalized recommendation information, the problem of poor user experience in existing AI interaction technologies is solved, and interaction accuracy and user satisfaction are improved.
Patent Information
- Application Number
- CN202411956090.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-12-28
AI Technical Summary
The existing AI interaction technology has poor user experience and is difficult to meet the diverse needs of users.
By building user portraits based on multi-source data, combining the associated information and user portrait of the first voice, personalized recommendation information is generated, and voice feedback is performed through wearable devices to improve interaction accuracy and user experience.
It has achieved an in-depth understanding of user characteristics, preferences and habits, provided individual suggestions, improved interaction accuracy and user satisfaction, and enhanced user trust between users and devices.
Smart Images

Figure CN119920249B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of human-computer interaction, and more specifically, relates to an AI interaction method and apparatus, a wearable device, and a storage medium. Background Art
[0002] With the development of technology, people's expectations for human-computer interaction methods have been increasing day by day. Traditional interaction methods such as keyboards and mice have gradually become difficult to meet the needs of diversification and convenience. People's demand for a more natural and convenient human-computer interaction has promoted the progress of AI interaction methods. In terms of voice interaction, acoustic models and language models have been evolving, gradually shifting from models based on traditional statistical methods to neural network models based on deep learning, such as deep neural networks, recurrent neural networks, and their variants, long short-term memory networks and gated recurrent units. The development of these technologies has continuously revolutionized speech recognition and synthesis technologies.
[0003] However, the current AI interaction technology still has problems such as poor user experience and difficulty in meeting the diverse needs of users. Summary of the Invention
[0004] The purpose of the present disclosure is to provide an AI interaction method and apparatus, a wearable device, and a storage medium to improve the user experience and meet the diverse needs of users.
[0005] In the first aspect of the embodiments of the present disclosure, an AI interaction method is provided, which is applied to a wearable device and includes:
[0006] Constructing a user profile based on multi-source data; the multi-source data is data of a target user, and the target user is a user wearing the wearable device;
[0007] In response to receiving a first voice of the target user, determining a target intent type based on the first voice; determining multi-modal perception data of the target user based on the target intent type; determining associated information corresponding to the first voice based on the multi-modal perception data; determining recommended information corresponding to the first voice based on the associated information and the user profile;
[0008] Generating a second voice based on the associated information and the recommended information; the second voice is the voice played by the AI in the wearable device.
[0009] In the second aspect of the embodiments of the present disclosure, an AI interaction apparatus is provided, which is applied to a wearable device and includes:
[0010] A first calculation module, configured to construct a user profile based on multi-source data; the multi-source data is data of a target user, and the target user is a user wearing the wearable device;
[0011] A second computing module, configured to determine associated information corresponding to the first voice upon receiving the first voice of the target user, and determine recommended information corresponding to the first voice based on the associated information and the user profile;
[0012] A third computing module, configured to generate a second voice based on the associated information and the recommended information; the second voice is the voice played by the AI in the wearable device.
[0013] In a third aspect of the embodiments of the present disclosure, a wearable device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor, where when the processor executes the computer program, the steps of the above AI interaction method are implemented.
[0014] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above AI interaction method are implemented.
[0015] The beneficial effects of the AI interaction method, device, wearable device, and storage medium provided by the embodiments of the present disclosure are as follows:
[0016] In the embodiments of the present disclosure, by constructing a user profile based on multi-source data, the user's characteristics, preferences, and habits can be deeply understood, and personalized suggestions can be provided for the user, improving the satisfaction of the user's use. The embodiments of the present disclosure combine the associated information of the first voice, including hidden information, context information, etc., and are not limited to the surface content of the voice, and can more accurately understand the user's true intention in a specific environment and state, avoid misunderstandings, effectively improve the accuracy of interaction, and enhance the trust between the user and the wearable device. In summary, the embodiments of the present disclosure enable the wearable device to better meet the diverse needs of users. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0018] Figure 1 It is a flowchart of the AI interaction method provided by an embodiment of the present disclosure;
[0019] Figure 2 It is a structural block diagram of the AI interaction device provided by an embodiment of the present disclosure;
[0020] Figure 3Schematic block diagram of a wearable device provided by an embodiment of the present disclosure. Detailed implementation manners
[0021] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from hindering the description of the present disclosure.
[0022] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the accompanying drawings.
[0023] Please refer to Figure 1 , Figure 1 Flow schematic diagram of an AI interaction method provided by an embodiment of the present disclosure. This method is applied to a wearable device and may include S101 to S103.
[0024] S101: Construct a user profile based on multi-source data. The multi-source data is data of the target user, and the target user is the user wearing the wearable device.
[0025] In this embodiment, the multi-source data includes application behavior data, historical interaction data, and user identity information.
[0026] Constructing a user profile based on multi-source data includes:
[0027] Construct a user profile based on application behavior data, historical interaction data, and user identity information.
[0028] In this embodiment, the wearable device may include smart glasses. Multi-source data refers to a data set from multiple different channels or aspects. The multi-source data may include data related to the target user and associated with the wearable device. The target user is an individual wearing a specific wearable device, and the data collected and analyzed by the device is all for the situation of this user.
[0029] The application behavior data may include the types of applications used, usage frequency, usage duration, operation paths within the application, etc. The historical interaction data may include the content involved in voice interaction, interaction time, interaction frequency, interaction methods, and feedback on the interaction results. The user identity information may include age, gender, occupation, health status, living area, etc. By collecting this multi-source data in different dimensions, mining the rules and associations therein, analyzing the user's behavior patterns, interests, living habits, and needs, a user profile can be outlined, providing a basis for personalized services.
[0030] Exemplarily, a dedicated data collection module is provided inside the wearable device to record application behavior data, historical interaction data in real time, and obtain user identity information when the user initially sets up the device. The collected data is classified and stored according to categories. For example, different database tables are established to store application behavior, historical interaction, and user identity information data respectively, facilitating subsequent processing.
[0031] Data analysis and machine learning techniques are adopted, such as clustering analysis, decision tree algorithms, etc. For application behavior data, features such as the usage frequency and duration distribution of different applications are analyzed. For historical interaction data, common interaction patterns, instruction preferences, etc. are extracted. For user identity information, the behavior differences of users with different identity attributes are examined in combination with the other two types of data. By integrating these analysis results, a user portrait is constructed in the form of vectors or structured data. For example, the user's preference degrees for different applications, common interaction behaviors, etc. are used as different dimensions of the portrait to represent.
[0032] The user portrait can be updated regularly, and the description of the user is continuously optimized and improved according to the newly collected data to make it more in line with the actual situation of the user.
[0033] S102: In response to receiving the first voice of the target user, determine the target intent type based on the first voice. Determine the multi-modal perception data of the target user based on the target intent type. Determine the associated information corresponding to the first voice based on the multi-modal perception data. Determine the recommended information corresponding to the first voice based on the associated information and the user portrait.
[0034] In this embodiment, the associated information corresponding to the first voice may include the hidden information corresponding to the first voice.
[0035] The first voice refers to the voice command issued by the wearable device user, with diverse contents, such as "I'm a bit tired and want to relax".
[0036] The associated information is the data related to the first voice, including the content directly expressed in the voice and the potential hidden information. The hidden information can be the user's emotion and potential needs inferred through voice intonation, specific words, etc. It can also include the context information when the voice is issued, such as the noise level of the environment where the device is located, the user's motion state, etc. It can also include the information hidden based on the translated text corresponding to the first voice. For example: The first voice is: "What's there to eat nearby", and the associated information (i.e., the hidden information) corresponding to the first voice may include the content corresponding to "nearby" and "eat", such as determining and generating the location of the target user, generating restaurants within a radius of a kilometers, determining the restaurant names, directions, and distances from the target user.
[0037] The user portrait is a description of the user's comprehensive features, including information such as the user's interests, hobbies, living habits, and health status.
[0038] Recommendation information is a suggestion given to meet the user's current needs based on associated information and user profiles, such as recommending a quiet coffee shop nearby or providing suggestions for playing soothing music.
[0039] Exemplarily, when the first voice is received, the voice is first analyzed and processed to obtain its key content and relevant context information, and the corresponding associated information is generated. Combining the pre-constructed user profile, which includes the user's preferences, habits, etc., recommended information that meets the user's needs and characteristics is screened out from the resource library.
[0040] Exemplarily, the first voice is converted into text content using speech recognition technology, and then the intent, key entities, etc. of the text are analyzed through natural language processing technology. At the same time, the current environment and user status information are obtained by combining the sensors of the wearable device to determine the associated information. According to the key content in the associated information, a matching search is performed in the target retrieval resource library determined based on the user profile to generate recommendation information.
[0041] In this embodiment, the target intent types include information query types and life service types. The multimodal perception data includes user physiological data and environmental perception data.
[0042] Determining the multimodal perception data of the target user based on the target intent type includes:
[0043] If the target intent type is an information query type, the multimodal perception data of the target user is determined to be environmental perception data.
[0044] If the target intent type is a life service type, the multimodal perception data of the target user is determined to be user physiological data and environmental perception data.
[0045] In this embodiment, the target intent type refers to the main purpose category expressed by the user through the first voice, which can be divided into information query types (such as asking for news, weather, etc. information) and life service types (such as ordering takeout, making an appointment for services, etc. related needs to improve life convenience).
[0046] Multimodal perception data refers to the acquired text, audio, image, etc. data, which is used to more comprehensively understand the user status and the surrounding environment.
[0047] User physiological data can include parameters reflecting the body condition such as the user's heart rate, blood pressure, body temperature, body posture, fatigue level, etc.
[0048] Environmental perception data can include information such as the temperature, humidity, light intensity, noise level, geographical location, and surrounding object recognition around the wearable device. The wearable device can be connected to a smart bracelet or other auxiliary electronic devices to obtain multimodal perception data.
[0049] Exemplarily, a large language model is used to perform text recognition and semantic analysis on the first voice. By training a neural network-based classifier, according to features such as keywords and grammatical structures in the voice content, it is classified into information query type or life service type.
[0050] Integrate appropriate sensors in the wearable device, such as a heart rate sensor, an accelerometer, etc. The heart rate sensor continuously monitors heart rate data, and the accelerometer can judge body posture and exercise intensity by analyzing the motion pattern. The collected data is filtered, calibrated and then stored. Or, connect the relevant sensors to the wearable device (smart glasses) for data transmission.
[0051] Equip with a temperature sensor, a humidity sensor, a microphone, a GPS module, etc. Each sensor collects environmental information in real time and performs data processing such as corresponding unit conversion and data fusion.
[0052] After determining the intention type, obtain the corresponding multi-modal perception data from the data acquisition module. If it is of the information query type, only extract the environmental perception data. If it is of the life service type, then extract the user physiological data and the environmental perception data. Then combine these data with the content of the first voice, and determine the associated information through a data analysis algorithm (such as association rule mining). For example, in the life service type, according to the situation that the user has a fast heart rate and the environmental temperature is high, recommend heatstroke prevention-related services or recommendations for the user.
[0053] In this embodiment, by distinguishing the target intention type and combining multi-modal perception data, the user's needs can be accurately understood. Considering the actual situations of the user and the environment, more context-aware associated information is provided, improving the accuracy of the recommendation. This embodiment synthesizes the user's physiological and environmental data, enabling a comprehensive understanding of the user. This avoids the one-sidedness of only focusing on a single factor, making the associated information reflect the user's true state, better meeting the user's needs, and enhancing the user experience. Accurate associated information helps to improve the interaction quality, making the user feel the intelligence and thoughtfulness of the device, and enhancing the user's trust and dependence on the wearable device.
[0054] S103: Generate a second voice based on the associated information and the recommended information. The second voice is the voice played by the AI in the wearable device.
[0055] In this embodiment, the recommended information is useful content for the user obtained based on the associated information and the user profile, such as a list of restaurants that meet the user's taste and distance requirements and their evaluations. The second voice is the response voice of the device to the user, and it feeds back content such as the recommended information to the user in voice form. The second voice can have a suitable intonation set according to the type of the recommended information. For example, a lively intonation is used for recommending entertainment activities, and a steady intonation is used for important notifications. The second voice can include the specific text content of the recommended information, prompts, etc., such as "Here are some restaurants recommended for you". The voice playback duration of the second voice can be determined according to the amount of content.
[0056] Exemplarily, based on the associated information and the recommended information, a natural language generation template or model is used to generate smooth and logical text content. A voice synthesis engine is used to select a suitable voice timbre, convert the generated text into the second voice, and at the same time adjust parameters such as the intonation and speech rate of the voice to make it more natural. Finally, it is played to the user through the speaker of the wearable device.
[0057] It can be concluded from the above that through the detailed user profile and in-depth analysis of the first voice in this embodiment, highly personalized recommended information can be provided for the user. This embodiment combines the hidden information in the associated information, such as the user's emotions, potential needs, and environmental and status information, to make the interaction more intelligent. It avoids simply responding to the surface content of the user's voice, can better understand the user's true intentions, accurately give suggestions that conform to the current situation, and reduce misunderstandings and ineffective interactions.
[0058] Screening the recommended information based on the user profile and the associated information can effectively utilize the data in the resource library. It avoids searching for and processing a large amount of irrelevant information, improves the operation efficiency, and can also provide valuable content for the user more quickly. Regularly updating the user profile can ensure that the interaction process always conforms to the user's latest behavior patterns and changing needs, thereby enhancing the user experience and meeting the diverse needs of the user.
[0059] In one embodiment of the present disclosure, a user profile is constructed based on application behavior data, historical interaction data, and user identity information, including:
[0060] The application behavior data is divided based on the application type to obtain application behavior data corresponding to multiple time periods, and each time period corresponds to one application type.
[0061] Based on the application type and data volume corresponding to each time period, a first local user profile corresponding to the application behavior data is determined.
[0062] A second local user profile is determined based on the historical interaction data.
[0063] A third local user profile is determined based on the user identity information.
[0064] Fuse the first partial user portrait, the second partial user portrait, and the third partial user portrait based on the weight coefficient matrix corresponding to the application behavior data, historical interaction data, and user identity information to obtain the user portrait.
[0065] In this embodiment, determining the first partial user portrait corresponding to the application behavior data based on the application type and data volume corresponding to each time period includes:
[0066] Determine the application type - data volume sequence based on the application type and data volume corresponding to each time period, and determine the first partial user portrait based on the application type - data volume sequence.
[0067] In this embodiment, the application type - data volume sequence includes an application type sequence and a data volume sequence corresponding to the application type sequence, and the application type sequence and the data volume sequence are mutually mapped in the dimension of time period.
[0068] In this embodiment, determining the second partial user portrait based on the historical interaction data includes:
[0069] Divide the historical interaction data based on the interaction type to obtain the historical interaction data corresponding to multiple time periods, and each time period corresponds to one interaction type.
[0070] Determine the interaction type - data volume sequence based on the interaction type and data volume corresponding to each time period, and determine the second partial user portrait based on the interaction type - data volume sequence.
[0071] In this embodiment, the interaction type - data volume sequence includes an interaction type sequence and a data volume sequence corresponding to the interaction type sequence, and the application type sequence and the data volume sequence are mutually mapped in the dimension of time period.
[0072] In this embodiment, determining the third partial user portrait based on the user identity information includes: determining the third partial user portrait from the user tag library based on the user identity information.
[0073] In this embodiment, the application type may include social applications, entertainment applications, health applications, office applications, etc. The identification of the application type facilitates the division of the application behavior data, and each subset of the application behavior data after division corresponds to a unique time interval and data volume. The data volume refers to the number of data records generated for a certain application type or interaction type within a certain time period.
[0074] The first part of the user portrait represents the portrait part that initially depicts user characteristics and habits from the perspective of the target user's application usage behavior. It mainly reflects the user's usage preferences, usage activity levels, etc. for different types of applications at different times, and is part of the content that constitutes the complete user portrait. The second part of the user portrait represents the partial portrait content of the target user in terms of the topics, demand directions, and frequency of interactions during the interaction with the system at different times. The third part of the user portrait is to match relevant tags that conform to the user from a pre-set set of typical description tags (i.e., the user tag library) that can represent different identity characteristics based on the user's identity information, so as to outline the partial portrait that reflects the user's characteristics, preferences, etc. from the basic attribute level of the user, and supplement the identity background-related feature descriptions for the complete portrait.
[0075] The weight coefficient matrix is a matrix form representation that measures the relative importance of application behavior data, historical interaction data, and user identity information. The elements in the weight coefficient matrix correspond to the weight values of each data type in constructing the user portrait. By reasonably setting these weights, it is possible to comprehensively coordinate the contribution degrees of different data to the final user portrait according to business requirements and actual application scenarios, and then achieve a more practical and targeted portrait fusion.
[0076] Exemplarily, the weight coefficient matrix of application behavior data, historical interaction data, and user identity information can be calculated based on the principal component analysis weighting method.
[0077] In this embodiment, the application behavior data is divided based on the application type, which can carefully present the user's usage preferences and activity levels for various applications at different times. Combining with the historical interaction data, it further clarifies the user's needs and interaction situations in different scenarios, making the portrait more targeted. This embodiment uses the user identity information to match relevant tags from the tag library, supplementing the user identity background characteristics and comprehensively enriching the content of the user portrait. By determining the weight coefficient matrix through the principal component analysis weighting method, it reasonably coordinates the contributions of various data types to the portrait and realizes a more practical and accurate portrait fusion, providing personalized services for users.
[0078] In an embodiment of the present disclosure, determining the recommended information corresponding to the first voice based on the association information and the user portrait includes:
[0079] Determining the target retrieval resource library based on the user portrait.
[0080] Determining the reply text corresponding to the first voice from the target retrieval resource library based on the association information.
[0081] Determining the recommended information corresponding to the first voice from the target retrieval resource library based on the reply text.
[0082] In this embodiment, determining the target retrieval resource library based on the user profile includes:
[0083] Determine user tags and temporal features based on the user profile.
[0084] Match the first retrieval resource library from the retrieval resources based on the user tags.
[0085] Determine the target retrieval resource library from the first retrieval resource library based on the temporal features.
[0086] In this embodiment, the target retrieval resource library is screened out from numerous retrieval resources, which highly matches the needs and characteristics of the current user and is used to find the reply text corresponding to the first voice and the recommended information in a specific database.
[0087] User tags are identifiers extracted from the user profile to describe the user's characteristics, preferences, behavior habits, etc., such as "sports enthusiast", "foodie", "night owl", etc., which facilitate quickly classifying users and locating relevant resources.
[0088] Temporal features include user characteristics related to the time sequence, which can include parameters such as the user's activity level in different time periods, behavior preferences in specific time periods, and the change of device usage frequency over time.
[0089] In this embodiment, the user tags and temporal features are analyzed based on the constructed user profile. Through the user tags, the first retrieval resource library that generally matches the user's interests and habits can be initially screened out, and this step is matched based on the relevance between the tags and the resources. Then, the temporal features are further used to accurately locate in the first retrieval resource library to find the target retrieval resource library that better fits the user's current time-related behavior habits and other situations. Finally, based on the associated information, the corresponding reply text is found in this accurate target retrieval resource library, and then the recommended information is determined, so that the recommended information not only conforms to the overall characteristics of the user but also meets the specific situational needs at present.
[0090] Exemplarily, the constructed user profile is parsed, and the user tags are identified through a feature extraction algorithm. For example, the tags are determined according to the application types frequently used by the user, the activities frequently participated in, etc. At the same time, a time series analysis algorithm is used to analyze the user's behavior data in different time periods, and the temporal features are extracted, such as calculating the mean value and variance of the usage frequency in a specific time period. A large retrieval resource library is established to store resource contents in various aspects such as knowledge Q&A, life services, and hobbies.
[0091] Based on user tags, use a text matching algorithm to screen out the first retrieval resource library from the retrieval resources, and gather the resources with a high degree of relevance to the user tags together. Based on temporal features, further determine the target retrieval resource library in the first retrieval resource library. For example, if the user is active at night and likes entertainment, select the target retrieval resource library with richer resources related to entertainment at night.
[0092] According to the key content in the associated information, such as user needs, current environment, etc., use text retrieval technology to find the reply text corresponding to the first voice in the target retrieval resource library. Then, further determine the recommended information corresponding to the first voice based on the guiding information, relevant resource links, etc. in the reply text. For example, if the reply text mentions that a certain restaurant is good, the recommended information can be the detailed address, preferential information, etc. of the restaurant.
[0093] This embodiment uses the tags and temporal features in the user portrait to screen out the target retrieval resource library, which can ensure that the recommended information fits the user's interests and current situation. By accurately positioning the target retrieval resource library, it avoids searching in a large number of irrelevant resources, reduces the consumption of computing resources and time costs, and quickly and accurately provides replies and recommendations for users. The recommended information that conforms to the overall user characteristics and current needs allows users to feel that the device understands them, enhances the interaction satisfaction between the user and the device, and improves user loyalty.
[0094] In an embodiment of the present disclosure, generating a second voice based on the associated information and the recommended information includes:
[0095] Generate a first reply text based on the associated information and a second reply text based on the recommended information.
[0096] Concatenate the first reply text and the second reply text to generate a target text, and synthesize the second voice based on the target text.
[0097] In this embodiment, the first reply text: the response content generated for the associated information, focusing on the feedback on the user's current situation or needs, such as "Understood, you are looking for a nearby restaurant. 15 nearby restaurants have been searched for you".
[0098] The second reply text: focuses on the elaboration and explanation of the recommended information, such as "Recommend restaurant A to you. It is 500 meters near you and has a very good evaluation. It specializes in the Sichuan cuisine you like" or "There are hot pot, barbecue, and grilled meat nearby. Which kind of cuisine do you want to eat?".
[0099] The target text: the complete text content formed by concatenating the first reply text and the second reply text, serving as the basis for synthesizing the second voice.
[0100] The second voice: the voice played by the AI in the wearable device to the user, used to convey the associated information and the recommended information.
[0101] For example, related information: the user asks "I feel a little tired and uncomfortable", and the wearable device recognizes "feeling" "tired" and "uncomfortable", and determines it as a life service. Wake up the relevant devices connected to the wearable device, obtain the user's physiological characteristics and environmental characteristics, and obtain data information about the user's fast heart rate, slightly high blood pressure, and high current ambient temperature and humidity. It is judged that the user may be in a state of fatigue and slight discomfort, and the location information is displayed at home.
[0102] Recommended information: It is recommended that users take a rest first, turn on the air conditioner to adjust the indoor temperature to a comfortable range, and remind users to replenish appropriate amount of water.
[0103] First reply text: "You may feel tired now due to a variety of factors."
[0104] Second reply text: "You can take a rest first, turn on the air conditioner to a comfortable temperature, and drink some water. Do you need me to call you?"
[0105] Second voice: "You may feel tired due to many factors. You can take a rest first, turn on the air conditioner to a comfortable temperature, and drink some water. Do you need me to call you?"
[0106] Exemplary, associated information: the user says "I want to go to a nearby shopping mall", the wearable device connects to the user's mobile phone, identifies "nearby" and "shopping mall", wakes up the map APP in the mobile phone, and obtains that the current user is in a busy traffic area, the surrounding noise is loud, and the user walks slowly.
[0107] This embodiment can fully cover user needs by generating a first reply text and a second reply text respectively and splicing them together, and provide users with complete information from need confirmation to specific suggestions, thereby avoiding information loss.
[0108] The second voice combines the understanding of user context and targeted recommendations, making the interaction more in line with human communication habits, making users feel like they are communicating with an understanding assistant, and improving the user experience. The voice form clearly and accurately conveys related information and recommended information, especially when the user is busy or inconvenient to view the screen, which can efficiently serve the user.
[0109] Corresponding to the AI interaction method in the above embodiment, Figure 2 This is a structural block diagram of an AI interaction device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 2 The AI interaction device 20 is applied to a wearable device, and includes: a first computing module 21, a second computing module 22 and a third computing module 23.
[0110] Among them, the first computing module 21 is used to construct a user profile based on multi-source data. The multi-source data is data of a target user, and the target user is a user wearing a wearable device.
[0111] The second computing module 22 is used to, in response to receiving the first voice of the target user, determine the target intent type based on the first voice, determine the multi-modal perception data of the target user based on the target intent type, determine the associated information corresponding to the first voice based on the multi-modal perception data, and determine the recommended information corresponding to the first voice based on the associated information and the user profile.
[0112] The third computing module 23 is used to generate a second voice based on the associated information and the recommended information. The second voice is the voice played by the AI in the wearable device.
[0113] In an embodiment of the present disclosure, the multi-source data includes application behavior data, historical interaction data, and user identity information. The first computing module 21 is specifically used to construct a user profile based on the application behavior data, historical interaction data, and user identity information.
[0114] In an embodiment of the present disclosure, the first computing module 21 is specifically further used to divide the application behavior data based on the application type to obtain application behavior data corresponding to multiple time periods, and each time period corresponds to an application type.
[0115] Determine the first local user profile corresponding to the application behavior data based on the application type and data volume corresponding to each time period.
[0116] Determine the second local user profile based on the historical interaction data.
[0117] Determine the third local user profile based on the user identity information.
[0118] Fuse the first local user profile, the second local user profile, and the third local user profile based on the weight coefficient matrix corresponding to the application behavior data, historical interaction data, and user identity information to obtain the user profile.
[0119] In an embodiment of the present disclosure, the target intent types include information query types and life service types. The multi-modal perception data includes user physiological data and environmental perception data.
[0120] The second computing module 22 is specifically further used to, if the target intent type is an information query type, determine that the multi-modal perception data of the target user is environmental perception data.
[0121] If the target intent type is a life service type, determine that the multi-modal perception data of the target user is user physiological data and environmental perception data.
[0122] In one embodiment of the present disclosure, the second computing module 22 is further specifically configured to determine a target retrieval resource library based on a user profile.
[0123] Determine a response text corresponding to the first voice from the target retrieval resource library based on the association information.
[0124] Determine recommendation information corresponding to the first voice from the target retrieval resource library based on the response text.
[0125] In one embodiment of the present disclosure, the second computing module 22 is further specifically configured to determine user tags and temporal features based on a user profile.
[0126] Match a first retrieval resource library from the retrieval resources based on the user tags.
[0127] Determine a target retrieval resource library from the first retrieval resource library based on the temporal features.
[0128] In one embodiment of the present disclosure, the third computing module 23 is specifically configured to generate a first response text based on the association information and generate a second response text based on the recommendation information.
[0129] Concatenate the first response text and the second response text to generate a target text, and synthesize a second voice based on the target text.
[0130] See Figure 3 , Figure 3 is a schematic block diagram of a wearable device provided in an embodiment of the present disclosure. As Figure 3 shown, the wearable device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store a computer program, and the computer program includes program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module in the above device embodiments, such as Figure 2 the functions of the modules 21 to 23 shown.
[0131] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0132] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.
[0133] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.
[0134] In specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first and second embodiments of the AI interaction method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the wearable device 300 described in the embodiments of the present disclosure, which will not be elaborated herein.
[0135] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing related hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0136] The computer-readable storage medium can be the internal storage unit of the wearable device in any of the foregoing embodiments, such as the hard disk or memory of the wearable device. The computer-readable storage medium can also be an external storage device of the wearable device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the wearable device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the wearable device. The computer-readable storage medium is used to store the computer program and other programs and data required by the wearable device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0137] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0138] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the wearable device and the unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described in detail here.
[0139] In several embodiments provided by this application, it should be understood that the disclosed wearable device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, or can also be in the form of electrical, mechanical or other connections.
[0140] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can also be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.
[0141] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0142] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. An AI interaction method, characterized in that, Applied to a wearable device, including: Construct a user profile based on multi-source data; the multi-source data is data of a target user, and the target user is a user wearing the wearable device; In response to receiving the first voice of the target user, determine the target intent type based on the first voice; determine the multi-modal perception data of the target user based on the target intent type; determine the associated information corresponding to the first voice based on the multi-modal perception data; determine the recommended information corresponding to the first voice based on the associated information and the user profile; Generate a second voice based on the associated information and the recommended information; the second voice is the voice played by the AI in the wearable device; The target intent types include information query types and life service types; the multi-modal perception data includes user physiological data and environmental perception data; The determining the multi-modal perception data of the target user based on the target intent type includes: If the target intent type is the information query type, determine the multi-modal perception data of the target user as the environmental perception data; If the target intent type is the life service type, determine the multi-modal perception data of the target user as the user physiological data and the environmental perception data.
2. The AI interaction method according to claim 1, wherein The multi-source data includes application behavior data, historical interaction data, and user identity information; The constructing a user profile based on multi-source data includes: Construct the user profile based on the application behavior data, the historical interaction data, and the user identity information.
3. The AI interaction method according to claim 2, wherein The constructing the user profile based on the application behavior data, the historical interaction data, and the user identity information includes: Divide the application behavior data based on the application type to obtain application behavior data corresponding to multiple time periods, and each time period corresponds to one application type; Determine the first local user profile corresponding to the application behavior data based on the application type and data volume corresponding to each time period; Determine the second local user profile based on the historical interaction data; Determine the third local user profile based on the user identity information; Fuse the first local user profile, the second local user profile, and the third local user profile based on the weight coefficient matrix corresponding to the application behavior data, the historical interaction data, and the user identity information to obtain the user profile.
4. The AI interaction method according to claim 1, wherein The determining the recommended information corresponding to the first voice based on the associated information and the user profile includes: Determine the target retrieval resource library based on the user profile; Determine the reply text corresponding to the first voice from the target retrieval resource library based on the associated information; Determine the recommended information corresponding to the first voice from the target retrieval resource library based on the reply text.
5. The AI interaction method according to claim 4, wherein The determining the target retrieval resource library based on the user profile includes: Determine user tags and temporal features based on the user profile; Match the first retrieval resource library from the retrieval resources based on the user tags; Determine the target retrieval resource library from the first retrieval resource library based on the temporal features.
6. The AI interaction method according to claim 1, wherein The generating a second voice based on the associated information and the recommended information includes: Generate a first response text based on the associated information, and generate a second response text based on the recommended information; Concatenate the first response text and the second response text to generate a target text, and synthesize the second voice based on the target text.
7. An AI interaction device, characterized in that, Applied to a wearable device, including: A first computing module for constructing a user profile based on multi-source data; the multi-source data is data of a target user, and the target user is the user wearing the wearable device; A second computing module for, in response to receiving the first voice of the target user, determining a target intent type based on the first voice; determining multi-modal perception data of the target user based on the target intent type; determining associated information corresponding to the first voice based on the multi-modal perception data; and determining recommended information corresponding to the first voice based on the associated information and the user profile; The target intent type includes information query type and life service type; the multi-modal perception data includes user physiological data and environmental perception data; The second computing module is specifically configured to, if the target intent type is the information query type, determine that the multi-modal perception data of the target user is the environmental perception data; if the target intent type is the life service type, determine that the multi-modal perception data of the target user is the user physiological data and the environmental perception data; A third computing module for generating a second voice based on the associated information and the recommended information; the second voice is the voice played by the AI in the wearable device.
8. A wearable device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Personalized information recommendation method and device, electronic equipment and medium
CN114417051A
Music recommendation method and device, computer storage medium and program product
CN118964733A