Interactive accompanying method and device, electronic equipment and storage medium
By introducing an interactive companionship system into smart speakers, using multiple monitoring devices to acquire and integrate activity data, and generate and execute interactive data, the problems of single interaction methods and limited data sources in the existing technology are solved, and a more efficient and diverse child education and growth companionship experience is achieved.
Patent Information
- Application Number
- CN202510077091.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-23
AI Technical Summary
Existing smart speakers have problems such as single interaction mode, limited data sources and difficulty in creating an immersive learning environment in terms of children's education and growth companionship.
By awakening the interactive companion system in the interactive companion system, the activity data of multiple monitoring devices (voice and visual devices) is obtained, the data is fused and the pre-trained matching model is called to determine the interaction conditions, the interactive data is generated and sent to multiple interactive devices for execution.
It realizes the acquisition of active data through multiple monitoring devices, improves the accuracy and diversity of interactive data, provides diversified interaction methods, and improves the user experience and system support.
Smart Images

Figure CN120030118A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an interactive companionship method, device, electronic device and storage medium. Background Art
[0002] With the development of smart technology, devices such as smart speakers have gradually entered households, but there are many problems in children's education and growth companionship.
[0003] Traditional smart speakers mainly rely on voice interaction, which is not intuitive enough for children, especially when learning complex concepts. Moreover, they usually simply answer questions and lack in-depth analysis and records of children's learning process, and cannot provide systematic support for children's long-term growth. In addition, existing smart devices are mostly limited to their own small screens or voice output in content display, making it difficult to create an immersive learning environment. Summary of the invention
[0004] In view of this, the embodiments of the present application provide an interactive companionship method, device, electronic device and storage medium to solve the problems in the prior art that the interactive companionship data source is limited and the companionship method is single.
[0005] A first aspect of an embodiment of the present application provides an interactive companionship method, comprising:
[0006] In response to receiving a wake-up instruction or the interactive companion system detecting an autonomous wake-up condition, waking up the interactive companion system;
[0007] In response to receiving the first activity data of the companion object sent by the first monitoring device, the interactive companion system calls the second monitoring device to obtain the second activity data of the companion object, the first monitoring device and the second monitoring device each including at least one monitoring device, and the monitoring device includes a voice monitoring device and a visual monitoring device;
[0008] In response to determining that the first activity data and the second activity data satisfy a preset interaction condition, generating interaction data;
[0009] The interactive data is sent to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
[0010] In some embodiments, determining whether the first activity data and the second activity data satisfy a preset interaction condition includes:
[0011] fusing the first activity data and the second activity data to obtain fused activity data;
[0012] The pre-trained matching model is called to determine whether the fused activity data meets the preset interaction conditions.
[0013] In some embodiments, the first activity data and the second activity data are fused to obtain fused activity data, including:
[0014] Respectively obtaining voice monitoring data in the first activity data and the second activity data, parsing the voice monitoring data, and obtaining voice key information;
[0015] Respectively obtaining visual monitoring data in the first activity data and the second activity data, parsing the visual monitoring data, and obtaining visual key information;
[0016] The voice key information and the visual key information are fused to obtain fused activity data.
[0017] In some embodiments, calling a pre-trained matching model to determine whether the fused activity data meets a preset interaction condition includes:
[0018] In response to determining that the voice key information satisfies a preset voice priority response interaction condition, determining that the number of companion objects satisfies a preset interaction condition;
[0019] In response to determining that the voice key information does not satisfy the preset voice priority response interaction condition, and the visual key information satisfies the preset visual priority response interaction condition, determining that the number of companion objects satisfies the preset interaction condition;
[0020] In response to determining that the voice key information does not meet the preset voice priority response interaction conditions, and the visual key information does not meet the preset visual priority response interaction conditions, the fused activity data is input into a pre-trained matching model to determine whether the fused activity data meets the preset interaction conditions.
[0021] In some embodiments, generating interaction data includes:
[0022] Determining the identity of the companion based on the companion's activity data;
[0023] Generate interaction data based on the identity of the companion object and preset interaction conditions.
[0024] In some embodiments, sending the interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data includes:
[0025] Determine location information and posture information of the accompanying object based on the activity data of the accompanying object;
[0026] determining at least one interactive device based on the position information and posture information of the accompanying object;
[0027] The interactive data is sent to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
[0028] In some embodiments, sending the interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data includes:
[0029] Obtaining first to Nth sub-interaction data in the interaction data, where the first to Nth sub-interaction data are related to each other, and N is an integer greater than 1;
[0030] Determine location information and posture information of the accompanying object based on the activity data of the accompanying object;
[0031] Determine N interactive devices based on the location information and posture information of the accompanying object;
[0032] The interactive data is sent to N interactive devices, so that each interactive device performs an interactive operation corresponding to one of the first to Nth sub-interactive data.
[0033] A second aspect of an embodiment of the present application provides an interactive companion device, including:
[0034] A wake-up module, configured to wake up the interactive companion system in response to receiving a wake-up instruction or the interactive companion system detecting an autonomous wake-up condition;
[0035] The monitoring module is configured to, in response to receiving first activity data of a companion object sent by a first monitoring device, the interactive companion system calls a second monitoring device to obtain second activity data of the companion object, wherein the first monitoring device and the second monitoring device each include at least one monitoring device, and the monitoring device includes a voice monitoring device and a visual monitoring device;
[0036] A processing module configured to generate interaction data in response to determining that the first activity data and the second activity data satisfy a preset interaction condition;
[0037] The interactive module is configured to send interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
[0038] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0039] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0040] Compared with the prior art, the embodiments of the present application have the following beneficial effects: after the interactive companion system is awakened, if the first activity data of the companion object sent by the first monitoring device is received, the embodiment of the present application calls the second monitoring device to obtain the second activity data of the companion object, and generates interaction data when it is determined that the first and second activity data meet the preset interaction conditions, and then sends the interaction data to at least one interaction device so that each interaction device executes the interaction operation corresponding to the interaction data. The activity data of the companion object can be obtained by linking multiple different monitoring devices, and the interaction data can be determined after the data obtained by the linked multiple devices are integrated, and the interaction data can be executed in multiple different interaction devices, which enriches the information source when obtaining interactive information, improves the accuracy of the generated interaction data, provides diversified interaction methods, and enhances user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0042] Figure 1 It is a flow chart of an interactive companionship method provided in an embodiment of the present application.
[0043] Figure 2 It is a flowchart of a method for determining whether first activity data and second activity data meet preset interaction conditions provided in an embodiment of the present application.
[0044] Figure 3 It is a flowchart of a method for fusing first activity data and second activity data to obtain fused activity data provided by an embodiment of the present application.
[0045] Figure 4 It is a flowchart of a method provided in an embodiment of the present application for calling a pre-trained matching model to determine whether fused activity data meets preset interaction conditions.
[0046] Figure 5 It is a flowchart of a method for generating interactive data provided in an embodiment of the present application.
[0047] Figure 6 It is a flowchart of a method provided in an embodiment of the present application for sending interactive data to at least one interactive device so that each interactive device executes an interactive operation corresponding to the interactive data.
[0048] Figure 7It is a flowchart diagram of another method provided by an embodiment of the present application for sending interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
[0049] Figure 8 It is a schematic diagram of an interactive companion device provided in an embodiment of the present application.
[0050] Fig. 9 It is a schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0052] An interactive companionship method and device according to an embodiment of the present application will be described in detail below with reference to the accompanying drawings.
[0053] As mentioned above, traditional smart speakers mainly rely on voice interaction, which is not intuitive enough for children, especially when learning some complex concepts. For example, when a child asks a question, the smart speaker understands the question through voice recognition technology, then searches for the answer from the built-in knowledge base and responds to the child in voice form. This interaction method can only achieve voice questions and answers. For children, a single voice feedback may not be intuitive enough and is not conducive to their better understanding of the answer.
[0054] At the same time, traditional smart speakers usually simply answer questions, lack in-depth analysis and recording of children's learning process, and cannot provide systematic support for children's long-term growth. At the same time, the content display is mostly limited to its own small screen or voice output, making it difficult to create an immersive learning environment.
[0055] In view of this, an embodiment of the present application provides an interactive companionship method. After the interactive companionship system is awakened, if the first activity data of the companionship object sent by the first monitoring device is received, the second monitoring device is called to obtain the second activity data of the companionship object, and when it is determined that the first and second activity data meet the preset interaction conditions, interaction data is generated, and then the interaction data is sent to at least one interactive device so that each interactive device executes the interaction operation corresponding to the interaction data. The activity data of the companionship object can be obtained by linking multiple different monitoring devices, and the interaction data can be determined after the data obtained by the linked multiple devices are integrated, and the interaction data can be executed in multiple different interactive devices, which enriches the information source when obtaining interactive information, improves the accuracy of the generated interaction data, provides diversified interaction methods, and enhances user experience.
[0056] Figure 1 is a flow chart of an interactive companionship method provided in an embodiment of the present application. Figure 1 As shown, the interactive companionship method includes the following steps:
[0057] In step S101, in response to receiving a wake-up instruction or the interactive companion system detecting an autonomous wake-up condition, the interactive companion system wakes up.
[0058] In step S102, in response to receiving the first activity data of the companion object sent by the first monitoring device, the interactive companion system calls the second monitoring device to obtain the second activity data of the companion object.
[0059] Wherein, the first monitoring device and the second monitoring device each include at least one monitoring device, and the monitoring device includes a voice monitoring device and a visual monitoring device.
[0060] In step S103 , in response to determining that the first activity data and the second activity data satisfy a preset interaction condition, interaction data is generated.
[0061] In step S104, the interactive data is sent to at least one interactive device, so that each interactive device performs an interactive operation corresponding to the interactive data.
[0062] In some embodiments of the present application, the method may be performed by an interactive companion system. In one example, the interactive companion system may be set on a server or a terminal device with a certain processing capability.
[0063] In another example, the interactive companionship system may include at least one monitoring device and at least one interactive device, or the interactive companionship system may communicate with at least one third-party monitoring device to obtain activity data of the companionship object, and communicate with at least one third-party interactive device to send interactive data to the third-party interactive device, and enable the third-party interactive device to perform interactive operations based on the interactive data.
[0064] In some embodiments of the present application, the interactive companion system can be awakened when receiving a wake-up instruction, or when an autonomous wake-up condition is detected. The wake-up instruction can be, for example, a preset wake-up voice instruction, a preset wake-up gesture instruction, or other wake-up instructions, which are not limited here.
[0065] The autonomous wake-up conditions can be set according to actual needs. For example, when the companion is an infant, crying can be used as an autonomous wake-up condition. For another example, when the companion is a student, the preset time can be used as an autonomous wake-up adjustment. For another example, when the companion is an elderly person, actions such as falling can be used as autonomous wake-up conditions. It is understandable that the setting of autonomous wake-up conditions can also be adjusted according to actual needs, and there is no restriction here.
[0066] In certain embodiments of the present application, after being awakened, the interactive companion system can communicate with at least one monitoring device to receive activity data of the companion object sent by the at least one monitoring device. The monitoring device can be a voice monitoring device or a visual monitoring device, and each monitoring device monitors the companion object of the interactive companion system and sends the activity data of the companion object obtained by monitoring to the interactive companion system.
[0067] When the interactive companion system receives the first activity data of the companion object sent by the first monitoring device, the interactive companion system can call the second monitoring device to obtain the second activity data of the companion object. The first monitoring device can be a voice monitoring device or a visual monitoring device, and the first monitoring device can include one or more monitoring devices.
[0068] The second monitoring device may also include one or more monitoring devices. Moreover, the second monitoring device may be a monitoring device of the same type as the first monitoring device. For example, if the first monitoring device is a voice monitoring device, the second monitoring device may also be a voice monitoring device. Alternatively, if the first monitoring device is a visual monitoring device, the second monitoring device may also be a visual monitoring device.
[0069] The second monitoring device may also be a monitoring device of a different type from the first monitoring device. For example, if the first monitoring device is a voice monitoring device, the second monitoring device may be a visual monitoring device. Alternatively, if the first monitoring device is a visual monitoring device, the second monitoring device may be a voice monitoring device.
[0070] The interactive companion system can determine whether a preset interactive condition is met based on the first activity data and the second activity data. If so, interactive data corresponding to the interactive condition is generated. The interactive companion system can also send the interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
[0071] The interactive device may be, for example, a voice interactive device, in which case the interactive operation may be voice interaction. Alternatively, the interactive device may be an interactive device with a display screen, in which case the interactive operation may be voice and visual interaction, wherein the visual interaction may be one or more of text interaction, visual interaction and video interaction, and the interactive device with a display screen may be, for example, a television, a computer, a tablet computer, a mobile phone, an interactive robot with a screen, etc.
[0072] According to the technical solution provided in the embodiment of the present application, after the interactive companion system is awakened, if the first activity data of the companion object sent by the first monitoring device is received, the second monitoring device is called to obtain the second activity data of the companion object, and when it is determined that the first and second activity data meet the preset interaction conditions, interaction data is generated, and then the interaction data is sent to at least one interactive device so that each interactive device executes the interaction operation corresponding to the interaction data. The activity data of the companion object can be obtained by linking multiple different monitoring devices, and the interaction data can be determined after the data obtained by the linked multiple devices are integrated, and the interaction data can be executed in multiple different interactive devices, thereby enriching the information source when obtaining interactive information, improving the accuracy of the generated interaction data, providing diversified interaction methods, and enhancing user experience.
[0073] That is to say, the interactive companion system can be equipped with multiple distributed voice devices, which are distributed in different areas where the companion is active, such as bedrooms, living rooms, game areas, etc. These voice devices can, for example, use microphone array technology to capture the voice of the companion in an all-round and highly sensitive manner, ensuring accurate reception of voice signals no matter where the companion speaks.
[0074] Multiple voice devices can be linked seamlessly and share voice data in real time through network communication technology. When a voice device detects the voice of the companion, other devices can work together to enhance the sound and locate the direction, improving the accuracy and stability of voice recognition. For example, when the companion makes a sound in the bedroom, the voice device in the living room can also receive the signal and work with the bedroom device to determine the companion's location and voice content, so that the system can respond more promptly and accurately.
[0075] On the other hand, the interactive companion system can also deploy multiple high-definition video cameras at different locations to achieve all-round visual monitoring of the companion. The video camera can be a camera with wide-angle shooting, intelligent zoom and automatic tracking functions, which can capture the companion's movements, expressions and environment in real time.
[0076] The cameras can be linked and collaborated through intelligent algorithms. When one camera detects the movement or specific behavior of the companion, other cameras can automatically adjust the angle and focal length to achieve seamless switching and relay tracking, ensuring that the companion is always within the monitoring range. At the same time, the camera can also perform visual fusion processing, integrating visual information from multiple perspectives to provide the system with more comprehensive and accurate information on the status of the companion.
[0077] For example, when the companion object walks from one room to another, the cameras in different rooms can work together to achieve continuous tracking and shooting, and then use methods such as cross-camera tracking to extract the companion object's visual activity data, and then determine the companion object's behavior trajectory and activity pattern.
[0078] On the other hand, the interactive companion system can be connected to an AI (Artificial Intelligence) big model. The interactive companion system can send the activity data of the companion object received to the big model for processing, and use the big model to determine whether the activity data of the companion object meets the preset interaction conditions. If the big model determines that the activity data of the companion object meets the preset interaction conditions, it can generate interaction data. The big model can send the interaction data back to the interactive companion system, and the interactive companion system sends it to at least one interactive device, so that each interactive device performs the interaction operation corresponding to the interaction data.
[0079] That is, the big model can serve as the intelligent center of the interactive companion system. The big model can be trained using the historical data of multiple companions, covering many aspects of knowledge such as the companion's language pattern, cognitive development, and emotional expression. The big model has powerful natural language processing capabilities and can accurately understand the voice input of the companion, whether it is simple daily speech or complex question inquiries, and can quickly parse its semantics and intentions. For example, when the companion says "I want to hear a story about dinosaurs", the big model can understand the needs of the companion and select dinosaur stories suitable for the companion's age and cognitive level from the rich story library.
[0080] In addition, the large model can not only process voice information, but also deeply integrate with other modules. It can comprehensively analyze the status and needs of the companion according to the image information collected by the camera, the sound characteristics obtained by the voice device, and the interactive data fed back by the screen, and provide more personalized and accurate responses. For example, by analyzing the expression and concentration of the companion when watching the animation, the large model can adjust the playback speed of the animation or recommend related interactive content.
[0081] On the other hand, the interactive companion system can be connected to multiple different types of screens, including but not limited to smart TV screens, tablet screens, children's learning machine screens, companion robot screens, etc. These screens can be flexibly linked and displayed according to the activity scenes and needs of the companion objects.
[0082] The screens can be interconnected through wireless or wired communication technology. For example, when the companion is watching an educational video in the living room, the system can simultaneously display the relevant interactive content on the tablet screen, making it easier for the companion to operate and interact. Or when the companion is doing learning activities, multiple screens can simultaneously display teaching content from different angles, such as one screen showing text explanations and another screen showing animation demonstrations, to enhance the learning effect. At the same time, the screen can also be adaptively adjusted according to the companion's line of sight and concentration, providing a more comfortable viewing experience.
[0083] Figure 2 is a flow chart of a method for determining whether first activity data and second activity data satisfy a preset interaction condition provided by an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:
[0084] In step S201 , the first activity data and the second activity data are fused to obtain fused activity data.
[0085] In step S202, a pre-trained matching model is called to determine whether the fused activity data meets the preset interaction conditions.
[0086] In some embodiments of the present application, it is possible to determine whether a preset interaction condition is met based on the first activity data and the second activity data. In one example, the first activity data and the second activity data may be first fused to obtain fused activity data, and then a pre-trained matching model may be called to determine whether the fused activity data meets the preset interaction condition.
[0087] In other words, the interactive companion system can receive and integrate multimodal data from voice devices, camera devices and screens, and fuse the data. During the fusion process, the voice data can be converted into text information, and the features of the companion object's expression, action, scene, etc. in the image data can be analyzed for correlation. At the same time, combined with the screen interaction data, such as click operations, browsing history, etc., a comprehensive companion object behavior and demand model can be constructed.
[0088] For example, by analyzing the voice questions, facial expressions, and on-screen operations of the accompanying subject while watching an educational video, we can determine the subject's level of understanding and interest in the video content, thereby adjusting subsequent teaching strategies or recommending more suitable content.
[0089] At the same time, advanced data analysis algorithms and machine learning techniques can be used to conduct in-depth mining of the integrated multimodal data to extract key information such as the behavior patterns, interest preferences, learning progress, etc. of the companionship objects, providing data support for personalized companionship and education.
[0090] For example, by analyzing the daily voice communication content and question frequency of the companion, we can draw the trend of the companion's interest in specific fields (such as science, art, stories, etc.). Or according to the completion time and accuracy of the companion in different learning tasks, we can evaluate the learning ability and knowledge mastery of the companion, and provide a basis for formulating personalized learning plans.
[0091] Figure 3 is a flow chart of a method for fusing first activity data and second activity data to obtain fused activity data provided by an embodiment of the present application. Figure 3 As shown, the method comprises the following steps:
[0092] In step S301, voice monitoring data in the first activity data and the second activity data are respectively obtained, and the voice monitoring data are analyzed to obtain voice key information.
[0093] In step S302, visual monitoring data in the first activity data and the second activity data are respectively obtained, and the visual monitoring data are analyzed to obtain visual key information.
[0094] In step S303, the voice key information and the visual key information are fused to obtain fused activity data.
[0095] In some embodiments of the present application, when the first activity data and the second activity data are fused, the voice monitoring data in the first activity data and the second activity data can be obtained respectively, and the voice monitoring data can be parsed to obtain voice key information. At the same time, the visual monitoring data in the first activity data and the second activity data can be obtained respectively, and the visual monitoring data can be parsed to obtain visual key information. Finally, the voice key information and the visual key information are fused to obtain the fused activity data.
[0096] The operation of parsing the first activity data and the second activity data can be implemented by the interactive companion system. For example, a voice recognition model and a visual recognition model can be set in the interactive companion system to respectively recognize the voice information and visual information in the activity data to obtain voice key information and visual key information.
[0097] In some other embodiments, the operation of parsing the first activity data and the second activity data can also be implemented by a large model. For example, a preprocessing model can be set in the large model, and the preprocessing model includes a speech recognition model and a visual recognition model, which are used to recognize the speech information and visual information in the activity data respectively to obtain speech key information and visual key information.
[0098] Figure 4 1 is a flow chart of a method for calling a pre-trained matching model to determine whether the fused activity data meets the preset interaction conditions provided in an embodiment of the present application. Figure 4 As shown, the method comprises the following steps:
[0099] In step S401, in response to determining that the voice key information satisfies a preset voice priority response interaction condition, it is determined that the number of companion objects satisfies a preset interaction condition.
[0100] In step S402, in response to determining that the voice key information does not satisfy the preset voice priority response interaction condition, and the visual key information satisfies the preset visual priority response interaction condition, it is determined that the number of companion objects satisfies the preset interaction condition.
[0101] In step S403, in response to determining that the voice key information does not meet the preset voice priority response interaction conditions, and the visual key information does not meet the preset visual priority response interaction conditions, the fused activity data is input into a pre-trained matching model to determine whether the fused activity data meets the preset interaction conditions.
[0102] In certain embodiments of the present application, calling the matching model to determine whether the fused interaction data satisfies the preset interaction conditions can be to first determine whether the voice key information satisfies the preset voice priority response interaction conditions. If so, it is determined that the preset interaction conditions are currently met, and the preset interaction conditions at this time are the preset voice priority response interaction conditions.
[0103] The preset voice priority response interaction condition can be set according to actual needs and can include one or more interaction conditions. For example, if the companion is an infant, crying can be set as the preset voice priority response interaction condition. If the companion is a student, the preset voice priority response interaction condition can be set as detecting voice key information containing the keyword "learning".
[0104] On the other hand, if it is determined that the voice key information does not meet the preset voice priority response interaction conditions, but the visual key information meets the preset visual priority response interaction conditions, it can also be determined that the preset interaction conditions are currently met, and the preset interaction conditions at this time are the preset visual priority response interaction conditions.
[0105] The preset visual priority response interaction condition can be set according to actual needs and can include one or more interaction conditions. For example, if the companion is an infant or an elderly person, the preset visual priority response interaction condition can be set to detecting visual key information containing the keyword "fall".
[0106] On the other hand, if it is determined that the voice key information does not meet the preset voice priority response interaction conditions, and the visual key information does not meet the preset visual priority response interaction conditions, the fused activity data can be input into a pre-trained matching model, and the matching model determines whether the fused activity data meets the preset interaction conditions.
[0107] In other words, the big model also has the ability to process data in real time and can respond quickly to the immediate behavior and needs of the companion. When the companion issues a voice command or performs a specific operation, the system can analyze and process the data within milliseconds and provide timely feedback, such as answering the companion's questions, adjusting the screen display content, or initiating related interactive activities.
[0108] Figure 5 is a flow chart of a method for generating interactive data provided in an embodiment of the present application. Figure 5 As shown, the method comprises the following steps:
[0109] In step S501, the identity of the companion object is determined based on the activity data of the companion object.
[0110] In step S502, interaction data is generated based on the identity of the companion object and preset interaction conditions.
[0111] In certain embodiments of the present application, different interaction data may be set for different companion objects for the same preset interaction condition. That is, when generating interaction data, the companion object's identity may be first determined based on the companion object's activity data, for example, by voiceprint recognition or face recognition. Then, the generation model corresponding to the companion object's identity is called in the large model, and the corresponding interaction data is generated using the preset interaction condition.
[0112] In other words, the big model can create user profiles for each companion based on the interests, learning and other needs of different companions obtained through historical data analysis. Then, after determining the preset interaction conditions, it can match and recommend personalized interaction data for the companion in the content library, or call the generation model corresponding to the companion's identity to generate the corresponding interaction data.
[0113] The interactive data may be, for example, learning resources, entertainment content, and interactive activities, etc. For example, if the companion is determined to be interested in painting through the user portrait, interactive conditions are preset for entertainment, and interactive data related to painting tutorial videos, painting games, and related art appreciation can be generated.
[0114] In addition, the interactive data can be dynamically adjusted according to the growth stage and learning progress of the companion. As the companion grows and accumulates knowledge, the system will gradually recommend more challenging and in-depth content to promote the continuous learning and development of the companion.
[0115] Taking children as an example, the interactive companion system can include a variety of interactive learning courses and games, aiming to stimulate the learning interest and cultivate various abilities of the companions in an interesting way. For example, in language learning games, the companions can improve their language expression ability by talking with virtual characters and imitating pronunciation; mathematical thinking games can exercise the logical thinking and spatial cognition ability of the companions through solving puzzles and jigsaw puzzles.
[0116] The interactive companion system can provide real-time feedback and encouragement based on the performance of the companion during the interaction, enhancing the companion’s sense of participation and self-confidence. At the same time, by recording the companion’s behavioral data during learning and games, the interactive experience and teaching strategies can be further optimized.
[0117] At the same time, the interactive companion object can also comprehensively record the companion object's growth process, including voice development, behavioral changes, learning progress, etc., and save every moment of interaction between the companion object and the system in the form of data to form a detailed growth file.
[0118] In addition, we regularly evaluate the growth of the accompanying object and generate visual growth reports to show parents the development trends and achievements of the accompanying object in various fields. The reports also provide professional educational suggestions and development forecasts to help parents better understand the growth of the accompanying object and formulate reasonable education plans.
[0119] Figure 6 1 is a flow chart of a method for sending interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data provided by an embodiment of the present application. Figure 6 As shown, the method comprises the following steps:
[0120] In step S601, position information and posture information of the companion object are determined based on the activity data of the companion object.
[0121] In step S602, at least one interactive device is determined based on the position information and posture information of the accompanying object.
[0122] In step S603, the interactive data is sent to at least one interactive device, so that each interactive device performs an interactive operation corresponding to the interactive data.
[0123] In some embodiments of the present application, after determining the interactive data, the interactive companion system can determine the interactive device. In one example, the location information and posture information of the companion object can be determined based on the activity data of the companion object. For example, the location information of the companion object can be determined by voiceprint positioning or visual positioning. For another example, the posture information of the companion object can be determined by image recognition and other methods.
[0124] At least one interactive device may be determined using the determined position information and posture information. The determined interactive device may be an interactive device whose distance from the accompanying object is less than a preset distance threshold. Furthermore, if the interactive data includes voice interactive data, the determined interactive device needs to include an interactive device with a voice playback function. If the interactive data includes at least one of text, image, and video data, the determined interactive device needs to include an interactive device with a display function, and the device directly in front of the accompanying object is preferentially determined as the interactive device.
[0125] After determining the interactive devices, the interactive companion system can send the interactive data to each interactive device so that it can perform interactive operations corresponding to the interactive data.
[0126] Figure 7 1 is a flow chart of another method for sending interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data provided by an embodiment of the present application. Figure 7 As shown, the method comprises the following steps:
[0127] In step S701, first to Nth sub-interaction data in the interaction data are obtained.
[0128] The first to Nth sub-interaction data are associated with each other, and N is an integer greater than 1.
[0129] In step S702, the position information and posture information of the companion object are determined based on the activity data of the companion object.
[0130] In step S703, N interactive devices are determined based on the position information and posture information of the companion object.
[0131] In step S704, the interactive data is sent to N interactive devices, so that each interactive device performs an interactive operation corresponding to one of the first to Nth sub-interactive data.
[0132] In some embodiments of the present application, after determining the interactive data, the interactive companion system can determine multiple interactive devices and display interactive operations corresponding to different sub-interactive data in multiple interactive devices, thereby providing diversified interactive methods.
[0133] In one example, the interactive data may include multiple sub-interactive data, such as the first to Nth sub-interactive data. The sub-interactive data are interrelated. For example, if the interactive data is for teaching a certain content, the sub-interactive data of the interactive data may be a knowledge explanation video related to the content, exercises related to the content, etc.
[0134] The interactive companion system can determine the location information and posture information of the companion object based on the activity data of the companion object, and determine N interactive devices based on the location information and posture information. That is, the number of interactive devices determined is the same as the number of sub-interactive data items in the interactive data. Then, the interactive companion system sends the interactive data to the N interactive devices, and displays one sub-interactive data item in each interactive device.
[0135] Among them, the corresponding sub-interaction data and interaction devices can be determined according to the type and attributes of the sub-interaction data, as well as the type and attributes of each interaction device. For example, if the interaction data includes three sub-interaction data, the type of the first sub-interaction data is audio data, the type of the second sub-interaction data is video type and the attribute is explanation video, and the type of the third sub-interaction data is video type and the attribute is interactive video. At the same time, the three interaction devices determined include a TV, a tablet computer and a speaker, respectively. Then, the first sub-interaction data can be sent to the speaker, the second sub-interaction data can be sent to the TV, and the third sub-interaction data can be sent to the tablet computer, thereby providing rich interactive companionship for the companion object.
[0136] Still taking the example of a child as the companion, after the interactive companion system is awakened, the voice device and visual device are always in working state, monitoring the surrounding environment and behavior of the companion in real time. When the companion makes a sound or moves, the voice device quickly captures the voice signal, and the camera starts and adjusts the angle to shoot.
[0137] Voice data can be transmitted in real time to the data processing and analysis module of the interactive companion system for voice recognition and semantic understanding. Image data collected by the camera is also synchronously transmitted to the data processing and analysis module for image analysis and target recognition to identify the expression, action, location and other information of the companion object.
[0138] The data processing and analysis module integrates voice and image data for analysis, and combines it with the intelligent processing of the large model to determine the needs and intentions of the companion. For example, if the companion says "I'm hungry" and the camera captures the companion near the kitchen, the interactive companion system can infer that the companion may need food and provide relevant responses, such as asking the companion what he wants to eat or recommending some healthy snack options.
[0139] Based on the analysis results, the interactive companionship system displays relevant information on the screen or initiates corresponding interactive activities. For example, if the companion expresses the desire to play a game, the interactive companionship system will display the game interface on the appropriate screen and guide the companion to operate through voice. At the same time, the system will continue to monitor the companion's interactive process and adjust the feedback content in real time to ensure that the companion has a good companionship experience.
[0140] In addition, the interactive companion system develops a personalized learning plan based on the age and development stage of the companion. When it is time to study, the interactive companion system will guide the companion to carry out learning activities through voice and screen prompts. For example, for the companion at the early childhood stage, simple cognitive courses may be arranged, such as recognizing numbers, colors, shapes, etc.
[0141] During the learning process, multiple screens are linked to display learning content, such as one screen playing teaching videos and another screen showing related interactive exercises. The voice device will simultaneously provide voice explanations and guidance to help the accompanying person better understand and master the knowledge.
[0142] The companion can interact with the system by answering questions by voice or operating on the screen. The interactive companion system will evaluate the learning effect in real time and adjust the teaching content and difficulty according to the companion's answers and operations. If the companion encounters difficulties with a certain knowledge point, the interactive companion system will provide more detailed explanations and examples, or switch to a simpler teaching method until the companion understands.
[0143] After the study is completed, the interactive companion system will generate a study report to record the learning process and results of the companion, including learning time, knowledge points mastered, error rate, etc. Parents can view the report to understand the learning situation of the companion and adjust the subsequent study plan according to the system's suggestions.
[0144] At the same time, during the entire process of interaction between the companion object and the system, the data processing and analysis module will record various data in real time, including voice content, image data, screen interaction operations, timestamps, etc. These data are stored in a safe and reliable database for subsequent analysis.
[0145] In some implementations, the accumulated data can also be analyzed in depth on a regular basis (e.g., daily, weekly, or monthly). First, the raw data can be cleaned and preprocessed to remove noise and abnormal data to ensure the accuracy and completeness of the data. Then, data analysis algorithms and machine learning models are used to mine and analyze the data to extract key information such as the behavior patterns, interest preferences, and learning progress of the companion object.
[0146] Finally, based on the analysis results, the personalized model of the companion is updated and the service strategy of the system is adjusted. For example, if the companion is found to have a strong interest in music recently, the interactive companion system will add more music-related resources to the subsequent content recommendations and optimize the setting of music learning courses to better meet the needs of the companion.
[0147] The results of data recording and analysis can also be used for long-term growth assessment and trend analysis. By comparing data from different time periods, we can observe the development and changes of the accompanying object in various aspects, provide parents with comprehensive growth reports and professional educational suggestions, and help parents better accompany the growth of the accompanying object.
[0148] The technical solution provided in the embodiment of the present application realizes all-round and multi-angle real-time monitoring and interaction of the companion object through the linkage of multiple voice devices, camera devices and screens, providing the companion object with a seamless companionship and education experience. This multi-device collaborative working mode can obtain the information of the companion object more comprehensively, improve the interactive companion system's ability to perceive the companion object's needs and behaviors, use multi-dimensional perception to improve perception accuracy, and increase response speed.
[0149] The specially trained large model is used to process the voice and image information of the companion object to achieve intelligent understanding, personalized response and content recommendation. The performance and adaptability of the large model directly affect the interactive quality and educational effect of the interactive companion system. Through in-depth data analysis, the interactive content is personalized and the interactive experience is enriched.
[0150] Based on the analysis of the multimodal data of the companions, we provide each companion with a customized learning plan, content recommendations and interactive activities to meet the individual needs and interest development of the companions. At the same time, by establishing a complete security and privacy protection mechanism, we ensure the safety and reliability of hardware devices, the encryption and stability of network communications, and the privacy protection of data.
[0151] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present application, which will not be described one by one here.
[0152] The following is an embodiment of the device of the present application, which can be used to execute the embodiment of the method of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the method of the present application.
[0153] Figure 8 Schematic diagram of an interactive companion device provided in an embodiment of the present application. Figure 8 As shown, the device comprises:
[0154] The wake-up module 801 is configured to wake up the interactive companion system in response to receiving a wake-up instruction or the interactive companion system detecting an autonomous wake-up condition.
[0155] The monitoring module 802 is configured to respond to receiving first activity data of the companion object sent by the first monitoring device, and the interactive companion system calls the second monitoring device to obtain the second activity data of the companion object. The first monitoring device and the second monitoring device each include at least one monitoring device, and the monitoring device includes a voice monitoring device and a visual monitoring device.
[0156] The processing module 803 is configured to generate interaction data in response to determining that the first activity data and the second activity data meet a preset interaction condition.
[0157] The interaction module 804 is configured to send interaction data to at least one interaction device, so that each interaction device performs an interaction operation corresponding to the interaction data.
[0158] According to the technical solution provided in the embodiment of the present application, after the interactive companion system is awakened, if the first activity data of the companion object sent by the first monitoring device is received, the second monitoring device is called to obtain the second activity data of the companion object, and when it is determined that the first and second activity data meet the preset interaction conditions, interaction data is generated, and then the interaction data is sent to at least one interactive device so that each interactive device executes the interaction operation corresponding to the interaction data. The activity data of the companion object can be obtained by linking multiple different monitoring devices, and the interaction data can be determined after the data obtained by the linked multiple devices are integrated, and the interaction data can be executed in multiple different interactive devices, thereby enriching the information source when obtaining interactive information, improving the accuracy of the generated interaction data, providing diversified interaction methods, and enhancing user experience.
[0159] In some embodiments, determining whether the first activity data and the second activity data meet the preset interaction condition includes: fusing the first activity data and the second activity data to obtain fused activity data; calling a pre-trained matching model to determine whether the fused activity data meets the preset interaction condition.
[0160] In some implementations, the first activity data and the second activity data are fused to obtain fused activity data, including: respectively obtaining voice monitoring data in the first activity data and the second activity data, parsing the voice monitoring data to obtain voice key information; respectively obtaining visual monitoring data in the first activity data and the second activity data, parsing the visual monitoring data to obtain visual key information; and fusing the voice key information and the visual key information to obtain fused activity data.
[0161] In some embodiments, a pre-trained matching model is called to determine whether the fused activity data meets the preset interaction conditions, including: in response to determining that the voice key information meets the preset voice priority response interaction conditions, determining that the preset interaction conditions are met; in response to determining that the voice key information does not meet the preset voice priority response interaction conditions, and the visual key information meets the preset visual priority response interaction conditions, determining that the preset interaction conditions are met; in response to determining that the voice key information does not meet the preset voice priority response interaction conditions, and the visual key information does not meet the preset visual priority response interaction conditions, inputting the fused activity data into the pre-trained matching model to determine whether the fused activity data meets the preset interaction conditions.
[0162] In some implementations, generating the interaction data includes: determining the identity of the companion object based on the activity data of the companion object; and generating the interaction data based on the identity of the companion object and preset interaction conditions.
[0163] In some embodiments, interactive data is sent to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data, including: determining the position information and posture information of the companion object based on the activity data of the companion object; determining at least one interactive device based on the position information and posture information of the companion object; sending interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
[0164] In some embodiments, interactive data is sent to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data, including: obtaining first to Nth sub-interaction data in the interactive data, the first to Nth sub-interaction data are mutually related, and N is an integer greater than 1; determining the position information and posture information of the companion object based on the activity data of the companion object; determining N interactive devices based on the position information and posture information of the companion object; sending the interactive data to N interactive devices so that each interactive device performs an interactive operation corresponding to one sub-interaction data in the first to Nth sub-interaction data.
[0165] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0166] Fig. 9 Schematic diagram of an electronic device provided in an embodiment of the present application. Fig. 9 As shown, the electronic device 9 of this embodiment includes: a processor 901, a memory 902, and a computer program 903 stored in the memory 902 and executable on the processor 901. When the processor 901 executes the computer program 903, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 901 executes the computer program 903, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0167] The electronic device 9 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 9 may include, but is not limited to, a processor 901 and a memory 902. Those skilled in the art will appreciate that Fig. 9 The electronic device 9 is merely an example and does not limit the electronic device 9 , and may include more or less components than those shown in the figure, or different components.
[0168] The processor 901 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0169] The memory 902 may be an internal storage unit of the electronic device 9, for example, a hard disk or memory of the electronic device 9. The memory 902 may also be an external storage device of the electronic device 9, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 9. The memory 902 may also include both an internal storage unit of the electronic device 9 and an external storage device. The memory 902 is used to store computer programs and other programs and data required by the electronic device.
[0170] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0171] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0172] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An interactive companionship method, characterized in that: include: In response to receiving a wake-up instruction or the interactive companion system detecting an autonomous wake-up condition, waking up the interactive companion system; In response to receiving first activity data of a companion object sent by a first monitoring device, the interactive companion system calls a second monitoring device to obtain second activity data of the companion object, wherein the first monitoring device and the second monitoring device each include at least one monitoring device, and the monitoring device includes a voice monitoring device and a visual monitoring device; In response to determining that the first activity data and the second activity data satisfy a preset interaction condition, generating interaction data; The interactive data is sent to at least one interactive device, so that each interactive device performs an interactive operation corresponding to the interactive data.
2. The method according to claim 1, characterized in that: Determining whether the first activity data and the second activity data meet a preset interaction condition includes: fusing the first activity data and the second activity data to obtain fused activity data; A pre-trained matching model is called to determine whether the fused activity data meets a preset interaction condition.
3. The method according to claim 2, characterized in that The fusing the first activity data and the second activity data to obtain the fused activity data includes: Respectively obtaining voice monitoring data in the first activity data and the second activity data, parsing the voice monitoring data, and obtaining voice key information; Respectively acquiring visual monitoring data in the first activity data and the second activity data, and analyzing the visual monitoring data to obtain key visual information; The voice key information and the visual key information are fused to obtain fused activity data.
4. The method according to claim 3, characterized in that The calling of the pre-trained matching model to determine whether the fused activity data meets the preset interaction condition includes: In response to determining that the voice key information satisfies a preset voice priority response interaction condition, determining that a preset interaction condition is satisfied; In response to determining that the voice key information does not satisfy the preset voice priority response interaction condition, and the visual key information satisfies the preset visual priority response interaction condition, determining that the preset interaction condition is satisfied; In response to determining that the voice key information does not meet the preset voice priority response interaction condition, and the visual key information does not meet the preset visual priority response interaction condition, the fused activity data is input into a pre-trained matching model to determine whether the fused activity data meets the preset interaction condition.
5. The method according to claim 1, characterized in that The generating of interactive data comprises: determining the identity of the companion object based on the activity data of the companion object; The interaction data is generated based on the identity of the companion object and the preset interaction conditions.
6. The method according to claim 1, characterized in that The sending the interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data includes: Determining the position information and posture information of the companion object based on the activity data of the companion object; Determine at least one interactive device based on the position information and posture information of the companion object; The interactive data is sent to at least one interactive device, so that each interactive device performs an interactive operation corresponding to the interactive data.
7. The method according to claim 1, characterized in that The sending the interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data includes: Acquire first to Nth sub-interaction data in the interaction data, where the first to Nth sub-interaction data are related to each other, and N is an integer greater than 1; Determining the position information and posture information of the companion object based on the activity data of the companion object; Determine N interactive devices based on the location information and posture information of the companion object; The interactive data is sent to the N interactive devices, so that each interactive device performs an interactive operation corresponding to one of the first to Nth sub-interactive data.
8. An interactive companion device, characterized in that: include: A wake-up module, configured to wake up the interactive companion system in response to receiving a wake-up instruction or the interactive companion system detecting an autonomous wake-up condition; The monitoring module is configured to, in response to receiving first activity data of a companion object sent by a first monitoring device, the interactive companion system calls a second monitoring device to obtain second activity data of the companion object, wherein the first monitoring device and the second monitoring device each include at least one monitoring device, and the monitoring device includes a voice monitoring device and a visual monitoring device; A processing module, configured to generate interaction data in response to determining that the first activity data and the second activity data satisfy a preset interaction condition; The interactive module is configured to send the interactive data to at least one interactive device so that each interactive device performs an interactive operation corresponding to the interactive data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Intelligent dialogue method and system for accompanying old people based on user portraits
CN120998202A