Multimedia operation service system and method based on digital avatar
Through the combination of digital clone intelligence and monitoring control, the problems of high threshold for use, large safety hazards and strong dependence for the elderly when using digital devices are solved, barrier-free and safe multimedia operation services are achieved, and the convenience and security of digital life of the elderly are improved.
Patent Information
- Application Number
- CN202510863870.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The elderly, teenagers, etc. have problems such as high barriers to use, high security risks, strong dependence on digital services, and susceptible to online fraud when using digital devices. The existing technology is difficult to provide barrier-free and secure multimedia operation services.
The multimedia operation service system based on digital clones is adopted, and through the digital clones and the monitoring control terminal combined with the cloud server, it provides three-dimensional virtual characters for voice interaction and operation agents, and integrates UI rendering, voice interaction, dialogue management, security control and data acquisition encryption modules to realize barrier-free interaction and security protection between the elderly and digital devices.
It lowers the threshold for the elderly to use digital devices, enhances the sense of security and trust, reduces the burden of remote assistance for guardians, and realizes multi-functional integrated services such as information accessibility, life services, entertainment care, health management and safety protection.
Smart Images

Figure CN120358367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimedia operation service, and more specifically, to a multimedia operation service system based on digital avatars. Background Art
[0002] In recent years, the industry has made many efforts to address the digital divide among the elderly. For example, in 2021, the Ministry of Industry and Information Technology launched the "Special Action for the Adaptation of Internet Applications to the Aging and Accessibility" and issued the "General Design Specifications for the Adaptation of Mobile Internet Applications to the Aging", requiring the adjustment of font size and line spacing in the elderly-friendly application interface, and strictly prohibiting interference such as pop-up ads. Major Internet applications have also launched "elder mode" or "care mode", which are characterized by enlarged fonts and buttons, simplified interfaces, and voice broadcasts and one-click operations to lower the threshold for use. These aging-friendly renovations make it easier for some elderly users to watch news, pay bills, call taxis, shop online, etc. However, these improvements are mainly focused on interface optimization and operation simplification, and are still not enough to fully address the needs of the elderly, young people, etc. in terms of trust, security, and comprehensive services.
[0003] In the cultural context of China, families play an important role in helping the elderly, young people, etc. to cross the digital divide. The phenomenon of "digital feeding back" shows that nearly 70% of the elderly were taught to use digital applications such as WeChat by their children and younger generations at home; when the elderly, young people, etc. encounter difficulties in using digital devices, nearly 90% of the first people to seek help are family members. For example, in order to prevent their parents from being deceived online, many children often choose to directly handle various online affairs on their behalf, rather than teaching their parents to operate themselves in a time-consuming and laborious manner. For example, children register online, buy tickets, take taxis, order meals, etc. for the elderly, young people, etc., and help their elders complete the necessary matters in digital life in an agent way. Although this approach is efficient, it also means that the elderly, young people, etc. rely on others for digital services and their autonomy is limited. In particular, when their children are not around, the elderly will feel at a loss. What is more serious is that the elderly and young people often face high risks of untrue telecommunications and untrue information after going online. Scammers and bad information often target the elderly group who are relatively lacking in media literacy. Data shows that in the fraud cases targeting middle-aged and elderly netizens in 2024, more than 97% of the victims, such as the elderly, eventually suffered property losses, which is much higher than that of young people. Some lawless elements take advantage of the lack of network knowledge and vigilance of the elderly and young people, pretending to be customer service, acquaintances, and even using AI to pretend to be the children of the elderly and young people to commit fraud, causing a large number of tragedies. These phenomena reflect that the elderly and young people have strong emotional needs in the digital world, but also have serious security risks and trust blind spots.
[0004] Therefore, there is an urgent need for a digital avatar-based multimedia operation service system and method that can access digital services without barriers and safely, reduce the usage threshold, enhance the sense of trust and security, and at the same time reduce the burden of remote assistance for guardians. Summary of the Invention
[0005] In view of the above problems, the purpose of the present invention is to provide a digital avatar-based multimedia operation service system to solve the problems that the elderly, teenagers, and disabled groups currently have a demand for multimedia functions such as payment, reading, and checking the weather on electronic products, but are highly dependent on guardians and are easily deceived by online fraud.
[0006] A digital avatar-based multimedia operation service system provided by the present invention includes a digital avatar intelligent agent deployed on a first client, a guardianship control terminal deployed on a second client, and a cloud server communicatively connected to the digital avatar intelligent agent and the guardianship control terminal; wherein, The digital avatar intelligent agent includes a three-dimensional virtual character communicatively connected to media software and the home, and a function module for providing function services for the three-dimensional virtual character; wherein, the function module is used to configure the three-dimensional virtual character and locally process the voice information and click request information received by the three-dimensional virtual character to obtain a request signal; The cloud server is used to generate a feedback signal for the request signal and instruct the three-dimensional virtual character to feedback the feedback signal to the digital avatar intelligent agent or the guardianship control terminal; The guardianship control terminal is used to monitor and remotely control the communication data of the three-dimensional virtual character in real time.
[0007] Preferably, the function module includes a UI rendering module, a voice interaction module, a dialogue management and local logic module, a security control module, and a data collection and encryption module; wherein, The UI rendering module is used to perform real-time rendering of a 3D model based on a graphics engine and a preset digital human image element library and interface element library to form a real-time changing digital human three-dimensional picture, and perform voice dubbing for the digital human three-dimensional picture based on pre-recorded video clips and pre-acquired voice clips to form a digital avatar intelligent agent; The voice interaction module is used to continuously monitor the user's daily voice by calling the microphone, summon the three-dimensional virtual character by comparing the daily voice with the wake-up word pre-stored locally, collect the command voice issued by the user to the three-dimensional virtual character, convert the command voice into a text to generate a request signal, send the request signal to the cloud server, receive the feedback signal returned by the cloud server, and convert the feedback signal into an output voice to instruct the three-dimensional virtual character to play the output voice, or display the output text corresponding to the output voice in large font on the display of the first client; wherein, the user is the user of the first client. The dialogue management and local logic module is used to match the local stored corpus according to the daily voice. If a stored corpus is matched, the local action corresponding to the stored corpus is executed. And, if the stored corpus is a to-do corpus, a dialogue is initiated to the guardianship control terminal, and the guardianship control terminal is instructed to handle business for the user. The security control module is used to filter the data received by the first client based on the preset control rules and the preset blacklist, and perform access restrictions on the first client based on the preset restriction instructions. The data collection and encryption module is used to encrypt the request signal and the feedback signal, capture sensitive data based on the preset collection rules, and directly transmit the encrypted sensitive data to the guardianship control terminal.
[0008] Preferably, in the digital avatar intelligent agent, the actions of the digital human three-dimensional picture are matched with the content of the dubbing in real time; wherein, The digital human three-dimensional picture is the AI image of the user's guardian; The dubbing configured for the AI image is the guardian voice preset by the guardian; The three-dimensional picture of the AI image is a 3D stereoscopic digital human picture generated according to the guardian's photo.
[0009] Preferably, the cloud server includes a pre-trained large model, a database and management platform, a digital human control module, and a monitoring module; wherein, The large model is used to convert the request signal into a request text, perform natural language recognition on the request text based on the natural language understanding model to parse the request intention, and determine the response purpose according to the request intention, the preference data stored in the database and management platform, or the monitoring logs stored in the log database; make a feedback signal according to the response purpose; the feedback signal is feedback data or a feedback call instruction. The digital human control module is used to feedback the feedback data to the voice interaction module, or remotely control the media software and home appliances according to the feedback call instruction; The monitoring module is used to perform risk analysis on all feedback data and feedback call instructions based on a preset risk analysis model to obtain the user's evaluation status data, perform health analysis on the user's physiological data transmitted by the first client to obtain health status data, and edit the evaluation status data and the health status data into a monitoring log and transmit it to the guardianship control terminal, and synchronously store the monitoring log in the log database; The database and management platform are used to perform big data analysis on the monitoring log to obtain preference texts, and summarize and store the preference texts and the set data preset by the user to form preference data.
[0010] Preferably, the feedback data is data adapted to the response purpose based on the media software and home appliances and the database based on the large model.
[0011] Preferably, the media software and home appliances include social software, weather software, news software, music software, e-commerce software, payment software, smart home cloud platform, and medical and health platform.
[0012] Preferably, the guardianship control terminal logs in based on an application or a web page; the function interface of the guardianship control terminal includes a real-time monitoring panel, a permission and filtering setting panel, and a communication and agency panel; among them, The real-time monitoring panel is used to receive the monitoring log to obtain the operation information of the user on the first client; The permission and filtering setting panel is used to set access permissions for the digital avatar intelligent agent to control the user's usage space on the first client, used to set a blacklist for the digital avatar intelligent agent to perform social filtering on the first client, and is also used to set a payment limit for the payment software in the first client to control the user's payment amount on the first client; The communication and agency panel is used to perform business agency for the user according to the agency corpus, and after the agency is completed, instruct the digital avatar intelligent agent to broadcast to the user or send a prompt message indicating the completion of the agency to the user.
[0013] The present invention also provides a multimedia operation service method based on a digital avatar. Among them, a multimedia service with targeted restrictions is performed based on the multimedia operation service system based on a digital avatar as described above, including: Configure a three-dimensional virtual character through a preset function module, and locally process the voice information and click request information received by the three-dimensional virtual character to obtain a request signal; the function module and the three-dimensional virtual character are integrated in a preset digital avatar intelligent agent; Make a feedback signal for the request signal through a preset cloud server, and instruct the three-dimensional virtual character to feedback the feedback signal to the digital avatar intelligent agent or a preset guardianship control terminal; Monitor and remotely control the communication data of the three-dimensional virtual character through the guardianship control terminal.
[0014] Preferably, configuring a three-dimensional virtual character and locally processing the voice information and click request information received by the three-dimensional virtual character to obtain a request signal includes: Perform real-time rendering of a 3D model based on a graphics engine and a preset digital human image element library and interface element library to form a digitally changing three-dimensional human picture, and dub the digitally changing three-dimensional human picture based on pre-recorded video clips and pre-acquired voice clips to form a digital avatar intelligent agent; Call the microphone to continuously monitor the daily voice of the user, summon the three-dimensional virtual character by comparing the daily voice with a locally pre-stored wake-up word, collect the command voice issued by the user to the three-dimensional virtual character, and perform text conversion on the command voice to generate a request signal; wherein, the user is the user of the digital avatar intelligent agent.
[0015] Preferably, making a feedback signal for the request signal and instructing the three-dimensional virtual character to feedback the feedback signal to the digital avatar intelligent agent or the guardianship control terminal includes: Send the request signal to the cloud server, so that the cloud server converts the request signal into a request text, performs natural language recognition on the request text based on a natural language understanding model to parse the request intention, and determines the response purpose according to the request intention, preference data stored in the database and management platform or monitoring logs stored in the log database; make a feedback signal according to the response purpose; the feedback signal is feedback data or a feedback call instruction; Feed the feedback data back to the voice interaction module of the digital avatar intelligent agent, or remotely control the media software and home appliances connected to the digital avatar intelligent agent according to the feedback call instruction; Based on a preset risk analysis model, perform risk analysis on all feedback data and feedback call instructions to obtain the user's evaluation status data, perform health analysis on the user's physiological data transmitted by the first client to obtain health status data, and edit the evaluation status data and the health status data into a monitoring log and transmit it to the guardianship control terminal, and synchronously store the monitoring log in the log database; Perform big data analysis on the monitoring log to obtain preference texts, and summarize and store the preference texts and the set data preset by the user to form preference data.
[0016] As can be seen from the above technical solutions, the multimedia operation service system and method based on digital avatars provided by the present invention include a digital avatar intelligent agent, which includes a three-dimensional virtual character that communicatively connects media software and a home, and a function module that provides functional services for the three-dimensional virtual character; among them, the three-dimensional virtual character is configured through the function module, and the voice information and click request information received by the three-dimensional virtual character are locally processed to obtain a request signal, and a feedback signal is made for the request signal through the cloud server, and the three-dimensional virtual character is instructed to feedback the feedback signal to the digital avatar intelligent agent or the guardianship control terminal.
[0017] In this way, the guardianship control terminal can perform real-time monitoring and remote control on the communication data of the three-dimensional virtual character, thereby breaking through the unfamiliarity of traditional machine assistants. The digital avatar image of the elderly's children is creatively used as an interaction agent, and the natural trust of the elderly in the parent-child relationship is utilized to solve the problems of low user stickiness and low trust, making the elderly more willing to accept and use digital technology services subjectively; Children can remotely configure and intervene in the behavior of the intelligent agent to achieve appropriate supervision of the digital life of the elderly, especially in preventing fraud and security, and timely intervene. This solution that integrates the guardian's authority into the intelligent assistant enables the digital assistant to fully exert its intelligence and always be within the controllable range; Moreover, integrating information accessibility, life services, entertainment care, health management, safety protection and other aspects into one system, multiple needs of the elderly can be met through a digital avatar interface, avoiding the cumbersome process of switching between numerous apps and devices. This all-round integrated service concept greatly improves convenience. Brief Description of the Drawings
[0018] By referring to the following description of the specification in conjunction with the drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the drawings: Figure 1 It is a logic block diagram of a multimedia operation service system based on digital avatars according to an embodiment of the present invention; Figure 2A scenario example of generating a digital twin agent for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 3 A detailed function example of a digital twin agent for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 4 A scenario example of querying the weather for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 5 A scenario example of speech recognition for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 6 A scenario example of anti-fraud for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 7 An example of an external device for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 8 A function example of emergency event handling for a multimedia operation service system based on digital twins according to an embodiment of the present invention; Figure 9 A flowchart of a multimedia operation service method based on digital twins according to an embodiment of the present invention. Detailed implementation manners
[0019] Currently, the elderly population, teenagers, and disabled groups have needs for multimedia functions such as payment, reading, and checking the weather on electronic products, but they are highly dependent on guardians and are easily deceived by online fraud.
[0020] In view of the above problems, the present invention provides a multimedia operation service system and method based on digital twins, and the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0021] To illustrate the multimedia operation service system and method based on digital twins provided by the present invention, Figure 1 - FIG. 9 gives an exemplary illustration of an embodiment of the present invention.
[0022] The following description of the exemplary embodiments is merely illustrative in nature and is in no way intended to limit the present invention, its application, or its use. Technologies and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies and devices should be considered as part of the specification.
[0023] As Figure 1 、 Figure 2As shown together, the present invention provides a multimedia operation service system 100 based on a digital avatar, which includes a digital avatar intelligent agent 110 deployed on a first client, a guardianship control terminal 120 deployed on a second client, and a cloud server 130 communicatively connected to the digital avatar intelligent agent and the guardianship control terminal; wherein, the first client can be an elderly mobile phone terminal, which can be implemented by an elderly mobile phone, and the elderly mobile phone can be an elderly smart phone or an ordinary smart phone; the second client is a children's terminal, and the children's terminal can also be referred to as a children's remote terminal; The digital avatar intelligent agent 110 includes a three-dimensional virtual character 111 communicatively connected to media software and a home, and a function module 112 for providing functional services for the three-dimensional virtual character; wherein, the function module 112 is used to configure the three-dimensional virtual character, and locally process the voice information and click request information received by the three-dimensional virtual character to obtain a request signal; The cloud server (cloud service) 130 is used to make a feedback signal for the request signal, and instruct the three-dimensional virtual character to feedback the feedback signal to the digital avatar intelligent agent 110 or the guardianship control terminal 120; The guardianship control terminal 120 is used to monitor and remotely control the communication data of the three-dimensional virtual character in real time.
[0024] Specifically, in Figure 1 、 Figure 2 In the embodiment shown together, the function module 112 includes a UI rendering module 1121, a voice interaction module 1122, a dialogue management and local logic module 1123, a security control module 1124, and a data collection and encryption module 1125; wherein, The UI rendering module 1121 is used to perform real-time rendering of a 3D model based on a graphics engine and a preset digital human image element library and interface element library to form a real-time changing digital human three-dimensional picture, and dub the digital human three-dimensional picture based on pre-recorded video clips and pre-acquired voice clips to form a digital avatar intelligent agent; The voice interaction module 1122 is used to continuously monitor the user's daily voice by calling a microphone, summon the three-dimensional virtual character by comparing the daily voice with a locally pre-stored wake-up word, collect the command voice issued by the user to the three-dimensional virtual character, perform text conversion on the command voice to generate a request signal, send the request signal to the cloud server, receive the feedback signal returned by the cloud server, and convert the feedback signal into output voice to instruct the three-dimensional virtual character to play the output voice, or display the output text corresponding to the output voice in large font on the display of the first client; wherein, the user is the user of the first client; The dialogue management and local logic module 1123 is used to match the local stored corpus according to the daily speech. If the stored corpus is matched, the corresponding local action is executed. Moreover, if the stored corpus is a to-do corpus, a dialogue is initiated to the guardianship control terminal, and the guardianship control terminal is instructed to handle business for the user; The security control module 1124 is used to filter the data received by the first client based on the preset control rules and the preset blacklist, and perform access restrictions on the first client based on the preset restriction instructions; The data acquisition and encryption module 1125 is used to encrypt the request signal and the feedback signal, capture sensitive data based on the preset acquisition rules, and directly transmit the encrypted sensitive data to the guardianship control terminal.
[0025] Among them, in this digital avatar intelligent agent, the actions of the digital human three-dimensional picture are matched with the content of the dubbing in real time; among them, The digital human three-dimensional picture is the AI image of the user's guardian; The dubbing configured for the AI image is the guardian voice preset by the guardian; The three-dimensional picture of the AI image is a 3D stereoscopic digital human picture generated according to the guardian's photo.
[0026] In a specific embodiment, the first client is the main place where the multimedia operation service system based on the digital avatar runs, including the display interface of the digital avatar intelligent agent and the local interaction logic. The digital avatar intelligent agent is integrated in the first client and is responsible for directly interacting with the user, including rendering the image (animation or video) of the child digital avatar, playing voice, listening to the voice commands of the elderly, and touch operations, etc. The digital avatar intelligent agent includes a communication-connected media software and a three-dimensional virtual character of the home, and a function module that provides function services for the three-dimensional virtual character; this function module has the following function modules: UI rendering module: Present the digital human image and related interface elements based on the graphics engine. It can adopt the method of real-time rendering of 3D models or combining pre-recorded video clips with voice synthesis. The interface layout is simple and clear, and necessary auxiliary graphics are provided (such as large icons with high contrast, guiding arrows).
[0027] Voice interaction module: Call the microphone to continuously listen to the elderly's speech, activate voice recognition through local wake-up word detection, and upload the voice command or inquiry to the cloud for recognition / understanding. It can also use the built-in offline instruction set in a weak network environment. After the recognized instruction is processed by the dialogue management, it is reported back by the text-to-speech TTS module in the voice of the child.
[0028] Dialogue Management and Local Logic Module: A lightweight dialogue manager is responsible for handling common instructions and maintaining the dialogue state. Commonly used corpora and rules are stored locally (for example, when the elderly person says "I want to see my child's photos", the photos of the children can be directly found in the local photo library and displayed), which improves the response speed. The local logic is also responsible for some agile responses, such as immediately triggering a pre-plan locally in case of an emergency (for example, when detecting an alarm from the elderly person's fall sensor, immediately ask by voice and notify the cloud).
[0029] Security Control Module: Basic information filtering is implemented on the mobile phone side, such as using the local phone number blacklist to intercept obvious fraud calls and using keywords to block sensitive text messages. At the same time, monitor system permission calls to prevent unauthorized applications from accessing sensitive data. Under the control of the policy rules issued by the cloud, this module enforces various permission constraints set for the children.
[0030] Data Collection and Encryption: After the mobile phone side collects the interaction data of the elderly person (voice content, operation logs, health sensor readings, etc.) and performs preliminary processing, it uploads them to the cloud for storage / analysis through an encrypted channel. Sensitive data (such as video call content) can be encrypted point-to-point and directly transmitted to the children's devices without passing through the cloud relay to ensure personal information security.
[0031] In this embodiment, the cloud server 130 includes a pre-trained large model 131, a database and management platform 132, a digital human control module 133, and a monitoring module 134; among them, The large model 131 is used to convert the request signal into a request text, perform natural language recognition on the request text based on a natural language understanding model to parse the request intention, and determine the response purpose according to the request intention, preference data stored in the database and management platform 132, or monitoring logs stored in the log database; make a feedback signal according to the response purpose; the feedback signal is feedback data or a feedback call instruction; The digital human control module 133 is used to feedback the feedback data to the voice interaction module, or remotely control the media software and home appliances according to the feedback call instruction; The monitoring module 134 is used to perform risk analysis on all feedback data and feedback call instructions based on a preset risk analysis model to obtain the user's evaluation status data, perform health analysis on the user's physiological data transmitted from the first client to obtain the health status data, and edit the evaluation status data and the health status data into a monitoring log and transmit it to the guardianship control end, and synchronously store the monitoring log in the log database; The database and management platform 132 is used to perform big data analysis on the monitoring log to obtain preference texts, and summarize and store the preference texts and the set data preset by the user to form preference data.
[0032] The feedback data is data adapted to the response purpose and made based on the media software and home and based on the database of the large model.
[0033] Media software and home include social software, weather software, news software, music software, e-commerce software, payment software, smart home cloud platform, medical and health platform, etc.
[0034] In other words, you can use voice through the digital avatar intelligent body to control social software, weather software, news software, music software, e-commerce software, payment software, smart home cloud platform, medical and health platform and many other functional software, as well as many home appliances in the home.
[0035] In a specific embodiment, the cloud serves as the brain and center of the system, taking on complex AI computing, data storage and remote communication functions. The cloud architecture is the cloud server, which can adopt a microservice approach, and the main modules include: Large model, including: speech recognition and natural language understanding model, that is, using the powerful AI model in the cloud to convert the uploaded elderly voice into text, and perform natural language understanding (NLU) to analyze the elderly's intentions. Combine the context and the elderly's personal preferences (stored in the user model database) to generate corresponding replies or execution instructions; and dialogue and decision-making AI engine, which is the brain of the digital avatar intelligent body. Based on the dialogue system and reasoning engine, it is responsible for deciding how to respond to the elderly's requests and what actions to take. The engine can be fine-tuned by a large pre-trained language model, with a certain open dialogue capability (for chatting and companionship), and a built-in business logic tree (for handling specific affairs). The engine will refer to the policies and permissions set by the children, and perform security checks before answering (such as if the elderly ask to buy an item, first check whether it is within the permitted scope).
[0036] The digital human control module, also known as the digital human rendering and speech synthesis service, is used for: When the cloud conversation engine generates a text reply or requires the digital avatar to express specific emotions and actions, the service is responsible for converting the text into a child's voice (TTS) and generating the corresponding digital human expression / lip animation data. If the mobile phone has sufficient performance, the TTS and expression generation model can be decentralized to the client for execution, otherwise the cloud generates audio and animation and then packages it for delivery. Through cloud generation, more complex models can be used to make the digital avatar's voice and facial expressions more natural and realistic.
[0037] The monitoring module, also known as the monitoring and anomaly detection module, is used for: continuously analyzing various data uploaded by the elderly terminal in the cloud, and identifying abnormal patterns through machine learning models. For example, real-time transcription analysis of call voices to check for fraud keywords, whether browsing behavior suddenly deviates from the norm, whether health data exceeds the standard, etc. Once a preset rule is triggered or the AI model determines a high risk, it immediately notifies the children's terminal through the message push service, and can instruct the digital avatar on the elderly terminal to take preliminary measures (such as suspending suspicious operations). The algorithm model of this module will comprehensively judge by combining the authoritative fraud number database, phishing website blacklist, and training data on common abnormal behaviors of elderly users, aiming for high accuracy and low false alarm rates.
[0038] Database and management platform: The cloud sets up a user database (storing basic information of the elderly, preferences, children's permission configurations, etc.), a log database (recording interactions and events), and a health record database, etc. The management platform is provided for children and maintenance personnel to view the status of the elderly, adjust configurations, and perform system operation and maintenance management. The platform has fine-grained permission control, and only authorized children's accounts can access the data of the corresponding elderly, and multi-factor authentication is used to ensure security.
[0039] Content service integration module: To implement various life service functions, the cloud integrates multiple third-party API interfaces (weather, news, music, e-commerce, payment, smart home cloud platform, medical and health platform, etc.). When the dialogue engine determines that an external service needs to be called (such as querying the weather, purchasing goods, adjusting the air conditioner at home), it interacts with the third party through this module to obtain or issue instructions, and then returns the results to the dialogue engine to organize the language and feedback to the elderly. The module can also make personalized recommendations based on the elderly's historical usage data (such as entertainment content).
[0040] In this embodiment, the guardianship control terminal 120 logs in based on an application or a web page; the function interface of the guardianship control terminal 120 includes a real-time monitoring panel 121, a permission and filtering setting panel 122, a communication agency panel 123, etc.; among them, The real-time monitoring panel 121 is used to receive the monitoring log to obtain the operation information of the user on the first client; The permission and filtering setting panel 122 is used to set access permissions for the digital avatar agent to control the usage space of the user on the first client, to set a blacklist for the digital avatar agent to perform social filtering on the first client, and to set a payment limit for the payment software in the first client to control the payment amount of the user on the first client; The communication agency panel 123 is used to perform business agency for the user according to the agency corpus, and after the agency is completed, it instructs the digital avatar agent to broadcast to the user or send a reminder message indicating that the agency is completed.
[0041] Specifically, children and the like can log in to the system through a dedicated mobile App, mobile mini-program, or web page as the child end (the second client) to remotely assist and manage the digital avatar agent of the elderly's mobile phone (the digital avatar agent of the first client). Its function interface includes: Real-time monitoring panel: Displays the key status of the elderly (online / offline, last activity, summary of health indicators, etc.) and any alarm notifications. Children can click on a specific alarm to view the details and choose to call the elderly with one click (dial video / audio to the elderly through the agent, appearing in the image of the child himself) or send instructions to the agent (such as asking the agent to temporarily disable a certain function).
[0042] Permissions and filtering settings panel: Children can adjust the permission policies of the agent on this interface, such as adding new blacklist numbers / words, modifying consumption limits, adding a list of trusted contacts (the agent needs to be particularly vigilant when receiving calls from outside this list), etc. All settings are synchronized to the cloud in real-time and sent to the elderly end to take effect.
[0043] Communication and agency panel: In addition to monitoring, the child end also allows for proactive initiation of conversations or operations. For example, children can input text / voice through the App and let the digital avatar of the elderly convey the message (display / announce it to the elderly); or remotely make an appointment for the elderly to see a doctor, purchase goods, and then the agent tells the elderly "I have helped you handle such and such things." This actually provides a semi-artificial agency model where children can pre-process some affairs for the elderly when they are free, and the agent executes and gives feedback, reducing the burden on the elderly.
[0044] Logs and records: Children can view historical logs, including which functions the elderly used every day, summary of topics talked with the agent, handling of abnormal alarms, etc. This helps to keep track of the parents' living dynamics. When it is found that the elderly's needs in a certain aspect increase (such as often asking a certain type of question), children can also provide targeted assistance or adjust the service content of the agent.
[0045] Implementation key points: The architecture of the entire multimedia operation service system based on digital avatars adopts a modular design and communicates through standard interfaces to ensure scalability and maintainability. The mobile phone terminal and the cloud are connected through an encrypted communication protocol, and a two-way confirmation mechanism is adopted for important instructions to prevent false instructions. When the system is initially deployed, the devices of children and the elderly need to complete the identity binding and digital avatar creation process. The model data of the digital avatar and the authorization policies of children are stored in the cloud, and the necessary parts are cached on the elderly's mobile phone terminal to support basic functions for offline use (such as locally viewing photo albums, making family calls, etc.). In terms of cross-platform, the mobile application is developed through mainstream frameworks, and at the same time, the accessibility interfaces of each platform are fully utilized (such as Android Accessibility Service and iOS Voice Over adaptation) to ensure compatibility with different models. The system is designed with progressive enhancement - if the performance of the elderly's mobile phone is limited, a simplified 2D avatar and voice mode can be selected; if the performance is good, a realistic 3D image and richer interactions can be provided. Through the above architecture and implementation methods, the system of the present invention can operate stably and efficiently in the real environment and meet the target functional requirements.
[0046] In addition, Figure 3 The specific achievable functions of the digital avatar agent (hereinafter referred to as the agent) on the elderly's mobile phone terminal are shown more briefly and comprehensively, including voice broadcast and graphic optimization through the information accessibility module; personalized recommendation and companion interaction through the intelligent entertainment module; device connection, data analysis and health reminder through the health monitoring module; family affection communication and psychological comfort through the emotional companion module; voice control of household appliances through the smart home module; automatic interception and warning through the anti-fraud module, etc.
[0047] More specifically, as Figures 4 - 8 shown, in a more specific embodiment, it can assist the elderly, teenagers, disabled groups, etc. in performing auxiliary functions in but not limited to the following aspects: Barrier-free information dissemination: To address the problem of the elderly having difficulty accessing written information, the intelligent agent acts as the elderly's information translator and announcer. For all written notices, messages, and news content, the intelligent agent will automatically convert them into voice (using the voices of their children) and read them aloud to the elderly, and can provide simplified explanations as needed. For example, when receiving a bank text message, the intelligent agent will explain the content of the text message to the elderly in simple language (such as "Your pension has been credited this month"), avoiding obscure financial terms. At the same time, the intelligent agent supports the elderly to obtain information by asking questions in voice, such as asking about the weather, news, calendar reminders, etc. The intelligent agent will answer in a conversational manner, without the elderly having to read a complex interface or search on their own. Visually, the intelligent agent interface minimizes large blocks of text and uses more graphics, icons, and concise labels to convey information, and can adjust the contrast and font size according to the elderly's eyesight. Through the multi-modal barrier-free presentation of information, the elderly can access various digital contents without barriers.
[0048] Voice and graphical assisted operation: The intelligent agent greatly reduces the complexity of mobile phone operations. The elderly can directly issue instructions in natural language and let the intelligent agent complete a series of cumbersome touch screen operations on their behalf. For example, when the elderly say "Transfer 100 yuan to my daughter for me", the intelligent agent will repeat and confirm, then automatically open the payment application, select the payee, fill in the amount, and finally only ask the elderly to perform the necessary authorization (such as entering a fingerprint or giving an oral confirmation). Almost no navigation of the interface by the elderly is required throughout the process. For some steps that the elderly need to complete manually, the intelligent agent will also provide graphical guidance: highlight the position of the button that should be clicked on the screen, or use animated arrows to indicate the operation sequence, allowing the elderly to complete it by following the instructions. For example, when teaching the elderly to switch the front camera during a video call, the intelligent agent can display an arrow pointing to the relevant icon and give a voice prompt "Click here to switch the camera". In addition, the intelligent agent has a voice error correction and repeat confirmation mechanism. When the elderly's instructions are unclear or the system does not hear clearly, it will politely request clarification in the tone of their children, avoiding risks caused by misoperations. Through the dual assistance of voice + graphics, the elderly can use various application functions with the least cognitive burden and complete tasks that were previously difficult to complete independently.
[0049] Intelligent Entertainment Recommendation: To enrich the spiritual and cultural life of the elderly, the intelligent device is built with an intelligent recommendation engine that provides personalized entertainment content based on the elderly's interests and preferences. The system regularly pushes short videos, songs, operas, cross talks, sketches, news, etc. that are healthy and positive to the elderly, and can continuously adjust the recommendation strategy according to the elderly's feedback. For example, if the system knows that the elderly like to listen to operas and storytelling, then after dinner every day, the intelligent device might say, "Dad, do you want to listen to a Peking opera? I'll play a section of 'Borrowing the East Wind' by Ma Lianliang for you." Or if it detects that the elderly like to watch short videos for entertainment in the afternoon, the intelligent device will prepare some humorous and positive-energy videos to play on the screen. In addition, the intelligent device can also accompany the elderly in some intellectual interactions, such as telling jokes, guessing riddles, chatting about past memories, etc., playing the role of spiritual comfort. All recommended content has been filtered and reviewed by the system to ensure that there is no vulgar or false information. If the elderly show disinterest in a certain type of content (such as frequently skipping a certain program), the system will reduce such recommendations; on the contrary, it will increase the weight for the content they like, forming an ever-adaptive entertainment supply. Through the push and company of the intelligent device, the digital life of the elderly will be more colorful and no longer limited to lonely moments of swiping the WeChat Moments or watching TV.
[0050] Health Monitoring and Reminder: Considering the needs of elderly health management, the multimedia operation service system based on digital avatars in this embodiment supports connecting various wearable and home medical devices. The digital avatar intelligent agent can monitor sensor data in real time to achieve round-the-clock health data monitoring and timely feedback. The intelligent agent application uses an open interface and can be docked with common health IoT devices on the market, such as blood pressure monitors, blood glucose meters, heart rate monitors, smartwatches, emergency call buttons, etc. Once these devices measure data, the results will be transmitted to the intelligent agent via Bluetooth or the network, and the intelligent agent will archive and analyze them. If an abnormal increase in blood pressure is detected, the intelligent agent will immediately remind the elderly to pay attention to rest or take medicine as prescribed by the doctor, and notify their children of the abnormal situation through the cloud. The intelligent agent will also regularly urge the elderly to measure their vital signs, take medicine on time, and exercise appropriately every day. For example: "Mom, it's 3 pm now. It's time to measure your blood sugar." "Dad, you haven't gone for a walk today. Do you want me to play some soothing music for you now so that you can move around?" In addition, combined with the positioning function of the mobile phone, and when an emergency is detected based on sensor data, the intelligent agent will quickly conduct preliminary handling. If the elderly respond normally, the event will be recorded in the cloud. If the elderly do not respond or confirm discomfort, the guardianship control terminal will be notified in time to inform their children, and the children can remotely intervene in the emergency treatment, etc.; the intelligent agent can ensure travel safety: when it detects that the elderly have walked out of the safe range or got lost, the system can notify their family members and guide the elderly back through voice. The health module also provides an emergency help function. The elderly only need to shout out a preset distress word (such as "I'm not feeling well") to the intelligent agent, and the system will immediately call the emergency contact or 120 for first aid and send out the location information. Through continuous health monitoring and considerate reminders, the intelligent agent becomes the personal health assistant of the elderly, helping them maintain good living habits and get timely assistance in case of accidents. Emotional Companionship and Communication: The digital avatar of this system is not only a tool but also acts as an emotional companion for the elderly. The intelligent agent has considerable conversation ability and can chat with the elderly to relieve their spiritual loneliness. When children cannot accompany the elderly often, the intelligent agent can comfort them in a timely manner: greet them actively every morning and evening ("Mom, how do you feel today"), send blessings on special days ("Dad, today is your birthday. My family and I wish you a long and healthy life"), and give comfort and guidance when detecting that the elderly are in a low mood (showing concern through voice and expressions). Since the intelligent agent uses the image and voice of the children, the elderly will feel as if their children are by their side asking about their well-being, thus obtaining a real sense of companionship. This emotional communication is not a simple preset but uses an AI dialogue system to achieve more natural interaction. The intelligent agent can recall things mentioned by the elderly (such as saying that the back hurt yesterday, and asking "Is your back better today" today), and can also make appropriate responses according to the elderly's mood and tone (if it hears the elderly's voice is depressed, the tone will be softer and more encouraging).Of course, the system will clearly inform the elderly that this is just a "digital avatar" of their children to avoid cognitive confusion, but this familiar communication method can still greatly relieve the loneliness of the elderly. The goal of the emotional companionship module is to make the intelligent agent a trustworthy family member for the elderly in the digital world and provide them with psychological support.
[0051] Smart home control: With the development of the Internet of Things, many elderly people's homes have also started to be equipped with smart home devices, such as smart light bulbs, TVs, air conditioners, door locks, security sensors, etc. However, for the elderly, the mobile phone apps for operating these devices may be too complex. This system integrates smart home control into the digital avatar intelligent agent, and home devices can be controlled through natural language. The elderly can directly tell the intelligent agent their needs. For example, "Help me turn on the living room light", and the intelligent agent will connect to the corresponding home platform API through the cloud to perform the light-on action and give a voice feedback "The light has been turned on for you". Another example is that at night, when the elderly get up and it is inconvenient to get the mobile phone, they can call the intelligent agent "I want to go to the bathroom", and the intelligent agent will immediately turn on the night light and prompt the ground obstacles. The intelligent agent can also automatically execute some home linkage scenarios according to the daily schedule. For example, at 10 pm, it reminds the elderly to rest and automatically turns off the TV, closes the smart curtains, etc. For the security system with a monitoring camera, the elderly can also let the intelligent agent call up the real-time picture on the mobile phone to check the situation at the door. All these control instructions are verified through the children's permissions to ensure safety (for example, operations involving security such as opening the door lock require prior authorization or reconfirmation by the children). By controlling the smart home by voice, the elderly do not need to learn various apps, and can truly achieve that with just one sentence, the whole house responds, which is both convenient and improves home safety.
[0052] Information Security Protection and Anti-Fraud: Ensuring the information security of the elderly in the digital life is a major feature module of this system. The intelligent agent constructs a firewall both at the front end and in the cloud: First, it has built-in call and message filtering functions. Drawing on various databases and number libraries, it automatically intercepts high-risk calls and text messages and processes them before the elderly are aware. For example, suspected fraud calls will be politely hung up by the intelligent agent or directly transferred to the intelligent agent for handling, preventing the elderly from directly contacting scammers. If it is a fraud text message, the intelligent agent will not remind the elderly to view it and will mark it in the children's log. Second, for those interactions that are not in the blacklist but are still suspicious, the intelligent agent will be vigilant during the interaction and monitor keywords at any time. Once it detects typical fraud situations such as the other party asking for a transfer or requesting a verification code, the intelligent agent will immediately trigger a secondary confirmation: asking the elderly in the image of their children "Are you sure? Maybe wait a while. Let me help you verify." and at the same time pushing an alarm notification to the children. Third, the intelligent agent will also refute rumors and correct the information obtained by the elderly on the Internet. Many elderly people are prone to believing online rumors and untrue health preservation information. Therefore, when the elderly relay a suspicious piece of information or ask relevant questions, the intelligent agent will check authoritative materials and make clarifications to guide the elderly to distinguish between truth and falsehood. For example, when the elderly say "I saw a folk remedy that can cure many diseases", the intelligent agent (if it determines that it is a rumor) will respond "These folk remedies have no scientific basis. It's more reliable to listen to the doctor." Finally, in the security module, all abnormal events (such as abnormal logins, device abnormalities, etc.) will be recorded and reported, and children can view the security report regularly. Through multi-level protection means, the intelligent agent is equivalent to hiring a personal security guard for the elderly's mobile phones, minimizing their chances of suffering digital risks.
[0053] It should be noted that the digital avatar intelligent agent, the guardianship control terminal in this embodiment, and the organic combination of the digital avatar intelligent agent and the cloud server enable the digital avatar intelligent agent of this system to fully meet the needs of the elderly in aspects such as information acquisition, operation execution, entertainment, health and safety, and emotional communication. Importantly, all these functions are presented under the interactive interface of the children's avatar. The operation method is always the same for the elderly - they can complete it by talking to the "child" or simply tapping. There is no need to learn the usage of different software. This integrated design greatly improves the usability and practical value of the system.
[0054] It should be noted that through the multimedia operation service system based on digital avatars in this embodiment, the following can be achieved: The digital avatar of the child as the only interaction interface: The assistant robot that the elderly users see and hear on their mobile phones will adopt the image and voice of their children. This digital avatar can be a 3D virtual image modeled after the real photos of the children, or a deep synthesis model of the real videos / voices of the children. All external digital interactions (such as obtaining information, operating applications, accessing online services, etc.) are completed through this avatar proxy. For example, when the elderly person wants to obtain the weather forecast, the "child" intelligent agent on the screen will inform the weather in a kind tone; when receiving a text message, the intelligent agent will read it out in the voice of the child and give an explanation. The familiar image greatly enhances the trust and intimacy of the elderly, reducing their wariness of machines and learning costs.
[0055] Remotely controllable permissions and content management: The system gives the children (real people) of the elderly a remote management port. The children can connect to the intelligent agent system used by the elderly through their own mobile phone App or web interface to manage the permissions and policies of the digital avatar. Specifically, the children can set the scope of applications and functions that the intelligent agent can access (for example, allowing it to help parents buy daily necessities online, but restricting it from transferring money on behalf of parents without confirmation); content that is suspicious to the children can be added to the blacklist library (such as certain high-incidence fraud numbers, keywords), and the intelligent agent will automatically intercept this information and prevent the elderly from seeing it. At the same time, the children's side can view the logs and monitoring data of the intelligent agent, such as what the elderly have recently asked the intelligent agent to do and whether there are any abnormal behaviors. Through this remote monitoring mechanism, the children actually play the role of a "background administrator", protecting the elderly without disturbing them. The children can also call or video chat with the elderly through the intelligent agent at any time, appearing in the image of the digital avatar to achieve "real-time filial piety online".
[0056] Intelligent identification of anomalies and real-time alerts: The system is built-in with advanced AI monitoring algorithms to detect anomalies in the interaction situation of the elderly's terminals. When a suspected risk event occurs, the system will immediately notify the children to intervene. Anomalous events include but are not limited to: The elderly are on the phone with a suspected fraud number (the intelligent agent can identify common fraud scripts based on the phone number or the content of the conversation voice); the elderly attempt to transfer a large amount of money; the elderly chat with strangers for a long time or browse suspicious websites; the intelligent agent detects that the elderly are in a low mood, with abnormal health data, etc. When these situations are detected, the intelligent agent will immediately send an alert to the children through the cloud, including the event description, time, possible risk level, etc., to facilitate the children to contact the elderly in time to verify the situation or directly intervene and stop through the intelligent agent. For example, if the intelligent agent determines that the elderly are answering a fraud call, it can first politely terminate the call on its own (saying in the voice of the child, "Don't worry, mom / dad. I'll help you check it out"), and send an emergency notice to the children's mobile phones, waiting for the children to decide the next step. Such a mechanism ensures that the children can intervene in time at critical moments and minimize the risks.
[0057] Multifunctional integrated intelligent assistant: The digital avatar intelligent body is not only a beautified interface, but also a one-stop assistant that integrates multiple capabilities such as barrier-free interaction, life services, entertainment, and health. Most of the digital functions required by the elderly in daily life can be completed through dialogue with the intelligent body or simple instructions, without switching between multiple apps. The intelligent body integrates AI capabilities such as speech recognition, natural language processing, and computer vision, and is connected with various service interfaces to achieve "one entrance, full network service". The specific functions will be introduced in detail in the "Main Function Modules" section below.
[0058] Cross-platform deployment, no special hardware dependency: This system is implemented in software form and is compatible with mainstream smartphone platforms (Android, iOS, etc.), without the need to customize special hardware. By using cross-platform development technologies (such as Flutter, ReactNative, etc.), the same set of code can run on mobile phones of different brands and models, facilitating large-scale promotion. The elderly do not need to buy new dedicated equipment, just install the application on existing smartphones or elderly smartphones (elderly phones). The system makes full use of the existing sensors of the mobile phone (such as microphones for voice interaction, cameras for video calls or emotion analysis, accelerometers for detecting falls, etc.) and network connectivity to minimize costs and usage barriers. The entire solution focuses on practicality and accessibility, with the goal of allowing most smartphones with ordinary configurations to run this system smoothly, thereby covering a wide range of elderly user groups.
[0059] Rapid creation and personalization of digital avatars: In terms of implementation, the creation of children's digital avatars will rely on today's mature AI digital human generation technology. Children only need to provide a small amount of material (such as a few minutes of video and audio samples), and the system can generate a highly realistic digital image in the cloud. Current existing technologies show that only 3 minutes of real-person oral video and about 100 voice samples are needed to train and generate a digital human model that is highly similar to the real person's portrait and voice. This system uses this technology to customize each family's exclusive digital avatar, so that it can restore the characteristics of children as much as possible in appearance, voice, and speaking style. The generated digital avatar is sent to the elderly's mobile phone after confirmation by the children, and the necessary model data is saved locally. During the interaction process, some simple dialogues and expressions are rendered in real time by the mobile phone; when encountering complex dialogues, the cloud AI model support can be requested. The entire personality model also supports a certain degree of personalized configuration. For example, children can set the name and tone of the intelligent body (whether it is humorous, etc.), and even define specific reminder sentences containing the caring words that children usually use, so that the elderly can feel the strong family affection.
[0060] Timely summary and communication between the digital avatar and the children: For the stipulated pre - plans, the digital avatar processes them independently, and regularly summarizes and analyzes the elderly's behaviors, then pushes the results to the children so that they can comprehensively and accurately understand the elderly's behaviors in a timely manner. For pre - plans beyond the regulations, the digital avatar notifies the children in real - time, and the children handle them manually. The digital avatar learns independently and forms pre - plans.
[0061] In summary, the core of the multimedia operation service system based on the digital avatar in this embodiment is to create a trustworthy "digital child" on the smart phone as the agent for the elderly to interact with the digital world. By talking to the "digital child", the elderly can obtain information and use services; while the children can ensure safety through remote supervision and endow the intelligent agent with enough intelligence to handle daily affairs and abnormal situations. This solution integrates artificial intelligence, digital simulation, and interpersonal emotion elements, and pioneeringly introduces the family trust relationship into the field of human - computer interaction, which can effectively improve the convenience and safety of the elderly integrating into digital life.
[0062] As described above, the multimedia operation service system based on the digital avatar provided in this application has the following differences from the existing technical solutions: Different from traditional age - friendly software: Currently, the "easy mode" or "elderly version" provided by many mobile phone manufacturers and applications mainly simplifies the interface, such as enlarging the font, streamlining options, and increasing contrast. Such transformations do not change the main body and logic of the interaction, and the elderly still need to learn and operate multiple applications by themselves. In contrast, this system introduces a digital intelligent agent as the main body of the interaction. The elderly only need to get used to talking to one "person", and the intelligent agent will execute specific tasks on their behalf, avoiding complex human - machine interface interactions. Therefore, we do not simply optimize the UI, but provide a new interaction paradigm, greatly reducing the cognitive threshold.
[0063] Different from general voice assistants (such as Siri, Xiaoai Tongxue, etc.): Common voice assistants can accept commands to complete some tasks, but they usually do not have personality settings, their voice feedback is mechanical, and they lack in - depth customization for the elderly in terms of functions. Our intelligent agent has a specific relative image and emotional interaction, and its conversation content and style are closer to the psychological needs of the elderly. At the same time, general voice assistants do not provide remote monitoring functions and do not have built - in anti - fraud mechanisms; this system has the participation of the children's end (the second client, or the children's remote end) and security strategies, making it more suitable for the elderly group in need of care. In addition, voice assistants often cannot perform complex operations across applications (or require the user to authorize each step), while the proxy execution ability of this system is more powerful and can complete complex tasks for the elderly across the application ecosystem, which is also a major difference.
[0064] Differences from remote control / assistance software: There are some tools or systems on the market for children to remotely assist their parents with their mobile phones (such as the built-in remote assistance of various mobile phone brands, third-party remote control apps, etc.). The basic idea of these solutions is for children to directly take over the elderly's devices for operation demonstrations or control. Essentially, it is still manual operation and cannot run autonomously. In contrast, our system emphasizes artificial intelligence autonomy. Usually, the intelligent agent automatically helps the elderly, and children are only involved in decision-making when there are abnormalities or specific needs. This reduces the burden on children as they don't need to monitor constantly. Moreover, remote control software generally does not consider the emotional and entertainment needs of the elderly, while our system provides rich companionship and care functions, offering a comprehensive solution.
[0065] Differences from elderly care robots / hardware devices: Some high-end elderly care products such as home companion robots and smart speakers also have functions like chatting and simple assistance. However, these hardware products are expensive and lack customization for individual family relationships, and their interaction images are mostly cartoon or machine-like. Our digital avatar solution only requires a mobile phone, without hardware costs, and the image can be that of the elderly's children, which is much more affectionate than a cold robot face. In terms of functionality, it is often difficult for companion robots to perform complex mobile application operations and access fine-grained network services, while the intelligent agent on the mobile phone can complete more tasks using the existing software and hardware conditions of the mobile phone. In short, we provide a solution that is closer to the family scenario at a lower cost.
[0066] Differences from single-function applications: Currently, apps for the elderly usually solve single problems separately. For example, some provide health reminders, some provide anti-loss positioning, and some provide chatting and making friends. However, it is impossible for the elderly to install and proficiently use numerous apps, and the scattered functions cannot work together effectively. The system of this invention integrates many functional modules and provides services through a unified interface, avoiding information silos and fragmented operations. Data sharing can also be achieved on the unified platform. For example, health data can affect entertainment recommendations, and the emotional state can trigger reminders for children to interact. These are collaborative effects that cannot be achieved by multiple independent apps. Therefore, the comprehensiveness and coordination of this system are not possessed by current fragmented applications.
[0067] Generally speaking, the multimedia operation service system based on digital avatars in this embodiment is different from any existing single software or hardware product. It integrates concepts such as digital human technology, intelligent assistants, remote monitoring, anti-fraud, and security assistance into a whole, forming a brand-new technology integration for the digital life of the elderly. This innovative combination gives the system unique advantages in user experience and functional depth, representing a major expansion and improvement of the existing technology.
[0068] The multimedia operation service system based on digital avatars in this embodiment can be applied to the following scenarios: The elderly living alone or in empty-nest families: For the elderly who live separately from their children, the digital avatar intelligent agent can make up for the lack of daily care and companionship, enabling them to "see" and "hear" their children's care at home, reducing their sense of loneliness, and having someone to guard them at any time when encountering network risks.
[0069] The elderly with mobility difficulties or semi-disabilities: Due to diseases and other conditions, some elderly people find it more difficult to operate mobile phones and are even unable to go out to handle affairs frequently. Through voice and proxy operations, the intelligent agent can help them obtain services without barriers, including remote registration for medical treatment, online shopping with home delivery of medicine, etc., improving their quality of life. For the elderly with dementia, the system's safety monitoring and reminder functions can also play a role (but more strict permission limitations are required from their children).
[0070] Ordinary elderly users with smartphones: Even if the elderly can basically use mobile phones, they can still benefit from this system. For example, many elderly people can use WeChat to chat but do not know other functions. The intelligent agent can guide them step by step to try new digital services and provide anti-fraud protection. This helps to expand the participation of the elderly in the digital world and enjoy more conveniences.
[0071] Community elderly care and nursing homes: The system can also be deployed in community elderly care service centers or nursing homes as part of an intelligent elderly care solution. The image of a caregiver can serve as the digital avatar for multiple elderly people, providing services for group users. In addition, the system can be docked with community hospitals and nursing home monitoring systems to achieve information sharing and improve the management efficiency of institutional elderly care.
[0072] Other people in need of digital assistance: Although this system is mainly designed for the elderly, its concept is also applicable to other digital disadvantaged groups, such as adults with low educational levels, visually impaired people, and intellectually disabled people. They can also designate a trusted person as their digital avatar image and obtain a barrier-free digital service experience through a similar mechanism. Therefore, the present invention also has application potential in a broader field of assistive technology.
[0073] The multimedia operation service system based on digital avatar in this embodiment has no hardware requirements, and tries to lower the hardware threshold so that the vast majority of smartphones and tablets can support it. The multimedia operation service system based on digital avatar needs a continuous network connection to access cloud AI services. Therefore, a mobile network of 4G or above, or a broadband WiFi environment is required. When the network is poor, the local functions of the system can still work limitedly, but the overall experience will decline. Therefore, it is necessary to ensure smooth network in the elderly's residence during actual use. In addition, to save traffic, the system will optimize audio and video transmission, such as preferentially using WiFi to download and update resources. If you need to use health monitoring and smart home functions, corresponding peripheral devices are required. For example, a Bluetooth medical device compatible with the market is needed to measure heart rate and blood pressure; smart home appliances need to support network control (WiFi or Bluetooth) and have open interfaces. Of course, these peripherals are not essential parts of the system, but are added according to the specific needs of the elderly. Through modular design, the system can gradually connect to new device categories. Even if the user does not have any additional devices, the core functions of companionship, assistant, and protection can still be provided by the mobile phone itself. If the mobile phone has a security chip (such as fingerprint or face recognition hardware), the system can call it for sensitive operation confirmation to improve security. However, this is not mandatory, and a password can be used as a substitute.
[0074] In summary, the multimedia operation service system based on digital avatar in this embodiment can be deployed and used without replacing existing devices. For the vast majority of elderly users with smartphones, they only need to install the application and perform a one-time setup to enjoy the digital elderly care service brought by the present invention. It can adapt to different application ranges from family individuals to institutional elderly care, and based on the high flexibility and low-cost deployment characteristics of the multimedia operation service system based on digital avatar, it helps to promote the popularization of smart elderly care and barrier-free technologies in the whole society, and truly realizes the benefit of the elderly with technology.
[0075] As Figure 9 shown, the present invention also provides a multimedia operation service method based on digital avatar, which performs targeted restricted multimedia services based on the multimedia operation service system as described above, including: S1: Configure a three-dimensional virtual character through a preset function module, and locally process the voice information and click request information received by the three-dimensional virtual character to obtain a request signal; the function module and the three-dimensional virtual character are integrated in a preset digital avatar intelligent body; S2: Make a feedback signal for the request signal through a preset cloud server, and instruct the three-dimensional virtual character to feedback the feedback signal to the digital avatar intelligent body or a preset guardianship control terminal; S3: Real-time monitor and remotely control the communication data of the three-dimensional virtual character through the guardianship control terminal.
[0076] In this embodiment, a three-dimensional virtual character is configured, and the voice information and click request information received by the three-dimensional virtual character are locally processed to obtain a request signal, including: S11: Based on a graphics engine, a preset digital human image element library, and an interface element library, real-time rendering of a 3D model is performed to form a real-time changing digital human three-dimensional picture, and based on pre-recorded video clips and pre-acquired voice clips, voiceovers are performed for the digital human three-dimensional picture to form a digital avatar intelligent agent; S12: The microphone is called to continuously monitor the daily voice of the user, the three-dimensional virtual character is summoned by comparing the daily voice with a wake-up word pre-stored locally, and the command voice issued by the user to the three-dimensional virtual character is collected, and the command voice is text-converted to generate a request signal; wherein, the user is the user of the digital avatar intelligent agent.
[0077] A feedback signal is made for the request signal, and the three-dimensional virtual character is instructed to feedback the feedback signal to the digital avatar intelligent agent or the guardianship control terminal, including: S21: The request signal is sent to a cloud server, so that the cloud server converts the request signal into a request text, performs natural language recognition on the request text based on a natural language understanding model to parse the request intention, and determines a response purpose according to the request intention, preference data stored in the database and management platform, or monitoring logs stored in the log database; a feedback signal is made according to the response purpose; the feedback signal is feedback data or a feedback call instruction; S22: The feedback data is fed back to the voice interaction module of the digital avatar intelligent agent, or the media software and home appliances connected to the digital avatar intelligent agent are remotely controlled according to the feedback call instruction; S23: Based on a preset risk analysis model, risk analysis is performed on all feedback data and feedback call instructions to obtain user evaluation status data, health analysis is performed on the user physiological data transmitted from the first client to obtain health status data, and the evaluation status data and the health status data are edited into a monitoring log and transmitted to the guardianship control terminal, and the monitoring log is synchronously stored in the log database; S24: Big data analysis is performed on the monitoring log to obtain preference text, and the preference text and setting data preset by the user are summarized and stored to form preference data.
[0078] For a more specific implementation manner, refer to the embodiment of the multimedia operation service system based on digital avatars above, and details are not described herein again.
[0079] As described above, the multimedia operation service method based on digital avatars provided by the present invention monitors and remotely controls the communication data of three-dimensional virtual characters in real time, thereby breaking through the sense of strangeness of traditional machine assistants. It creatively uses the digital avatar image of the elderly's children as an interaction agent, taking advantage of the elderly's natural trust in the parent-child relationship to solve the problems of low user stickiness and trust. This makes the elderly more willing to accept and use digital technology services subjectively; children can remotely configure and intervene in the behavior of the intelligent agent to achieve appropriate supervision of the elderly's digital life, especially in timely intervention in anti-fraud and security aspects. This solution that integrates the guardian's permissions into the intelligent assistant enables the digital assistant to give full play to its intelligence while remaining within the controllable range; integrating information accessibility, life services, entertainment care, health management, security protection and other aspects into a system, multiple needs of the elderly can be met through a digital avatar interface, avoiding the cumbersome process of switching between numerous apps and devices. This all-round integrated service concept greatly improves convenience. Moreover, the intelligent agent (digital avatar intelligent agent) gains the trust of the elderly through the image of the children and secretly executes security strategies, achieving a protection effect without any burden on the elderly. This subtle security design greatly improves the success rate of fraud prevention. And it integrates real-time voice semantic analysis and a security knowledge base, which can greatly improve the accuracy of intelligent agent applications.
[0080] As described above by way of example with reference to the accompanying drawings, the multimedia operation service system and method based on digital avatars proposed according to the present invention have been described. However, those skilled in the art should understand that various improvements can be made to the above-mentioned multimedia operation service system and method based on digital avatars proposed by the present invention without departing from the content of the present invention. Therefore, the protection scope of the present invention should be determined by the content of the appended claims.
Claims
1. A multimedia operation service system based on digital twins, characterized in that, It includes a digital avatar intelligent agent deployed on the first client, a guardianship control terminal deployed on the second client, and a cloud server communicatively connected to the digital avatar intelligent agent and the guardianship control terminal; wherein, the digital avatar intelligent agent includes a three-dimensional virtual character communicatively connected to media software and a home, and a function module for providing functional services to the three-dimensional virtual character; wherein, the function module is used to configure the three-dimensional virtual character and locally process the voice information and click request information received by the three-dimensional virtual character to obtain a request signal; the cloud server is used to generate a feedback signal for the request signal and instruct the three-dimensional virtual character to feedback the feedback signal to the digital avatar intelligent agent or the guardianship control terminal; the guardianship control terminal is used to monitor and remotely control the communication data of the three-dimensional virtual character in real time.
2. The multimedia operation service system based on digital avatars according to claim 1, wherein, The function module includes a UI rendering module, a voice interaction module, a dialogue management and local logic module, a security control module, and a data collection and encryption module; wherein, the UI rendering module is used to perform real-time rendering of a 3D model based on a graphics engine and a preset digital human image element library and interface element library to form a real-time changing three-dimensional digital human picture, and dub the three-dimensional digital human picture based on pre-recorded video clips and pre-obtained voice clips to form a digital avatar intelligent agent; the voice interaction module is used to continuously monitor the user's daily voice by calling a microphone, summon the three-dimensional virtual character by comparing the daily voice with a locally pre-stored wake-up word, collect the command voice sent by the user to the three-dimensional virtual character, perform text conversion on the command voice to generate a request signal, send the request signal to the cloud server, receive the feedback signal returned by the cloud server, and convert the feedback signal into output voice to instruct the three-dimensional virtual character to play the output voice, or display the output text corresponding to the output voice in large font on the display of the first client; wherein, the user is the user of the first client; the dialogue management and local logic module is used to match the local stored corpus according to the daily voice. If the stored corpus is matched, the corresponding local action is executed. And if the stored corpus is an agency corpus, a dialogue is initiated to the guardianship control terminal and the guardianship control terminal is instructed to handle business for the user; the security control module is used to filter the data received by the first client based on preset control rules and a preset blacklist, and perform access restriction on the first client based on preset restriction instructions; the data collection and encryption module is used to encrypt the request signal and the feedback signal, and capture sensitive data based on preset collection rules, and directly transmit the encrypted sensitive data to the guardianship control terminal.
3. The multimedia operation service system based on digital avatars as claimed in claim 2, wherein, In the digital avatar intelligent agent, the actions of the three-dimensional digital human picture are matched with the content of the dubbing in real time; wherein, the three-dimensional digital human picture is the AI image of the user's guardian. The voiceover configured for the AI image is the guardian voice preset by the guardian; The three-dimensional image of the AI image is a 3D stereoscopic digital human image generated based on the guardian's photo.
4. The multimedia operation service system based on digital avatars according to claim 3, characterized in that The cloud server includes a pre-trained large model, a database and management platform, a digital human control module, and a monitoring module; among them, The large model is used to convert the request signal into a request text, perform natural language recognition on the request text based on a natural language understanding model to parse the request intention, and determine the response purpose according to the request intention, preference data stored in the database and management platform, or monitoring logs stored in the log database; make a feedback signal according to the response purpose; the feedback signal is feedback data or a feedback call instruction; The digital human control module is used to feedback the feedback data to the voice interaction module, or remotely control the media software and home appliances according to the feedback call instruction; The monitoring module is used to perform risk analysis on all feedback data and feedback call instructions based on a preset risk analysis model to obtain the user's evaluation status data, perform health analysis on the user physiological data transmitted by the first client to obtain the health status data, and edit the evaluation status data and the health status data into a monitoring log and transmit it to the guardianship control terminal, and synchronously store the monitoring log in the log database; The database and management platform is used to perform big data analysis on the monitoring log to obtain preference texts, and summarize and store the preference texts and the setting data preset by the user to form preference data.
5. The multimedia operation service system based on digital avatars according to claim 4, characterized in that The feedback data is data adapted to the response purpose based on the media software and home appliances and the database of the large model.
6. The multimedia operation service system based on digital avatars according to claim 5, characterized in that The media software and home appliances include social software, weather software, news software, music software, e-commerce software, payment software, smart home cloud platform, and medical and health platform.
7. The multimedia operation service system based on the digital twin as claimed in claim 6, wherein The guardianship control terminal logs in based on an application or a web page; the function interface of the guardianship control terminal includes a real-time monitoring panel, a permission and filtering setting panel, and a communication and agency panel; among them, The real-time monitoring panel is used to receive the monitoring log to obtain the operation information of the user on the first client; The permission and filtering setting panel is used to set access permissions for the digital avatar agent to control the user's usage space on the first client, used to set a blacklist for the digital avatar agent to perform social filtering on the first client, and is also used to set a payment limit for the payment software in the first client to control the user's payment amount on the first client; The communication and agency panel is used to perform business agency for the user according to the agency corpus, and after the agency is completed, instruct the digital avatar agent to broadcast to the user or send a reminder message indicating that the agency is completed to the user.
8. A multimedia operation service method based on digital avatars, characterized in that, A multimedia service with targeted restrictions based on the digital twin-based multimedia operation service system according to any one of claims 1 to 7, including: Configuring a three-dimensional virtual character through a preset function module, and locally processing the voice information and click request information received by the three-dimensional virtual character to obtain a request signal; the function module and the three-dimensional virtual character are integrated in a preset digital twin intelligent agent; Making a feedback signal for the request signal through a preset cloud server, and instructing the three-dimensional virtual character to feedback the feedback signal to the digital twin intelligent agent or a preset guardianship control terminal; Real-time monitoring and remote control of the communication data of the three-dimensional virtual character through the guardianship control terminal.
9. The method for multimedia operation service based on digital twin according to claim 8, wherein Configuring a three-dimensional virtual character, and locally processing the voice information and click request information received by the three-dimensional virtual character to obtain a request signal, including: Performing real-time rendering of a 3D model based on a graphics engine and a preset digital human image element library and interface element library to form a real-time changing digital human three-dimensional picture, and dubbing the digital human three-dimensional picture based on pre-recorded video segments and pre-acquired voice segments to form a digital twin intelligent agent; Calling a microphone to continuously monitor the user's daily voice, summoning the three-dimensional virtual character by comparing the daily voice with a locally pre-stored wake-up word, collecting the command voice issued by the user to the three-dimensional virtual character, and performing text conversion on the command voice to generate a request signal; wherein, the user is the user of the digital twin intelligent agent.
10. The multimedia operation service method based on digital avatar according to claim 9, wherein, Making a feedback signal for the request signal, and instructing the three-dimensional virtual character to feedback the feedback signal to the digital twin intelligent agent or the guardianship control terminal, including: Sending the request signal to a cloud server, so that the cloud server converts the request signal into a request text, performing natural language recognition on the request text based on a natural language understanding model to parse the request intention, and determining a response purpose according to the request intention, preference data stored in a database and management platform, or monitoring logs stored in a log database; making a feedback signal according to the response purpose; the feedback signal is feedback data or a feedback call instruction; Feeding back the feedback data to the voice interaction module of the digital twin intelligent agent, or remotely controlling the media software and home appliances connected to the digital twin intelligent agent according to the feedback call instruction; Performing risk analysis on all feedback data and feedback call instructions based on a preset risk analysis model to obtain user evaluation status data, performing health analysis on the user physiological data transmitted by the first client to obtain health status data, and editing the evaluation status data and the health status data into a monitoring log and transmitting it to the guardianship control terminal, and synchronously storing the monitoring log in the log database; Performing big data analysis on the monitoring log to obtain preference texts, and summarizing and storing the preference texts and preset data set by the user to form preference data.
Citation Information
Patent Citations
Information monitoring method, equipment and information monitoring system
CN105163330A
Target function processing method and device, terminal and storage medium
CN112163862A
Nursing dialogue system based on artificial intelligence
CN115017918A
Video generation method and device, electronic equipment and computer program product
CN119922351A
Old person health service system based on health service robot
CN204971271U