Control system and control method for interactive object, and computer program
An AI-driven control system addresses risks in virtual communities by determining user literacy and using NPCs for parental control, creating a safe and engaging environment adaptable to diverse cultural and religious contexts.
Patent Information
- Application Number
- PCT/JP2025/003311
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-01-31
- Publication Date
- 2025-10-02
AI Technical Summary
Existing virtual community environments face risks due to cultural, religious, and ethical diversity, with rule-based algorithms struggling to adapt and human expert assistance being costly and discontinuous, necessitating a comprehensive risk mitigation solution.
An AI-driven control system that determines user internet literacy and social community risk levels, performing parental control through non-player characters (NPCs) to manage interactions and mitigate risks.
Provides a safe virtual community environment that reduces anxiety and fosters engagement by ensuring ethical security and individual expression, adapting to diverse cultural and religious contexts.
Smart Images

Figure JP2025003311_02102025_PF_FP_ABST
Abstract
Description
Interactive object control system, control method, and computer program
[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to a control system and control method for interactive objects used in a virtual space or a social community, and a computer program.
[0002] Recently, with the development of remote environments and the democratization of XR environments, geographical constraints on communication have been eliminated, and expectations are rising for the potential for a variety of entertainment businesses and improved engagement, in addition to the use of remote work and social networking services (SNS). Furthermore, expectations are high for business possibilities arising from the opportunities for sharing and connection brought about by the spread of the metaverse. Online games can be said to be at the forefront of this trend. Relatively young age groups are also participating in online gaming platforms, and there are high expectations that this will be a booming business area in the future.
[0003] However, the development of remote environments makes it easy to communicate across religious, national, and cultural boundaries, raising concerns about unexpected risks from various perspectives, including ethical, cultural, customary, linguistic, and religious. For example, cultures vary in the degree to which religious teachings are faithfully observed. Currently, the only way to determine which areas require caution is to rely on real-time advice from experts, and this service has not yet been established. Furthermore, rule-based algorithms have difficulty keeping up with changes in virtual community environments and dealing with complex conditions. Furthermore, risks may arise due to the influence of international relations, necessitating the assistance of experts such as political scientists. While relying on human operations such as experts is possible, there are limitations in terms of continuity and cost. Furthermore, there is a movement in Europe toward a European AI regulation bill that also focuses on respecting and protecting human rights. Businesses subject to the regulations must make every effort to develop and use AI systems or underlying models in accordance with general principles of diversity, non-discrimination, and fairness. AI systems are required to be developed and used in a way that is inclusive of diverse actors and promotes equal access, gender equality, and cultural diversity, while avoiding the risk of discriminatory effects and unfair bias, which are prohibited by law.
[0004] At present, responding to these risks is left to the literacy of each individual user, which is a particular challenge for younger generations. Services that comprehensively avoid these risks have yet to be established for various businesses operating in remote or XR environments.
[0005] Zhe Cao, Gines Hidalgo, Tomas Simon, Shih-En Wei and Yaser Sheikh, "OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", arXiv, 30 May 2019
[0006] An object of the present disclosure is to provide a control system and control method for interactive objects, and a computer program, while eliminating risks that may lurk in a virtual space or a social community.
[0007] The present disclosure has been made in consideration of the above-mentioned problems, and a first aspect thereof is an interactive object control system used in a virtual space or a social community, comprising: a control unit; and an AI determination unit that determines the internet literacy of a user who uses the interactive object or determines the risk of the social community based on historical information of the user, and the control unit performs parental control on the interactive object according to the internet literacy of the user or the risk level of the social community.
[0008] However, the term "system" used here refers to a logical collection of multiple devices (or functional modules that realize specific functions), regardless of whether each device or functional module is contained within a single housing. In other words, both a single device consisting of multiple parts or functional modules and a collection of multiple devices are considered "systems."
[0009] The user history information includes at least one of an internet usage history, a game history, a purchase site history, a behavior history in the social community, and a conversation history.
[0010] The interactive object may be an avatar. The control system according to the first aspect may include a generation AI unit that generates a non-player character (NPC) that is not operated by a user. The control unit uses the non-player character to perform the parental control based on the content of a conversation between the avatar and another avatar.
[0011] Furthermore, a second aspect of the present disclosure is a method for controlling an interactive object used in a virtual space or a social community, the method including: a determination step of determining the internet literacy of a user who uses the interactive object based on historical information of the user, or determining the risk of the social community using an AI model; and a control step of performing parental control on the interactive object depending on the internet literacy of the user or the risk level of the social community.
[0012] Furthermore, a third aspect of the present disclosure is a computer program written in a computer-readable format to be executed on a computer to control interactive objects used in a virtual space or a social community, causing the computer to function as: an AI determination unit that determines the Internet literacy of a user who uses the interactive object or determines the risk of the social community based on historical information of the user; and a control unit that performs parental control on the interactive object depending on the Internet literacy of the user or the risk level of the social community.
[0013] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or via a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative action is exerted on the computer, and the same operational effects as those of the control system according to the first aspect of the present disclosure can be obtained.
[0014] According to the present disclosure, it is possible to provide a control system and control method for interactive objects, as well as a computer program, by utilizing AI technology while eliminating risks that may lurk in virtual spaces or social communities.
[0015] It should be noted that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited to these. Furthermore, the present disclosure may also bring about additional effects in addition to the effects described above.
[0016] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description based on the embodiments and accompanying drawings.
[0017] FIG. 1 illustrates the configuration of system 100. FIG. 2 illustrates the configuration of system 200 for adding individuality to an avatar's nonverbal expressions. FIG. 3 illustrates the main modules involved in avatar dancing. FIG. 4 illustrates the basic configuration of system 400. FIG. 5 illustrates the basic configuration of video generation system 500. FIG. 6 illustrates a functional block diagram for extracting movements from a dance video. FIG. 7 illustrates a functional block diagram for generating personal motion data. FIG. 8 illustrates a functional block diagram for extracting personal features by body part. FIG. 9 illustrates a system configuration for analyzing and determining user features. FIG. 10 illustrates a mechanism for generating reference motion data. FIG. 11 illustrates a mechanism for feature recommendation and self-feature correction by the evaluation and supervision unit 509. FIG. 12 illustrates an example of an individual service implemented on a communication platform. FIG. 13 illustrates an example of the basic operation of sharing content between users' information terminals on a communication platform. FIG. 14 illustrates a mechanism for switching avatars depending on the trust and intimacy between interlocutors. FIG. 15 is a diagram illustrating a mechanism for linking avatar generation using AI generation with existing avatar generation services. FIG. 16 is a diagram illustrating an example in which a communication platform has built a virtual community environment that spans culturally and religiously diverse groups and countries. FIG. 17 is a diagram illustrating an example in which an object that poses a religious risk is replaced with another object that does not pose a religious risk in the virtual space of another religious sphere. FIG. 18 is a diagram illustrating a functional configuration for processing an object that poses a religious risk in a shared virtual space. FIG. 19A is a diagram illustrating a specific method for acquiring feature data. FIG. 19B is a diagram illustrating a specific method for acquiring feature data. FIG. 19C is a diagram illustrating a specific method for acquiring feature data. FIG. 20 is a diagram illustrating a processing flow implemented by a communication platform. FIG. 21 is a diagram illustrating a mechanism for passing through a security gate before entering a closed space. FIG. 22 is a diagram illustrating an example in which voice and a camera are used for information input.FIG. 23 is a diagram showing an information processing flow centered on the audio information processing block. FIG. 24 is a diagram showing an information processing flow centered on image processing. FIG. 25 is a diagram showing an information processing flow centered on SNS history processing. FIG. 26 is a diagram showing the functional configuration of a parental control function according to the present disclosure. FIG. 27 is a diagram showing the overall configuration of a parental control function according to the present disclosure. FIG. 28 is a diagram showing an example of the internal configuration of the total risk assessment unit 2705. FIG. 29 is a diagram showing an example of generating a risk simulation. FIG. 30 is a diagram showing a flow for calculating a total risk. FIG. 31 is a diagram showing a specific example of determining a risk level in the total risk assessment unit 2705. FIG. 32 is a diagram showing a mechanism for determining a user's internet literacy using a literacy assessment AI model. FIG. 33 is a diagram showing an example of the relationship between internet literacy, risk level, and weighting coefficient. FIG. 34 is a diagram showing a mechanism for linking a parental control function according to the present disclosure with an app. FIG. 35 is a flowchart showing a processing procedure for controlling a user's (child's) internet access. FIG. 36 is a flowchart showing a processing procedure for training a literacy determination AI model. FIG. 37 is a flowchart showing a processing procedure for determining a user's (child's) internet literacy. FIG. 38 is a flowchart showing a detailed processing procedure for controlling a user's (child's) internet access. FIG. 39A is a diagram showing an example of UI operation on a child's smartphone and a parent's smartphone. FIG. 39B is a diagram showing internal operation during a period in which an app is continuously used on a child's smartphone. FIG. 40 is a diagram showing an example of the hardware configuration of an information processing device 2000. FIG. 41 is a diagram showing an example of parental control using an NPC.
[0018] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.
[0019] A. Overview B. System Configuration C. Non-verbal Expression Using Avatars C-1. Overview C-2. Background C-3. Mechanism for Imparting Individuality to Avatar Non-verbal Expressions C-3-1. Analysis Module C-3-2. Proposal Module C-3-3. Critic Module C-3-4. Other Platform Deployment Module C-4. Avatar Dance Generation System C-5. Video Generation System Adding Personal Features C-5-1. Adjusting the Amount of Personal Features Added C-5-2. Extracting Movements from Dance Videos C-5-3. Extracting Personal Features C-5-5. Extracting Features of User's Dance Performance C-5-6. Explanation of Reference Motion Data C-5-7. Improving Avatar Dance D. Communication Platform D-1. Basic Configuration D-2. Basic Movements D-3. Avatar Switching According to Reliability D-4. Linking with External Services D-5. D-6. Risk avoidance in virtual community environments D-7. AI-based proxy solutions D-8. Information processing procedures D-9. Mechanism for passing through security gates D-10. Examples E. Parental control E-1. Overview E-1-1. Basic functions E-1-2. Effects E-2. Overall structure E-3. Risk assessment E-3-1. Overall structure E-3-2. Risk simulation E-3-3. Calculation of total risk E-3-4. Assessment of internet literacy E-3-5. Relationship between internet literacy, risk level, and weighting coefficient E-4. Parental control app integration E-5. Flow E-5-1. Flow for controlling internet access E-5-2. Flow for learning the literacy assessment AI model E-5-3. Flow for assessing internet literacy E-5-4. Flow for controlling internet access F. Configuration of information processing device
[0020] A. Overview While the development of a remote environment is expected to lead to the expansion of online businesses such as communication and entertainment, there are concerns that unexpected risks may arise from various perspectives, including ethical, cultural, customary, linguistic, and religious.
[0021] Therefore, this disclosure proposes a technology that creates a platform consisting of a closed space based on an AI (Artificial Intelligence) engine, constructs a real-time, always-on virtual environment, and autonomously performs information filtering, conversion, adjustment, etc. using generating AI, with the main purpose of avoiding risks in virtual community environments.
[0022] According to the present disclosure, it is possible to provide a safe virtual community environment that eliminates anxiety and risk factors even when connected at all times. The safe virtual community environment provided by the present disclosure is expected to foster cooperation and empathy, increasing the opportunities for creativity and activity. Furthermore, engagement in the virtual community space provided by the present disclosure will increase, increasing opportunities for business creation within this space.
[0023] B. System Configuration This disclosure realizes a platform that allows players to express their individuality in social games, increasing user engagement while enjoying games in a safe and secure environment. This disclosure builds a platform that provides ethical security and reduces risks for young people, and adds individual user individuality to games, providing a safe and enjoyable environment for many users.
[0024] 1 is a schematic diagram showing the configuration of a system 100 to which the present disclosure is applied. The illustrated system 100 comprises a communication platform unit 110 and an application unit 150.
[0025] The application unit 150 includes a social game application 151 and various other services 152 related to communication in a virtual community environment. In the present disclosure, the social game application 151 includes a generation AI.
[0026] On the other hand, the communication platform unit 110 provides each service of the application unit 150 with a safe virtual community environment that eliminates anxiety and risk factors from various perspectives, such as ethics, culture, customs, language, and religion, even when connected at all times.
[0027] The communication platform unit 110 processes each data element handled in the virtual community environment. Data elements include video, audio, documents, web, etc., and the communication platform unit 110 performs image processing, audio processing, text processing, web processing, and various other sensor processing on these data elements. The communication platform unit 110 operates various functional modules, such as an information detection unit 111 that detects information from the processed data elements, an analysis unit 112 that analyzes each detected piece of information, a discrimination unit 113 that distinguishes each piece of information based on the analysis results, and a generation unit 114 that generates each element of the virtual community environment, such as video content (e.g., avatars), based on the discrimination results. An avatar is a type of content used in a virtual space and is generally a character used as a user's avatar in remote work, communication via social networking services, etc.
[0028] Within the communication platform unit 110, data processing such as image processing, voice processing, text processing, web processing, various sensor processing, etc. is performed on data elements such as video, audio, documents, web, etc.
[0029] The information detection unit 111 extracts virtual community information such as human body characteristics, environmental images, environmental sounds, SNS information, URL (Uniform Resource Locator) information, and movement information from the processed data elements. The analysis unit 112 analyzes the conversations, speech characteristics, background noise, background, tones / facial expressions, and movement characteristics in the virtual community based on the information extracted by the information detection unit 111. The discrimination unit 113 discriminates user identification, emotions, preferences, behavior, speech, personal authentication, and the like in the virtual community based on the analysis results by the analysis unit 112. The generation unit 114 generates each element of the virtual community environment based on the discrimination results by the discrimination unit 113, such as content such as still images, videos, and audio, for example, videos of the avatar itself, virtual backgrounds and background noises in the space where the avatar exists, content other than the avatar, and supplementary information such as advertisements (including advertisements using background signs and other supplementary information). The generation unit 114 also converts the language spoken in the virtual community and converts risky vocabulary into non-risky vocabulary. The communication platform unit 110 includes a history management database 115 that manages the operation history of each of the function modules 111 to 114.
[0030] The communication platform unit 110 operates functional modules such as an authentication unit 116, a parental control unit 117, and a content adjustment unit 118 that cooperate with the services of the application unit 150.
[0031] In the system 100 shown in FIG. 1 , the content adjustment unit 118 in the communication platform unit 114 adjusts each piece of content generated by the generation unit 114 so that it is compatible with the application 150. For example, the generation unit 114 can generate an avatar, and the content adjustment unit 118 can further adjust the avatar so that it operates in the social game application 151. The content adjustment unit 118 also performs processing to add "personality" to the avatar when the avatar is used to perform non-verbal physical expressions such as facial expressions, limb movements, and dancing. Details of non-verbal expressions using the avatar will be described in Section C below. In addition, the personal authentication unit 116 and the parental control unit 117 can manage personal authentication and behavior using functional modules such as the discrimination unit 113 operating in the communication platform unit 110. The parental control unit 117 also performs parental control on the avatar's behavior. The parental control unit 117 can also implement parental control without user operation by using non-player characters (NPCs). For example, parental control can be implemented based on the content of conversations between avatars through NPCs by having NPCs appear in the social community in which the user participates, or by changing the appearance, speech, or items used by the avatars. Details of the parental control function will be described in Section E below.
[0032] The communication platform will be described in more detail in Section D below. It should be noted that the communication platform unit 110 can also add necessary functional modules other than those shown in the figure that work with applications and services. Furthermore, at least two or more modules of the information detection unit 111, analysis unit 112, discrimination unit 113, generation unit 114, parental control unit 117, and content adjustment unit 118 in the communication platform unit 110 can be integrated into one large AI model. The large AI model can be, for example, a large language model (LLM) or a platform model.
[0033] C. Non-verbal Expression Using Avatars C-1. Overview Avatars are a type of content used in virtual spaces, and are generally characters used as the user's alter ego in situations such as remote work and communication via social media. When attempting to use an avatar to express non-verbal movements using the body, such as facial expressions, limb movements, or dancing, there is no way to add a "personal touch" to the avatar. This creates the problem that all avatars end up with uniform movements. For example, even if you instruct a generation AI to generate unique avatar movements, without input of "personal touch," the result will simply be an increased variation of imitating someone else.
[0034] First, the user needs to input their "personal body expressions" into the AI generator, but unlike language, it is difficult for the user to know how to provide non-verbal body expressions. Therefore, the AI needs to provide instructions to the user on body expressions that can be expressed by the user's body and that can be read by an interface between the AI and humans.
[0035] Therefore, in the present disclosure, the generation AI instructs the acquisition of the user's facial expressions and motions according to the user's environment (device used, such as a personal computer (PC), smartphone, etc.), estimates the user's physical "habits" and "characteristics" based on a comparison of the acquired user's facial expressions and motions with preset standard values, and presents avatar movements that reflect the user's habits and characteristics.
[0036] Here, in order to acquire the user's facial expressions more naturally, for example, a video that encourages joy, anger, sadness, and happiness is presented to the user, thereby acquiring the user's facial expressions in a natural way. Furthermore, when acquiring the user's motion, for example, a simple exercise is presented to the user, thereby acquiring the user's bone and motion data. Furthermore, the standard value to be compared with the acquired user's facial expressions and motion is specifically the average value of physical expressions. The generation AI then grasps deviations and fluctuations from the average value of physical expressions, understands them as the user's physical "habits" and "characteristics," and presents avatar movements that reflect the user's habits and characteristics. Note that the "average value of physical expressions" takes user attributes such as "cultural background," "gender," and "age" into full consideration.
[0037] The AI also presents the user with avatar movements that reflect the user's habits and characteristics. The AI then provides feedback to the AI about what the user likes and dislikes through multimodal interaction, such as voice, images, videos, text, illustrations, and gestures, repeatedly providing instructions for modifying the avatar's movements. By incorporating the user's own physical habits and characteristics into the avatar in this way, an avatar that expresses the user's unique movements can be generated, allowing for unique physical expression that is neither uniform nor imitation. Furthermore, when the user modifies the avatar's movements or dance motions, this feedback can be effectively collected and trained into the AI model. For example, feedback on whether the user likes certain movements or gestures can be collected, and the AI can learn based on this information to reflect or change the behavior and dance motions of the avatar it generates.
[0038] By converting the "personality" created here into a safe and secure cross-platform format, it can also be used for avatars in other applications. In business applications, this allows users to participate in meetings using avatars and communicate their intentions without speaking, improving the efficiency of consensus building in meetings. In entertainment applications, users can express their individuality through their avatars when uploading their own avatar expressions (such as dancing) to social networking sites like video viewing platforms. Furthermore, when multiple users simultaneously participate in multiplayer games by controlling avatars, the users' "personality" is reflected in the avatars they control in the game, improving engagement with the game.
[0039] C-2. Background Here, the background to the non-verbal expression using avatars in this disclosure will be explained.
[0040] 1 , the communication platform unit 110 provides a safe platform that eliminates anxiety and risk factors even when connected at all times to the social game application 151. Motivations for participating in a social game on the communication platform according to the present disclosure include the following:
[0041] ・I want to be seen as amazing (desire for recognition). ・I want to get along with friends. ・I want to leave a record. ・I want to play in a safe and secure place.
[0042] When making an avatar perform non-verbal expressions such as dancing in a social game, the present disclosure has the following features.
[0043] ・Since individuality cannot be expressed by dancing according to a model or example (reference), it is possible to give each user their own individuality (characteristics). ・Dances can be changed to reflect the opinions and criticisms of third parties (evaluators) (including evaluations by AI models). ・Dances can be automatically changed based on familiarity with other participants. ・Users can extract their own individuality (facial expressions, motion features) with simple movements and interfaces.
[0044] It is expected that the following functions will be added to non-verbal expressions such as avatar dancing according to the present disclosure.
[0045] - A function to adjust the avatar's personality features depending on whether they are expressing individuality or anonymity. - A function to adjust emotions (emotional features) depending on whether a dancer expresses individuality or anonymity. Emotional features are adjusted based on the user's mood (feeling) as determined by AI. - A function to adjust whether the user's emotion is positive or negative based on the presence or absence of an audience. Changes in the user's emotions are determined from the sound and sharpness of the movements. - A function to adjust the appearance of the dance based on the familiarity of the audience. When an unfamiliar audience is present, the expression of the user's features is suppressed to make the dance appear more normal. - A function to generate personal features. This adjusts the ratio between the dance model or model (reference) and the user's features, and also adjusts the reflection ratio of the user's features based on the familiarity of the shared space. - As a parental control function, the reflection ratio of personal features is adjusted based on the safety level of the shared space. This can be applied not only to dance but also to the reflection ratio of the personal features of the avatar's face. The reflection ratio of personal features is controlled using AI. - A function to modify the dance based on the judges' judgment.
[0046] C-3. Mechanism for Imparting Individuality to an Avatar's Nonverbal Expressions Figure 2 schematically illustrates the configuration of a system 200 for imparting individuality to an avatar's nonverbal expressions. A specific example of an "avatar's nonverbal expression" is the avatar's dancing. The illustrated system 200 includes an analysis module 210, a proposal module 220, a reviewer module 230, and a multi-platform (PF) deployment module 240. For ease of explanation, Figure 2 illustrates the system divided into modules by function, but at least two or more of modules 210-240 can be integrated into a single large AI model, such as an LLM or a platform model.
[0047] C-3-1. Analysis Module The analysis module 210 is a module that analyzes the user's "uniqueness" by analyzing past data and data acquired from the user. The analysis module 210 basically analyzes accumulated past data, but if the data is insufficient, it acquires further necessary data from the user and performs the analysis. Furthermore, the analysis of "uniqueness" is basically performed based on a comparison with a preset standard value. Below, we will explain each of the "analysis of past data," "acquisition of necessary data," and "comparison with standard values."
[0048] Analysis of Past Data: A user's social networking site contains a large amount of data (including profile data) that describes the user's characteristics. For example, a photo / video sharing social networking site may contain videos of the user's daily dance practice, or scenes of various sports such as baseball and swimming. The analysis module 210 allows any existing data uploaded by the user to the social networking site to be input as learning data for the AI.
[0049] The avatars used by users do not necessarily match the attributes of the users themselves. For example, a male user may use a female avatar, or an avatar of a different age or race. Therefore, the analysis module 210 also analyzes the context of the avatar. For example, in the case of a female avatar, the analysis module 210 analyzes characteristic movements from animations and videos and footage posted via the avatar.
[0050] Acquisition of necessary data: If the analysis module 210 does not have enough previously accumulated data, it uses the AI model to attempt to acquire the necessary data from the user. Specifically, the AI model directly or indirectly instructs the user on movements to acquire information about the user's movements. A direct instruction is an instruction in a form that the user can understand, such as "Show me your back." An indirect instruction is, for example, the AI model generating model or exemplary guide content and presenting it to the user to encourage the user to imitate the model or exemplary. Specifically, the AI model automatically generates funny video guide content to help the user acquire a natural smiling expression and presents it to the user to encourage the user to laugh naturally; automatically generates sad video guide content to help the user acquire a sad expression and presents it to the user to encourage the user to make a sad expression; or generates model or exemplary guide content for exercises and presents it to the user to encourage the user to exercise.
[0051] Comparison with standard values: The standard values to be compared with the acquired facial expressions and motions of the user are specifically the average values of physical expressions. The AI model then grasps deviations and fluctuations from the average values of physical expressions and understands them as the user's physical "habits" or "characteristics." Note that the "average values of physical expressions" take user attributes such as "cultural background," "gender," and "age" into full consideration. Specifically, the AI model learns the "average values of physical expressions" by classifying the user's cultural background, user attributes such as gender, content genre, etc. from tag information such as videos uploaded to a video-sharing SNS and accompanying comments.
[0052] C-3-2. Suggestion Module Based on the information obtained by the analysis module 210, multiple avatar movements that reflect the user's physical characteristics are generated. The suggestion module 220 suggests the generated multiple (e.g., two or three) avatar movements to the user as options. For the sake of convenience, the functional module that generates the avatar movements and functional modules related to the generation of avatar movements are not shown in FIG. 2.
[0053] The user selects the avatar movement they like from among multiple options proposed by the suggestion module 220. To further refine their selection, the user provides feedback on what they like and dislike about the selected avatar movement. The user can provide feedback in a highly abstract form, such as "like ____." The suggestion module 220 repeatedly proposes modified avatar movements based on the feedback, thereby creating avatar movements that capture the user's intentions.
[0054] C-3-3. Critic Module The critic module 230 has a group of critic AI models with various personalities discuss the avatar movements created by the proposal module 220 and change them to better movements. The group of AI models used here may include multiple AI models with different viewpoints for evaluating the avatar movements, such as an AI model that emphasizes expressiveness, an AI model that emphasizes rhythmic sense, an AI model that emphasizes originality, or even an AI model that is critical of the avatar movements.
[0055] C-3-4. Other Platform Deployment Module The avatar movements created through the critic module 230 are created for the avatar owned by the user and input into the analysis module 210. The other platform deployment module 240 extracts the features of the created avatar movements so that the movements can be reflected in avatars with different mechanisms, such as on other platforms, and obtains and exports information that can be deployed to other platforms.
[0056] C-4. Avatar Dance Generation System Figure 3 shows the main modules involved in avatar dance as a representative example of an avatar's non-verbal expression. The main functional modules involved in avatar dance are an analysis module 301, a generation module 302, a video generation module 303, a control module 304, a critic module 305, a proposal module 306, and a multi-platform deployment module 307. For ease of explanation, Figure 3 shows modules divided by function, but at least two or more of modules 301 to 307 can be combined and constructed into a single large AI model, such as an LLM or a base model.
[0057] The analysis module 301 is a functional module that analyzes the user's past videos to analyze physical features related to the user's "uniqueness." The generation module 302 is a functional module that allows the user to generate an avatar using an avatar creation tool.
[0058] The video generation module 303 is a functional module that generates a video of the avatar generated by the generation module 302 dancing. The control module 304 is a functional module that controls the avatar generated by the generation module 302. Specifically, the control module 304 controls the dance movements of the avatar so that the dance movements of the avatar generated by the video generation module 303 reflect the physical characteristics of the user obtained by the analysis module 301.
[0059] The critic module 305 has a group of critic AI models with various personalities, and has the group of critic AI models discuss the avatar dance video generated by the video generation module 303, and feeds back the results of the discussion to the control module 304. The control module 304 modifies the avatar dance to improve the movements based on the results of the discussion by the group of AI models. The group of AI models used here may include multiple AI models with different perspectives for evaluating the avatar's movements, such as an AI model that emphasizes expressiveness, an AI model that emphasizes rhythmic sense, an AI model that emphasizes originality, and even an AI model that is critical of the avatar's movements.
[0060] The suggestion module 306 suggests a plurality of avatar movements to the user as options based on the results of discussions by the critic AI models in the critic module 305. The user selects a favorite from the plurality of avatar movement options suggested by the suggestion module 306 and provides feedback of requests in a highly abstract form. The control module 304 controls the video generation module 303 to generate an avatar dance video that meets the user requests fed back from the avatar dance suggestion module 306.
[0061] The other platform deployment module 307 extracts the characteristics of the completed avatar's movement under the control of the control module 304, acquires information that can be deployed to other platforms, and enables it to be exported so that the movement of the avatar completed under the control of the control module 304 can be reflected in avatars with different mechanisms on other platforms.
[0062] FIG. 4 shows the basic configuration of a system 400 for moving an avatar.
[0063] A user can create a design for their own avatar using a production tool 401. The production tool 401 is an authoring device used to create content, or an information processing device on which authoring software runs. The avatar design created by the production tool 401 is registered in an avatar design database 402. In addition, a user can call up an avatar registered in the avatar design database 402 into the production tool 401 at any time and edit it.
[0064] The generation unit 403 generates content. Specifically, the generation unit 403 calls up an avatar design from the avatar design database 402 and generates an avatar to which a dance movement is to be assigned. The controller 404 controls the content generated by the generation unit 403. Specifically, the controller 404 causes the avatar generated by the generation unit 403 to perform movements such as dancing based on external instructions. The rendering engine 405, under the control of the controller 404, renders 3D animation of the content generated by the generation unit 403. Specifically, under the control of the controller 404, the rendering engine 405 renders 3D animation of the dance video of the avatar generated by the generation unit 403.
[0065] In the controller 404, the basic motion model for reproducing the user's motion with an avatar based on motion tracking is as follows.
[0066] (1) Motion Capture (Tracking): The motion capture unit 406 uses a motion capture system to capture the physical movements of an actual person. The motion capture system uses sensors and cameras to capture physical movements, facial expressions, and the like in real time. (2) Data Processing: The motion capture unit 406 processes the data obtained from the motion capture system to generate motion data including parameters such as the position, angle, and speed of each joint. (3) Animation Creation: The controller 404 uses the motion data to control the movement of the avatar. The controller 404 sets the character's posture and movements based on the motion data to achieve smooth avatar movement. (4) Real-Time Movement: To control the avatar in real time, a system is required that reflects the motion data obtained by motion capturing in the avatar animation in real time. This makes it possible for the avatar to move in accordance with the user's movements.
[0067] In the above-described operation model, the controller 404 uses motion data obtained by motion capture to make the avatar dance. In this case, the user's dancing skills are directly expressed by the avatar, so there is a problem that if a user only has average dancing skills, the avatar's dancing will be awkward (i.e., the avatar will not dance well in real time).
[0068] In contrast, if the basic movement data is created from model data, the controller 404 can easily make the avatar dance. Model data can be, for example, bio-kinetic data, videos of dance instructors, or videos posted on social media. However, in this case, there is no way to add "personality" to the avatar, that is, individuality, and the dance movements of all users end up being uniform.
[0069] Furthermore, it is also possible to reproduce the user's facial expressions on an avatar based on face tracking by the movement acquisition unit 406, using the same motion model and example data as above.
[0070] C-5. Video Creation System Adding Personal Characteristics Figure 5 shows the configuration of a video creation system 500 that creates videos such as avatar dancing. The main feature of video creation system 500 is that it further includes a function for adding personal characteristics to the avatar dancing, in addition to the basic configuration of the avatar dancing system shown in Figure 4. First, the operation of video creation system 500 will be generally described.
[0071] The motion acquisition unit 501 performs motion tracking and face tracking on a dance video of an actual person to capture body movements, facial expressions, and the like in real time, and performs data processing such as spatial mapping of motion vectors and feature data formatting to generate personal motion data of the person. The motion acquisition unit 501 may also perform similar processing on dance videos accumulated in the past (for example, acquired from a video sharing platform) instead of real-time videos of an actual person to generate personal motion data of the person. The personal motion data generated by the motion acquisition unit 501 is stored in a non-volatile manner in a predetermined storage unit (not shown).
[0072] The reference motion data generation unit 512 generates reference motion data, which is average or exemplary motion data. The reference motion data generation unit 512 generates reference motion data by capturing videos of dance instructors or dance videos posted on social media and analyzing the dance motion. The motion data generated by the reference motion data generation unit 512 is stored in a non-volatile manner in a specified storage unit (not shown).
[0073] The individual feature extraction unit 502 extracts, as individual features, the difference between the individual motion data extracted by the movement acquisition unit 501 and the reference motion data obtained from the reference motion data generation unit 512. The individual feature extraction unit 502 can capture individual features from the direction of movement of body parts (head, hands, feet, each joint) and their speed (the time from movement to stopping, or vice versa). The individual feature extraction unit 502 can also generate individual features from data such as the jump (height) and rotation (angular velocity) of the body.
[0074] A user designs their own avatar using the production and control tool 503. The avatar designed by the production and control tool 503 is registered in the avatar design database 504. The user can also use the production and control tool 503 to instruct the avatar's dance sequence. The user here may be the user who uses the avatar and the subject of the dance video. Alternatively, the user of the production and control tool 503 may be a user who specializes in designing avatars and instructing dance sequences, different from the user who uses the avatar. The production and control tool 503 is an authoring device used to create content, or an information processing device on which authoring software runs.
[0075] The generation unit 505 generates content. Specifically, the generation unit 505 generates an avatar based on an avatar design retrieved from the avatar design database 504. The generation unit 505 may cause a trained model (generative AI) to generate the avatar. The controller 506 then retrieves the avatar from the avatar database 504 and controls the dance movements of the avatar based on the dance sequence instructed by the production and control tool 503.
[0076] When avatar dance faithfully follows a dance sequence, the movements of all users' avatars end up being uniform. Therefore, the controller 506 adds individuality to the avatar dance using the personal features extracted by the personal feature extraction unit 502. The controller 506 also controls the amount of personal features added to the avatar dance. For example, the controller 506 controls the amount of personal features added to the avatar dance based on emotions, the presence or absence of an audience, the level of intimacy with the audience, parental control, etc.
[0077] The controller 506 performs parental control on the avatar according to the user's internet literacy and the risk of the social community. The user's internet literacy is determined based on the user's behavior or conversations on the internet, in games, on purchasing sites, and in social communities, and the risk of the social community is determined according to the user's literacy. Furthermore, the controller 506 may control the generation unit 505 to change the avatar generated in order to eliminate anxiety or risk factors from various perspectives, such as ethical, cultural, customary, linguistic, and religious, in the remote environment where the avatar is displayed.
[0078] The rendering engine 507, under the control of the controller 506, renders a 3D animation of the content generated by the generation unit 505. Specifically, under the control of the controller 506, the rendering engine 507 renders a 3D animation of the avatar generated by the generation unit 505 dancing.
[0079] The video generation system 500 further includes a function for evaluating the avatar dance generated by the rendering engine 507 and a function for modifying the dance based on the evaluation results. A feature extraction unit 508 analyzes the dance and other movements of the avatar generated by the rendering engine 507 and extracts feature amounts. An evaluation and supervision unit 509 then evaluates the avatar dance using the feature amounts of the avatar's dance movements extracted by the feature extraction unit 508.
[0080] The evaluation and supervision unit 509 acts as a "supervisor," so to speak, using an AI model to evaluate the generated 3D avatar dance animation. The evaluation and supervision unit 509 has a group of critic AI models with various personalities discuss the avatar dance and output an evaluation of the avatar dance. The group of AI models used here may include multiple AI models with different perspectives for evaluating the avatar's movements, such as an AI model that emphasizes expressiveness, an AI model that emphasizes rhythmic sense, and an AI model that emphasizes originality.
[0081] The evaluation and supervision unit 509 generates and outputs feature correction data for improving the avatar dance based on the evaluation results from the AI model. The evaluation and supervision unit 509 may generate the feature correction data by appropriately using the feature database 510. After evaluating the avatar dance, the evaluation and supervision unit 509 may generate appropriate comments and advice, recommend them to the user, and present them to the user via a UI (User Interface) unit 511 to prompt the user to make a selection. The user selects the advice they wish to adopt on the GUI of the UI unit 511. The UI unit 511 generates and outputs feature correction data for improving the avatar dance based on the advice selected by the user. The UI unit 511 may also generate feature correction data for improving the avatar dance based on information obtained through multimodal interaction with the user. Methods for improving the avatar dance include a method of modifying an avatar dance that has already been generated and a method of regenerating the avatar dance.
[0082] The individual feature extraction unit 502 receives feature correction data from the evaluation and supervision unit 509 or the UI unit 511 and corrects the individual features to improve the avatar dance. The controller 506 uses the corrected individual features to control the rendering engine 507 so as to add individuality to the avatar dance and improve it.
[0083] The following describes in detail the characteristic operations of each functional module of the video generation system 500.
[0084] C-5-1. Adjustment of the Amount of Added Personal Characteristics The controller 506 controls the amount of added personal characteristics to the avatar dance when adding individuality to the avatar dance. For example, the controller 506 controls the amount of added personal characteristics to the avatar dance based on emotions, the presence or absence of an audience, the level of intimacy of the audience, parental control (described below), etc.
[0085] For example, when dancing with strangers, a user may want to minimize personal characteristics. Also, when dancing with only close friends, it is possible to increase the additive level of personal characteristics to perform a more distinctive dance. The relationship level with fellow dancers or audience members can be set, for example, by operating an app on a smartphone. Of course, the relationship level may also be set automatically based on past history data. Furthermore, the relationship level of an avatar may also be set automatically based on the frequency of meetings, using the ID or name of the other person's avatar.
[0086] The controller 506 may assign an addition level of the personal feature amount according to the relationship level with a person. For example, the relationship level with a person may be expressed as a value from 1 to 10, with relationship level 1 being "loose" and the intimacy increasing as the level value increases, with relationship level 10 being "close." Furthermore, the addition level of the personal feature amount may be expressed as a value from 0 to 10, with addition level 0 meaning that no personal feature amount is added, and the amount of addition increasing as the level value increases, with addition level 10 being defined as the maximum amount of addition of the personal feature amount.
[0087] C-5-2. Extracting Movements from Dance Videos Here, we will explain the process of extracting movements from dance videos in the movement acquisition unit 501. Figure 6 shows a functional block diagram for extracting movements from a user's dance video in the movement acquisition unit 501.
[0088] The posture determination unit 601 captures images of each frame from the input dance video and determines the dancer's posture from the captured images. For example, a known skeletal information acquisition technology such as Openpose (see Non-Patent Document 1) may be used to determine posture from the captured images. Using Openpose, it is possible to detect human joint points simply by inputting a still image, and by displaying the joint positions as points, it is possible to extract skeletal data such as knee angles and elbow throws.
[0089] The skeleton point determination unit 602 determines the skeleton points by taking the following features and inputs them to the feature vector extraction processing unit 603. The feature vector extraction processing unit 603 is configured with a trained model such as an autoencoder. The encoder 603A creates individual features from the image with the determined skeleton points. The decoder 603B decodes the individual features to generate motion data and outputs it as dance control data.
[0090] C-5-3. Extraction of Personal Features The personal feature extraction unit 502 extracts personal features by comparing the personal motion data extracted by the movement acquisition unit 501 with reference motion data. Fig. 7 shows a functional block that allows the personal feature extraction unit 502 to compare the personal motion data with the reference motion data and generate personal motion data. In the example shown in Fig. 7, the personal feature extraction unit 502 is configured to extract personal features using a trained model such as an autoencoder.
[0091] The encoder 701 receives the personal motion data and the reference motion data and creates personal features that are the difference between the two. The decoder 702 decodes the personal features and outputs the motion data of the personal features to the controller 506.
[0092] The individual feature extraction unit 502 can also extract individual features on a part-by-part basis. FIG. 8 shows a functional block diagram for extracting individual features on a part-by-part basis. In the illustrated example, the individual motion data extracted by the movement acquisition unit 501 is divided into parts such as the face, right arm, left arm, body, right leg, and left leg, and input to an encoder corresponding to each part. The encoder 701 in FIG. 7 includes a face encoder 801, a right arm encoder 802, a left arm encoder 803, a body encoder 804, a right leg encoder 805, and a left leg encoder 806. The face encoder 801 extracts face motion features, the right arm encoder 802 extracts right arm motion features, the left arm encoder 803 extracts left arm motion features, the body encoder 804 extracts body motion features, the right leg encoder 805 extracts right leg motion features, and the left leg encoder 806 extracts left leg motion features.
[0093] Since motion features are extracted for each body part in this way, correction can be made for each body part when the decoder 810 (corresponding to the decoder 702 in FIG. 7) decodes the data to generate motion data. In the example shown in FIG. 8, hand motion correction data is input to the decoder, and motion data with the hand motion corrected is output. The hand motion correction data is correction data generated by the evaluation and supervision unit 509 based on the evaluation results using an AI model, or correction data generated by the UI unit 511 based on advice selected by the user.
[0094] C-5-5. Extracting Features of User's Dance Performance In order to extract user's features, it is effective to have the user perform a task and extract features from their movements.
[0095] Fig. 9 shows the configuration of a system for analyzing and determining user features. In the system shown in Fig. 9, the user is basically instructed to assume various poses to extract features, or interesting content is presented in the form of images or audio, and the user's poses and facial expressions at those times are captured and the features are analyzed. The user's poses and facial expressions are then compared with standard data, and personal features are determined based on the differences.
[0096] Task elements such as the poses and facial expressions to be adopted by the user are instructed via the production and control tool 503. For example, routine tasks for acquiring personal features, such as "Raise your hand" or "Perform a simple dance," are presented to the user on screen or by voice. The user takes the instructed pose in front of the camera, and the movement is captured to obtain motion data.
[0097] Furthermore, if the instructed task is not a physical movement such as a pose or dance but rather the acquisition of facial expression features resulting from emotions such as a smile, it is difficult for the user to create a natural smile or other facial expression even if the user is instructed to "smile" via screen or voice. Even if facial motion data obtained by capturing an unnatural smile or other facial expression such as a forced smile is used, it is not possible to extract the individual features of the user's natural smile. Therefore, when the instructed task is a facial expression rather than a physical movement or pose, a method may be used in which guide content is presented to the user to guide the user to easily create the instructed facial expression, instead of directly instructing the user on the task.
[0098] In the system shown in FIG. 9 , AI is utilized in the process of presenting a user with guide content that guides the user to easily create a specified facial expression. Examples of the AI used include a large-scale language model (LLM) and generative AI. By utilizing these AIs, it is possible to generate "texts and images that make the user laugh (naturally)" in response to a task instruction such as "laugh." For example, by inputting a prompt such as "texts and images that make the user laugh (naturally)" into the large-scale language model 901, the large-scale language model 901 can present the user with funny content searched for on the Internet. Furthermore, it is possible to present the user with an "image that makes the user laugh (naturally)" generated from the text "laugh" using a Text2Image model 902, which is one of the generative AIs.
[0099] The movement acquisition unit 501 captures the user's movements and facial expressions in response to the presentation of a task or content, and generates motion data of the movements and facial expressions. The individual feature extraction unit 502 then uses an autoencoder to extract motion data of the instructed movements and facial expressions, and compares the extracted motion data of the movements and facial expressions with standard feature data to extract individual features of the instructed movements and facial expressions.
[0100] C-5-6. Description of Reference Motion Data The individual feature extraction unit 502 uses average or exemplary reference motion data to extract individual features from the individual motion data. The reference motion data generation unit 512 analyzes movements from many videos and finds commonalities, thereby generating standard reference motion data suitable for extracting individual features.
[0101] FIG. 10 shows how the reference motion data generating unit 512 generates reference motion data representing standard dance movements from dance video.
[0102] The video classification unit 1001 classifies model data such as videos of dance instructors and dance videos published on social media into groups of similar dance videos based on music, tempo, and posture. The video capture unit 1002 then captures the dance videos for each group.
[0103] The posture determination unit 1003 determines the posture of the dancer from the captured dance video using posture determination technology such as Openpose. The skeleton data extraction unit 1004 then extracts skeleton data of the dancer based on the posture and generates dance motion data.
[0104] Dance motion data is input to the autoencoder 1005, and dance features are extracted by the previous encoder. The autoencoder 1005 compares two videos and determines that the movements are the same if the closest point in one video frame and the closest point in the other video frame are deemed to be in the same category. Furthermore, an averaged movement is extracted by selecting the closest point that minimizes a loss function using the nearest neighbor method. The similar video motion extraction unit 1006 then extracts motions of similar videos from the output of the decoder downstream of the autoencoder 1005, and the similar video feature averaging unit 1007 calculates the average of the features of similar videos output by the previous encoder to create reference motion data. The created reference motion data is stored in the dance feature database 1008.
[0105] C-5-7. Avatar Dance Improvement The evaluation and supervision unit 509 evaluates the 3D animation of the avatar dance using an AI model and generates appropriate comments and advice. Fig. 11 shows a mechanism for improving the avatar dance in the video generation system 500 shown in Fig. 5, i.e., a mechanism for the evaluation and supervision unit 509 to recommend avatar dance features and correct the avatar's own features.
[0106] The individual feature extraction unit 502 extracts individual features by comparing the individual motion data extracted by the motion acquisition unit 501 with the reference motion data generated by the reference motion data generation unit 512. The controller 506 uses the individual features extracted by the individual feature extraction unit 502 to control the dance movements of the avatar so as to add individuality to the avatar dance.
[0107] The rendering engine 507 renders a 3D animation of the avatar dancing under the control of the controller 506. The feature extraction unit 508 analyzes the dance movements of the avatar generated by the rendering engine 507 and extracts dance feature amounts. The evaluation and supervision unit 509 then evaluates the avatar dance using the avatar dance feature amounts extracted by the feature extraction unit 508. The evaluation and supervision unit 509 has a group of critic AI models with various personalities discuss the avatar dance and output an evaluation of the avatar dance.
[0108] The evaluation and supervision unit 509 generates and outputs feature correction data for improving the avatar dance based on the evaluation results from the AI model. After evaluating the avatar dance, the evaluation and supervision unit 509 may also generate appropriate comments and advice and present them to the user via the UI unit 511. In this case, the user can select the advice they want to adopt via the GUI screen of the UI unit 511. In the example shown in FIG. 11 , the evaluation and supervision unit 509 outputs hand movement correction data to the individual feature extraction unit 502 based on the feature of "right hand movement" extracted from the dance feature database 1008. Alternatively, the UI unit 511 presents a UI screen for the user to select whether or not to adopt the advice generated by the evaluation and supervision unit 509, "Your hand movements should be a little sharper," and outputs hand movement correction data to the individual feature extraction unit 502 based on the user's selection. The individual feature extraction unit 502 then inputs feature correction data from the evaluation and supervision unit 509 or UI unit 511 and corrects the individual feature in order to improve the avatar dance.
[0109] The individual feature extraction unit 502 corrects the individual features to improve the avatar dance when feature correction data is input from the evaluation and supervision unit 509 or the UI unit 511. The controller 506 uses the corrected individual features to control the rendering engine 507 to add individuality to the avatar dance and improve it.
[0110] D. Communication Platform The communication platform proposed in this disclosure serves as a platform for virtual community environments, including SNS, by creating a closed space based on an AI model, constructing a real-time, always-on virtual environment, and autonomously filtering, converting, and adjusting information using generating AI.
[0111] A communication platform can create a safe virtual community environment for each service, eliminating anxieties and risk factors from various perspectives, such as ethics, culture, customs, language, and religion, even when connected at all times. Providing such a virtual community environment is expected to foster cooperation and empathy, increasing the scope for creativity and activity. It also increases engagement in the virtual community space, increasing opportunities for business creation within this space.
[0112] The communication platform will be positioned as a constantly connected platform, and will implement an AI-based safety management system. The communication platform will also function as a basic service like an operating system, allowing various services to be deployed on this platform.
[0113] The AI-based platform layer can also be configured to perform basic ethical risk control and security management. This service can be operated borderlessly, with cloud services uploading data as needed, and incident information can be collected borderlessly and applied to updating learning data.
[0114] D-1. Basic Configuration Figure 12 shows an example of individual services implemented on a communication platform. The communication platform layer is configured as a platform for virtually always-on connection and controls engagement and risk in communication using AI. The communication platform includes functional modules such as a renderer (REN) that generates avatars (as described above). In addition, a layer that controls UI and interaction is interposed between the communication platform and the service layer. The upper service layer includes a content presentation layer and an extended service layer. The content presentation layer includes media services that share UGC (User Generated Content). The extended service layer consists of extensible applications and performs existing media capture and collaborative services. The service layer can provide services while tracking the sense of distance in real time according to the connection distance determined by the communication platform.
[0115] D-2. Basic Operation Figure 13 schematically illustrates an example of a basic operation for sharing content between users' information devices (e.g., smartphones) on a communication platform. The illustrated basic operation provides a constantly connected environment via a network, allowing users to flexibly capture themselves and input speech using information devices such as smartphones to communicate with other smartphone users. Based on an AI model, the communication platform operates basic services such as person identification, person tracking, background identification, background sound separation, and person speech separation in communication between users, for example, via their smartphones. As described with reference to Figure 1, the communication platform generates avatars by performing information detection, analysis, and discrimination on each imported data element. Specifically, the communication platform uses AI services to identify people, track people, distinguish backgrounds, and detect and separate background noise and speech, collecting and analyzing basic information for the virtual community environment and forming the basis for the real-time environment. The other party in the conversation quantifies the relationship between the conversing people and stores it as data in a profile. The default setting may be set to the lowest level of relationship. In this case, the other party will also have the most general avatar (i.e., no individuality added) and the user's spoken voice will be converted into conversational voice by AI and played back.
[0116] D-3. Switching Avatars According to Trust Levels Figure 14 shows a mechanism for switching avatars according to the trust and intimacy between interlocutors. User profiles are constantly updated on the communication platform. When trust levels exceed a certain value, avatars may be switched in stages and discontinuously. Specifically, a realistic person's appearance is projected onto the avatar's appearance. Furthermore, discontinuous switching can be controlled by adjusting the distance on a vector plane between the avatar's speech and the real voice. Avatars and speech can be generated using generation AI. Avatars are generated based on the situation, person emotions, authenticity determination, and rights protection determined by person identification, person tracking, background discrimination, and background sound separation performed by AI services.
[0117] To control the ratio of avatars to real images, it is important to determine the trust level of participating communities and between participants. In virtual communities with multiple conversation partners, avatars are generated and sent in a state (ratio) that matches each other's individual trust level. For example, if a partner is participating only with an avatar and generated voice, communication can be maintained while maintaining a similar distance. While it is necessary to measure the trust level of each partner and manage risks, one method is to use a phased approach, such as not disclosing a participant's detailed information to other participants until a certain amount of data has been collected. Trust levels are expected to be determined using analysis of speech, eye movements, facial expressions, and conversation context. AI can be used to further enhance environmental analysis.
[0118] D-4. Collaboration with External Services Figure 15 shows a mechanism for linking avatar generation using generation AI with existing avatar generation services. The avatar generated by AI on the communication platform and an avatar provided by another service may be switched between manually or automatically by the user. Users may create their own avatars or may be able to bring in avatars already in use on other services. The communication platform may support OpenXR to reflect features of glTF (GL Transmission Format), USD (Universal Scene Description), and interaction. The communication platform's extensibility, which enables collaboration with external parties, can be expanded to include value-added businesses such as advertising.
[0119] D-5. Risk Avoidance in Virtual Community Environments Figure 16 shows an example of a communication platform that builds a virtual community environment that spans culturally and religiously diverse groups and countries. Specifically, the sphere of religion A and language A and the sphere of religion B and language B are constantly connected via the communication platform. In such cases, various issues, such as diverse diversity issues and differences in ethical standards, can lead to unexpected risks, a topic that has been frequently discussed in metaverse discussions. The communication platform disclosed herein can identify risky terms, sentences, and contexts according to language and religious profiles, and can remove and prevent communication of terms and behaviors that are acceptable in one sphere but prohibited in another. Furthermore, according to the present disclosure, in a virtual community environment connected to a sphere to which the European AI regulatory bill, which focuses on respect and protection of human rights, applies, terms and behaviors that may pose a risk of discriminatory effects or unfair bias prohibited by law can be removed and prevented from being communicated to the other party. A specific method for removing prohibited terms and behaviors is to build a database on the service platform, update it, and determine a matching state by combining multiple factors. It is also possible to use a certain threshold as a criterion for judgment. Recent AI-based judgments have been effective. This prevents conversational accidents due to cultural and religious differences across borders. For example, forbidden words and behaviors can be filtered out as noise to prevent them from being conveyed to the other party. These processes are performed in real time. To prevent computational overhead, instead of changing the avatar for each user, the output side can identify (filter) specific religious groups or dangerous destinations in advance and convert only those. For example, information on the other party's nationality and region can be obtained, and switching can be performed in advance based on language, etc. However, if the generated avatars cross borders (in other words, if there are many avatars), changing the appearance or speech content of multiple avatars could cause the system to experience computational overhead.To avoid such problems, it is desirable to specify in advance the country or region to which the output will be sent, and, for example, only when the country or region to which the output will be sent is determined to be a high-risk country designated in advance or when AI or other means determine that the country or region is a high-risk country, change the avatar's appearance, speech, behavior, etc. in accordance with the ethics, culture, customs, language, religion, or rules of that country.
[0120] Due to differences in not only terminology and behavior but also common sense between spheres, it may be preferable from the perspective of risk avoidance to make some of the environmental information present in the shared virtual space invisible to other spheres or to subtly hide it so that it is only visible to certain spheres. Figure 17 shows an example in which a religious object 1701 that is visible in the virtual space of religious sphere A is replaced with another object 1702 that does not pose a religious risk and displayed in the virtual space of religious sphere B.
[0121] FIG. 18 shows a schematic functional configuration for processing religiously risky objects in a shared virtual space.
[0122] The image input unit 1801 inputs an image of a virtual space from one of the areas. Then, the feature extraction unit 1802 extracts features from the input image. The feature extraction unit 1802 can also obtain vector values from the image contours and use them as feature data. Next, the object discrimination unit 1803 discriminates objects included in the image based on the feature data consisting of the vector values of the image contours. The object discrimination unit 1803 may discriminate objects from the image (virtual space) using an AI model (not shown).
[0123] The prohibited object identification unit 1804 identifies objects that are prohibited in the other party's virtual space from among the objects detected in the image (virtual space) by the object discrimination unit 1803. Specifically, a "prohibited" object is an object that, due to cross-national ethical, cultural, customary, linguistic, or religious differences, could pose a risk of causing a conversation accident if it appears in the other party's virtual space.
[0124] The prohibited object identification unit 1804 identifies objects that are prohibited in the other party's virtual space using the AI engine 1805 and the participant's personal profile data 1806. The AI engine 1805 is assumed to have previously learned teacher data that are culturally prohibited images, teacher data that are religiously prohibited images, teacher data that are ethically prohibited images, and teacher data that are prohibited images from other perspectives.
[0125] The AI engine 1805 inputs the participant's personal profile data and determines the risk of each object determined by the object determination unit 1803 from the perspectives of culture, religion, and ethics in the participant's sphere. A certain threshold value may be used to determine the risk. The prohibited object identification unit 1804 then identifies prohibited objects in the virtual space of other spheres based on the risk determination results for each object by the AI engine 1805. In the input image denoted by reference numeral 1811, the object surrounded by a dashed ellipse is the prohibited object identified by the prohibited object identification unit 1804.
[0126] The image correction unit 1807 performs a correction process to replace the identified prohibited object in the image with another object that is not prohibited (non-risky). The image correction unit 1807 may apply AI to the image correction. The AI engine 1808 estimates a correction policy based on the participant profile data 1806 of the other group. For example, the AI engine 1808 estimates an object that is less risky and can be replaced. Then, the image correction unit 1807 performs an image correction process to replace the risky prohibited object with a safe object using the object estimated by the AI engine 1808, and outputs the corrected image. In the output image indicated by the reference numeral 1812, the object surrounded by a dashed ellipse is the object replaced from the prohibited object by the image correction unit 1807.
[0127] A specific method by which the feature extraction unit 1802 acquires feature data will be described with reference to FIG. 19. First, as shown in FIG. 19A, a 16 x 16 region around a keypoint in the input image is extracted. Next, as shown in FIG. 19B, the image is rotated based on the "representative orientation" of the keypoint. Then, as shown in FIG. 19C, the image is divided into 4 x 4 blocks to create 16 small blocks. A gradient histogram with bins in eight directions is created for each small block. This results in a total of 128 values, which are expressed as vectors and used as feature data for the image.
[0128] The functional configuration shown in FIG. 18 may be used to automatically adjust risk avoidance based on the age of participants. For example, if a community includes adults or younger, the virtual community environment may be presented to the non-adult participants after images inappropriate for non-adults in the virtual community environment have been removed or modified to images appropriate for non-adults. The AI engine 1805 can estimate objects inappropriate for non-adults by referencing the participant profile data. The AI engine 1808 can also estimate images appropriate for non-adults by referencing the participant profile data. In addition to image modification, the communication platform can also filter audio based on environmental sounds and speech context.
[0129] Images and audio captured from the real world may contain high-risk information, such as personal information. In such cases, a functional configuration similar to that shown in FIG. 18 may be used to identify and separate the high-risk information and delete or replace it with other information before transmission. Risk assessment may be performed using a certain threshold value with reference to participant profile data, or an AI model may be used to automatically assess and delete the information.
[0130] D-6. AI-Based Proxy Solution While it is possible to build a virtual community environment on a communication platform based on the present disclosure, time lag issues arise when providing a borderless service. To address this issue, an AI-based proxy solution can be provided, with avatars that have individual characteristics. The communication platform provides seamlessly transformable avatars, as shown in Figure 14. By giving this function an AI proxy role, communication can be performed on behalf of the user. Conversations and interactions are recorded, and key points can be automatically extracted and reviewed in time-shift mode. In this case, the reliability of the AI function is extremely important. Safety features can be included to ensure that individual profiles do not exceed unique characteristics and to temporarily suspend risky conversations and interactions. Meanwhile, a support function for suspending responses can be implemented by automatically generating and providing multiple possible answers, enabling rapid responses.
[0131] In conventional SNS, once a message is sent, the other party's actions are left to their own devices until they respond. Individuals make the delicate decision of whether or not they want to continue the conversation, which is reflected in the response time and text. When using an avatar as a proxy, it is possible to reflect various changes in physical condition in real-time updates to the profile, or, if you want to maintain a sense of distance, you can set the system to fix it to a mechanical proxy and monitor the situation. These settings can also be controlled by the real-time characteristic update function.
[0132] D-7. Application In this disclosure, personal data is aggregated in the functions of a communication platform such as the cloud and utilized for analysis and the generation of response actions. Therefore, in addition to viewing it as personal information, analyzing a variety of borderless events can also lead to risk assessment functions and contribute to safety. Personal data can also be managed using AI.
[0133] The communication platform according to the present disclosure can generate personalized avatars using AI, and can therefore be an example of a method for implementing an avatar with agentic capabilities.
[0134] A variety of applications and business models can be developed and expanded on the always-on communication platform disclosed herein. Various developments are possible as functions of avatars with agent-like characteristics. Examples of applications of the communication platform include:
[0135] ・Presenting advertising media in virtual space ・Creating and sharing content on communication platforms ・Sharing games and attractions on communication platforms ・Developing educational content services ・Sharing travel-related information, guidance, and sharing empathy
[0136] D-8. Information Processing Procedure The functional configuration of the communication platform according to the present disclosure is also shown in Figure 1. As shown in Figure 20, the processing flow performed by the communication platform can be roughly divided into four layers: a detection layer by the information detection unit 111, an analysis layer by the analysis unit 112, a discrimination layer by the discrimination unit 113, and a generation layer by the generation unit 114.
[0137] In the detection layer, the information detection unit 111 collects information using various detection means and sends it to the analysis layer. For example, images captured by a camera equipped in a device such as a user's smartphone, sensing information (personal characteristics, environment, etc.) by other sensors, and SNS information acquired via the smartphone are collected.
[0138] In the analysis layer, the analysis unit 112 analyzes conversations, speech characteristics, background noise, background, facial expressions, movement characteristics, etc. in the virtual community based on the information collected in the detection layer. AI may be used for the analysis.
[0139] The analysis results generally need to be processed by ranking the status, categorizing, cultural grouping, etc. In the discrimination layer, the discrimination unit 113 performs person identification, emotion, preference, behavior prediction, etc. of users in the virtual community based on the analysis results.
[0140] In the generation layer, the generation unit 114 generates images, audio, and context based on the discrimination results. Each generation task may be broken down into detailed sections and processed in parallel. The output information from the generation layer changes in real time and is linked to the generation of the avatar. As a result, the distance between the components of the avatar and the characteristics of a specific person changes adaptively in vector space, and the similarity of the generated avatar changes. The data used may be registered in a database as profile data for each participant and updated sequentially.
[0141] D-9. Mechanism for passing through security gates As shown in Figure 21, the communication platform provides a mechanism for passing through a security gate before entering a closed space.
[0142] The security gate applies AI services and authentication. Authentication may be performed in real time. Considering cases where the interface is exposed to the outside, such as on smartphones, if an inconsistency is detected during real-time security authentication, a change in the avatar communicated to the authenticating party may be coordinated. For example, if an authentication error occurs for a party, the avatar may be immediately reverted to the default setting, and the person's characteristics may be removed from the avatar.
[0143] For identity authentication, a system that uses biometric authentication to exclude actions by anyone other than the person in question may be adopted. For example, when incorporating face authentication, instead of sending a facial image, which is a heavy load, to an authentication server, the device may perform face detection for authentication and send the acquired face authentication metadata to the authentication server, thereby reducing the amount of data transferred. Furthermore, biometric authentication can also be replaced by fingerprint authentication, which is constantly sensing.
[0144] D-10. Example Fig. 22 shows an outline of an example in which voice and a camera are used for information input. The input information further includes SNS history and location information.
[0145] The audio and image information is converted into information data in the respective feature extraction blocks and sent to the information processing block. In the information processing block, audio processing involves identifying conversation sentences, identifying speech features, and separating background noise. In image processing, background images are separated. The results of the information processing are generated as information data representing multiple states and sent to the discrimination block. In the discrimination block, information processing is performed using the state data from the audio and video to estimate a person's emotions, profile, preference information, and behavioral pattern discrimination.
[0146] The risk determination block can also perform risk determination (or risk prediction) regarding behavioral inference based on information from a risk determination database that is stored separately from the judgment. Specifically, the risk determination block determines the user's Internet literacy based on the user's behavior or conversations on the Internet, games, purchase sites, and social communities, and determines the risk of the social community according to the user's literacy. Risk determination will be described in detail in the following section E-3.
[0147] The basic element generation block uses AI to generate an avatar based on the environmental information and inferred information obtained from the discrimination block, and also generates a virtual background (environment) and background noise. In addition, the application generation block can also generate optimized advertisements and service content as applications, and can also convert language and risky vocabulary in this process.
[0148] Figure 23 shows in detail the information processing flow, focusing on the voice information processing block, of the embodiment shown in Figure 22. In this processing flow, an avatar is generated that reflects behavioral features predicted from voice input, for example, via the microphone of a smartphone used by a user. In Figure 23, the part surrounded by a dashed line can be realized using an AI model.
[0149] First, a process is performed to separate the input speech from background noise. Then, the separated speech is subjected to speech-to-text conversion and conversation extraction processes, as well as speech feature extraction and speech feature identification processes. Then, an emotion analysis of the person who spoke is performed based on the resulting conversation and speech feature information. In addition, characteristics of the background sound are extracted after noise has been removed. The background sound characteristics include information about attractions such as TV and music playback. Preference information of the person is obtained and identified from the characteristics of the background sound. Person identification is performed based on the emotion analysis and preference information of the person obtained as described above.
[0150] The person identification results obtained in the audio information processing block are then mixed with similar person characteristic data obtained from the image information processing block (described later) and the SNS activity analysis block (described later) to comprehensively predict the person's behavioral characteristics. The behavioral characteristics match the person profile. The results of this person behavioral characteristic prediction process are reflected in the risk judgment in the risk judgment processing block and then sent to the avatar generation block. The risk judgment processing block performs risk judgment taking into account third party and environmental information in addition to the predicted person behavioral characteristics. In the risk judgment process, parameters indicating the degree to which realism and behavioral characteristics are reflected in the avatar are continuously changed and reflected in the avatar. The avatar generation block generates an avatar taking into account the predicted person behavioral characteristics and the risk judgment results.
[0151] Figure 24 shows in detail the information processing flow, focusing on image processing, of the embodiment shown in Figure 22. In this processing flow, an avatar is generated that reflects behavioral features predicted from an image taken with the camera of a smartphone used by the user, for example. The input image is a moving image. In Figure 24, the part surrounded by the dashed line can be realized using an AI model.
[0152] First, the input image is separated into images of people and other images. Person separation is performed based on personal data such as facial and physical information (volumetric features, etc.). Person separation can be performed using AI matting processing. Clothing analysis, behavioral pattern analysis, and facial expression analysis are performed on the separated person images. In behavioral pattern analysis, for example, volumetric capture is used to separate the person's movements using vector information, etc., to determine the characteristics of the behavioral pattern (features such as dancing, standing still, signaling, etc.). Emotion classification and determination are also performed using face recognition processing using AI that learns from facial expression changes due to emotions.
[0153] Then, by using the characteristic information about these people, the data is narrowed down to person status data in real time. For example, by having the AI judge the state of a person as "warm-colored casual clothing, a bright expression, and a large-moving gait," it is determined that the person is in a positive, active mental state. This process is performed by time sampling, and by continuously observing the data, the transition of the person's status changes can be obtained over time.
[0154] In parallel with the person analysis process, image feature extraction is performed on the remaining image after subtracting the person, followed by background image analysis. Background identification and object identification are performed in parallel as part of the background image analysis. This is also performed using AI matting. Based on the results of background identification, background features are classified, i.e., environmental information such as the color of the background image, the wall pattern, and whether or not it is indoors is determined. Objects identified through object identification are then classified to identify surrounding furniture, tableware, electrical appliances, paintings, ornaments, etc. By setting up and processing AI judgment processes for each object category, such as furniture and electrical appliances, in parallel, it is possible to quickly classify each identified object. Then, a person's objective preference trends are analyzed from multiple information sources, namely background features and object classification. For example, a monotone color preference can be classified as a preference for simplicity.
[0155] By integrating the real-time status data of a person, consisting of information about the person (person status data) acquired from the image as described above and information about the surrounding area other than the person (person preference tendency), with similar person feature data obtained from the audio information processing block (described above) and the SNS activity analysis block (described below), the behavioral features of the person can be predicted with higher accuracy. For example, the image information processing block determines that the person is close to dancing to music. The results of this behavioral feature prediction process are reflected in the risk assessment process and then sent to the avatar generation block. The risk assessment process block performs risk assessment by taking into account the predicted person behavioral features, as well as third-party person and environmental information. In the risk assessment process, parameters indicating the degree to which realism and behavioral features are reflected in the avatar are continuously changed and reflected in the avatar. The avatar generation block generates an avatar by taking the risk assessment results into account for the predicted person behavioral features. Therefore, the results of the person's behavioral feature prediction are reflected in the generation of the person's avatar. For example, it is possible to control the avatar of a person who is determined to have negative preferences so that the distance between the avatar and a specified personal characteristic is a fixed distance and the avatar is not affected by real-time changes in emotions.
[0156] Figure 25 shows in detail the information processing flow of the embodiment shown in Figure 22, focusing on processing of SNS history. In this processing flow, an avatar is generated that reflects behavioral characteristics predicted from the user's SNS viewing history and online purchase history, for example. The part surrounded by the dashed line in Figure 25 can be realized using an AI model.
[0157] Online activities such as viewing social media and making online purchases using smartphones have become commonplace. First, multiple pieces of historical information related to these online activities are acquired using an API (Application Programming Interface) or a means of inter-app integration individually prepared. This data is sampled at a predetermined interval, such as every half day, and updated in real time. While shorter update intervals allow for more realistic data, the sampling interval can also be determined by balancing communication frequency and processing load. The acquired data is then formatted for communication and inter-app interoperability. The processed data is stored in an activity history database. As a result, the activity history database is updated at regular intervals.
[0158] From this activity history database, the person's activity characteristics are extracted. The activity characteristics are based on preferences such as entertainment preferences, purchasing tendencies, activities (SNS posting categories (play, work, etc.), social outreach), and personality analysis. Next, a personality analysis (favorite things, things, culture, etc.) is performed based on the extracted activity characteristic data and the activity characteristics extracted from the activity history information. Personality analysis can also be applied to mental health status and personality classification, and previous research examples can be used. The results of the personality analysis are used for real-time emotion classification and assessment, and real-time person status data based on real-time activities is obtained. By integrating this real-time person status data with the output results of the audio information processing block and image information processing block, the person's behavioral characteristics are predicted comprehensively and with higher accuracy. The results of this behavioral characteristic prediction process are reflected in the risk assessment process and sent to the avatar generation block. The risk assessment processing block performs risk assessment by taking into account the predicted person's behavioral characteristics as well as third-party and environmental information. In the risk assessment process, parameters that indicate the degree to which realism and behavioral characteristics are reflected in the avatar are continuously changed and reflected in the avatar. The avatar generation block generates an avatar by taking into account the predicted human behavioral characteristics and the risk assessment result.
[0159] E. Parental Control E-1. Overview With the widespread use of smartphones and the spread of online and remote learning, children's opportunities to use the internet are increasing. These trends have made parental control increasingly important in order to protect children from the risks inherent in using the internet and online games. In recent years, the internet and online games have become more complex, and children's responses to risks vary greatly depending on their internet literacy. Therefore, simple age restrictions are no longer an effective way to protect children from risks. Children with high literacy levels may find ways to remove restrictions, such as finding loopholes, even if excessive restrictions are imposed.
[0160] Furthermore, as the risks of the Internet have become more diverse, parents may not have enough knowledge to determine effective countermeasures, leading to the issue of not being able to protect their children from these risks. Some parents are preventing their children from using the Internet at all, including high-quality content, because they do not fully understand the risks due to the complexity of the Internet. For children, not having access to the Internet means they cannot play online games with friends or join in on online gaming discussions at school, which can lead to social isolation, isolation, and bullying, even if they are able to avoid the risks of the Internet, which can lead to real-world risks.
[0161] Therefore, the communication platform according to the present disclosure provides appropriate parental control using an AI model such as generative AI. Specifically, the communication platform according to the present disclosure uses an AI model to assess a child's internet literacy and the risks inherent in a child's participation in a social community based on the child's history information, and then implements parental control based on the assessment results. The child's history information here includes, for example, the child's internet usage history, game history, purchase site history, social community behavior history, and conversation history. Furthermore, as part of parental control, an AI model is used to monitor a child's online behavior.
[0162] Examples of using AI models to monitor children's online behavior include content filtering, cyber monitoring, time limit settings, payment monitoring and restrictions, limiting personal information leaks, and behavior monitoring and restrictions using NPCs. (1) Content filtering: Using AI models, you can limit the content children can access. For example, you can automatically block content deemed inappropriate for children, such as adult or violent content. (2) Cyber monitoring: You can use AI models to monitor children's online activities and detect inappropriate behavior or dangerous situations. For example, you can detect risky behavior such as bullying, stalking, and suicide threats and notify parents and other relevant parties. (3) Time limit settings: You can use AI models to limit the time children can access the network. For example, you can restrict network access to certain hours to prioritize study or sleep. (4) Payment monitoring and restrictions: If a game includes paid items, you can prohibit them or notify parents to make a decision on whether to use them and manage the amount of restrictions. (5) Limiting personal information leaks: You can restrict the careless posting of names, addresses, passwords, and other family information. (6) Behavior monitoring and restriction using NPCs: Parental control is performed based on the content of conversations between avatars by introducing NPCs into the social community in which the user is participating. Parental control is also performed by changing the appearance, speech content, and items used by the avatars.
[0163] Of the above, the examples of "content filtering," "cyber monitoring," and "setting time limits" are initiatives that utilize AI models to improve children's online safety. However, sufficient consideration must also be given to protecting children's privacy and personal information. It is important for parents and educational institutions to create a safe online environment for children by combining appropriate monitoring and education.
[0164] The parental control function disclosed herein uses an AI model pre-trained based on various risk information to assess a child's internet literacy and the risk of the social community the child uses, making it possible to determine risks that cannot be determined by conventional parameter-based filtering. By using an AI model, the parental control function disclosed herein can suppress excessive blocking and notify (alert) parents of their child's risks. Furthermore, in multi-user play environments such as online games, the parental control function disclosed herein can realize multi-parental control (a mechanism for monitoring multiple users in groups) that takes into account the level of each individual user and monitors all users at a safe level.
[0165] It is important for parents and children to consult and decide on rules before starting to use the Internet, and to continue to consult and decide on rules whenever the child makes new requests. However, there are problems such as parents' lack of understanding of the Internet, children starting new things without consulting their parents, and not being able to find time to consult with their parents. Furthermore, even parents who have a certain level of understanding of the Internet can find it difficult to keep up with the constantly changing environment of the Internet and online games.
[0166] In contrast, according to the parental control function disclosed herein, after the parent and child consult with each other to decide on rules and the child begins using the Internet, parental control is implemented by using an AI model to assess the child's Internet literacy and the risks of the social communities used by the child.
[0167] E-1-1. Basic Functions Fig. 26 shows a schematic functional configuration of the parental control function according to the present disclosure. The parental control function according to the present disclosure is realized by the cooperative operation of a parental control unit 2601, an internet usage monitoring unit 2602, and a risk assessment unit 2603.
[0168] The internet usage monitoring unit 2602 monitors the child's internet usage using a smartphone based on the child's profile and the child's literacy acquired from the child's behavioral history and internet literacy database. Specifically, the internet usage monitoring unit 2602 monitors the child's internet usage history, game history, purchase site history, behavioral history in social communities, conversation history, etc. The internet usage monitoring unit 2602 also records information obtained by monitoring the child's smartphone in the child's behavioral history and internet literacy database. Specifically, the internet usage monitoring unit 2602 is comprised of an app for monitoring internet usage that runs on the child's smartphone.
[0169] The risk assessment unit 2603 assesses the user's internet literacy based on the user's behavior or conversations on the internet, in games, on purchasing sites, and in social communities monitored by the internet usage monitoring unit 2602, and assesses the risk of the social community according to the user's literacy. The risk assessment unit 2603 uses multiple risk assessment AI models to assess risk for each category. Each risk assessment AI model is pre-trained using a training dataset of risk cases in the corresponding category. The risk assessment unit 2603 then assesses the risk level for each category using each risk assessment AI model for the child's internet usage status notified by the internet usage monitoring unit 2602, and outputs the assessment result to the parental control unit 2601.
[0170] The parental control unit 2601 monitors the child's internet usage so as to avoid or reduce risks in categories with increased risk levels, based on the child's internet usage status notified by the internet usage monitoring unit 2602 and the risk level determination results for each category by the risk determination unit 2603. Furthermore, the parental control unit 2601 notifies (warns) the parent's smartphone when the child's risk level in any category exceeds a predetermined threshold.
[0171] In the example shown in Fig. 26, the parental control unit 2601 performs parental control using an NPC that is not operated by the user. On the child's smartphone, the NPC issues a voice message such as "That's dangerous!", "Let's stop!", or "Talk to your mom!" to warn the child or urge the child to consult with a parent (to obtain consent). On the parent's smartphone, the NPC notifies the parent of the child's smartphone usage status by issuing a voice message such as "Taro is participating in a game. It's been 15 minutes since he started."
[0172] Figure 41 shows another example of parental control using NPCs that are not operated by the user. This figure assumes that multiple NPCs appear in addition to the user's avatar in a virtual space such as an online game or social community.
[0173] While a user or their avatar is active in the virtual space, the risk assessment unit 2603 uses an AI model in the background to assess the child's internet literacy and the risks inherent in participating in a social community based on the child's history information. The parental control unit 2601 then implements parental control using an NPC based on the assessment result of the risk assessment unit 2603. The NPC performs parental control based on the content of the dialogue between avatars. The parental control is also implemented by changing the NPC's appearance, speech content, and items used by the avatar. In the example shown in FIG. 41 , when a user's avatar receives an invitation from another avatar, "Would you like to join our group?", the risk assessment unit 2603 assesses the risk, and the parental control unit 2601 has the NPC in the same space say, "No! Check with your mother!", issuing a warning not to accept the invitation easily and avoiding risk to the user.
[0174] The basic functions of the parental control function utilizing the AI model according to the present disclosure are as follows:
[0175] Basic function 1: Setting the child's control level (1) AI determines the child's internet literacy based on their operation history and conversations. If their internet literacy is lower than the standard value, the risk value increases and restrictions are strengthened. (2) Parents can also set parameters such as usage time restrictions.
[0176] Basic Function 2: Child Monitoring (1) Risk Assessment of the Internet and Services (Content) Preventing malicious site access. Since UGC is expected to often not have control codes attached, AI assesses the risk. (2) Purchase Assessment Decisions for online shopping, billing, and app downloads. AI determines whether to prohibit, set a price limit, and determine categories of apps that are allowed to be downloaded depending on the level. (3) Monitoring Posting of Personal Information It is necessary to use at least one of text and images to determine whether a child is sending confidential information to another party via an electronic messaging service. It is necessary to detect whether an image sent as part of an electronic message shows, among other things, a bank card, social security card, or identification card. When such a situation is detected, in some embodiments, a notification is automatically sent to a third party (e.g., parent, teacher, etc.). (4) Setting the risk assessment level based on the child's operation history
[0177] Basic function 3: Functions of social media monitoring app Avoid conversations that involve general bullying, slander, libel, discrimination, etc.
[0178] Basic function 4: Communication (avatar generation), defamation monitoring (1) Introduce non-personal characters (NPCs) to the social community in which the user is participating, and monitor the content of conversations between avatars through the NPCs. (2) Based on the monitoring results, parental control is performed by changing the appearance, shape, and content of speech of the avatars, and the items used by the avatars.
[0179] Basic function 5: Communication (reporting) of child usage status to parent device, remote control
[0180] E-1-2. Effects Examples of effects brought about by the parental control function utilizing an AI model according to the present disclosure include the following.
[0181] Effect 1: Parental controls that take children's internet literacy into consideration allow for optimal restrictions with minimal risk to many children, preventing excessive usage restrictions and reducing stress for children. Effect 2: Parents can achieve appropriate parental control simply by using the communication platform disclosed herein, eliminating the need to decide not to allow their children to access the internet or use online games because they don't understand it. Many children will be able to use the internet appropriately. Effect 3: With the parental control function disclosed herein, the AI model is constantly learning and updating, allowing for appropriate control that follows children's growth and the changing times, leading to continued use.
[0182] The parental control function utilizing the AI model according to the present disclosure can detect the following situations:
[0183] ・When parents have strongly advised children not to give out personal information: Restrict the output of personal information on the internet or in online games. ・When parental permission is required to download an app: An AI model is used to evaluate apps from app download sites and SNS, etc., to determine whether or not to download them. ・Chatting in online games is prohibited: An AI model is used to provide support, detect dangerous words such as slander and bullying, and restrict their input. ・A child pretends that their online game friends are school friends: An AI model is used to ensure safety while taking the child's psychology into consideration. Restrict parents' access to their children's personal information depending on their internet literacy.
[0184] The parental control function according to the present disclosure is not the same as a parent, but utilizes an AI model to watch over children as a third party. It is like a friendly neighborhood uncle who watches over the community in the real world. The parental control function according to the present disclosure may be a pet character in a virtual community environment. The parental control function according to the present disclosure may also watch over children as an NPC in an online game or metaverse. The parental control function according to the present disclosure may display an icon or the like to indicate that the parental control function is watching over children, thereby providing a deterrent effect.
[0185] The parental control function according to the present disclosure can protect a child's personal information according to the child's literacy level. Furthermore, the parental control function according to the present disclosure can dynamically adjust the alert level according to the risk determined by AI. For example, when the risk level increases, the number of NPCs for monitoring purposes may be increased, or an ordinary man who appears for monitoring purposes may be transformed into a police officer.
[0186] E-2. Overall Configuration Figure 27 shows the overall configuration of the parental control function according to the present disclosure. Figure 27 basically illustrates the functions that operate on the smartphone of a child (i.e., the person being monitored). The parental control function according to the present disclosure includes a parental control unit 2701, a usage history management unit 2702, a user profile management unit 2703, a literacy assessment unit 2704, a total risk assessment unit 2705, and a usage time management unit 2706.
[0187] The parental control unit 2701 monitors the user's (child's) internet usage using a smartphone. The parental control unit 2701 then passes data on access to SNSs and online games made on the smartphone to the total risk assessment unit 2705, and based on the risk assessment data returned from the total risk assessment unit 2705, restricts the user's access to SNSs and online games, or implements interactions using the smartphone's UI to help the user avoid risks. The parental control unit 2701 also stores the user's internet usage history obtained from the smartphone in the usage history management unit 2702. The parental control unit 2701 also obtains information on the user's smartphone usage time from the usage time management unit 2706, and restricts access so that usage time does not exceed a predetermined time limit, or implements interactions using the smartphone's UI to ensure that usage time does not exceed a predetermined time limit. For example, when the user's smartphone usage time approaches the time limit, a message such as "It's about time to wrap it up" may be displayed on the smartphone's UI screen (or an audio guidance such as "It's about time to wrap it up" may be output).
[0188] The user profile management unit 2703 manages the user profile, including information on the user's smartphone usage history managed by the usage history management unit 2702. The literacy determination unit 2704 determines the user's Internet literacy based on the user profile managed by the user profile management unit 2703.
[0189] When the total risk assessment unit 2705 receives data on accesses to SNS or online games made by the user on the smartphone, it assesses the risk that this access data poses to the user for each category in consideration of the user's Internet literacy rank assessed by the literacy assessment unit 2704, and outputs a total risk that integrates the risk assessment results for each category. The total risk assessment unit 2705 then returns the risk assessment data to the parental control unit 2701. The total risk assessment unit 2705 analyzes text, etc., by section, and outputs the risk level for each section and the overall risk level to the parental control unit 2701. Furthermore, when the child's risk level exceeds a predetermined threshold, the total risk assessment unit 2705 sends a notification (warning) to the parent's smartphone (the parental control unit within the smartphone).
[0190] The parental control unit 2701 restricts the user's use of apps (access to SNS or online games) or uses the smartphone's UI to implement interactions for the user to avoid risks based on the total risk level returned from the total risk determination unit 2705. Furthermore, if the total risk level exceeds a predetermined threshold, the parental control unit 2701 sends a notification (warning) to the parent's smartphone (parental control unit within the smartphone).
[0191] The total risk determination unit 2705 uses multiple risk determination AI models to determine the risk for each category. In the example shown in Figure 27, seven categories of risk are assumed: "fraud," "rumor," "personal information leakage," "libel," "sexual content," "violent content," and "copyright and portrait right infringement." The total risk determination unit 2705 determines the risk level of the user's access data for each category using a "fraud risk determination AI model," a "rumor risk determination AI model," a "personal information leakage risk determination AI model," a "libel and slander risk determination AI model," a "sexual risk determination AI model," a "violent risk determination AI model," and a "copyright and portrait right infringement risk determination AI model," which respectively determine the risk for each of these categories, and outputs the total risk by integrating the risks for each category.
[0192] Each risk determination AI model used by the total risk determination unit 2705 is a trained AI model that has been pre-trained using risk cases of each category, namely, "fraud case training data," "hoax case training data," "personal information leakage case training data," "defamation case training data," "sexual content case training data," "violence content case training data," and "copyright and portrait right infringement case training data." Because an AI model is used to determine the risk of each category, a rule-based judgment may be used to integrate the risks of each category and determine the total risk.
[0193] E-3. Risk Assessment E-3-1. Overall Configuration Fig. 28 shows an example of the internal configuration of the total risk assessment unit 2705. In the example shown, the total risk assessment unit 2705 includes a data interface unit 2801, a data capture unit 2802, a risk simulation unit 2803, an individual risk assessment unit 2804, and an integration unit 2805.
[0194] The data interface unit 2801 receives access data from the parental control unit 2701. The access data includes data of various modalities such as images (moving images, still images), audio, text, and web URLs.
[0195] The data capture unit 2802 captures data of each modality included in the access data. For example, in the case of access data for an online game application, video capture is performed to capture moving images of characters appearing in the game, sound capture is performed to capture sound such as game sound effects, text information related to the game application such as the game title, genre, publisher, and release date, and the web URL for accessing the online game.
[0196] The risk simulation unit 2803 generates possible risks by simulation from the data of each modality captured by the data capture unit 2802. In the example shown in FIG. 28 , the risk simulation unit 2803 includes a video risk simulator 2803-1 that simulates risks from video captured by the data capture unit 2802, a sound risk simulator 2803-2 that simulates risks from sound captured by the data capture unit 2802, and a text risk simulator 2803-3 that simulates risks from text information captured by the data capture unit 2802. Each of the risk simulators 2803-1, ... can be configured as a risk simulation AI using an autoencoder. An encoder at a front stage of the autoencoder extracts features of the input modality data, and a decoder at a rear stage generates and outputs modality data that assumes risks based on the extracted features.
[0197] The individual risk determination unit 2804 is equipped with individual risk determination AI models such as a "fraud risk determination AI model," a "hoax risk determination AI model," a "personal information leakage risk determination AI model," a "defamation risk determination AI model," a "sexual risk determination AI model," a "violence risk determination AI model," and a "copyright / portrait right infringement risk determination AI model" in order to individually determine the risk for each category. These individual risk determination AI models determine the risk levels of the "fraud risk," "hoax risk," "personal information leakage risk," "defamation risk," "sexual content risk," "violence content risk," and "copyright / portrait right infringement risk" posed by the access data, based on the data generated by the risk simulation unit 2803 assuming the risks of each modality included in the access data.
[0198] The integration unit 2805 integrates the risk levels determined for each category by the individual risk determination unit 2804 and outputs the total risk of the access data. Since AI is used to determine the risk of each category performed by the individual risk determination unit 2804, the integration unit 2805 may determine the total risk on a rule-based basis.
[0199] E-3-2. Risk Simulation The risk simulation unit 2803 generates, by simulation, possible risks from the data of each modality captured by the data capture unit 2802. The risk simulation unit 2803 can be configured as a risk simulation AI using an autoencoder that simulates risks for each modality.
[0200] 29 shows an example of generating a risk simulation in the risk simulation unit 2803. The risk simulation unit 2803 receives text and images captured from access data by the data capture unit 2802. Therefore, a video risk simulator (image generation model) 2803-1, which simulates risk from video, generates an image that assumes a risk from the captured video. Also, a risk simulator (document generation model) 2803-3, which performs a text simulation of risk from text information, generates a sentence that assumes a risk from the captured text.
[0201] E-3-3. Calculation of Total Risk Figure 30 shows a schematic diagram of the flow for calculating total risk. The individual risk assessment unit 2804 is provided with a risk assessment AI model for each category in order to assess the risk for each category. In the example shown in Figure 30, for the sake of simplicity, it is assumed that the unit is provided with an "a risk assessment AI model," a "b risk assessment AI model," ..., and an "n risk assessment AI model" corresponding to each of categories a, b, ..., n.
[0202] The data interface unit 2801 receives website content, game application content, web URLs, etc. as access data. Website content and game application content include data of various modalities such as images (moving images, still images) and audio. Web URLs are basically text data.
[0203] The data capture unit 2802 captures data for each modality, such as video, audio, text, etc., from this access data. The risk simulation unit 2803 then extracts the feature quantities of the captured data for each modality, and generates and outputs modality data that assumes risks.
[0204] The individual risk assessment unit 2804 assesses the risk level of each category using a risk assessment AI model for each category. In the example shown in FIG. 30, the risk assessment AI model a assesses the risk level R of category a. aThe risk determination AI model b determines the risk level R of category b. b ..., the n risk determination AI model determines the risk level R of category n n Determine the following.
[0205] The integration unit 2805 integrates the risk level R determined for each category by the individual risk determination unit 2804. a , R b , ..., R n The total risk R total For example, the integration unit 2805 calculates the risk level R a , R b , ..., R n and weight coefficient k a , k b , ..., k n The total risk level R is calculated as shown in the following formula. total Calculate.
[0206] R total =R a ×k a +R b ×k b +...+R n ×k n
[0207] Risk level R for each category a , R b , ..., R n and the weighting coefficient k for each category a , k b , ..., k n The user's (child's) internet literacy may be determined and changed based on the user's (child's) internet literacy. The user's internet literacy is managed by the literacy determination unit 2704. Details of internet literacy determination will be described later.
[0208] Then, the total risk determination unit 2705 calculates the calculated total risk level R total Specifically, the total risk determination unit 2705 determines whether the content being accessed by the user (child) is safe based on the total risk level R total is a predetermined threshold R ref Compared with the total risk R totalis the threshold R ref (i.e., R total >R ref ), it determines whether the content the user (child) is accessing is safe and sends a notification (warning) to the parent's smartphone.
[0209] FIG. 31 shows a specific example of how the total risk determination unit 2705 determines the risk level.
[0210] The fraud risk determination AI model, hoax risk determination AI model, personal information leakage risk determination AI model, slander risk determination AI model, sexual risk determination AI model, violence risk determination AI model, and copyright / portrait right infringement risk determination AI model in the individual risk determination unit 2804 determine the risk level that the user's access data poses to each category. The risk level is a value between 0 and 1. In the example shown in Figure 31, each risk determination AI model determines a fraud risk level of 0.1, a hoax risk level of 0.1, a personal information leakage risk level of 0, a slander risk level of 0.4, a sexual risk level of 0.3, a violence risk level of 0.5, and a copyright / portrait right infringement risk level of 0.3, respectively.
[0211] Meanwhile, based on the internet literacy assessed for the user, a weight k = 1 is assigned to the fraud risk, a weight k = 1 to the hoax risk, a weight k = 1 to the personal information leakage risk, a weight k = 3 to the slander risk, a weight k = 3 to the sexual risk, a weight k = 3 to the violence risk, and a weight k = 0 to the copyright / portrait right infringement risk. Therefore, after each weighting, the fraud risk level is 0.1, the hoax risk level is 0.1, the personal information leakage risk level is 0, the slander risk level is 1.2, the sexual risk level is 0.9, the violence risk level is 1.5, and the copyright / portrait right infringement risk level is 0.
[0212] The integration unit 2805 multiplies the risk level of each category by a weighting coefficient and adds them together to obtain a total risk level R total = 3.8, where the threshold value R ref If the total risk level R is set to 3, total is the threshold R ref(i.e., R total >R ref ), the total risk determination unit 2705 determines that the content being accessed by the user (child) is risky. In this case, the parental control unit 2701 notifies (warns) the parent's smartphone that the child is accessing risky content.
[0213] E-3-4. Determination of Internet Literacy The literacy determination unit 2704 first initializes the user's Internet literacy, and then updates the Internet literacy based on the user's Internet usage history.
[0214] The literacy assessment unit 2704 initially sets the user's Internet literacy based on whether the user can correctly configure pre-prepared Internet access settings. Specifically, one or more tasks are performed to assess the user's Internet literacy, and the user's Internet literacy is initially set based on whether the user was able to correctly perform the tasks. Examples of tasks for assessing Internet literacy include the following. However, the following are merely examples, and it is not necessary to try all of the tasks. Furthermore, tasks not listed below may be added to assess Internet literacy.
[0215] - Launch a browser. - Search for the word "XXX" on an XXX search site. - Check tomorrow's weather forecast. - Find related information from today's news on a news media site. - Copy the URL. - Bookmark the web page. - Open a new tab in your browser. - Clear the cache and cookies from your browser.
[0216] Instead of having the user actually perform the above tasks, as in the case of the initial setup of Internet literacy, questions may be posed to the user when setting up the application, and the initial setup may be performed based on the answers from the user. In this case, the questions and answers may be conducted via a UI screen or in a conversational format with a voice agent.
[0217] The literacy assessment unit 2704 then assesses the user's Internet literacy based on the user history. The user history is managed by the user profile management unit 2703. Examples of user history relevant to assessing Internet literacy include websites used and games used. The use of websites and games may involve risks. The user history of websites used, such as usage time (or usage time zone) and usage status (whether access is only, whether posts are made, whether charges are made, whether there are multiple participants, etc.), is used to assess Internet literacy. The user history of games used, such as usage time (or usage time zone) and usage status (whether access is only, whether posts are made, whether charges are made, whether there are multiple participants, etc.), is used to assess Internet literacy.
[0218] Of course, even if it is not the web sites and games used, the literacy determination unit 2704 may determine the user's Internet literacy by further including other user history where access or use may be risky.
[0219] The literacy assessment unit 2704 assesses the user's internet literacy using a literacy assessment AI model. FIG. 32 shows a mechanism for assessing a user's internet literacy using the literacy assessment AI model. The literacy assessment AI model is configured using, for example, an autoencoder. The literacy assessment AI model inputs the results of the literacy assessment task and the user's history of web sites used, games used, etc., and assesses the user's internet literacy rank as a value from 1 to 9.
[0220] For example, by monitoring the operation history of a user (child), if it is confirmed that the user does not access risky sites, always adheres to usage time limits, or has gained experience, the Internet literacy rank of the user (child) is changed. In addition, a rule may be applied such that the Internet literacy rank of the user (child) is changed if the parent's Internet literacy is sufficiently high or if the user (child) is getting older.
[0221] The literacy determination unit 2704 may perform a literacy determination task in the initial settings to determine the user's initial Internet literacy, or may update the Internet literacy based on the user's subsequent history.
[0222] E-3-5. Relationship between Internet literacy, risk level, and weighting coefficient As described above, the literacy determination unit 2704 determines the rank of the user's Internet literacy as a value from 1 to 9. Then, the total risk determination unit 2705 calculates the risk level R for each category. a , R b , ..., R n and the weighting coefficient k for each category a , k b , ..., k n This is determined based on the user's (child's) internet literacy.
[0223] Specifically, the total risk assessment unit 2705 assigns a risk level and a weighting coefficient to each rank of Internet literacy for each risk category. Figure 33 shows an example of the relationship between Internet literacy rank, risk level, and weighting coefficient for the risk of the "XXX" category.
[0224] E-4. Parental Control App Cooperation Fig. 34 shows a mechanism for cooperation between the parental control function and an app according to the present disclosure.
[0225] The communication platform acquires access data for apps used by users. The apps used by users include, for example, social networking services (SNS) and games. The access data for these apps includes data in various modalities, such as images, audio, and text.
[0226] Within the communication platform, the total risk assessment unit 2705 acquires access data (images, audio, text) of the apps used by the user via the parental control unit 2701. The total risk assessment unit 2705 then uses AI to assess the total risk of the apps used by the user and outputs the assessment result to the parental control unit 2701.
[0227] The parental control unit 2701 restricts the user's use of apps (access to SNS or online games) and implements interactions for the user to avoid risks using the smartphone UI based on the total risk level returned from the total risk determination unit 2705. Furthermore, the parental control unit 2701 sends a notification (warning) to the parent's smartphone when the risk level of the total risk exceeds a predetermined threshold.
[0228] E-5. Flow E-5-1. Flow for controlling Internet access Figure 35 shows the processing procedure for controlling a user's (child's) Internet access in the form of a flowchart. This processing operation is mainly performed by the parental control unit running on the child's smartphone.
[0229] First, the child launches an application (SNS, game, etc.) on the smartphone and starts accessing a website (step S3501).
[0230] In response to the application launch, the parental control unit requests approval of the application launch from the parent's smartphone (step S3502). If approval is granted on the parent's smartphone (step S3503), the application launches on the child's smartphone (step S3504), enabling access to the website. The child then begins using the application on their smartphone (step S3505).
[0231] Also, when approval is given on the parent's smartphone (step S3503), and the parent approves the acquisition of the child's smartphone operation log (user history) (step S3507), the parental control unit begins acquiring the child's operation log (step S3508).
[0232] The literacy assessment unit updates the child's Internet literacy based on the user's operation log (step S3509). The total risk assessment unit 2705 calculates the total risk level of the child's app usage (Internet access) and determines whether or not there is a risk by threshold assessment or the like (step S3510). In step S3510, the total risk assessment unit 2705 calculates the total risk level based on the Internet literacy updated in the preceding step S3909.
[0233] If the total risk level exceeds the threshold and it is determined that the child's use of the app (Internet access) is risky, the parental control unit stops the app (step S3506), restricts Internet access, and ends this process.The parental control unit also notifies (warns) the parent's smartphone that the child is accessing risky content (step S3511).
[0234] E-5-2. Flow for Learning the Literacy Determination AI Model Figure 36 shows, in the form of a flowchart, the processing procedure for learning the literacy determination AI model used in the literacy determination unit 2704.
[0235] First, learning data consisting of user history data that poses a risk to the user (child), the user's Internet literacy rank, and the user's age is acquired (step S3601). Next, a literacy assessment AI model is trained based on the judgment of the parental control unit (step S3602). Then, a human determines the judgment result of Internet access restrictions (step S3603). By repeating the above process, the literacy assessment AI model is trained.
[0236] E-5-3. Flow of Determining Internet Literacy Figure 37 shows in the form of a flowchart the processing procedure by which the literacy determination unit 2704 determines the Internet literacy of a user (child).
[0237] The literacy determination unit 2704 acquires whether or not the user can correctly configure the pre-prepared Internet access settings as a determination factor for the user's Internet literacy (step S3701).
[0238] Furthermore, the literacy determination unit 2704 acquires the user's history of the web sites and games used (step S3702).
[0239] Next, the literacy assessment unit 2704 inputs the judgment factors acquired in step S3701 and the user history acquired in step S3702 into the literacy standard generation AI model to generate literacy assessment standards (step S3703).Then, the internet literacy assessment AI model determines the user's internet literacy from the user's history information acquired in step S3702 based on the literacy assessment standards generated in step S3703 (step S3704).
[0240] E-5-4. Flow for Controlling Internet Access Figure 38 shows in flowchart form the detailed processing procedure for controlling the user's (child's) Internet access. Figure 35 shows the general processing procedure for controlling the user's (child's) Internet access, but Figure 38 shows the processing procedure in which the parental control unit controls the child's Internet access using the agent app's linking function.
[0241] First, the same agent app related to parental control is installed on each smartphone of the child and the parent, and these smartphones are linked (step S3801). If the child and the parent use the same smartphone, it is sufficient to install the agent app on only one of the smartphones (step S3821).
[0242] When a child launches an app (such as a social networking site or game) on their smartphone and begins accessing a website (step S3802), the agent app's linking function sends a request for approval from the child's smartphone to the parent's smartphone to launch the app (step S3803). If the parent's smartphone approves the request (step S3804), the app launches on the child's smartphone, allowing access to the website, and the child begins using the app (step S3805). If the child and parent use the same smartphone, the parent's approval (e.g., fingerprint authentication) on the same smartphone (step S3822) is required before the child can begin using the app (step S3805).
[0243] The agent app monitors the time spent using apps on the child's smartphone (step S3811). The agent app also acquires an operation log on the child's smartphone (step S3813), sets the child's Internet literacy (step S3816), and determines and changes the child's Internet literacy as needed (step S3817). Based on the Internet literacy determined in step S3816, the agent app then checks the risks of app launch or website access (step S3812) and determines whether there are any risks in the child's use of the app or website access (step S3818). The agent app also uses a risk-determination AI model to determine the risks of apps (e.g., games) and websites the child is accessing (step S3814) and registers the determination results in the app and website database (step S3815).
[0244] If it is determined in step S3818 that there is no risk in the child's use of the app or access to the website, the child continues to use the app or access the website on his or her smartphone (step S3806).Then, the child autonomously ends the use of the app or the access to the website (step S3807).
[0245] On the other hand, if it is determined in step S3818 that there is a risk in the child's use of the app or access to the website, the agent app stops the use of the app on the child's smartphone or prohibits access to the website (step S3819).The agent app also uses its linking function to notify (warn) the parent's smartphone that there is a risk in the child's use of the app or access to the website (step S3820), and forcibly terminates the child's use of the app or access to the website (step S3807).
[0246] 39A shows an example of UI operation on a child's smartphone and a parent's smartphone when a child's Internet access is controlled using the agent app's collaboration function. When a child wants to use a game app, he or she first checks with the parent (SEQ 3901). After obtaining the parent's approval (SEQ 3902), the child launches the game app and begins playing. While the child continues using the game app, the agent app monitors the child's use until the usage time allowed by the parent is reached, and notifies the parent's smartphone that use is continuing (SEQ 3903).
[0247] Figure 39B shows the internal operation during the period when the app is being used on the child's smartphone. During the period when the app is being used on the child's smartphone, the parental control unit uses the agent app to acquire the operation history from the child's smartphone and analyzes the operation. Based on the analysis results, the parental control unit determines the child's Internet literacy rank. The risk assessment AI model then determines the risk of the child's app use and website access based on the child's current Internet literacy rank.
[0248] To protect children's personal information, the parental control unit does not report to the parent's smartphone every detail of a child's game app play (for example, the content of conversations within the game), but it does monitor risks. The parental control unit can also change its behavior depending on the child's Internet literacy level.
[0249] F. Configuration of Information Processing Device Fig. 40 shows an example of the hardware configuration of an information processing device 2000 that can operate as the system 100 to which the present disclosure is applied. The information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by a personal computer, but some of its functions may be configured by an information terminal such as a tablet or smartphone.
[0250] The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. When performing computationally intensive processes such as learning various AIs (described above) on the information processing device 2000, it is desirable that the CPU 2001 be a multi-core CPU (e.g., Apple M1 Max, etc.), or that the information processing device 2000 further include a multi-core processor (e.g., NVIDIA's "Quadro A6000") such as a GPU (Graphics Processing Unit) or a GPGPU (General-purpose computing on graphics processing unit) in addition to the CPU 2001. However, hereinafter, for convenience, these will be collectively referred to simply as the CPU 2001.
[0251] The ROM 2002 stores in a nonvolatile manner programs (such as a basic input / output system) and calculation parameters used by the CPU 2001. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as working data that change as appropriate during program execution. Programs loaded into the RAM 2003 and executed by the CPU 2001 include, for example, various application programs and an operating system (OS).
[0252] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which includes a CPU bus and other components. The CPU 2001 executes various application programs in an execution environment provided by an OS through the cooperative operation of the ROM 2002 and RAM 2003, thereby enabling various functions and services. If the information processing device 2000 is a personal computer, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. Note that the application program or some of the modules in the application program may use an existing library that is stored, shared, or made public through, for example, a source code management service. For example, a program for operating as the communication platform unit 110 or the social application unit 150 shown in FIG. 1 is executed on the information processing device 2000.
[0253] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).
[0254] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 40 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.
[0255] The input unit 2008 is composed of an input control circuit that generates an input signal based on an input from a user and outputs the signal to the CPU 2001. When the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, a mouse, a touch panel, a camera, and a microphone. The output unit 2009 includes display devices such as a liquid crystal display (LCD) device, an organic electroluminescence (EL) display device, and an LED (light emitting diode), as well as an audio output device such as a speaker. The input unit 2008 is used to input a video to be processed, and the output unit 2009 is used to display a GUI screen, etc.
[0256] The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.
[0257] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 2012. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.
[0258] The communication unit 2013 is a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark), and cellular communication networks such as 4G and 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) and a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, etc. Programs executed on the information processing device 2000 are installed from the outside, for example, via the communication unit 2013.
[0259] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.
[0260] Although the present specification has focused on embodiments in which online games are run on the communication platform according to the present disclosure and avatars are used in virtual community environments such as online games, the gist of the present disclosure is not limited thereto. The communication platform according to the present disclosure can also provide a safe virtual community environment for various applications and services other than online games, even when always connected, that is free from anxieties and risk factors from various perspectives, such as ethics, culture, customs, language, and religion.
[0261] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.
[0262] For convenience, this specification and drawings have been described as dividing the modules into functional modules, but at least two or more modules can be combined into one large AI model, such as an LLM or a foundation model.
[0263] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.
[0264] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc (registered trademark)), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to the realization of the present disclosure can be provided as so-called package software.
[0265] The program may also be transferred wirelessly or via a wire from a download site to a computer via a network such as a wide area network (WAN) typified by cellular, a local area network (LAN), the Internet, etc. The computer can receive the program transferred in this manner and install it in a large-capacity storage device such as an HDD or SSD within the computer.
[0266] The present disclosure may also be configured as follows.
[0267] (1) A control system for interactive objects used in a virtual space or a social community, comprising: a control unit; and an AI judgment unit that judges the internet literacy of a user who uses the interactive object or judges the risk of the social community based on historical information of the user, and the control unit performs parental control on the interactive object depending on the internet literacy of the user or the risk level of the social community.
[0268] (2) The interactive object control system described in (1) above, wherein the user's history information includes at least one of internet usage history, game history, purchase site history, behavioral history in the social community, and conversation history.
[0269] (3) The interactive object control system according to any one of (1) or (2), wherein the interactive object is an avatar.
[0270] (4) The interactive object control system according to (3) above, further comprising a generation AI unit that generates a non-player character (NPC) that is not operated by the user, and the control unit uses the non-player character to perform the parental control based on the content of a conversation between the avatar and another avatar.
[0271] (5) An interactive object control system as described in (3) above, comprising a generation AI unit that generates at least one of the avatar's external shape, speech content, and items used, and the generation AI unit changes the avatar's external shape, speech content, or items used depending on the user's Internet literacy or the level of the social community.
[0272] (6) The interactive object control system according to (4), wherein the control unit performs the parental control by notifying the parent of the user of consent confirmation based on the content of a conversation between the avatar and another avatar.
[0273] (7) A method for controlling an interactive object used in a virtual space or a social community, comprising: a determination step of determining the Internet literacy of a user who uses the interactive object based on historical information of the user, or determining the risk of the social community using an AI model; and a control step of performing parental control on the interactive object depending on the Internet literacy of the user or the risk level of the social community.
[0274] (8) A computer program written in a computer-readable format to be executed on a computer to control interactive objects used in a virtual space or a social community, the computer program causing the computer to function as: an AI judgment unit that judges the Internet literacy of a user using the interactive object or judges the risk of the social community based on the user's historical information; and a control unit that performs parental control on the interactive object depending on the user's Internet literacy or the risk level of the social community.
[0275] 100...system, 110...communication platform unit 111...information detection unit, 112...analysis unit, 113...discrimination unit 114...generation unit, 115...history management database, 116...authentication unit 117...parental control unit, 118...content adjustment unit 200...system, 210...analysis module 220...suggestion module, 230...critic module 240...other platform deployment module 301...analysis module, 302...generation module 303...video generation module, 304...control module 305...critic module, 306...suggestion module 307...other platform deployment module 400...system, 401...production tool 402...avatar design database, 403...generation unit 404...controller, 405...rendering engine 406...motion acquisition unit 500...video generation system, 501...motion acquisition unit 502...Personal feature extraction unit, 503...Production and control tool 504...Avatar design database, 505...Generation unit 506...Controller, 507...Rendering engine 508...Feature extraction unit, 509...Evaluation and supervision unit 510...Feature database, 511...UI unit 512...Reference motion data generation unit 601...Posture determination unit, 602...Skeleton point determination unit 603...Feature vector extraction processing unit, 603A...Encoder 603B...Decoder, 701...Encoder, 702...Decoder 801...Face encoder, 802...Right arm encoder 803...Left arm encoder, 804...Body encoder 805...Right foot encoder, 806...Left foot encoder, 810...Decoder 901...Large-scale language model, 902...Text2Image model 1001...Video classification unit, 1002...Video capture unit 1003... Posture determination unit, 1004... Skeleton data extraction unit, 1005... Autoencoder, 1006... Similar video motion extraction unit, 1007... Similar video feature averaging unit, 1008... Dance feature database, 1701... Religious object in virtual space of religious sphere A, 1702... Object replaced from religious object A, 1801... Image input unit, 1802... Feature extraction unit, 1803... Object discrimination unit, 1804... Prohibited object identification unit, 1805... AI engine1806...Participant individual profile data, 1807...Image correction unit 1808...AI engine 2000...Information processing device, 2001...CPU, 2002...ROM 2003...RAM, 2004...Host bus, 2005...Bridge 2006...Expansion bus, 2007...Interface unit 2008...Input unit, 2009...Output unit, 2010...Storage unit 2011...Drive, 2012...Removable recording medium 2013...Communication unit 2601...Parental control unit, 2602...Internet usage monitoring unit 2603...Risk assessment unit 2701...Parental control unit, 2702...Usage history management unit 2703...User profile management unit 2704...Literacy assessment unit, 2705...Total risk assessment unit 706...Usage time management unit 2801...Data interface unit, 2802...Data capture unit 2803: Risk simulation unit, 2804: Individual risk assessment unit, 2805: Integration unit
Claims
1. A control system for interactive objects used in a virtual space or a social community, comprising: a control unit; and an AI determination unit that determines the internet literacy of a user who uses the interactive object or determines the risk of the social community based on historical information of the user, wherein the control unit performs parental control on the interactive object according to the internet literacy of the user or the risk level of the social community.
2. The interactive object control system according to claim 1, wherein the user's history information includes at least one of internet usage history, game history, purchase site history, behavior history in the social community, and conversation history.
3. The interactive object control system according to claim 1 or 2, wherein the interactive object is an avatar.
4. The interactive object control system according to claim 3, further comprising a generation AI unit that generates non-player characters (NPCs) that are not operated by the user, and the control unit uses the non-player characters to perform the parental control based on the content of conversations between the avatar and other avatars.
5. An interactive object control system as described in claim 3, comprising a generation AI unit that generates at least one of the avatar's external shape, speech content, and items used, and the generation AI unit changes the avatar's external shape, speech content, or items used depending on the user's Internet literacy or the level of the social community.
6. The interactive object control system according to claim 4, wherein the control unit performs the parental control by notifying the parent of the user of consent confirmation based on the content of the conversation between the avatar and other avatars.
7. A method for controlling an interactive object used in a virtual space or a social community, comprising: a determination step of determining the internet literacy of a user who uses the interactive object based on historical information of the user, or determining the risk of the social community using an AI model; and a control step of performing parental control on the interactive object depending on the internet literacy of the user or the risk level of the social community.
8. A computer program written in a computer-readable format to be executed on a computer to control interactive objects used in a virtual space or a social community, causing the computer to function as: an AI judgment unit that judges the Internet literacy of a user or judges the risk of the social community based on the user's history of using the interactive object; and a control unit that performs parental control on the interactive object depending on the user's Internet literacy or the risk level of the social community.
Citation Information
Patent Citations
Content reproducing device, input method, program, recording medium, and television receiver
JP2010176210A
Filtering and parental control methods to limit visual effects on head-mounted displays
JP2018517444A
Generating method, generating program and information processing device
JP2023119197A
Measurement program, measurement method, and measurement device
JP2023176640A