System, method
A device predicts user behavior through supervised learning on acquired data, addressing the lack of user behavior prediction in existing systems and improving interaction and consultation services.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- 井泽 佑斗
- Filing Date
- 2024-12-24
- Publication Date
- 2026-07-06
AI Technical Summary
Existing systems lack a method for predicting user behavior regarding certain problems, particularly in interactions involving multiple users.
A device that acquires information on user behavior over multiple days, performs supervised learning on this data, and predicts user behavior based on specified training data, allowing for personalized interactions and discussions.
This configuration enables the generation of ideas and stimulates discussions by accurately predicting user behavior based on individual characteristics, enhancing user interaction and consultation services.
Abstract
Description
Technical Field
[0001] The present invention relates to a device or a prediction device or a prediction method.
Background Art
[0002] The statements in this section only provide background information related to the present disclosure and do not necessarily constitute prior art.
[0003] Patent Document 1 discloses a conversation support device.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the inventor recognized that at least in the above embodiment, there is a drawback that there is no method for predicting the user's behavior regarding a certain problem.
Means for Solving the Problems
[0006] At least one disclosure of the present invention is In a state where at least two or more users exist, for each user, Acquire information regarding the behavior of the user on at least two or more days, Based on the information, request the AI to perform supervised learning, Separate and store the learning data, Based on the specified learning data, request the AI to predict the behavior of the corresponding user regarding a certain problem, Device To provide.
Effects of the Invention
[0007] This configuration has the advantage of at least being able to generate ideas or stimulate discussion by predicting user behavior based on user characteristics.
[0008] These and other aspects, features, and advantages of the Disclosure will become apparent from the following detailed written description of preferred embodiments and aspects taken in conjunction with the following drawings, but variations and modifications thereof may be implemented without departing from the spirit and scope of the novel concepts of the Disclosure. An aspect of one embodiment in the Disclosure may be combined with or replaced by one or more aspects of another embodiment disclosed herein, insofar as they do not conflict. [Modes for carrying out the invention]
[0009] The following disclosure provides many different embodiments and examples for carrying out different features of the presented subject matter. For the sake of simplicity, specific examples of components and arrangements are disclosed below. Of course, these are merely examples and are not intended to be limiting. For example, a structure in which a first feature is covered by or in contact with a second feature subsequently disclosed may include embodiments in which an additional feature is formed between the first and second features so that they do not come into direct contact, as well as embodiments in which the first and second features are formed so that they do not come into direct contact. Furthermore, the disclosure may repeat reference numbers and / or letters in various examples. This repetition is for the sake of brevity and clarity and does not require in itself to be related to the various embodiments and / or configurations described. Furthermore, when describing the first element as being "connected" or "joined" to the second element, such description includes embodiments in which the first and second elements are directly connected or joined to each other, as well as embodiments in which the first and second elements are indirectly connected or joined to each other by having one or more other elements interposed between them.
[0010] As used herein, the phrase "at least one of" encompasses all the exemplary variations. For example, the phrase "comprises at least one of A, B, or C" is synonymous with "consisting of A, B, C and combinations thereof," and encompasses all conceivable variations of A, B, C, A+B, A+C, B+C, and A+B+C.
[0011] In this disclosure, disclosures involving the use of machines, electronic operators, or computers may include embodiments of methods, recording media, apparatus, or programs. Any statement used herein that “A is B” may be replaced with “A includes B” unless otherwise stated herein, to the extent that it does not conflict with the statement.
[0012] The terms used in this disclosure, including those used in the claims, may be interpreted in light of the descriptions and drawings in the specification, and, to the extent that they do not contradict the implications of this disclosure, they may also be interpreted in relation to what one or more citizens have called, represented, understood or practiced, or may have done in the past, present, or future, as such. The following embodiments can be used for the operating method in at least one embodiment. The following will be explained by reference to the description in JP6456303, which describes at least one embodiment in detail (beginning of reference).
[0013] As used herein, the term “computer” schematically includes, as known in the art, a processor, memory, at least one information storage / retrieval device such as a hard drive, disk drive or flash drive or memory stick, or other non-temporary computer-readable medium or non-temporary storage device, at least one input device such as a keyboard, mouse, point and touch device, touchscreen or microphone, and a display structure such as a well-known computer screen. In addition, a computer may include one or more network connections, such as wired or wireless connections. As known in the art, such a computer or computer system may more or less include, but is not limited to, for example, tablet computers or smart devices, other electronic media and electronic devices.
[0014] As used herein, the terms “cloud” or “cloud computing” refer to centralized and virtualized computing facilities where all computing resources are shared. For application systems and subsystems, since they all reside within the “cloud,” it is no longer possible to refer to specific machines.
[0015] As used herein, the term “Distributed Internet Services System” refers to a distributed internet services platform that translates internet applications for execution in various computing environments. A DIS system distributes internet applications, including content, data, and logic, to any number and type of device, to whatever extent appropriate, and along a network, via Component Distribution Servers / Asset Distribution Servers. Through a DIS, internet applications can be hosted and centrally managed as a service tailored to each user's needs, and can be locally cached and executed on the user's device or nearby locations while maintaining their integrity. Web-enabled computing devices can be upgraded with DIS software to become DIS-enabled, capable of enjoying and running distributed internet services. The Distributed Internet Service System is fully described in any one of the following patent families: U.S. Patents No. 7,136857, 7150015, 7181731, 7209921, 7430610, 7685183, 7685577, 7752214, 8326883, 8386525, 8443035, 8458142, 8458222, 8473468, 8527545, and 8650226, and U.S. Patent Publications 20120005205 and 20130091252, all of which, like the present invention, are jointly owned by OP40 Holdings, Inc., and all of which are incorporated by reference. (End of quote)
[0016] Regarding the operating method used in at least one embodiment, the following embodiments can be taken for conventional internet systems that do not use a distributed internet. The following will be explained by reference to the description in JP7113047, which describes at least one embodiment in detail (beginning of reference).
[0017] Embodiments including those specifically disclosed herein can provide an automated response system based on artificial intelligence that is implemented in a manner that mimics actual human conversation, thereby enabling more natural communication with users while quickly and conveniently handling inquiries, reservations, delivery orders, and more.
[0018] The multiple electronic devices 110, 120, 130, and 140 may be fixed or mobile terminals implemented by a computer system. Examples of the multiple electronic devices 110, 120, 130, and 140 include AI speakers, smartphones, mobile phones, navigation systems, PCs (personal computers), notebook PCs, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablets, game consoles, wearable devices, IoT (Internet of Things) devices, VR (virtual reality) devices, and AR (augmented reality) devices. As an example, Figure 1 shows an AI speaker as electronic device 110, but in embodiments of the present invention, electronic device 110 may mean one of a variety of physical computer systems that can communicate with other electronic devices 120, 130, 140 and / or servers 150, 160 via a network 170 using substantially wireless or wired communication methods.
[0019] The communication method is not limited, and may include not only communication methods that utilize communication networks that can be included in network 170 (for example, mobile communication networks, wired internet, wireless internet, broadcasting networks, satellite networks, etc.), but also short-range wireless communication between devices. For example, network 170 may include one or more arbitrary networks such as PAN (personal area network), LAN (local area network), CAN (campus area network), MAN (metropolitan area network), WAN (wide area network), BBN (broadband network), and the Internet. Furthermore, network 170 may include, but is not limited to, one or more network topologies, including bus networks, star networks, ring networks, mesh networks, starbus networks, tree or hierarchical networks.
[0020] Servers 150 and 160 may each be implemented by one or more computer devices that communicate with a plurality of electronic devices 110, 120, 130, 140 via a network 170 to provide instructions, code, files, content, services, and the like. For example, server 150 may be a system that provides a first service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170, and server 160 may also be a system that provides a second service to a plurality of electronic devices 110, 120, 130, 140 connected via network 170. As a more specific example, server 150 may provide, as the first service, a service (such as an automatic response service, for example) targeted by a corresponding application to a plurality of electronic devices 110, 120, 130, 140 through an application that is a computer program installed and executed on the plurality of electronic devices 110, 120, 130, 140. As another example, server 160 may provide, as the second service, a service that distributes files for installation and execution of the above-described application to a plurality of electronic devices 110, 120, 130, 140.
[0021] FIG. 2 is a block diagram for explaining the internal configurations of an electronic device and a server in an embodiment of the present invention. In FIG. 2, the internal configuration of electronic device 110 and the internal configuration of server 150 are described as examples for an electronic device. Also, the other electronic devices 120, 130, 140 and server 160 may have the same or similar internal configurations as the above-described electronic device 110 or server 150.
[0022] The electronic device 110 and the server 150 may include memory 211, 221, processors 212, 222, communication modules 213, 223, and input / output interfaces 214, 224. The memory 211, 221 may be a non-temporary computer-readable recording medium and may include non-temporary mass storage devices such as RAM (random access memory), ROM (read-only memory), disk drives, SSDs (solid-state drives), and flash memory. Here, non-temporary mass storage devices such as ROM, SSDs, flash memory, and disk drives may be included in the electronic device 110 and the server 150 as separate non-temporary storage devices distinct from the memory 211, 221. The memory 211, 221 may also store an operating system and at least one program code (for example, code for a browser installed and run on the electronic device 110, or code for an application installed on the electronic device 110 to provide a specific service). Such software components may be loaded from a computer-readable recording medium separate from the memory 211, 221. Such other computer-readable recording media may include computer-readable recording media such as floppy® drives, disks, tapes, DVD / CD-ROM drives, and memory cards. In other embodiments, software components may be loaded into memories 211, 221 through communication modules 213, 223 which are not computer-readable recording media. For example, at least one program may be loaded into memories 211, 221 based on a computer program (for example, the application described above) that is installed by a file provided over the network 170 by a file distribution system (for example, the server 160 described above) that distributes installation files for a developer or application.
[0023] Processors 212 and 222 may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. The instructions may be provided to processors 212 and 222 by memory 211, 221 or communication modules 213, 223. For example, processors 212 and 222 may be configured to execute instructions received according to program code recorded in a recording device such as memory 211, 221.
[0024] Communication modules 213 and 223 may provide functions for the electronic device 110 and the server 150 to communicate with each other via the network 170, or may provide functions for the electronic device 110 and / or the server 150 to communicate with other electronic devices (for example, the electronic device 120) or other servers (for example, the server 160). As an example, a request generated by the processor 212 of the electronic device 110 according to program code recorded in a recording device such as memory 211 may be transmitted to the server 150 via the network 170 under the control of the communication module 213. Conversely, control signals, instructions, contents, files, etc. provided under the control of the processor 222 of the server 150 may be received by the electronic device 110 through the communication module 213 of the electronic device 110 via the communication module 223 and the network 170. For example, control signals, instructions, contents, files, etc. received through the communication module 213 from the server 150 may be transmitted to the processor 212 or the memory 211, and the contents, files, etc. may be recorded in a recording medium (the non-temporary recording device described above) that the electronic device 110 may further include.
[0025] The input / output interface 214 may be a means for interface with an input / output device 215. For example, an input device may include a keyboard, mouse, microphone, camera, etc., and an output device may include a display, speaker, haptic feedback device, etc. As another example, the input / output interface 214 may be a means for interface with a device that integrates input and output functions into one, such as a touchscreen. The input / output device 215 may consist of the electronic device 110 and one other device. Also, the input / output interface 224 of the server 150 may be a means for interface with an input or output device (not shown) that connects to or can be included in the server 150. As a more specific example, when the processor 212 of the electronic device 110 processes instructions for a computer program loaded into memory 211, a service screen or content configured using data provided by the server 150 or electronic device 120 may be displayed on the display via the input / output interface 214.
[0026] Furthermore, in other embodiments, the electronic device 110 and the server 150 may include more components than those shown in Figure 2. However, it is not necessary to explicitly show most of the conventional components in the figure. For example, the electronic device 110 may be implemented to include at least some of the input / output devices 215 described above, and may further include other components such as transceivers, cameras, various sensors, and databases. As a more specific example, if the electronic device 110 is an AI speaker, the electronic device 110 may be implemented to further include a variety of components that are generally included in an AI speaker, such as various sensors, camera modules, various physical buttons, buttons using a touch panel, input / output ports, and vibrators for vibration. (End of quote)
[0027] The following disclosure concerns a machine. According to at least one embodiment, the user terminal comprises a control unit, RAM, storage unit, graphics processing unit, communication interface, and interface unit, each connected by an internal bus.
[0028] According to at least one embodiment, the control unit consists of a CPU and ROM. The control unit executes programs stored in the storage unit and controls the user terminal. RAM is the work area of the control unit. The storage unit is a memory area for saving programs and data. The control unit reads programs and data from RAM and processes them. By processing the programs and data loaded into RAM, the control unit outputs drawing commands to the graphics processing unit.
[0029] According to at least one embodiment, the graphics processing unit is connected to the display unit. The display unit has a display screen. When the control unit outputs a drawing command to the graphics processing unit, the graphics processing unit outputs a video signal for displaying an image on the display screen. Here, the display unit may be a touch panel equipped with a touch sensor. The touch panel of this display unit functions as an input unit.
[0030] According to at least one embodiment, the communication interface can be connected to a communication network wirelessly or via a wired connection, and can send and receive data with a server device via the communication network. The data received via the communication interface is loaded into RAM and processed by the control unit. External memory (e.g., an SD card) is connected to the interface unit.
[0031] According to at least one embodiment, the user terminal is not particularly limited as long as it is a computer device having a display screen and an input unit. Examples of user terminals include conventional mobile phones, tablet devices, smartphones, and desktop or notebook personal computers. A VR goggle, i.e., a screen (or two display panels, one for each eye) attached to a frame (or headset) that is fixed or attached to the head with a strap, may also be used. The user terminal has an audio output unit.
[0032] According to at least one embodiment, a user terminal can communicate with a server device via a communication network. It can transmit or receive information by establishing a communication connection via the communication network.
[0033] According to at least one embodiment, the server device comprises at least a control unit, RAM, a storage unit, and a communication interface, each connected by an internal bus.
[0034] According to at least one embodiment, the control unit consists of a CPU and ROM, executes programs stored in the storage unit, and controls the server device. The control unit also has an internal timer for timing. RAM is the work area of the control unit. The storage unit is a memory area for saving programs and data. The control unit reads programs and data from RAM and performs program execution processing based on information received from the user terminal, etc.
[0035] This document discloses AI. According to at least one embodiment, artificial intelligence includes machine learning, deep learning, generative AI, large-scale language models (LLMs), foundational models, and generative AI. Generative AI uses transformers and employs numerous attention mechanisms. It uses self-supervised learning and Extract Prediction. In this case, the AI can predict the next word. Given a sentence, it predicts the next word from the sentence up to that point. It generates a large number of supervised learning problems. These enable an AI that can predict the next word. Generative AI can predict grammatical structure, topic connections, and what kind of sentences people with a certain writing style are likely to write. Furthermore, by simply predicting the next sentence, generative AI can learn the underlying structure, causal relationships, and knowledge. Generative AI scales quickly, and its accuracy improves as the number of parameters increases. Ordinary statistics and machine learning overfit if the model parameters are too large compared to the data sample size. LLMs become more accurate as the number of parameters increases. One generative AI has 175 billion parameters. Generative AI is overlaid with supervised learning to facilitate smooth conversation. They are taught not to say anything strange. They write essays and act as call center operators.
[0036] According to at least one embodiment, Large Language Models (LLMs) are, in a non-inclusive sense, machine learning natural language processing models built using large datasets and deep learning techniques. Generally, they are adapted to various natural language processing (NLP) tasks such as text classification and generation, sentiment analysis, text summarization, and question answering using a technique called "fine-tuning," which involves training them on specific tasks. According to at least one embodiment, self-supervised learning is close to intrinsic human intelligence. When humans act, they are always predicting what will happen next and predicting the next input. In the process, they can learn the structure of the external world. Predicting the next word is considered intrinsic intelligence and is close to what is done in the cerebral cortex. According to at least one embodiment, Large Language Models memorize all the input information but generalize it to the extent necessary to predict the next word. They do not try to generalize all the information from the beginning. Large Language Models require capacity to remember information, and therefore require parameters. According to at least one embodiment, a large-scale language model incorporates eight models, each with 175 billion parameters or 220 billion parameters.
[0037] According to at least one embodiment, videos and images are represented as a collection of visual patches, which are small data units similar to text tokens in LLMs. Patches effectively represent models of visual data and are used as highly scalable and effective representations for training generative models on various types of videos and images. Videos are transformed into patches by first compressing the video into a low-dimensional latent space and then decomposing the representation into spatiotemporal patches.
[0038] According to at least one embodiment, a video compression network is a network that reduces the dimensionality of visual data, taking raw video as input and outputting a temporally and spatially compressed latent representation. An AI is trained in this compressed latent space and then generates video within this compressed latent space.
[0039] According to at least one embodiment, Spacetime Latent Patches extract a series of spatiotemporal patches that function as transformer tokens when given a compressed input video. The patch-based representation allows Sora to be trained on videos and images of varying resolutions, lengths, and aspect ratios, and controls the size of the resulting video by arranging randomly initialized patches into a grid of appropriate size during inference.
[0040] According to at least one embodiment, the AI is a diffusion model that, given a noisy patch (and conditioning information such as text prompts) as input, is trained to predict the original "clean" patch. The AI is a diffusion transformer, which exhibits remarkable scaling properties in various domains, including language modeling, computer vision, and image generation. Diffusion transformers are also effective as video generation models. The AI's sample quality improves significantly as the amount of training computation increases.
[0041] According to at least one embodiment, the AI applies caption regeneration technology to train a highly descriptive caption model, which is then used to generate text captions for all videos in the training set. Training highly descriptive captions improves not only the overall quality of the generated videos but also the fidelity of the text. GPT is used to convert short user prompts into long, detailed captions, which are then sent to the model. This allows the AI to generate high-quality videos that precisely follow the user prompts.
[0042] According to at least one embodiment, AI can perform vectorization in natural language processing according to the following flow. First, as a preprocessing step, the given text is cleaned. In the cleaning process, unnecessary words such as JavaScript code and HTML tags contained in the text are removed. These codes are used to display on the internet and are therefore not generally used in natural language processing. Next, the text is divided into words using morphological analysis. Morphological analysis is the classification of natural language sentences written in characters into the smallest meaningful linguistic units. "MeCab," "JUMAN," and "JANOME" can be used as morphological analysis tools. In normalization, words with the same meaning, such as variations in spelling, are unified into a single word. Stop words are words that are excluded from processing for reasons such as not being usable in natural language processing. Examples of stop words include particles and auxiliary verbs, which do not have meaning on their own. When calculating vectors, these may be removed, and only meaningful words may be targeted. Vectorization may also be performed without removing these stop words. Vectorization is the process of converting strings of words into vectors. Vectorization transforms word data into numerical data. When converting words to vectors, methods such as Bag of Words and distributed representations are used. Bag of Words is a method of vectorizing a given text using the number of occurrences of each word. It focuses on how often each word appears in the text, and does not consider the order of words or sentences. Distributed representation is a method of vectorizing by focusing on the meaning of words. By vectorizing the meaning of words, it is possible to assign similar vectors to words with similar meanings or usages, and the relationships between words can also be represented by vectors. With vector representation, it is possible to add and subtract the meanings of words. In applied processing, natural language converted into numerical data can be used as input for machine learning. Specifically, vectorized natural language is fed into a classifier to perform text classification.Tools used here include TensorFlow, scikit-learn, and PyTorch.
[0043] The present invention discloses a state in which at least two or more users exist. In at least one embodiment, users include general consumers and traders. Users use user terminals or computers.
[0044] Disclosure is made on a per-user basis. In at least one embodiment, the device performs predetermined operations for each user. Data is stored for each user.
[0045] This document discloses embodiments for obtaining information about a user's speech and actions over at least two days. In at least one embodiment, user speech and actions include the user's language activities. Language activities include all intellectual activities performed through language, such as speaking, listening, writing, and reading. These activities include not only performing these actions, but also one or more elements such as language for solving problems, language necessary for social life, activities that go beyond mere information transmission to organizing one's thoughts and communicating them to others, and language for a deeper understanding of the learning content of each subject. Language activities reflect not only the acquisition of knowledge, but also abilities that are important for living in the coming era, such as thinking ability, judgment ability, and expression ability. Information about a user's speech and actions includes data on the user's language activities. For example, this includes textual data of the user's language and supplementary data such as the category of emotion inferred from the facial muscles used when the user speaks. This also includes data of the user's language activities that can be processed by the device. Methods of data conversion include, if it is the user's language, acquiring the user's voice and transcribing the voice into text. It also includes acquiring data if the user's language is already posted on the user's social media. Other methods include capturing the user's facial expressions using a video acquisition device and requesting an AI, which has undergone supervised learning about human facial expressions, to determine what emotions the user's facial expressions reflect and categorize them (including categories such as joy, anger, sadness, and happiness).
[0046] This document discloses embodiments of which an AI is required to perform supervised learning based on such information. In at least one embodiment, such information includes information about the user's speech and behavior. The device requires the AI to perform supervised learning with the information about the user's speech and behavior. The AI includes a natural language processing model. The AI is further fine-tuned with the information about the user's speech and behavior. This method enables the AI to perform text classification and generation, sentiment analysis, text summarization, and question answering that reflect the user's features and characteristics. For example, the AI learns words specific to the user and the flow of sentences from information about the user's language. As already described, given a sentence, the AI guesses the next word from the sentence up to that point. A large number of supervised learning problems are generated. This allows for an AI that can guess the next word based on the user's characteristics. The generative AI can predict grammatical structure, topic connections, and what kind of sentences a person with a certain writing style is likely to write. In at least one embodiment, the device performs supervised learning for each user. That is, the device changes the information about the user's speech and behavior that should be used for supervised learning for each user. AI can predict what kind of sentence user A would write, and what kind of sentence user B would write.
[0047] This invention discloses an embodiment for storing the learning data separately. In at least one embodiment, the device stores the AI's learning data separately. For example, it stores learning data for user A and learning data for user B separately. This embodiment offers the convenience and industrial applicability of being able to store the AI's personality for at least one user and accumulate the results of that personality's learning.
[0048] This invention discloses an embodiment in which an AI is required to predict a user's corresponding verbal and physical behavior regarding a given task, based on specified training data. In at least one embodiment, the device acquires information about a task. The task includes tasks that can be responded to as verbal activities. For example, a question. For example, this includes an abstract question such as, "What do you think about the global warming problem?" Other examples include a specific question, such as, "Company A's sales are X billion yen, and it has product X1, product X2, etc. Should it sell product X3?", with specific data attached as preconditions for the question. For these tasks, the user in question can create the task in the information about the user's verbal and physical behavior. Alternatively, someone other than the user can create the task. The content of the task can be anything. It may be a task that only the user would know, or it may be a general task. The device requests the AI to predict a user's corresponding verbal and physical behavior regarding a given task, based on specified training data. The AI predicts the user's verbal and physical behavior regarding the task based on the stored training data. For example, based on an abstract question like, "What do you think about the global warming problem?", the AI predicts grammatical structure, topic connections, and what kind of sentences a person with a certain writing style is likely to write. For instance, based on user A's training data, the AI might generate a predicted sentence like, "To combat global warming, measures X1 and X2 are necessary..." Based on user B's training data, it might generate a predicted sentence like, "There is no scientific basis for global warming, so there is no need to take measures."
[0049] This embodiment offers at least the following advantages and industrial applicability: A persona can be saved on the device. Past social media activity can be scanned to discover the user's thoughts. Answering tasks (sometimes synonymous with questions) allows for the selection of the user's thought base to create a clone AI of the user. In this process, the user's personality traits can be extracted. Consulting and AI representations of famous CEOs can be rented to assist with daily life and business management. If the user is a designer or a famous business executive, the AI associated with that user can participate in internal meetings and offer suggestions on issues. This can be helpful for consulting and business management.
[0050] Information regarding the user's behavior is disclosed to the device as the user's SNS posting data. In at least one embodiment, the device references the user's SNS posting data. The device obtains proof information that the user has given permission. The user generates permission information that allows the AI to learn the posting data related to the SNS account owned by the user. As an example, an application implementing some or all of this embodiment performs user authentication and confirms that the user has given permission for the AI to learn the posting data related to the SNS account owned by the user. The verification method includes the user entering the password for that SNS account, an authentication method linked to that SNS (including, for example, entering a password or biometric information for a Google® account), or a combination thereof. When the device obtains proof information that the user has given permission for the AI to learn the posting data related to the SNS account owned by the user, it requests the AI to learn based on the user's SNS posting data. According to this embodiment, there is the convenience and industrial applicability of preventing the user's SNS posting data from being used for learning in a form that the user does not desire.
[0051] Disclosed is a device that acquires a user's voice or image, requests an AI to converse with the user or ask questions to the user, and uses information related to such conversation or questions as information related to the user's speech or actions. In at least one embodiment, the device has voice acquisition means or image acquisition means. For example, these may be provided in the user terminal. For example, this includes the microphone and camera of a smartphone or tablet terminal as the user terminal. A separate device may be made capable of communicating with the device. The device requests an AI capable of natural language processing to converse with the user or ask questions to the AI. The device stores initial question information to be asked first and initially asks the user this information. Such question information is determined by the business operator providing the application that implements some or all of this embodiment. It may also be determined arbitrarily by the user. Based on the initial question information, the AI converses with or asks questions to the user. In this method, the content of the conversation or questions is displayed as text on the user terminal or output as audio. Based on the user's response obtained from the user terminal (including methods of inputting as text or analyzing and transcribing audio input), the AI further engages in conversation or asks questions. The AI continues the conversation and asks questions based on the user's input responses. The business operator providing the application that implements part or all of this embodiment determines the question information to be asked regarding that conversation or conversation. The device treats the information related to the conversation or questions as information related to the user's behavior. This embodiment offers the convenience and industrial applicability of efficiently acquiring at least the user's characteristics, or acquiring user characteristics that cannot be understood from SNS posts.
[0052] Disclosed is an apparatus that obtains authentication information to prove that a user has authenticated and provides the authentication information or information relating to the authentication information to the training data corresponding to the user. In at least one embodiment, a user generates authentication information to prove that they have authenticated the prediction of the user's behavior in relation to a given task, based on the training data or the training data. Authentication includes the user authenticating that the training data based on the user or the AI that predicts the user's behavior in relation to a given task based on the training data reflects the user's characteristics. For example, a user may, through an application that implements some or all of this embodiment, refer to their training data or examine the content generated by the AI based on that training data and determine whether it reflects the user's characteristics. For example, user A The user reviews text generated by an AI with user-based training data and determines whether it resembles or does not resemble the user. The user provides proof information when they believe the AI reflects their characteristics. This proof information includes confirmation that the user has authenticated the AI, permission to use the training data, the AI's similarity score, and combinations thereof. The user inputs the similarity score, for example, "80% similar AI," "40% similar AI," or "not very similar AI." The expression can be set in any way. The device presents the user with user-based training data or an AI that predicts the user's behavior on a given task based on the training data, and the user authenticates the training data or the AI that predicts the user's behavior on a given task based on the training data. The device obtains proof information to prove that authentication has been made. For example, this proof information may include digital signatures, confirmation that a specified password has been entered, biometric authentication, or other available authentication methods. Information regarding the proof information includes the date of authentication, the form of authentication (digital signature, biometric authentication, etc.), the AI's similarity score, and one or more combinations thereof. The device provides the proof information or information regarding the proof information to the training data corresponding to the user. For example, this includes associating and storing certification information with the learning data, or attaching a timestamp that proves the certification information has not been tampered with. According to this embodiment, it is possible to determine whether or not a user has authenticated at least their learning data or the AI based on that learning data. This improves the ease of detecting misuse of a user's learning data, or enhances the trustworthiness of the learning data and the AI based on it, offering both convenience and industrial applicability.
[0053] Disclosed is a device that displays authentication graphics on a display screen when displaying content relating to training data that has been assigned authentication information or information relating to authentication information. In at least one embodiment, when the device activates the AI, it checks whether the training data has authentication information or information relating to authentication information. If either of these exists, the device displays authentication graphics on the display screen of the user who is using the content relating to the training data or the AI based on the training data. The term "user" here includes not only the user on which the training data is based, but also general consumers or traders who use the training data or the AI. The display screen includes the display screen of the user terminal when these users access, view, or use the training data or the AI based on the training data. The authentication graphics include information indicating that the user has been authenticated, authentication information, information relating to castle name information, user similarity, and one or more combinations thereof. This includes methods of displaying this information in text or displaying predetermined graphics corresponding to this information. The device stores the information of these graphics and displays the corresponding graphics when the above conditions are met. For example, when using the AI of user A that has authenticated user A's authentication information, the display screen of the user terminal using it displays graphics of "Authenticated," "Verified," or stamps indicating these. Of course, these are merely examples. In at least one embodiment, any indication that user A has authenticated is considered to fall under the category of "authentication-related graphics." This embodiment offers the convenience and industrial applicability of improving the ease of detecting misuse of user training data, or improving the reliability of training data and the AI based thereon.
[0054] Disclosed is a device that requests an AI to predict the behavior of two or more users based on the training data, and requests the AI to converse with these two or more users on a given task. In at least one embodiment, the device stores two or more training data for each user. The device requests the AI to predict the behavior of a corresponding user on a given task based on the two or more training data for each user. The device requests the AI to converse with the two or more users on a given task. The device initially stores question information and initially asks this information to the AI based on one of the users. Such question information is determined by the business operator providing the application that implements some or all of this embodiment. It can also be determined arbitrarily by the user. For example, the AI based on user A is asked a question, the AI predicts the behavior of user A, and the device obtains the predicted language. Next, based on that language, the AI predicts the language of user B and obtains that predicted language. Furthermore, based on that language, the AI predicts the behavior of user A. Through this repetition, the AI causes the AI based on two or more personalities to converse or discuss. According to this embodiment, there is the convenience and industrial applicability of automatically generating excellent ideas or judgments by having AIs based on at least two excellent users converse with each other.
[0055] The following is an overview of the embodiments described above.
[0056] With at least two users present, for each user, We obtain information about the user's words and actions over at least two days. Based on that information, the AI is instructed to perform supervised learning. The learning data is stored, The AI is asked to predict a user's behavior in response to a given task, based on specified training data. Device.
[0057] The above-mentioned device, Information regarding users' behavior is derived from their social media posts. Device.
[0058] The above-mentioned device, Acquire the user's voice or image, The AI is asked to converse with the user or ask the user questions. The information regarding the conversation or question will be treated as information regarding the user's words and actions. Device.
[0059] The above-mentioned device, Obtaining authentication information to prove that the user has authenticated, Provide proof information or information related to proof information to the learning data corresponding to the user. Device.
[0060] The above-mentioned device, Obtaining authentication information to prove that the user has authenticated, Provide proof information or information related to proof information to the learning data corresponding to the user in question. When displaying on the screen content related to training data or AI based on training data that has been assigned certification information or information related to certification information, a graphic related to authentication is displayed on the display screen. Device.
[0061] Based on the above training data, the AI is asked to predict the words and actions of two or more users. The AI is asked to converse with two or more users regarding a certain issue. Device.
[0062] In at least one embodiment, the apparatus is as described above, Acquire the user's voice or image, The information regarding the audio or image will be considered as information regarding the user's words and actions. Device.
[0063] In at least one embodiment, the device acquires the user's voice or image. The user speaks to the device on their own, even without receiving a conversation or question from the AI. The device acquires the user's voice or image using voice acquisition means or image acquisition means. The device uses at least the information regarding the voice or image as information regarding the user's speech or actions. According to this embodiment, at least the user has the convenience and industrial applicability of being able to spontaneously work to improve the accuracy and similarity of their own AI.
[0064] In at least one embodiment, information regarding the user's words and actions is video or photo data owned by the user. Video or photo data owned by the user includes, for example, video or photo data stored in the user's cloud. Google Photos® is one example. Of course, this is just an example; the format in which the data is stored is not limited to video or photo data owned by the user. For example, the photo or video may be taken or created by the user. For example, the photo or video may further include illustrations or graphics created by the user. The device acquires authentication data that authorizes the user to access or use the video or photo. After acquiring the authentication data, the device requests the AI to perform supervised learning based on the video or photo. According to this embodiment, there is the convenience and industrial applicability of being able to provide an AI based on data that at least reflects the user's preferences and characteristics. As a result of good faith consideration by the inventors, compared to performing supervised learning using only the user's SNS posting data, including the user's videos or photos, illustrations, and graphics in the supervised learning process can create an AI that is more similar to the user. This configuration includes at least some novel aspects that have not been previously revealed, offering the convenience and industrial applicability of creating AI or training data that is more similar to the user.
[0065] The invention disclosed herein only needs to achieve at least one of the effects described above.
Claims
1. With at least two users present, for each user, We obtain information about the user's words and actions over at least two days. Based on that information, the AI is instructed to perform supervised learning. The learning data is stored, The AI is asked to predict a user's behavior in response to a given task, based on specified training data. Device.
2. The apparatus according to claim 1, Information regarding users' behavior is derived from their social media posts. Device.
3. The apparatus according to claim 1, Acquire the user's voice or image, The information regarding the audio or image will be considered as information regarding the user's words and actions. Device.
4. The apparatus according to claim 1, Acquire the user's voice or image, The AI is asked to converse with the user or ask the user questions. The information regarding the conversation or question will be treated as information regarding the user's words and actions. Device.
5. The apparatus according to claim 1, Obtaining authentication information to prove that the user has authenticated, Provide proof information or information related to proof information to the learning data corresponding to the user. Device.
6. The apparatus according to claim 1, Obtaining authentication information to prove that the user has authenticated, Provide proof information or information related to proof information to the learning data corresponding to the user in question. When displaying on the screen content related to training data or AI based on training data that has been assigned certification information or information related to certification information, a graphic related to authentication is displayed on the display screen. Device.
7. The AI is requested to predict the words and actions of two or more users based on the learning data described in claim 1. The AI is asked to converse with two or more users regarding a certain issue. Device.