Interaction system for assisting with use of medium by user

The interactive system uses an AI server to provide contextual responses to children's questions about book content, addressing the limitations of sound pens by integrating voice input with book context, thereby enhancing engagement and understanding.

WO2025178442A1PCT designated stage Publication Date: 2025-08-28NEO LAB CONVERGENCE INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099343
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-11
Filing Date
2025-02-07
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Sound pens used by young children to read books do not allow for interactive conversations about the content, limiting their ability to engage and understand the material deeply.

Method used

An interactive system that utilizes an AI server to generate responses based on both the child's voice input and contextual information from the book, such as printed content and user interactions, using a large language model like Chat GPT to facilitate conversation.

Benefits of technology

Enables children to engage in interactive discussions about the book content, enhancing their understanding and interest in reading through AI-generated responses tailored to the book's context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099343_28082025_PF_FP_ABST
    Figure KR2025099343_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an interaction system for assisting with use of a medium by a user. More specifically, the present invention relates to an interaction system which, when a user reads a medium or performs a predetermined activity through the medium, resolves user's questions related to the content of the medium, helps the user understand the medium, or assists the user to smoothly perform the activity.
Need to check novelty before this filing date? Find Prior Art

Description

An interactive system to assist users in using the medium.

[0001] The present disclosure relates to an interactive system for assisting users in using a medium. More specifically, the present disclosure relates to a method and a system for performing the same, which resolves users' questions related to the content of a medium, enhances their understanding of the medium, or facilitates smooth execution of the activity when the user reads the medium or performs a designated activity through the medium.

[0002] Recently, sound pens have been actively used by young children (e.g., under 7 years old) to read storybooks or picture books on their own. When a child touches the book with the sound pen, the pen converts the text in the touched area into spoken words and outputs them.

[0003] These sound pens have the advantage of allowing children to read books without a guardian, but there is a problem in that children can only listen to the sound output from the sound pen unilaterally and cannot have a conversation about the content of the book like when a guardian reads a book.

[0004] The applicant developed the invention of the present disclosure by considering a method for resolving a child's curiosity about the content of a book while reading it, or for making the child understand the content of the book more clearly and become interested in the act of reading.

[0005] The task we are trying to solve is to provide responses to what the child says when he or she is reading a book on his or her own.

[0006] The task we are trying to solve is to encourage children to understand the content of books more deeply when they read books on their own.

[0007] The task we are trying to solve is to convey information and encourage learning to children through cards.

[0008] The problems to be solved are not limited to the problems described above, and problems not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from this specification and the attached drawings.

[0009] According to one embodiment, a server for communicating with an electronic device and an artificial intelligence server, comprising: a memory; a communication unit; and a control unit for controlling the memory and the communication unit;, wherein the control unit, [1] obtains reference information encoded in code data obtained by the electronic device, wherein the reference information is received from the electronic device through the communication unit or obtained by the control unit decoding the code data after receiving the code data from the electronic device through the communication unit, and the reference information relates to a medium with which a user interacts using the electronic device, and the reference information includes ID information of the medium, page information indicated by the electronic device, and location information within the page indicated by the electronic device, [2] based on the reference information, obtains content of interest about a part of the content printed on the medium and / or printed on the medium in which the user is interested from the memory to generate medium context information, and [3] obtains a voice text corresponding to the user's voice obtained by the electronic device, wherein the voice text is encoded by decoding the voice. An electronic device and an artificial intelligence server are provided that are obtained by converting STT (Speech-To-Text) -, [4] generate a prompt based on the obtained medium context information and the voice text, wherein the prompt includes the medium context information and the voice text -, [5] transmit the prompt to the artificial intelligence server through the communication unit, [6] receive a response from the artificial intelligence server to the transmitted prompt through the communication unit, [7] obtain a response voice corresponding to the received response, and [8] transmit the response voice to the electronic device through the communication unit.

[0010] The means of solving the problem are not limited to the above-described means of solving the problem, and means of solving the problem that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from this specification and the attached drawings.

[0011] In one embodiment, a child can read a book alone and converse with an artificial intelligence model (e.g., chat GPT) about the book's content.

[0012] In one embodiment, the responses of the artificial intelligence model provided to the child may be related to content printed in the book.

[0013] In one embodiment, the responses of the artificial intelligence model provided to the child may be related to parts of the book printed that the child is interested in.

[0014] In one embodiment, a child may develop greater understanding or interest in the content of a book than simply reading it.

[0015] In one embodiment, a child can be provided with various activity experiences through cards.

[0016] The effects of the invention are not limited to the effects described above, and effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from this specification and the attached drawings.

[0017] FIG. 1 is a diagram illustrating an interaction system according to one embodiment.

[0018] Figure 2 is a drawing showing the configuration of an electronic device according to one embodiment.

[0019] FIG. 3 is a drawing illustrating a code printed on a medium according to one embodiment.

[0020] Figure 4 is a diagram showing the configuration of a main server according to one embodiment.

[0021] Figure 5 is a diagram showing the configuration of an artificial intelligence server according to one embodiment.

[0022] FIG. 6 is a drawing showing a block diagram of an electronic pen service system using an interactive artificial intelligence service according to one embodiment.

[0023] FIG. 7 is a drawing showing a block diagram of an electronic pen service system using an interactive artificial intelligence service according to another embodiment.

[0024] FIG. 8 is a diagram showing information stored in a database according to one embodiment.

[0025] FIG. 9 is a diagram illustrating an aspect in which areas are distinguished in a medium according to one embodiment.

[0026] FIG. 10 is a drawing showing an aspect in which areas are distinguished in a medium according to another embodiment.

[0027] FIG. 11 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to the first embodiment.

[0028] FIG. 12 is a flowchart illustrating a method for assisting a user's use of a medium by using artificial context information according to a second embodiment.

[0029] FIG. 13 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a third embodiment.

[0030] FIG. 14 is a flowchart illustrating a method for assisting a user's use of a medium by using artificial context information according to a fourth embodiment.

[0031] Fig. 15 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a fifth embodiment.

[0032] Fig. 16 is a diagram showing a process of obtaining artificial context information according to the fifth embodiment.

[0033] Fig. 17 is a diagram showing a prompt being generated using artificial context information according to the fifth embodiment.

[0034] Figure 18 is a flowchart illustrating a method for generating a follow-up prompt according to the fifth embodiment.

[0035] FIG. 19 is a diagram showing a subsequent prompt being generated using artificial context information according to the fifth embodiment.

[0036] Figure 20 is a diagram showing a process of obtaining a pre-stored prompt according to the sixth embodiment.

[0037] Fig. 21 is a drawing showing a pre-stored prompt according to the sixth embodiment.

[0038] FIG. 22 is a diagram showing a subsequent prompt being generated using artificial context information according to the sixth embodiment.

[0039] FIG. 23 is a flowchart illustrating a method for assisting a user's use of a medium by using medium context information according to the seventh embodiment.

[0040] Figure 24 is a diagram showing a process for obtaining medium context information according to the seventh embodiment.

[0041] FIG. 25 is a diagram showing a prompt being generated using medium context information according to the seventh embodiment.

[0042] Figure 26 is a flowchart illustrating a method for generating a follow-up prompt according to the seventh embodiment.

[0043] FIG. 27 is a flowchart illustrating a method for assisting a user's use of a medium by using artificial context information and medium context information according to the eighth embodiment.

[0044] Figure 28 is a diagram showing a process for obtaining artificial context information and medium context information according to the eighth embodiment.

[0045] FIG. 29 is a diagram showing a prompt being generated using artificial context information and medium context information according to the eighth embodiment.

[0046] Fig. 30 is a flowchart illustrating a method for assisting a user's use of a medium using contextual information according to the ninth embodiment.

[0047] According to one embodiment, a server for communicating with an electronic device and an artificial intelligence server, comprising: a memory; a communication unit; and a control unit for controlling the memory and the communication unit;, wherein the control unit, [1] obtains reference information encoded in code data obtained by the electronic device, wherein the reference information is received from the electronic device through the communication unit or obtained by the control unit decoding the code data after receiving the code data from the electronic device through the communication unit, and the reference information relates to a medium with which a user interacts using the electronic device, and the reference information includes ID information of the medium, page information indicated by the electronic device, and location information within the page indicated by the electronic device, [2] based on the reference information, obtains content of interest about a part of the content printed on the medium and / or printed on the medium in which the user is interested from the memory to generate medium context information, and [3] obtains a voice text corresponding to the user's voice obtained by the electronic device, wherein the voice text is encoded by decoding the voice. An electronic device and an artificial intelligence server are provided that are obtained by converting STT (Speech-To-Text) -, [4] generate a prompt based on the obtained medium context information and the voice text, wherein the prompt includes the medium context information and the voice text -, [5] transmit the prompt to the artificial intelligence server through the communication unit, [6] receive a response from the artificial intelligence server to the transmitted prompt through the communication unit, [7] obtain a response voice corresponding to the received response, and [8] transmit the response voice to the electronic device through the communication unit.

[0048] The above memory stores a plurality of medium texts, each medium text corresponding to at least a part of a text printed on a specific medium, and the control unit loads a medium text corresponding to the medium ID information of the reference information from among the plurality of medium texts from the memory to obtain the medium context information.

[0049] The memory stores a plurality of page-specific texts, each page-specific text corresponding to at least a portion of text printed on a specific page of a specific medium, and the control unit loads page-specific text corresponding to the medium ID information and the page information of the reference information from the memory to obtain the medium context information.

[0050] The memory stores a plurality of position-specific texts, each position-specific text corresponding to at least a portion of text printed at a specific position within a specific page of a specific medium, and the control unit loads position-specific text corresponding to the medium ID information, the page information, and the position information of the reference information from the memory to obtain the medium context information.

[0051] The above prompt is generated by concatenating the spoken text with the above medium context information.

[0052] The above prompt further includes descriptive text about the medium context information.

[0053] The above prompt further includes pre-stored user information, wherein the user information includes at least one of age and gender for the user.

[0054] The above prompt further includes pre-stored guide text, which includes matters for the artificial intelligence server to consider when generating the response.

[0055] The control unit obtains a prompt form from the memory, modifies the prompt form using the medium context information and the spoken text, and generates the prompt.

[0056] The above prompt comprises at least a portion of a past prompt transmitted to the artificial intelligence server prior to obtaining the code data from the electronic device.

[0057] The control unit transmits the past prompt to the artificial intelligence server through the communication unit and receives a past response to the past prompt from the artificial intelligence server, wherein the prompt includes at least a part of the past response.

[0058] The reference information further includes area identification information regarding a question or instruction printed on a medium with which the user is interacting using the electronic device, the area identification information identifying a preset area within the medium, and the control unit obtains artificial context information corresponding to the question or instruction printed on the preset area from the memory based on the reference information, and the prompt further includes the artificial context information.

[0059] The control unit obtains a follow-up voice text corresponding to the follow-up voice of the user obtained by the electronic device, wherein the follow-up voice text is obtained by converting the follow-up voice into STT, generates a follow-up prompt based on the follow-up voice text, transmits the follow-up prompt to the artificial intelligence server through the communication unit, receives a follow-up response of the artificial intelligence server to the transmitted follow-up prompt through the communication unit, obtains a follow-up response voice corresponding to the received follow-up response, and transmits the follow-up response voice to the electronic device through the communication unit.

[0060] The above follow-up prompt further includes the medium context information.

[0061] The above artificial intelligence server includes a Large Language Model (LLM).

[0062] The electronic device includes an image sensor and is configured to photograph an area of ​​the medium by the user's operation and obtain the code data from the photographed image.

[0063] The electronic device has a pen shape having one end and the other end, and when the user points one area of ​​the medium with one end of the electronic device, the electronic device is configured to capture an image of the medium using the image sensor arranged at the one end.

[0064] According to another embodiment, a server for communicating with an electronic device and an artificial intelligence server, comprising: a memory; a communication unit; And a control unit that controls the memory and the communication unit; wherein the control unit [1] obtains reference information encoded in code data obtained by the electronic device, wherein the reference information is received from the electronic device through the communication unit or obtained by the control unit decoding the code data after receiving the code data from the electronic device through the communication unit, and the reference information relates to a medium with which a user interacts using the electronic device, and the reference information includes ID information of the medium, page information indicated by the electronic device, and location information within the page indicated by the electronic device, [2] based on the reference information, obtains content of interest about a part of the content printed on the medium and / or the content printed on the medium that the user is interested in from the memory to generate medium context information, [3] obtains voice data corresponding to the voice of the user obtained by the electronic device, and [4] A server is provided that communicates with an electronic device and an artificial intelligence server, which generates a prompt using the acquired medium context information and the voice data, [5] transmits the prompt to the artificial intelligence server through the communication unit, [6] receives a response from the artificial intelligence server to the transmitted prompt through the communication unit, and [7] transmits the received response or a response voice corresponding to the response to the electronic device through the communication unit.

[0065] The data type of the above medium context information is any of text, sound, image, or video.

[0066] If the data type of the above medium context information is text and the data type supported by the above artificial intelligence server is sound, the control unit converts the medium context information so that the data type of the above medium context information becomes sound before generating the prompt.

[0067] According to another embodiment, a server for communicating with an electronic device and an artificial intelligence server, comprising: a memory; a communication unit; And a control unit that controls the memory and the communication unit; wherein the control unit [1] obtains reference information encoded in code data obtained by the electronic device - wherein the reference information is received from the electronic device through the communication unit or obtained by the control unit decoding the code data after receiving the code data from the electronic device through the communication unit, and the reference information includes area identification information for identifying a preset area within a medium with which the user is interacting using the electronic device, and at least one of a symbol, a question, and an instruction is printed in the preset area - , [2] based on the reference information, obtains artificial context information corresponding to the preset area from the memory, [3] obtains voice text corresponding to the user's voice obtained by the electronic device - wherein the voice text is obtained by converting the user's voice into STT (Speech-To-Text) - , [4] obtains the obtained A server is provided that communicates with an electronic device and an artificial intelligence server, which generates a prompt based on artificial context information and the spoken text, wherein the prompt includes the artificial context information and the spoken text, [5] transmits the generated prompt to the artificial intelligence server through the communication unit, [6] receives a response of the artificial intelligence server to the transmitted prompt from the server, [7] obtains a response voice corresponding to the received response, and [8] transmits the obtained response voice to the electronic device through the communication unit.

[0068] The memory stores a plurality of questions - each question corresponding to a question printed on a specific medium - and the artificial context information is a question corresponding to the area identification information of the reference information among the plurality of questions.

[0069] The above memory stores a plurality of instructions - each instruction corresponding to an instruction printed on a specific medium - and the artificial context information is an instruction corresponding to the area identification information of the reference information among the plurality of instructions.

[0070] The above reference information further includes a medium ID, and the artificial context information is information corresponding to the medium ID and the area identification information among the information stored in the memory.

[0071] The above reference information further includes page information, and the artificial context information is information corresponding to the medium ID, the page information, and the area identification information among the information stored in the memory.

[0072] The above prompt is generated by concatenating the spoken text with the artificial context information.

[0073] The above prompt further includes descriptive text for the artificial context information.

[0074] The above prompt further includes pre-stored user information, wherein the user information includes at least one of age and gender for the user.

[0075] The above prompt further includes pre-stored guide text, which includes matters for the artificial intelligence server to consider when generating the response.

[0076] The control unit obtains a prompt form from the memory, modifies the prompt form using the artificial context information and the spoken text, and generates the prompt.

[0077] The above prompt comprises at least a portion of a past prompt transmitted to the artificial intelligence server prior to obtaining the code data from the electronic device.

[0078] The control unit transmits the past prompt to the artificial intelligence server through the communication unit and receives a past response to the past prompt from the artificial intelligence server, wherein the prompt includes at least a part of the past response.

[0079] The reference information further includes at least one of ID information of the medium, page information indicated by the electronic device, and location information within the page indicated by the electronic device, and the control unit, based on the reference information, acquires content of interest about a part of the content printed on the medium and / or printed on the medium that the user is interested in from the memory to generate medium context information, and the prompt further includes the medium context information.

[0080] The control unit obtains a follow-up voice text corresponding to the user's follow-up voice obtained by the electronic device, wherein the follow-up voice text is obtained by converting the follow-up voice into STT, generates a follow-up prompt based on the follow-up voice text and the artificial context information, transmits the follow-up prompt to the artificial intelligence server through the communication unit, receives a follow-up response of the artificial intelligence server to the transmitted follow-up prompt through the communication unit, obtains a follow-up response voice corresponding to the received follow-up response, and transmits the follow-up response voice to the electronic device through the communication unit.

[0081] The above follow-up prompt further includes the medium context information.

[0082] The above artificial intelligence server includes a Large Language Model (LLM).

[0083] The electronic device includes an image sensor and is configured to photograph an area of ​​the medium by the user's operation and obtain the code data from the photographed image.

[0084] The electronic device has a pen shape having one end and the other end, and when the user points one area of ​​the medium with one end of the electronic device, the electronic device is configured to capture an image of the medium using the image sensor arranged at the one end.

[0085] According to another embodiment, a server for communicating with an electronic device and an artificial intelligence server, comprising: a memory; a communication unit; And a control unit that controls the memory and the communication unit; wherein the control unit [1] obtains reference information encoded in code data obtained by the electronic device, wherein the reference information is received from the electronic device through the communication unit or obtained by the control unit decoding the code data after receiving the code data from the electronic device through the communication unit, and the reference information relates to a question or instruction printed on a medium with which the user is interacting using the electronic device, and the reference information includes area identification information for identifying a preset area within the medium, [2] based on the reference information, obtains artificial context information stored in advance corresponding to the preset area from the memory, [3] based on the obtained artificial context information, generates a prompt, wherein the prompt includes the artificial context information, [4] transmits the generated prompt to the artificial intelligence server through the communication unit, and [5] transmits the generated prompt from the server to the An electronic device and a server communicating with an artificial intelligence server are provided, which receive a response from the artificial intelligence server for a transmitted prompt, [6] obtain a response voice corresponding to the received response, and [7] transmit the obtained response voice to the electronic device through the communication unit.

[0086] The control unit obtains a voice text corresponding to the user's voice obtained by the electronic device - wherein the user's voice is obtained after the response voice is transmitted to and output by the electronic device, and the voice text is obtained by converting the subsequent voice into STT -, generates a subsequent prompt based on the voice text, transmits the subsequent prompt to the artificial intelligence server through the communication unit, receives a subsequent response of the artificial intelligence server to the transmitted subsequent prompt through the communication unit, obtains a subsequent response voice corresponding to the received subsequent response, and transmits the subsequent response voice to the electronic device through the communication unit.

[0087] According to another embodiment, a server for communicating with an electronic device and an artificial intelligence server, comprising: a memory; a communication unit; And a control unit that controls the memory and the communication unit; wherein the control unit [1] obtains reference information encoded in code data acquired by the electronic device - wherein the reference information is received from the electronic device through the communication unit or obtained by the control unit decoding the code data after receiving the code data from the electronic device through the communication unit, and the reference information includes area identification information for identifying a preset area within a medium with which the user is interacting using the electronic device, and at least one of a symbol, a question, and an instruction is printed in the preset area - , [2] based on the reference information, obtains artificial context information stored in advance corresponding to the preset area from the memory, [3] obtains voice data corresponding to the user's voice acquired by the electronic device, [4] generates a prompt using the acquired artificial context information and the voice data, and [5] transmits the prompt to the artificial intelligence server through the communication unit. A server is provided that communicates with an electronic device and an artificial intelligence server, which transmits, [6] receives a response from the artificial intelligence server to the transmitted prompt through the communication unit, and [7] transmits the received response or a response voice corresponding to the response to the electronic device through the communication unit.

[0088] The data type of the above artificial context information is any of text, sound, image, or video.

[0089] If the data type of the artificial context information is text and the data type supported by the artificial intelligence server is sound, the control unit converts the artificial context information so that the data type of the artificial context information becomes sound before generating the prompt.

[0090] The above-described purposes, features, and advantages will become more apparent through the following detailed description taken in conjunction with the accompanying drawings. However, the present invention is susceptible to various modifications and various embodiments. Therefore, specific embodiments will be illustrated in the drawings and described in detail below.

[0091] In the drawings, the thicknesses of layers and regions are exaggerated for clarity, and when an element or layer is referred to as "on" or "on" another element or layer, this includes not only the case where the element or layer is directly above the other element or layer, but also the case where another layer or other element is interposed. In principle, the same reference numerals represent the same elements throughout the specification. In addition, elements that have the same function within the scope of the same idea shown in the drawings of each embodiment are described using the same reference numerals, and redundant descriptions thereof will be omitted.

[0092] The numbers used in the description of this specification (e.g., first, second, etc.) are merely identifiers to distinguish one component from another.

[0093] In addition, the suffixes "module" and "part" for components used in the following examples are given or used interchangeably only for the convenience of writing the specification, and do not have distinct meanings or roles in themselves.

[0094] In the examples below, singular expressions include plural expressions unless the context clearly indicates otherwise.

[0095] In the following examples, terms such as “include” or “have” mean that a feature or component described in the specification is present, and do not preclude the possibility that one or more other features or components may be added.

[0096] For convenience of explanation, the sizes of components in the drawings may be exaggerated or reduced. For example, the sizes and thicknesses of each component shown in the drawings are arbitrarily shown for convenience of explanation, and the present invention is not necessarily limited to what is shown.

[0097] In some embodiments, where implementations are otherwise feasible, specific process sequences may be performed in a different order than described. For example, two processes described in succession may be performed substantially simultaneously, or in a reverse order from the described order.

[0098] In the following examples, when it is said that a film, region, component, etc. are connected, it includes not only cases where the films, regions, and components are directly connected, but also cases where other films, regions, and components are interposed between the films, regions, and components and are indirectly connected.

[0099] For example, when it is said in this specification that a film, region, component, etc. are electrically connected, it includes not only cases where the film, region, component, etc. are directly electrically connected, but also cases where another film, region, component, etc. is interposed and is indirectly electrically connected.

[0100] Unless specifically stated or clear from context, the term "about" in relation to a numerical value shall be understood to mean the numerical value stated plus or minus 10% of that numerical value, and the term "about" in relation to a numerical range shall be understood to mean a range from 10% below the lower limit of the numerical range to 10% above the upper limit of the numerical range.

[0101]

[0102] 1. Overview

[0103] As described in the background technology of the invention, to address the issue of sound pens that can only read books and not converse with or answer children's questions, one solution is to consider utilizing an AI server. Specifically, a system could be considered that records a child's voice, inputs the recorded voice (or the converted text) into a large language model (LLM) such as Chat GPT (Generative Pre-trained Transformer), generates a response, and then reads it back to the child.

[0104] At this time, the child is young and has relatively low language skills, making it difficult for him or her to use sentences with clear meanings. To compensate for this lack of language skills, the child frequently uses demonstrative pronouns such as 'this' or 'that' with his or her fingers while reading a book. This characteristic must be taken into account.

[0105] In other words, the large language model can only receive the child's voice (or text converted from the voice) as input, and it is difficult to receive information such as the child's pointing behavior, which causes the problem of not being able to know exactly what the child is trying to say or what he or she is curious about.

[0106] For example, if a child is reading a Snow White storybook and asks questions such as “Is this where the princess lives?” or “Why is he bothering the princess?”, and if these questions are passed to a large language model without any context (or without the context of the Snow White story), the large language model will have difficulty understanding the child’s intention, and as a result, it will output a response that has little relevance to the Snow White storybook.

[0107]

[0108] To solve such a problem, the present disclosure proposes a method of extracting and utilizing the content related to a user's voice when the user's voice (e.g., a user's question, etc.) is input.

[0109] At this time, in order to specify the content related to the user's voice, a code that can be identified by a camera or the like is printed on the medium together with the content, and a code that is recognized by an electronic device used by the user is used.

[0110] Accordingly, according to the method provided by the present disclosure, when a user simply converses with a generative and conversational AI model such as GPT, the user can utilize not only the user's voice, but also information about content printed in a medium of interest to the user. Utilizing content printed in a medium of interest to the user in a conversation with the AI ​​model helps the AI ​​model form a basic 'context' for conversing with the user. In other words, the core concept of the method provided by the present disclosure for solving the conventional problem is to additionally generate information about the 'context' in the conversation between the user and the AI ​​model and utilize this contextual information in the conversation.

[0111] Below, a specific system for implementing the method provided by the present disclosure is described, and various embodiments of the method provided by the present disclosure are further described. In particular, specific methods for generating contextual information and transmitting this contextual information to an artificial intelligence model will be clearly understood through the following description.

[0112]

[0113] 2. Interaction System

[0114] Below, the configuration and operation method of the interaction system are described with reference to Fig. 1.

[0115] FIG. 1 is a diagram illustrating an interaction system (10) according to one embodiment.

[0116] Referring to FIG. 1, the interaction system (100) may include an electronic device (1000), a main server (2000), and an artificial intelligence server (3000).

[0117] When using a medium, a user can interact with the medium using an interaction system (100).

[0118] Here, a medium refers to a physical medium on which information or content is recorded. For example, the information or content recorded on a medium may be a story, description, information, questions, activities, or instructions. Furthermore, a physical medium may not only be a book, paper, or film, but also a portion of the surface of an object, a portion of the surface of a sculpture, a portion of the surface of a piece of furniture, a portion of the surface of an electronic device, a portion of the surface of a building, etc. In this case, the method by which information is recorded on a medium does not refer to an "electronic" method; rather, it refers to a method by which a user can visually perceive the information recorded on the medium. For example, information or content such as images or text may be printed on a medium. Alternatively, a medium may be a panel, such as an LCD, OLED, or LED, or a display including a panel, and information or content such as images or text may be electronically output on a medium.

[0119] Examples of a user's use of Medium include the user reading a story or information recorded on Medium, or the user performing an activity or instruction recorded on Medium.

[0120] A user can interact with the medium using the interaction system (100).

[0121] Typically, interaction refers to a conversation between the interaction system (100) and the medium. Specifically, a user can ask questions and receive answers through the electronic device (1000) while using the medium. For example, the interaction may occur when a user asks a question or expresses an opinion using their voice while reading a book, and the electronic device (1000) outputs a response to the question or opinion in voice. In another example, the electronic device (1000) may output activities or instructions printed on a card in voice, and the user may interact by performing the activity or responding to the instructions.

[0122] The process by which a user interacts with a medium using an electronic device (1000) can be performed as follows.

[0123] First, the electronic device (1000) can acquire image data regarding the medium. The image data refers to an image obtained by the electronic device (1000) by photographing a portion of the medium through a user's operation. For example, the electronic device (1000) can acquire image data by photographing a portion of the medium while the user places the electronic device (1000) in contact with or in close proximity to a portion of the medium. Information regarding the medium may be encoded in the image data. The electronic device (1000) can transmit the image data to the main server (2000).

[0124] The electronic device (1000) can obtain the user's voice data. The electronic device (1000) can record the user's voice to obtain the user's voice data. The electronic device (1000) can transmit the user's voice data to the main server (2000).

[0125] An electronic device (1000) can transmit an electronic device ID to a main server (2000). The electronic device ID is information for identifying the electronic device (1000) and can be used to generate session information described below.

[0126] The main server (2000) can generate a prompt using image data and voice data. Specifically, the main server (2000) can analyze image data to obtain context information and generate a prompt using the context information and voice data.

[0127] The main server (2000) can transmit a prompt to the artificial intelligence server (3000). The artificial intelligence server (3000) includes a large language model (LLM) and can receive a prompt and output a response. The artificial intelligence server (3000) transmits the response to the main server (2000), and the main server can generate response voice data for the response and transmit it to the electronic device (1000). The electronic device (1000) can output the received response voice data.

[0128] When a user interacts with a medium using the interaction system (100), the aforementioned contextual information can be used to specify the context of the interaction. For example, if a user expresses a question or opinion while reading (or pointing to) a portion of a book, the content of the book or information about that portion of the book can serve as the context for the interaction. In this case, information about the portion the user read (or pointed to) while expressing the question or opinion is acquired as contextual information, and the acquired contextual information can be used to generate a prompt. In this case, the response generated by the artificial intelligence server (3000) can also take contextual information into account, and the response voice data output through the electronic device (1000) includes content corresponding to the intent of the user's question or expressed opinion.

[0129] The context information may include at least one of medium context information and artificial context information.

[0130] Medium context information refers to information about a medium that constitutes the context of the medium. For example, if the medium is a book containing a specific story, the context of the medium refers to the specific story, and the medium context information may refer to information about a portion of the specific story, such as words, sentences, paragraphs, or images that constitute the specific story. The medium context information may refer to information about a portion of the medium with which a user interacts. For example, the medium context information may refer to information about a word, sentence, paragraph, or image that a user points to or touches using an electronic device (1000) within the context of the medium. The method by which the medium context information is acquired will be described later.

[0131] Artificial contextual information refers to information about a medium that does not constitute the context of the medium. For example, artificial contextual information may include instructions, questions, or activities related to the medium. More specifically, if the medium is a book containing a specific story, the context of the medium may refer to the specific story, and the contextual information of the medium may refer to questions, instructions, or sentences that encourage specific actions related to the story.

[0132] Below, each component of the interaction system (100) is described in detail with reference to FIGS. 2 to 5.

[0133] FIG. 2 is a drawing showing the configuration of an electronic device (1000) according to one embodiment.

[0134] FIG. 4 is a diagram showing the configuration of a main server (2000) according to one embodiment.

[0135] FIG. 5 is a diagram showing the configuration of an artificial intelligence server (3000) according to one embodiment.

[0136] Referring to FIG. 2, the electronic device (1000) may include a sensing unit (1100), an electronic device memory (1200), an electronic device input unit (1300), an electronic device output unit (1400), an electronic device communication unit (1500), and an electronic device control unit (1600).

[0137] The electronic device (1000) can capture content printed on a medium through a sensing unit (1100). For example, the sensing unit (1100) includes an image sensor such as a camera, and according to a user's operation, the electronic device (1000) can capture at least a portion of the medium to obtain an image.

[0138] An image acquired through the sensing unit (1100) may include a code. Specifically, the code may be printed on a medium according to preset rules, and an image acquired by photographing the medium may include the code. In this case, the ink used when printing the code on the medium may be different from the ink used when printing content on the medium. For example, the code may be printed on the medium using ink that absorbs infrared rays, and the sensing unit (1100) of the electronic device (1000) may include an infrared image sensor, so that the image captured by the sensing unit (1100) may include a code rather than content printed on the medium (or may include both content and a code).

[0139] As described below, the electronic device (1000) can capture a code image by photographing a medium on which a code is printed, and the electronic device (1000) or the main server (2000) can analyze the acquired code image to obtain reference information indicating information about the medium or information about the medium.

[0140] Below, the code printed on the medium is described in detail with reference to Fig. 3.

[0141] FIG. 3 is a diagram illustrating a code printed on a medium according to one embodiment. FIG. 3 (a) is a diagram illustrating unit cells constituting the code. FIG. 3 (b) is a diagram illustrating a method for encoding information on an information segment.

[0142] Referring to Figure 3, the code can be implemented in a manner in which a single unit cell is repeatedly printed in two dimensions to fit a medium size. Each unit cell includes a plurality of line segments (or dots) arranged according to a specific rule, and the specific rule can be determined based on the information encoded in the unit cell.

[0143] Referring to (a) of Fig. 3, the line segments included in the unit cell are divided into a reference line segment and an information line segment, and the reference line segment is composed of a horizontal reference line segment, a vertical reference line segment, and an intersection reference line segment.

[0144] When an electronic device (1000) or a main server (2000) analyzes a code image, the area of ​​a unit cell can be divided through a reference line segment, and information encoded in the unit cell can be obtained through an information line segment.

[0145] Meanwhile, within a unit cell, virtual lines can be defined, and the virtual lines are composed of a plurality of horizontal virtual lines extending in the horizontal direction and a plurality of vertical virtual lines extending in the vertical direction. Virtual lines generally refer to lines that are not printed on the medium, but this is not necessarily the case and may be printed on the medium.

[0146] Multiple horizontal and vertical virtual lines intersect to form multiple intersections. These intersections are hereinafter referred to as "virtual reference points." Virtual reference points are generally not printed on the medium, but this is not necessarily the case and may be printed on the medium.

[0147] Referring to FIG. 3, one unit cell is illustrated as having a total of seven horizontal virtual lines and seven vertical virtual lines, and among them, horizontal reference line segments are arranged on the horizontal virtual line arranged at the top, and the vertical reference line segments are arranged on the vertical virtual line arranged at the leftmost.

[0148] Depending on how the information segments are arranged, a specific number of binary digits can be expressed. For convenience, the following description assumes that the information segments can express four binary digits (i.e., 2 bits).

[0149] For the sake of clarity, assume that the aforementioned virtual reference point is the origin and that an orthogonal coordinate system is defined in the directions of horizontal reference lines and vertical reference lines.

[0150] When the information to be encoded through the information segment is '00', the information segment can be arranged so that one end of the information segment (the end located closer to the virtual reference point) is located in the third quadrant of the rectangular coordinate system, and the other end (the end located farther from the virtual reference point) is located in the first quadrant of the rectangular coordinate system (see the leftmost end in (b) of Fig. 3).

[0151] If the information to be encoded through the information segment is '01', the information segment can be arranged so that one end of the information segment (the end located closer to the virtual reference point) is located in the fourth quadrant of the rectangular coordinate system, and the other end (the end located farther from the virtual reference point) is located in the second quadrant of the rectangular coordinate system (see the second from the left in (b) of Fig. 3).

[0152] In addition, when the information to be encoded through the information line segment is '10', the information line segment can be arranged so that one end of the information line segment (the end located closer to the virtual reference point) is located in the first quadrant of the rectangular coordinate system, and the other end (the end located farther from the virtual reference point) is located in the third quadrant of the rectangular coordinate system (see the third from the left in (b) of Fig. 3).

[0153] In addition, when the information to be encoded through the information segment is '11', the information segment can be arranged so that one end of the information segment (the end located closer to the virtual reference point) is located in the second quadrant of the rectangular coordinate system, and the other end (the end located farther from the virtual reference point) is located in the fourth quadrant of the rectangular coordinate system (see the fourth from the left in (b) of Fig. 3).

[0154] If, as shown in (b) of FIG. 3, 2 bits of information are encoded in an information line segment and, as shown in (a) of FIG. 3, 36 information lines are included in one unit cell, the information that can be encoded in one unit cell is 2 bits*36=72 bits.

[0155] However, since the method of arranging information segments is to position one end of the information segment relative to the other, it can also be designed by tilting the information segment at an angle or by separating the information segment from a reference point. In this case, the information that can be encoded in the information segment is not limited to 2 bits, and can be 3 bits or more.

[0156] Meanwhile, the code can be implemented in various ways other than the aforementioned methods. For example, the code can take the form of an image encoding specific information, such as a QR code.

[0157] The sensing unit (1100) acquires an image (hereinafter, “code image”) of some of the multiple codes pre-printed on the medium. The electronic device (1000) analyzes the acquired code image to acquire information encoded in the code image.

[0158] According to some embodiments, the codes pre-printed on the medium may encode information such as the title of the medium, the page of the medium, and the X-coordinate and Y-coordinate within a page of the medium. For example, a unit of codes may encode (book title, page, X-coordinate, Y-coordinate). In this case, when a user indicates a specific location of the medium through the electronic device (1000), the electronic device (1000) can recognize the specific location of the medium indicated by the user.

[0159] In some other embodiments, codes pre-printed on the medium may encode predetermined area identification information. In this case, a user can use the electronic device (1000) to identify the type of content (e.g., text, images, or icons) printed on a specific area of ​​the medium.

[0160] According to some embodiments, the electronic device (1000) can transmit information (hereinafter, code data) obtained through analysis of a code image to the main server (2000), and the main server (2000) can use the received code data to identify the text or image (i.e., a part of the printed content) printed at a specific location of the medium indicated by the user.

[0161] According to some other embodiments, the electronic device (1000) can transmit the acquired code image to the main server (2000), in which case the main server (2000) can analyze the code image to acquire the aforementioned code data, and as previously described, can specify what text or image (i.e., a part of the printed content) was printed at a specific location of the medium indicated by the user.

[0162] Referring back to FIG. 2, the electronic device memory (1200) can store information processed by the electronic device (1000), programs executed, etc. For example, the electronic device memory (1200) can store image data acquired by the sensing unit (1100). As another example, the electronic device memory (1200) can store code data acquired by analyzing image data. As yet another example, the electronic device memory (1200) can store data regarding a medium acquired using code data.

[0163] The electronic device memory (1200) can be implemented in hardware in the form of various storage devices such as ROM, RAM, EPROM, flash drive, or hard drive.

[0164] The electronic device input unit (1300) may include a microphone. The electronic device input unit (1300) may record the user's voice to generate a voice signal. The voice signal may be stored as voice data in the electronic device memory (1200).

[0165] The electronic device input unit (1300) may include a trigger button. When a user manipulates the trigger button (e.g., pressing or touching it), a microphone included in the electronic device input unit (1300) may be activated, thereby initiating voice recording. The user may press the trigger button and ask a question or say something, and the electronic device input unit (1300) may record the user's voice and generate a voice signal.

[0166] The electronic device input unit (1300) may be, but is not limited to, a keyboard, a button, a mouse, a microphone, a camera, a sensor, a touch screen, or a combination thereof.

[0167] The electronic device output unit (1400) may include a speaker. The electronic device output unit (1400) may output response voice data obtained from the main server (2000).

[0168] The electronic device output unit (1400) may be, but is not limited to, a display, a speaker, an indicator, or a combination thereof.

[0169] The electronic device communication unit (1500) can perform data communication between the electronic device (1000) and an external device. For example, the electronic device communication unit (1500) can transmit a code image (or code data, context information, etc.) to the main server (2000). As another example, the electronic device communication unit (1500) can receive response voice data from the main server (2000).

[0170] The electronic device communication unit (1500) may be, for example, a wired / wireless LAN (Local Area Network) module, a WAN module, an Ethernet module, a Bluetooth module, a Zigbee module, a USB (Universal Serial Bus) module, an IEEE 1394 module, a Wi-Fi module, a mobile communication module, a satellite communication module, or a combination thereof, but is not limited thereto.

[0171] The electronic device control unit (1600) can control the components of the electronic device (1000) or execute a program stored in the electronic device memory (1200).

[0172] The electronic device control unit (1600) may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), a state machine, an application specific integrated circuit (ASIC), a radio-frequency integrated circuit (RFIC), or a combination thereof, but is not limited thereto.

[0173] Hereinafter, for convenience of explanation, it is described that the electronic device (1000) is a device having a pen shape, and when a user contacts or places the end of the electronic device (1000) adjacent to a medium on which a code is printed, the sensing unit (1100) captures the medium to obtain a code image, the electronic device (1000) analyzes the code image to obtain code data, and the electronic device (1000) transmits the code data to the main server (2000). However, the technical idea of ​​the present disclosure is not limited thereto. For example, the electronic device (1000) may be a smartphone or a tablet, and in this case, an application for performing the functions of the electronic device (1000) described above (e.g., capturing a code image, obtaining, analyzing, and transmitting code data, etc.) may be stored in the smartphone or tablet, etc.

[0174] The main server (2000) may include a main server memory (2100), a main server communication unit (2200), and a main server control unit (2300).

[0175] The main server memory (2100) can store information processed by the main server (2000), programs executed, etc. Referring to FIG. 4, at least a STT (Speech to Text) model (2110), a TTS (Text to Speech) model (2130), and a database (2150) can be stored in the main server memory (2100).

[0176] The STT model (2110) refers to a program that converts voice data (or voice signal) into text data (or voice text). The main server (2000) can convert the user's voice data received from the electronic device (1000) into text data using the STT model. The main server (2000) can use the converted text data to generate the prompt described below.

[0177] Meanwhile, the main server (2000) may transmit voice data to an external server, where the voice data may be converted into text data and provided to the main server (2000). Alternatively, the main server (2000) may generate a prompt using voice data without converting the voice data into text data. In this case, the main server (2000) may not include an STT model (2110).

[0178] The TTS model (2130) refers to a program that converts text data into voice data. The main server (2000) can convert response text data obtained from the artificial intelligence server (3000) into response voice data. The main server (2000) can transmit the response voice data to the electronic device (1000).

[0179] Meanwhile, the main server (2000) may transmit response text data to an external server, and the response text data may be converted into response voice data by the external server and provided to the main server (2000). Alternatively, the main server (2000) may receive response voice data from the artificial intelligence server (3000). In this case, the main server (2000) may not include a TTS model (2130).

[0180] The database (2150) stores information about various contents printed on multiple media. For example, the database (2150) may store all text printed on multiple media and / or images printed on multiple media by mapping them to each medium.

[0181] The database (2150) may further store information about the pages on which text and images are printed for each medium. For example, the database (2150) may store necessary information to distinguish between the content printed on the first page and the content printed on the second page among the entire contents of a specific medium.

[0182] The database (2150) may further store, for each page, information regarding the location within the page where text and images are printed. For example, the database (2150) may store necessary information to distinguish between content printed at a first location (first coordinate) and content printed at a second location (second coordinate) among content printed on a specific page.

[0183] Meanwhile, the database (2150) may store user information (e.g., the user's unique identification number, name, age, gender, personality, or family relationships). User information may be recorded and stored by an administrator or guardian. User information may be used to generate prompts, as described below. By utilizing user information in generating prompts, customized responses to the prompts may be generated. Furthermore, user information may also be used to generate session information, as described below.

[0184] The main server memory (2100) can be implemented in hardware in the form of various storage devices such as ROM, RAM, EPROM, flash drive, or hard drive.

[0185] The main server communication unit (2200) can perform data communication between the main server (2000) and an external device. For example, the main server communication unit (2200) can receive a code image (or code data, context information, etc.) from the electronic device (1000). As another example, the main server communication unit (2200) can receive response data (response text data or response voice data) from the artificial intelligence server (3000). As yet another example, the main server communication unit (2200) can transmit response voice data to the electronic device (1000).

[0186] The main server communication unit (2200) may be, for example, a wired / wireless LAN (Local Area Network) module, a WAN module, an Ethernet module, a Bluetooth module, a Zigbee module, a USB (Universal Serial Bus) module, an IEEE 1394 module, a Wi-Fi module, a mobile communication module, a satellite communication module, or a combination thereof, but is not limited thereto.

[0187] The main server control unit (2300) can control the configurations of the main server (2000) or execute a program stored in the main server memory (2100).

[0188] The main server control unit (2300) may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), a state machine, an application specific integrated circuit (ASIC), a radio-frequency integrated circuit (RFIC), or a combination thereof, but is not limited thereto.

[0189] Although not shown in FIG. 4, the main server (2000) may include a main server input unit and a main server output unit. The main server input unit may receive input from an administrator who manages the main server (2000), and may be, but is not limited to, a keyboard, a button, a mouse, a microphone, a camera, a sensor, a touchscreen, or a combination thereof. The main server output unit may output information processed in the main server (2000) to the administrator, and may be, but is not limited to, a display, a speaker, an indicator, or a combination thereof.

[0190] The artificial intelligence server (3000) may include an artificial intelligence server memory (3100), an artificial intelligence server communication unit (3200), and an artificial intelligence server control unit (3300).

[0191] The artificial intelligence server memory (3100) can store information processed by the artificial intelligence server (3000), programs executed, etc. The artificial intelligence server (3000) can generate response data to a user's questions or speech, and for this purpose, a generative model can be stored in the artificial intelligence server memory (3100).

[0192] Here, the generative model can be a natural language processing model (NLP) or a large language model (LLM) trained with a large corpus of data. The generative model can be created by training a Generative Pre-trained Transformer (GPT) model that utilizes the structure of a transformer model. Examples of generative models include ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot, and these models can be stored in the AI ​​server memory (3100).

[0193] Meanwhile, the generative model may vary depending on the type of data to be generated by the artificial intelligence server (3000). For example, if the artificial intelligence server (3000) receives a voice prompt from the main server (2000) and generates voice response data, the generative model may be a voice-generating model, such as WaveNet. For another example, if the main server (2000) provides image data to the artificial intelligence server (3000) and the artificial intelligence server (3000) generates image response data, the generative model may be a model to generate images, such as a Generative Adversarial Network (GAN) or a Diffusion Model. In other words, the generative model may generate response data of the same type as the input data type (e.g., text, sound, image, or video), such as by inputting text data and generating text response data, or by inputting voice data and generating voice response data.

[0194] Alternatively, a generative model may generate response data of a different type from the input data type, such as by inputting text data and generating voice response data, or by inputting voice data and generating text response data.

[0195] Alternatively, a generative model can accept multiple data types at once and generate response data of one of the data types. For example, a generative model can accept data containing text and sound and output text response data or voice response data.

[0196] The artificial intelligence server memory (3100) can be implemented in hardware in the form of various storage devices such as ROM, RAM, EPROM, flash drive, or hard drive.

[0197] The artificial intelligence server communication unit (3200) can perform data communication between the artificial intelligence server (3000) and an external device. For example, the artificial intelligence server communication unit (3200) can receive a prompt from the main server (2000). As another example, the artificial intelligence server communication unit (3200) can transmit response data (response text data or response voice data) to the main server (2000).

[0198] The artificial intelligence server communication unit (3200) may be, for example, a wired / wireless LAN (Local Area Network) module, a WAN module, an Ethernet module, a Bluetooth module, a Zigbee module, a USB (Universal Serial Bus) module, an IEEE 1394 module, a Wi-Fi module, a mobile communication module, a satellite communication module, or a combination thereof, but is not limited thereto.

[0199] The artificial intelligence server control unit (3300) can control the configurations of the artificial intelligence server (3000) or execute a program stored in the artificial intelligence server memory (3100).

[0200] The artificial intelligence server control unit (3300) may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), a state machine, an application specific integrated circuit (ASIC), a radio-frequency integrated circuit (RFIC), or a combination thereof, but is not limited thereto.

[0201] Although not illustrated in FIG. 5, the artificial intelligence server (3000) may include an artificial intelligence server input unit and an artificial intelligence server output unit. The artificial intelligence server input unit may receive input from an administrator who manages the artificial intelligence server (3000), and may be, but is not limited to, a keyboard, a button, a mouse, a microphone, a camera, a sensor, a touchscreen, or a combination thereof. The artificial intelligence server output unit may output information processed in the artificial intelligence server (3000) to the administrator, and may be, but is not limited to, a display, a speaker, an indicator, or a combination thereof.

[0202] Below, another embodiment of the above-described interaction system (100) is described with reference to FIGS. 6 and 7.

[0203] FIG. 6 is a drawing showing a block diagram of an electronic pen service system using an interactive artificial intelligence service according to one embodiment.

[0204] Referring to FIG. 6, the electronic pen service system (101) using an interactive artificial intelligence service includes an electronic pen (110), a first server (120), and an interactive artificial intelligence server (130). Here, the electronic pen service system (101) corresponds to the interaction system (100) described through FIGS. 1 to 5, the electronic pen (110) corresponds to an electronic device (1000), the first server (120) corresponds to a main server (2000), and the interactive artificial intelligence server (130) corresponds to an artificial intelligence server (3000).

[0205] The electronic pen (110) includes a first communication unit (111), a code recognition unit (112), a first control unit (113), a TTS unit (114), and a speaker unit (115).

[0206] The first communication unit (111) can communicate with the first server (120) using a wired or wireless interface. The first communication unit (111) has substantially the same configuration as the electronic device communication unit (1500) described in FIG. 2.

[0207] The code recognition unit (112) recognizes a predetermined code from a printed matter on which a predetermined code is printed. The predetermined code includes location information, i.e., coordinate information, on the printed matter. The code recognition unit (112) includes a camera, and when the camera approaches the printed matter on which the predetermined code is printed within a predetermined distance, the camera captures the predetermined code within an area that the camera can recognize. Thereafter, the code recognition unit (112) transmits the image captured by the camera to the first control unit (113). The code recognition unit (112) has substantially the same configuration as the sensing unit (1100) described in FIG. 2.

[0208] The first control unit (113) reads coordinate information from a predetermined code of an image received from the code recognition unit (112) (coordinate information corresponding to a predetermined code of the image is retrieved from a pre-stored database, etc.). The first control unit (113) has substantially the same configuration as the electronic device control unit (1600) described in FIG. 2.

[0209] After that, the first control unit (113) transmits the read coordinate information to the first server (120) through the first communication unit (111).

[0210] As will be described later, the first communication unit (111) receives the result value for the coordinate information transmitted from the first server (120).

[0211] Meanwhile, the first control unit (1130) can also read area identification information from a predetermined code of an image received from the code recognition unit (112).

[0212] Region identification information is information used to identify a region within a medium. For example, a medium may include at least one page, each of which may contain printed content. Region identification information may be set for a specific region within the page. In this case, the corresponding region identification information may be encoded into a code image obtained by photographing the specific region. Here, the specific region may be an area printed with a specific shape, an area printed with text, or an area printed with an image.

[0213] Area identification information can be used in place of coordinate information. For example, the first control unit (113) can transmit the area identification information read out through the first communication unit (111) to the first server (120). In this case, the first server (120) can transmit a prompt corresponding to the area identification information to the interactive AI server (130), and the interactive AI server (130) can generate a result value for the received prompt and transmit it to the first server (120).

[0214] The TTS unit (114) is a component equipped with a TTS (Text To Speech) function and performs the function of converting text into voice. The TTS unit (114) converts the result value received by the first communication unit (111) into voice, and the converted voice is output to the user of the electronic pen (110) through the speaker unit (115). The TTS unit (114) has a configuration corresponding to the TTS model (2130) described in FIG. 4. The TTS unit (114) may be included in the first server (120) rather than the electronic pen (110). In this case, the electronic pen (110) can receive the voice converted from the result value from the first server (120) and output the converted voice through the speaker unit (115).

[0215] In an additional embodiment, the electronic pen (110) may further include a microphone (116) and an STT unit (117).

[0216] The electronic pen (110) receives the user's voice input through the microphone (116).

[0217] The STT unit (117) is a component equipped with STT (Speech To Text) capability and performs the function of converting voice into text. The STT unit (117) converts voice information received by the microphone (116) into text information, and the converted text information is transmitted to the first server (120). Thereafter, the first communication unit (111) receives a result value for the transmitted text as text from the first server (120), and the text received by the first communication unit (111) is converted into voice and output to the user of the electronic pen (110) through the speaker unit (115). The STT unit (117) is a configuration corresponding to the STT model (2110) described in FIG. 4. The STT unit (117) may also be included in the first server (120) rather than the electronic pen (110). In this case, the electronic pen (110) transmits a voice signal to the first server (120), and the first server (120) can convert the voice signal into text information using the STT unit (117).

[0218] Additionally, in another embodiment, the electronic pen (110) may further include a storage unit (118).

[0219] The storage unit (118) contains content specified in the coordinate information. The storage unit (118) may be included in the first server (120) rather than the electronic pen (110).

[0220] The first control unit (113) can read coordinate information from a predetermined code captured from an image received from a code recognition unit (112), extract content stored in the read coordinate information, and output the extracted content to the user through the speaker unit (115).

[0221] In one embodiment of the present invention, the ID information of the electronic pen (110) user and the ID information of the electronic pen (110) are included together with the coordinate information and character information transmitted to the first server (120).

[0222] The first server (120) includes a second communication unit (121), a first database (122), and a second control unit (123).

[0223] When the second communication unit (121) receives coordinate information from the electronic pen (110), the second control unit (123) stores it in the first database (122). The second communication unit (121) has a configuration substantially identical to the main server communication unit (2200) described in FIG. 4.

[0224] The first database (122) may have a prompt corresponding to coordinate information pre-stored. The prompt refers to input data that issues a command or instruction to the interactive artificial intelligence server (130). The first database (122) has a configuration corresponding to the database (2150) described in FIG. 4. The first database (122) may store information regarding a medium. For example, the first database (122) may store at least one of a medium ID for each of a plurality of mediums, a medium title, all text printed on a medium, text printed on a specific page of a medium, and text printed at a specific coordinate of a medium. In addition, the first database (122) may store a predetermined prompt and / or prompt form.

[0225] The second control unit (123) extracts a prompt corresponding to the received coordinate information and transmits the extracted prompt to the interactive artificial intelligence server (130). The second control unit (123) has substantially the same configuration as the main server control unit (2300) described in FIG. 4. The second control unit (133) can load a prompt from the first database (122). Alternatively, the second control unit (123) can load information regarding a prompt form and medium from the first database (122) to generate a prompt.

[0226] When the second communication unit (121) receives character information from the electronic pen (110), the second control unit (123) stores the received character information in the first database (122) and transmits it to the interactive artificial intelligence server (130).

[0227] A conversational artificial intelligence server (130) is a server with an embedded language processing model for utilizing artificial intelligence, and refers to a server that provides a conversational artificial intelligence service.

[0228] Examples of conversational AI services include OpenAI's ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot.

[0229] The interactive artificial intelligence server (130) extracts a result value using a language processing model by inputting a prompt or character information from the first server (120) and transmits the extracted result value to the first server (120). The result value is composed of string data.

[0230] The second communication unit (121) receives the result value from the interactive artificial intelligence server (130) and transmits the result value to the electronic pen (110).

[0231] A session refers to a series of processes of transmitting text information or a prompt to an interactive artificial intelligence server (130) and receiving a corresponding result value. When one session is terminated, the second control unit (123) generates session information by linking the received result value with the session ID, the text information or prompt transmitted to the interactive artificial intelligence server (130), and the ID of the user who transmitted the code information that became the basis of the text information or prompt, and the ID of the electronic pen (110), and stores this in the first database (122). The second control unit (123) stores this information in the form of a JSON file, and may further store time information for the session ID by adding it thereto.

[0232] The session information may include text converted from a user's voice signal, a prompt generated by the first server (120), and a result value generated by the interactive artificial intelligence server (130). The second control unit (123) extracts at least one session information stored in the first database (122) based on a user ID or an electronic pen ID for a predetermined period of time. The second control unit (123) transmits predetermined session prompt information and at least one extracted session information to the interactive artificial intelligence server (130) through the second communication unit (121). The predetermined session prompt may include content related to a summary of at least one session information. In addition, the predetermined session prompt may include content related to an analysis of at least one session information.

[0233] The conversational artificial intelligence server (130) extracts a result value using a language processing model using the session prompt received from the first server (120) as an input value, and transmits the extracted result value to the first server (120).

[0234] More specifically, the first server (120) may generate a session prompt using session information and transmit the generated session prompt to the interactive artificial intelligence server (130). The interactive artificial intelligence server (130) may receive the session prompt and output at least one of a conversation log, a conversation summary, a pronunciation evaluation, keywords frequently used by the user, the user's medium usage time, the user's medium type-specific usage time, and the user's areas of interest. The interactive artificial intelligence server (130) may provide the output analysis information to the first server (120).

[0235] The first server (120) stores the result value (or analysis information) for the received session prompt in the first database (122). When a user requests, the first server (120) provides the user with at least one session information and a result value based on the session prompt.

[0236] Specifically, the first server (120) can transmit the analysis information to a terminal (e.g., a laptop, desktop, or smartphone) of a guardian (e.g., the child's parent if the user is a child). The guardian can check the analysis information through his / her terminal.

[0237] Meanwhile, the first server (120) can generate analysis information using session information without using the interactive artificial intelligence server (130). The analysis information may include at least one of a conversation log representing a conversation between a user and an electronic pen (110), a conversation summary summarizing a conversation between a user and an electronic pen (110), pronunciation evaluation information representing an evaluation of the user's pronunciation, keywords frequently used by the user, the user's medium usage time, the user's medium type usage time, and the user's area of ​​interest. Fig. 7 is a drawing showing a block diagram of an electronic pen service system using an interactive artificial intelligence service according to another embodiment.

[0238] Referring to FIG. 7, an electronic pen service system (101) using an interactive artificial intelligence service includes an electronic pen (110), a first server (120), an interactive artificial intelligence server (130), and a second server (140). Here, the electronic pen service system (101) corresponds to the interaction system (100), the electronic pen (110) corresponds to the electronic device (1000), the first server (120) corresponds to the main server (2000), and the interactive artificial intelligence server (130) corresponds to the artificial intelligence server (3000). The difference between FIG. 6 and FIG. 7 is that in FIG. 7, a second server (140) is additionally provided. The second server (140) can perform the aforementioned session information acquisition, session prompt generation, and analysis information acquisition.

[0239] The electronic pen (110) includes a first communication unit (111), a code recognition unit (112), a first control unit (113), a TTS unit (114), and a speaker unit (115).

[0240] The first communication unit (111) can communicate with the first server (120) using a wired or wireless interface. The first communication unit (111) has substantially the same configuration as the electronic device communication unit (1500) described in FIG. 2.

[0241] The code recognition unit (112) recognizes a predetermined code from a printed matter on which a predetermined code is printed. The predetermined code includes location information, i.e., coordinate information, on the printed matter. The code recognition unit (112) includes a camera, and when the camera approaches the printed matter on which the predetermined code is printed within a predetermined distance, the camera captures the predetermined code within an area that the camera can recognize. Thereafter, the code recognition unit (112) transmits the image captured by the camera to the first control unit (113). The code recognition unit (112) has substantially the same configuration as the sensing unit (1100) described in FIG. 2.

[0242] The first control unit (113) reads coordinate information from a predetermined code of an image received from the code recognition unit (112). The first control unit (113) has substantially the same configuration as the electronic device control unit (1600) described in FIG. 2.

[0243] After that, the first control unit (113) transmits the read coordinate information to the first server (120) through the first communication unit (111).

[0244] As will be described later, the first communication unit (111) receives the result value for the coordinate information transmitted from the first server (120).

[0245] Meanwhile, the aforementioned area identification information may be used in place of coordinate information. For example, the first control unit (113) may transmit the area identification information read out via the first communication unit (111) to the first server (120). In this case, the first server (120) may transmit a prompt corresponding to the area identification information to the interactive AI server (130), and the interactive AI server (130) may generate a result value for the received prompt and transmit it to the first server (120).

[0246] The TTS unit (114) is a component equipped with TTS (Text To Speech) capability and performs the function of converting text into voice. The TTS unit (114) converts the result value received by the first communication unit (111) into voice, and the converted voice is output to the user of the electronic pen (110) through the speaker unit (115). The TTS unit (114) has a configuration corresponding to the TTS model (2130) described in FIG. 4. The TTS unit (114) may be included in the first server (120) rather than the electronic pen (110). In this case, the electronic pen (110) can receive the voice converted from the result value from the first server (120) and output the converted voice through the speaker unit (115).

[0247] In an additional embodiment, the electronic pen (110) may further include a microphone (116) and an STT unit (117).

[0248] The electronic pen (110) receives the user's voice input through the microphone (116).

[0249] The STT unit (117) is a component equipped with STT (Speech To Text) capability and performs the function of converting voice into text. The STT unit (117) converts voice information received by the microphone (116) into text information, and the converted text information is transmitted to the first server (120). Thereafter, the first communication unit (111) receives a result value for the transmitted text as text from the first server (120), and the text received by the first communication unit (111) is converted into voice and output to the user of the electronic pen (110) through the speaker unit (115). The STT unit (117) is a configuration corresponding to the STT model (2110) described in FIG. 4. The STT unit (117) may also be included in the first server (120) rather than the electronic pen (110). In this case, the electronic pen (110) transmits a voice signal to the first server (120), and the first server (120) can convert the voice signal into text information using the STT unit (117).

[0250] Additionally, in another embodiment, the electronic pen (110) may further include a storage unit (118).

[0251] The storage unit (118) contains content specified in the coordinate information. The storage unit (118) may be included in the first server (120) rather than the electronic pen (110).

[0252] The first control unit (113) can read coordinate information from a predetermined code captured from an image received from a code recognition unit (112), extract content stored in the read coordinate information, and output the extracted content to the user through the speaker unit (115).

[0253] In one embodiment of the present invention, the ID information of the electronic pen (110) user and the ID information of the electronic pen (110) are included together with the coordinate information and character information transmitted to the first server (120).

[0254] The first server (120) includes a second communication unit (121), a first database (122), and a second control unit (123).

[0255] When the second communication unit (121) receives coordinate information from the electronic pen (110), the second control unit (123) stores it in the first database (122). The second communication unit (121) has a configuration substantially identical to the main server communication unit (2200) described in FIG. 4.

[0256] The first database (122) may have a prompt corresponding to coordinate information pre-stored. The prompt refers to input data that issues a command or instruction to the interactive artificial intelligence server (130). The first database (122) has a configuration corresponding to the database (2150) described in FIG. 4. The first database (122) may store information regarding a medium. For example, the first database (122) may store at least one of a medium ID for each of a plurality of mediums, a medium title, all text printed on a medium, text printed on a specific page of a medium, and text printed at a specific coordinate of a medium. In addition, the database (2150) may store a predetermined prompt and / or prompt form.

[0257] The second control unit (123) extracts a prompt corresponding to the received coordinate information and transmits the extracted prompt to the interactive artificial intelligence server (130). The second control unit (123) has substantially the same configuration as the main server control unit (2300) described in FIG. 4. The second control unit (133) can load a prompt from the first database (122). Alternatively, the second control unit (123) can load information regarding a prompt form and medium from the first database (122) to generate a prompt.

[0258] When the second communication unit (121) receives character information from the electronic pen (110), the second control unit (123) stores the received character information in the first database (122) and transmits it to the interactive artificial intelligence server (130).

[0259] A conversational artificial intelligence server (130) is a server with an embedded language processing model for utilizing artificial intelligence, and refers to a server that provides a conversational artificial intelligence service.

[0260] Examples of conversational AI services include OpenAI's ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot.

[0261] The interactive artificial intelligence server (130) extracts a result value using a language processing model by inputting a prompt or character information from the first server (120) and transmits the extracted result value to the first server (120). The result value is composed of string data.

[0262] The second communication unit (121) receives the result value from the interactive artificial intelligence server (130) and transmits the result value to the electronic pen (110).

[0263] Additionally, the second communication unit (121) transmits the result value to the second server (140).

[0264] The second server (140) includes a third communication unit (141), a second database (142), and a third control unit (143).

[0265] The third communication unit (141) receives the result value from the first server (120) and stores it in the second database (142).

[0266] A session refers to a series of processes of transmitting text information or a prompt to an interactive artificial intelligence server (130) and receiving a corresponding result value. When one session is terminated, the third control unit (143) generates session information by linking the received result value with the session ID, the text information or prompt transmitted to the interactive artificial intelligence server (130), and the ID of the user who transmitted the code information that became the basis of the text information or prompt, and the ID of the electronic pen (110), and stores this in the second database (142). The third control unit (143) stores this information in the form of a JSON file, and may further store time information for the session ID by adding it thereto.

[0267] The third control unit (143) extracts at least one session information stored in the second database (142) based on the user ID or electronic pen ID for a predetermined period of time.

[0268] The third control unit (143) transmits predetermined session prompt information and at least one piece of extracted session information to the interactive artificial intelligence server (130) via the third communication unit (141). The predetermined session prompt may include content related to a summary of at least one piece of session information. Additionally, the predetermined session prompt may include content related to an analysis of at least one piece of session information.

[0269] The conversational artificial intelligence server (130) extracts a result value using a language processing model using a session prompt received from a second server (140) as an input value, and transmits the extracted result value to the second server (140).

[0270] The second server (140) stores the result value for the received session prompt in the second database (142). When there is a user request, the second server (140) provides the user with at least one session information and the result value according to the session prompt.

[0271]

[0272] 3. Information used in the interaction system

[0273] Below, with reference to FIG. 8, information stored in a database (2150) in an interaction system (100) is described.

[0274] FIG. 8 is a diagram illustrating information stored in a database (2150) according to one embodiment. The information stored in the database (2150) can be utilized by the main server (2000), and in particular, can be utilized by the main server (2000) to generate a prompt.

[0275] Referring to FIG. 8, the database (2150) may store information about a medium (information about a first medium to information about an n-th medium), prompts (a first prompt to an n-th prompt), and prompt forms (a first prompt form to an n-th prompt form).

[0276] Information about the medium can be used to generate prompts, as described below. Information about the medium can be understood as information that helps determine what medium the user uses or what areas the user is interested in.

[0277] Information about the first medium refers to information about the first medium among multiple mediums. Information about the first medium may include a first medium ID, a first medium type, a first medium name, a first medium summary, first medium text, first medium page-specific text, first medium area-specific text, first medium page-specific images, first medium area-specific images, first medium questions, and first medium instructions.

[0278] The first medium ID is information used to identify the first medium among the mediums. The first medium may be assigned a first unique identification number, and the first unique identification number may be the first medium ID.

[0279] The first type of medium refers to the form of the medium, which can be books, cards, and maps.

[0280] The First Medium Name is the information that identifies the First Medium. If the First Medium is a book, the First Medium Name is the book's title. If the First Medium is a card, the First Medium Name is the card's name.

[0281] The First Medium Summary is a summary of the content of the First Medium. If the First Medium is a book, the First Medium Summary is a summary of the story printed in the book. If the First Medium is a card, the First Medium Summary is a summary of the content printed on the card, either omitted or omitted.

[0282] A first medium text is information about the content of a first medium. If the first medium is a book, then a first medium text refers to the text of all or part of the text (or story) printed in the first medium.

[0283] The First Medium Page-by-Page Text is the printed text on each page of a book, if the First Medium is a book. If the First Medium is a card, the First Medium Page-by-Page Text is the image printed on the front or back of the card.

[0284] The text for the first medium region is text printed in a specific area on a specific page (or side) of the first medium. If the first medium is a book, the text for the first medium region is text (paragraphs or sentences) printed in a specific area on a specific page of the first medium. If the first medium is a card, the text for the first medium region is text printed in a specific area on the front or back of the card. Here, the specific area can be preset when storing information for each medium and can have various shapes, such as a polygon, circle, ellipse, or any arbitrary shape.

[0285] The First Medium Page Image is the printed image for each page of a book, if the First Medium is a book. If the First Medium is a card, the First Medium Page Image is the image printed on the front or back of the card.

[0286] A first medium region-specific image is an image printed in a specific area on a specific page (or side) of a first medium. If the first medium is a book, the first medium region-specific image is an image printed in a specific area on a specific page of the first medium. If the first medium is a card, the first medium region-specific image is an image printed in a specific area on the front or back of the card. Here, the specific area can be preset when storing information for each medium and can have various shapes, such as a polygon, circle, ellipse, or any arbitrary shape.

[0287] The text per first medium page, the text per first medium area, the image per first medium page, and the image per first medium area can be understood as medium context information. For example, as described below, when a user speaks or asks a question while touching or pointing to a specific page of the first medium or a text or image within a specific page using an electronic device (1000), the main server (2000) can obtain at least one of the text per first medium page, the text per first medium area, the image per first medium page, and the image per first medium area as medium context information, and can generate a prompt using the medium context information.

[0288] The first medium question is the question printed on the first medium.

[0289] If the first medium is a book, the first medium questionnaire may include questions about the story in the first medium. For example, the first medium questionnaire may include a quiz about the story in the first medium or a question asking about the protagonist's feelings. The first medium may have different questions printed on each page. Specifically, the first medium may have questions about the text printed on each page.

[0290] If the first medium is a card, the first medium question may include questions related to the content of the first medium. For example, if the content of the first medium is "Introduce Yourself," the first medium question may include "What is your name?" and / or "What are your hobbies?"

[0291] The First Medium Instructions are the instructions printed on the First Medium.

[0292] If the first medium is a book, the first medium instructions may include instructions regarding the story in the first medium. For example, the first medium instructions may include "expressing the user's thoughts" and "telling the user's experience" regarding the story in the first medium. Different instructions may be printed on each page of the first medium. Specifically, the first medium may print instructions regarding the text printed on each page.

[0293] If the first medium is a card, the first medium instructions may include instructions related to the content of the first medium. For example, if the content of the first medium is "future aspirations," the first medium instructions may include "Tell me about a person you want to be like," "Tell me about something you're good at," etc.

[0294] The first medium question or first medium instruction can be understood as artificial contextual information. For example, as described below, when a user uses an electronic device (1000) to touch or point to an area on the first medium where a question is printed, the main server (2000) may acquire at least one of the first medium question or first medium instruction as artificial contextual information and generate a prompt using the artificial contextual information.

[0295] Meanwhile, information about a medium may include, in addition to text, audio, images, or video. For example, the database (2150) may store audio for a summary of the medium, audio for text printed on the medium (printed text by page or region), and audio for questions or instructions printed on the medium. Furthermore, the database (2150) may store video related to the medium, video related to each page of the medium, video related to a region of the medium or a printed image, and the like. Accordingly, the contextual information described below may also include at least one of text, audio, image, or video.

[0296] A prompt refers to information input into the AI ​​server (3000). The AI ​​server (3000) can receive the prompt and generate a response. The prompt may include contextual information. The prompt stored in the database (2150) may include contextual information regarding a specific medium.

[0297] Specific examples of prompts will be described later.

[0298] A prompt form may include content for generating a prompt. Specifically, the prompt form may include sentences or paragraphs to provide context to the AI ​​server (3000). The prompt form may be modified (e.g., contextual information may be added) using additional information, such as contextual information, to become a prompt input to the AI ​​server (3000).

[0299] The main server (2000) can generate a prompt using a prompt form and text data converted from the user's voice data. Alternatively, the main server (2000) can generate a prompt using a prompt form and medium context information. Alternatively, the main server (2000) can generate a prompt using a prompt form and artificial context information. Alternatively, the main server (2000) can generate a prompt using a prompt form, artificial context information, and medium context information.

[0300] Specific examples of prompt formats will be provided later.

[0301] Meanwhile, the database (2150) may store other information in addition to the aforementioned information. Furthermore, the format in which information is stored in the database (2150) is not limited to the aforementioned format, and the manner in which information is stored (e.g., categories, relationships between pieces of information, etc.) may vary.

[0302] Below, the process of acquiring contextual information is described. As described above, contextual information refers to information about the area in which a user interacts with the medium using an electronic device (1000). Contextual information can be used to generate prompts.

[0303] When a user interacts with a medium using an electronic device (1000), the user touches or points to a specific area within the medium using the electronic device (1000). For example, when the user touches or places the electronic device (1000) in proximity to a specific area within the medium, the sensing unit (1100) of the electronic device (1000) can capture a picture of the specific area, thereby obtaining a code image.

[0304] The code image can be analyzed by an electronic device (1000) or a main server (2000), and code data can be obtained accordingly.

[0305] The main server (2000) can obtain reference information from code data. Here, the reference information refers to information for loading necessary information from the database (2150). For example, the reference information may include at least one of a medium ID, page information, and location information (coordinate information and / or area identification information).

[0306] The main server (2000) can retrieve necessary information from the database (2150) using reference information to obtain contextual information. The contextual information may include information corresponding to at least a portion of the text or images printed on the medium. The main server (2000) can generate a prompt using at least the contextual information.

[0307] Information included in the contextual information may be determined based on the area within the medium where the electronic device (1000) is in contact or positioned. In other words, when a user positions the electronic device (1000) in different areas within the medium, the contextual information ultimately acquired by the main server (2000) must be different. This means that different codes are printed in each area within the medium, and thus the code images acquired by the electronic device (1000) also differ.

[0308] Below, with reference to FIGS. 9 and 10, the manner in which areas within a medium are distinguished is described.

[0309] FIG. 9 is a diagram illustrating an aspect in which areas are distinguished in a medium according to one embodiment.

[0310] Referring to Figure 9, the medium may include a picture area, a text area, a question area, page numbers, and symbols. The codes printed in each area may have different patterns. For example, by photographing the codes printed in each area and decoding the resulting code data, different area identification information can be obtained.

[0311] The picture area is an area where a specific picture is printed. When an electronic device (1000) comes into contact with the picture area and decodes the obtained code data, at least a medium ID, page information, and area identification information indicating the picture area can be obtained. The main server (2000) can use the obtained medium ID, page information, and area identification information to obtain an image corresponding to the picture printed in the corresponding area from the database (2150).

[0312] The text area is an area where text is printed. When an electronic device (1000) comes into contact with the text area and decodes the obtained code data, at least a medium ID, page information, and area identification information indicating the text area can be obtained. The main server (2000) can use the obtained medium ID, page information, and area identification information to obtain text corresponding to the text printed in the corresponding area from the database (2150).

[0313] Meanwhile, text areas can be set for each sentence that makes up the text. In this case, multiple text areas can be set on a single page.

[0314] The question area is an area where a question is printed. When an electronic device (1000) comes into contact with the question area and decodes the obtained code data, at least a medium ID, page information, and area identification information indicating the question area can be obtained. The main server (2000) can use the obtained medium ID, page information, and area identification information to obtain text corresponding to the question printed in the corresponding area from the database (2150).

[0315] Meanwhile, the question area may be printed with at least one of a symbol, a question, and an instruction, and may be named differently depending on the printed content. For example, if instructions are printed, it may be named the instruction area, or if only symbols are printed, it may be named the symbol area.

[0316] The areas distinguished in the medium are not limited to the areas described above, and the units and rules for distinguishing areas can be determined in various ways.

[0317] FIG. 10 is a drawing showing an aspect in which areas are distinguished in a medium according to another embodiment.

[0318] Referring to Figure 10, the medium can be divided into a question area and a description area. An image can be printed in the question area, and text can be printed in the description area. The codes printed in each area can have different patterns. For example, by photographing the codes printed in each area and decoding the resulting code data, different area identification information can be obtained.

[0319] When an electronic device (1000) comes into contact with a question area and decodes the obtained code data, at least a medium ID and area identification information indicating the question area can be obtained. The main server (2000) can use the obtained medium ID and area identification information to obtain text corresponding to a question regarding the medium from the database (2150).

[0320] When an electronic device (1000) comes into contact with a description area and decodes the obtained code data, at least a medium ID and area identification information indicating the description area can be obtained. The main server (2000) can use the obtained medium ID and area identification information to obtain text corresponding to the description of the medium from the database (2150).

[0321] The areas distinguished in the medium are not limited to the areas described above, and the types of areas, the shapes of the shapes that specify the areas, and the patterns of the codes printed for the areas can be determined in various ways.

[0322]

[0323] 4. Interaction using artificial context information

[0324] Hereinafter, with reference to FIGS. 11 to 22, a method for an interaction system (100) to assist a user in using a medium by using artificial context information is described.

[0325] Medium's primary purpose is to convey specific stories or information to users, such as fairy tales, nature, animals, or science. However, beyond simply conveying specific stories or information, Medium can be designed to include additional content, such as questions, instructions, exercises, or exploration activities, to help users deeply understand the content.

[0326] In this case, without assistance from a guardian or other person, the user can only answer questions or follow instructions, making further learning or exploration difficult. Therefore, assisting the user in using Medium on behalf of their guardian can deepen their understanding and increase their interest in the stories and content printed on Medium.

[0327] Below, an interaction system (100) that assists a user in using a medium on behalf of a guardian is described, with particular emphasis on how the interaction system (100) generates prompts for interaction when additional questions or instructions are printed along with a story or information on the medium.

[0328] The following scenario is an example of a method for assisting a user's use of a medium using artificial contextual information. A child is reading a storybook without a guardian. The storybook contains not only the story but also quizzes and additional questions about the story. The child wants to use the interaction system (100) to discuss the quizzes and additional questions.

[0329] FIG. 11 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a first embodiment. Hereinafter, as described in FIG. 6, the electronic device (1000) will be referred to as an electronic pen (110), the main server (2000) as a first server (120), and the artificial intelligence server (3000) as an interactive artificial intelligence server (130). The descriptions in FIG. 6 can be equally applied to each configuration.

[0330] Referring to FIG. 11, in step 301, the electronic pen recognizes a predetermined code from a printed matter printed with the predetermined code and extracts coordinate information. The predetermined code includes location information, i.e., coordinate information, on the printed matter. The electronic pen includes a camera, and when the camera approaches the printed matter printed with the predetermined code within a predetermined distance, the camera captures the predetermined code within an area that the camera can recognize. Thereafter, the electronic pen reads the coordinate information from the predetermined code in the captured image.

[0331] In step 302, the electronic pen transmits the extracted coordinate information to the first server. Alternatively, the electronic pen transmits the captured image to the first server, which analyzes the captured image to obtain code data and may obtain at least the medium ID and coordinate information from the obtained code data. At this time, additional page information may be obtained from the code data.

[0332] The electronic pen can further transmit the ID information of the electronic pen user and the ID information of the electronic pen to the first server.

[0333] The first server associates and stores the received coordinate information, the electronic pen user's ID information, and the electronic pen's ID information.

[0334] In step 303, the first server extracts a prompt corresponding to the coordinate information.

[0335] The first server may pre-store prompts corresponding to coordinate information. A prompt is input data that issues commands or instructions to the interactive AI server. The first server may obtain a prompt corresponding to a medium ID, a prompt corresponding to coordinate information, or a prompt corresponding to both a medium ID and coordinate information from the first database.

[0336] Alternatively, the first server can retrieve information corresponding to the medium ID and coordinate information from the first database to obtain contextual information. In this case, the first server may further consider page information when obtaining contextual information.

[0337] The first server can generate a prompt using contextual information. For example, the first server can generate a prompt by adding content to the contextual information. In another example, the first server can obtain a prompt form corresponding to at least one of a medium ID, page information, and coordinate information, and generate a prompt using the contextual information and the prompt form.

[0338] In step 304, the first server transmits the extracted prompt to the conversational artificial intelligence server.

[0339] A conversational AI server is a server embedded with a language processing model for utilizing AI, providing conversational AI services. Examples of conversational AI services include OpenAI's ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot.

[0340] In step 305, the conversational artificial intelligence server uses a language processing model with the prompt received from the first server as an input value to generate and extract a result value.

[0341] The result value consists of string data.

[0342] In step 306, the conversational artificial intelligence server transmits the extracted result value to the first server.

[0343] In step 307, the first server transmits the received result value to the electronic pen.

[0344] In step 308, the electronic pen converts the received result value into voice using the text-to-speech function.

[0345] At step 309, the electronic pen outputs the converted voice to the user of the electronic pen.

[0346] In step 310, the first server creates session information.

[0347] A session refers to a series of processes involving sending prompts to an interactive AI server and receiving corresponding results. When a session ends, the first server generates session information by linking the received result values ​​to the session ID, the prompt sent to the interactive AI server, and the IDs of the user and electronic pen that sent the code information that formed the basis of the prompt. The first server then stores this information. The first server stores this information in a JSON file and can also add time information related to the session ID to further store it.

[0348] In step 311, the first server extracts at least one session information stored based on a user ID or electronic pen ID for a predetermined period of time and transmits it to the interactive artificial intelligence server along with predetermined session prompt information.

[0349] The predetermined session prompt may include content related to a summary of at least one session information. Additionally, the predetermined session prompt may include content related to an analysis of at least one session information.

[0350] In step 312, the conversational artificial intelligence server uses a language processing model to generate a result value using the session prompt received from the first server as an input value.

[0351] In step 313, the conversational artificial intelligence server transmits the generated result value to the first server.

[0352] The first server stores the result values ​​for the received session prompts. Upon user request, the first server provides the user with at least one session information and the result values ​​for the session prompts.

[0353] FIG. 12 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a second embodiment of the present invention. Hereinafter, as described in FIG. 6 , the electronic device (1000) will be referred to as an electronic pen (110), the main server (2000) as a first server (120), and the artificial intelligence server (3000) as an interactive artificial intelligence server (130). The descriptions in FIG. 6 can be equally applied to each component.

[0354] Referring to Figure 12, in step 501, the electronic pen receives the user's voice input through a microphone and converts the voice information into text information.

[0355] In step 502, the electronic pen transmits the converted character information to the first server.

[0356] The first server associates and stores the received text information, the electronic pen user's ID information, and the electronic pen's ID information.

[0357] The first server can generate a prompt using text information. For example, the first server can obtain a prompt form corresponding to a medium ID from the first database and generate a prompt using the prompt form and text information. For another example, the first server can generate a prompt by adding content to or modifying text information. For another example, the first server can generate a new prompt using a prompt and text information generated prior to obtaining the text information.

[0358] In step 503, the first server transmits text information to the interactive artificial intelligence server. The first server may transmit a prompt to the artificial intelligence server.

[0359] A conversational AI server is a server embedded with a language processing model for utilizing AI, providing conversational AI services. Examples of conversational AI services include OpenAI's ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot.

[0360] In step 504, the conversational AI server generates and extracts a result value using a language processing model with the character information received from the first server as input. Alternatively, the conversational AI server may generate a result value using the prompt received from the first server as input.

[0361] The result value consists of string data.

[0362] In step 505, the conversational artificial intelligence server transmits the extracted result value to the first server.

[0363] In step 506, the first server transmits the received result value to the electronic pen.

[0364] In step 507, the electronic pen converts the received result value into voice using the text-to-speech function.

[0365] At step 508, the electronic pen outputs the converted voice to the user of the electronic pen.

[0366] In step 509, the first server creates session information.

[0367] A session refers to a series of processes that involve sending text information to an interactive AI server and receiving corresponding results. When a session ends, the first server generates session information by linking the received result value to the session ID, the text information sent to the interactive AI server, and the user and electronic pen IDs that formed the basis of the text information. The first server then stores this information. The first server stores this information in a JSON file and can also add time information related to the session ID to further store it.

[0368] In step 510, the first server extracts at least one session information stored based on a user ID or electronic pen ID for a predetermined period of time and transmits it to the interactive artificial intelligence server along with predetermined session prompt information.

[0369] The predetermined session prompt may include content related to a summary of at least one session information. Additionally, the predetermined session prompt may include content related to an analysis of at least one session information.

[0370] In step 511, the conversational artificial intelligence server uses a language processing model to generate a result value using the session prompt received from the first server as an input value.

[0371] In step 512, the conversational artificial intelligence server transmits the generated result value to the first server.

[0372] The first server stores the result values ​​for the received session prompts. Upon user request, the first server provides the user with at least one session information and the result values ​​for the session prompts.

[0373] FIG. 13 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a third embodiment of the present invention. Hereinafter, as described in FIG. 7, the electronic device (1000) will be referred to as an electronic pen (110), the main server (2000) as a first server (120), and the artificial intelligence server (3000) as an interactive artificial intelligence server (130). The descriptions in FIG. 7 can be equally applied to each configuration.

[0374] Referring to Figure 13, in step 401, the electronic pen recognizes a predetermined code from a printed matter having a predetermined code printed thereon and extracts coordinate information. The predetermined code includes location information, i.e., coordinate information, on the printed matter. The electronic pen includes a camera, and when the camera approaches the printed matter having the predetermined code printed thereon within a predetermined distance, the camera captures the predetermined code within an area that the camera can recognize. Thereafter, the electronic pen reads the coordinate information from the predetermined code in the captured image.

[0375] In step 402, the electronic pen transmits the read coordinate information to the first server.

[0376] The electronic pen can further transmit the ID information of the electronic pen user and the ID information of the electronic pen to the first server.

[0377] The first server associates and stores the received coordinate information, the electronic pen user's ID information, and the electronic pen's ID information.

[0378] In step 403, the first server extracts a prompt corresponding to the coordinate information.

[0379] The first server pre-stores prompts corresponding to coordinate information. Prompts are input data that issue commands or instructions to the interactive AI server.

[0380] In step 404, the first server sends the extracted prompt to the conversational artificial intelligence server.

[0381] A conversational AI server is a server embedded with a language processing model for utilizing AI, providing conversational AI services. Examples of conversational AI services include OpenAI's ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot.

[0382] In step 405, the conversational artificial intelligence server uses a language processing model to generate and extract a result value using the prompt received from the first server as an input value.

[0383] The result value consists of string data.

[0384] In step 406, the conversational artificial intelligence server transmits the extracted result value to the first server.

[0385] In step 407, the first server transmits the received result value to the electronic pen.

[0386] In step 408, the electronic pen converts the received result value into voice using the text-to-speech function.

[0387] At step 409, the electronic pen outputs the converted voice to the user of the electronic pen.

[0388] In step 410, the first server transmits the received result value to the second server.

[0389] At step 411, the second server creates session information.

[0390] A session refers to a series of processes involving sending prompts to an interactive AI server and receiving corresponding results. When a session ends, the first server generates session information by linking the received result values ​​to the session ID, the prompt sent to the interactive AI server, and the IDs of the user and electronic pen that sent the code information that formed the basis of the prompt. The first server then stores this information. The first server stores this information in a JSON file and can also add time information related to the session ID to further store it.

[0391] In step 412, the second server extracts at least one session information stored based on the user ID or electronic pen ID for a predetermined period of time and transmits it to the interactive artificial intelligence server along with predetermined session prompt information.

[0392] The predetermined session prompt may include content related to a summary of at least one session information. Additionally, the predetermined session prompt may include content related to an analysis of at least one session information.

[0393] In step 413, the conversational artificial intelligence server uses a language processing model to generate a result value using the session prompt received from the second server as an input value.

[0394] In step 414, the conversational artificial intelligence server transmits the generated result value to the second server.

[0395] The second server stores the result values ​​for the received session prompts. Upon user request, the second server provides the user with at least one session information and the result values ​​for the session prompts.

[0396] FIG. 14 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a fourth embodiment. Hereinafter, as described in FIG. 7, the electronic device (1000) will be referred to as an electronic pen (110), the main server (2000) as a first server (120), and the artificial intelligence server (3000) as an interactive artificial intelligence server (130). The descriptions in FIG. 7 can be equally applied to each configuration.

[0397] In step 601, the electronic pen receives the user's voice input through a microphone and converts the voice information into text information.

[0398] In step 602, the electronic pen transmits the converted character information to the first server.

[0399] The first server associates and stores the received text information, the electronic pen user's ID information, and the electronic pen's ID information.

[0400] In step 603, the first server transmits text information to the interactive artificial intelligence server.

[0401] A conversational AI server is a server embedded with a language processing model for utilizing AI, providing conversational AI services. Examples of conversational AI services include OpenAI's ChatGPT, Google's Bard or Gemini, and Microsoft's Copilot.

[0402] In step 604, the conversational artificial intelligence server uses the character information received from the first server as an input value to generate and extract a result value using a language processing model.

[0403] The result value consists of string data.

[0404] In step 605, the conversational artificial intelligence server transmits the extracted result value to the first server.

[0405] In step 606, the first server transmits the received result value to the electronic pen.

[0406] In step 607, the electronic pen converts the received result value into voice using the text-to-speech function.

[0407] At step 608, the electronic pen outputs the converted voice to the user of the electronic pen.

[0408] In step 609, the first server transmits the received result value to the second server.

[0409] At step 610, the second server generates session information.

[0410] A session refers to a series of processes that involve sending text information or prompts to an interactive AI server and receiving corresponding results. When a session ends, the first server generates session information by linking the received result value to the session ID, the text information sent to the interactive AI server, and the IDs of the user and electronic pen that sent the voice information that formed the basis of the text information. The second server then stores this information. The second server stores this information in the form of a Jason file and may also store additional time information related to the session ID.

[0411] In step 611, the second server extracts at least one session information stored based on a user ID or electronic pen ID for a predetermined period of time and transmits it to the interactive artificial intelligence server along with predetermined session prompt information.

[0412] The predetermined session prompt may include content related to a summary of at least one session information. Additionally, the predetermined session prompt may include content related to an analysis of at least one session information.

[0413] In step 612, the conversational artificial intelligence server uses a language processing model to generate a result value using the session prompt received from the second server as an input value.

[0414] In step 613, the conversational artificial intelligence server transmits the generated result value to the second server.

[0415] The second server stores the result values ​​for the received session prompts. Upon user request, the second server provides the user with at least one session information and the result values ​​for the session prompts.

[0416] Fig. 15 is a flowchart illustrating a method for assisting a user's use of a medium using artificial context information according to a fifth embodiment.

[0417] Figure 16 is a diagram showing a process of obtaining artificial context information according to the fifth embodiment.

[0418] Fig. 17 is a diagram showing a prompt being generated using artificial context information according to the fifth embodiment.

[0419] Referring to FIG. 15, the medium use assistance method may include a step of obtaining a code image (S1100), a step of analyzing the code image to obtain code data (S1200), a step of obtaining reference information from the code data (S1300), a step of obtaining artificial context information based on the reference information (S1400), a step of generating a prompt using the artificial context information (S1500), a step of generating response data using the prompt (S1600), a step of obtaining response voice data corresponding to the response data (S1700), and a step of outputting the response voice data (S1800).

[0420] Each step is described in detail below.

[0421] A code image can be acquired (S1100). Specifically, a code is printed on a medium along with content (text or image), and when a user brings an electronic device (1000) into contact with or places it close to an area of ​​interest in the medium, the electronic device (1000) can capture at least a portion of the area of ​​interest using a sensing unit (1100) to acquire a code image. At this time, the sensing unit (1100) can be periodically activated to acquire an image. Alternatively, the sensing unit (1100) can be activated to acquire an image only when the user operates a button on the electronic device (1000).

[0422] Referring to FIG. 16, as an example, a user can use an electronic device (1000) to touch a part of a question area (or a part of the question area with a shape) on page 14 of a book, and the electronic device (1000) can use a sensing unit (1100) to obtain a code image for a code printed in the touched area.

[0423] Code data can be obtained by analyzing a code image (S1200). Specifically, an electronic device (1000) or a main server (2000) can obtain code data by analyzing a code image. For example, the electronic device (1000) can transmit a code image to the main server (2000), and the main server (2000) can obtain code data based on the arrangement of a plurality of fine dots included in the code image. As another example, the electronic device (1000) can obtain code data from the code image and transmit the obtained code data to the main server (2000). Here, the code data may mean data in which specific information among information about a medium is encoded.

[0424] Reference information can be obtained from code data (S1300). Specifically, referring to FIG. 16, the main server (2000) can decode the code data to obtain at least one of medium type, medium ID, page information, and location information (coordinate information and / or area identification information). For example, the reference information obtained in FIG. 16 may include {Book (medium type), BOOK000 (medium ID), 14 (page information), A1 (area identification information)}.

[0425] Artificial contextual information can be acquired based on reference information (S1400). Specifically, referring to FIG. 16, the main server (2000) can retrieve information corresponding to the reference information among information about the medium from the database (2150) to acquire the medium title and the question. As described above, the question is information with little relevance to the flow of the story printed on the medium, and is added by the content creator to increase the user's understanding or interest in the story, and thus can be understood as artificial contextual information. For example, the medium title acquired in FIG. 16 is "Snow White and the Seven Dwarfs," and the question is "Why is the queen jealous of the princess?"

[0426] A prompt can be generated using artificial context information (S1500). Specifically, the main server (2000) can generate a prompt using at least artificial context information. For example, the main server (2000) can load a prompt form from the database (2150) and modify the prompt form using artificial context information to generate a prompt. The main server (2000) can load a prompt form corresponding to reference information or context information from the database (2150). Meanwhile, the main server (2000) can use other information, such as context information or user information, in addition to artificial context information when modifying the prompt form.

[0427] For example, referring to FIG. 17, the main server (2000) can load a first prompt form from the database (2150). The first prompt form is text that includes a medium title, artificial context information, and a part where additional information is inserted. The main server (2000) can modify the first prompt form using the previously acquired medium title 'Snow White and the Seven Dwarfs', the question 'Why is the queen jealous of the princess?', and the pre-stored user information 'The child is 5 years old' (in addition, various user information such as the user's gender, personality, or family relationships can be used). The main server (2000) can modify the first prompt form to generate the first prompt.

[0428] As another example, the main server (2000) may generate a prompt by modifying or adding artificial contextual information. As another example, the main server (2000) may generate a prompt that only includes artificial contextual information.

[0429] The prompt may include descriptive text describing the artificial contextual information. Additionally, the prompt may further include guidance text describing what the AI ​​server (3000) should consider when generating a response.

[0430] Response data can be generated using a prompt (S1600). Specifically, the main server (2000) transmits the prompt generated in step S1500 to the artificial intelligence server (3000), and the artificial intelligence server (3000) can input the received prompt into a large language model to generate response data. For example, the artificial intelligence server (3000) can receive the first prompt illustrated in FIG. 17 and generate response data such as "Why is the queen jealous of the princess?" or "I think the queen is jealous of Snow White. Why is the queen jealous of the princess?"

[0431] The response data can be understood as data generated by the artificial intelligence server (2000) by considering artificial contextual information. That is, when a user uses the medium, the interaction system (100) can induce a conversation with the user, and the topic of the conversation can be related to artificial contextual information.

[0432] The response data may be text data or voice data. Specifically, the large language model included in the artificial intelligence server (3000) may be implemented in a text-based or voice-based manner, and if the input prompt is in text format, it may output text-based response data, and if the input prompt is in the form of a voice signal, it may output voice response data.

[0433] Response voice data corresponding to the response data can be obtained (S1700). Specifically, the main server (2000) can receive the response data from the artificial intelligence server (3000) and convert the response data into response voice data using the TTS model (2130). Meanwhile, the main server (2000) may not include the TTS model (2130), in which case the main server (2000) can obtain response voice data corresponding to the response data using an external text-to-speech conversion service. In addition, when the main server (2000) receives voice response data from the artificial intelligence server (3000), the main server (2000) can use the data as is without conversion.

[0434] Response voice data can be output (S1800). Specifically, the main server (2000) transmits voice response data to the electronic device (1000), and the electronic device (1000) can output the voice response data through the electronic device output unit (1400).

[0435] Hereinafter, a method for generating a follow-up prompt will be described with reference to FIGS. 18 and 19. A follow-up prompt can be generated when response data generated by considering previously artificial contextual information is delivered to the user and the user responds to it. The main server (2000) generates a follow-up prompt and transmits it to the artificial intelligence server (3000), and the artificial intelligence server (3000) outputs follow-up response data for the follow-up prompt. The follow-up response data is then delivered to the user, allowing a follow-up conversation to continue.

[0436] Figure 18 is a flowchart illustrating a method for generating a follow-up prompt according to the fifth embodiment.

[0437] FIG. 19 is a diagram showing a subsequent prompt being generated using artificial context information according to the fifth embodiment.

[0438] Referring to FIG. 18, a method for generating a follow-up prompt may include a step of obtaining a voice text corresponding to a user's voice (S2100), a step of generating a follow-up prompt using artificial context information and a voice text (S2200), a step of generating follow-up response data using the follow-up prompt (S2300), a step of obtaining follow-up response voice data corresponding to the follow-up response data (S2400), and a step of outputting the follow-up response voice data (S2500).

[0439] Each step is described in detail below.

[0440] A voice text corresponding to a user's voice can be obtained (S2100). Specifically, the electronic device (1000) can record the user's voice by activating the microphone of the electronic device input unit (1300) through the user's button operation. The electronic device (1000) transmits the recorded user voice data to the main server (2000), and the main server (2000) can convert the user voice data into voice text using the STT model (2110). Meanwhile, the main server (2000) may not include the STT model (2110), in which case the main server (2000) can obtain voice text corresponding to the user voice data using an external voice-to-text conversion service.

[0441] A follow-up prompt can be generated using artificial context information and spoken text (S2200). Specifically, the main server (2000) can generate a follow-up prompt using at least artificial context information and spoken text.

[0442] For example, the main server (2000) can generate a prompt by loading a follow-up prompt form from the database (2150) and modifying the follow-up prompt form using artificial context information and spoken text. Here, the main server (2000) can load a follow-up prompt form corresponding to reference information or context information from the database (2150). Meanwhile, the main server (2000) can utilize other information, such as context information or user information, in addition to artificial context information, when modifying the follow-up prompt form.

[0443] For example, referring to FIG. 19, the main server (2000) may load a first follow-up prompt form from the database (2150). The first follow-up prompt form may be text that includes artificial contextual information and a portion into which user voice information is inserted. The main server (2000) may modify the first follow-up prompt form using the question obtained in step S1400, “Why is the queen jealous of the princess?” and the voice text obtained in step S2100, “Because the mirror said the princess was prettier.” The main server (2000) may modify the first follow-up prompt form to generate the first follow-up prompt.

[0444] As another example, the main server (2000) may generate a follow-up prompt by modifying or adding content to the spoken text. As another example, the main server (2000) may generate a follow-up prompt that only includes the spoken text.

[0445] The follow-up prompt may include descriptive text describing the spoken text. Additionally, the follow-up prompt may include additional guidance text describing considerations for the AI ​​server (3000) when generating a response.

[0446] Follow-up response data can be generated using the follow-up prompt (S2300). Specifically, the main server (2000) transmits the follow-up prompt generated in step S2200 to the artificial intelligence server (3000), and the artificial intelligence server (3000) can input the received follow-up prompt into a large language model to generate follow-up response data. For example, the artificial intelligence server (3000) can receive the first follow-up prompt illustrated in FIG. 19 and generate follow-up response data such as "That's right, the queen is jealous of the princess because of the mirror" or "That's right, the queen is jealous of the princess because of the mirror. Have you ever felt jealous of someone?"

[0447] The follow-up response data can be understood as data generated by the artificial intelligence server (2000) considering artificial contextual information. That is, after the interaction system (100) induces a conversation with the user, the topic of the induced conversation may be continuously related to the artificial contextual information.

[0448] The follow-up response data may be text data or voice data. Specifically, the large language model included in the artificial intelligence server (3000) may be implemented in a text-based or voice-based manner, and if the input follow-up prompt is in text format, follow-up response data in text format may be output, and if the input follow-up prompt is in the form of a voice signal, follow-up voice response data may be output.

[0449] Follow-up response voice data corresponding to follow-up response data can be obtained (S2400). Specifically, the main server (2000) can receive follow-up response data from the artificial intelligence server (3000) and convert the follow-up response data into follow-up response voice data using the TTS model (2130). Meanwhile, the main server (2000) may not include the TTS model (2130), in which case the main server (2000) can obtain follow-up response voice data corresponding to the follow-up response data using an external text-to-speech conversion service. In addition, when the main server (2000) receives follow-up voice response data from the artificial intelligence server (3000), the main server (2000) can use the data as is without conversion.

[0450] Follow-up response voice data may be output (S2500). Specifically, the main server (2000) transmits the follow-up voice response data to the electronic device (1000), and the electronic device (1000) may output the follow-up voice response data through the electronic device output unit (1400).

[0451] After the subsequent response voice data is output, the subsequent prompt generation method may be performed again. Specifically, when a user records a voice using an electronic device (1000), a subsequent prompt is generated using the acquired voice text and artificial context information, and subsequent voice response data for the subsequent prompt is generated and output to the user (at this time, the format of the subsequent prompt may vary). This means that the conversation between the interaction system (100) and the user continues.

[0452] Meanwhile, if the artificial context information being used changes (e.g., if the user acquires different artificial context information by touching or pointing to a different question area using the electronic device (1000), the session may end and a new session may begin. In this case, a session refers to a conversation about a single artificial context information, and the beginning of a new session means that the method of assisting the user's use of the medium using the artificial context information illustrated in FIG. 15 is performed again.

[0453] The main server (2000) can store the prompt and response data generated in the session as session information. The session information can be used to generate analysis information as described above. For example, the main server (2000) or the artificial intelligence server (3000) can use the session information to generate at least one of a conversation log representing a conversation between a user and an electronic device (1000), a conversation summary summarizing a conversation between a user and an electronic device (1000), a pronunciation evaluation representing an evaluation of the user's pronunciation, keywords frequently used by the user, the user's medium usage time, the user's medium usage time by type, and the user's areas of interest.

[0454] Hereinafter, with reference to FIGS. 20 to 22, a method of assisting a user in using a medium by using a pre-stored prompt among methods of assisting a user in using a medium by using artificial context information by an interaction system (100) is described.

[0455] Figure 20 is a diagram showing a process of obtaining a pre-stored prompt according to the sixth embodiment.

[0456] Fig. 21 is a drawing showing a pre-stored prompt according to the sixth embodiment.

[0457] FIG. 22 is a diagram showing a subsequent prompt being generated using artificial context information according to the sixth embodiment.

[0458] A method for assisting a user's use of a medium using pre-saved prompts may include the steps described in FIGS. 15 and 18. Accordingly, any overlapping portions with the previously described content will be omitted, and the content described in FIGS. 15 and 18 may be applied equally.

[0459] Referring to Figure 20, a medium is a card with specific information printed on it, and a user can touch or position an electronic device (1000) on a specific area of ​​the card. The medium can be divided into a question area and an explanation area.

[0460] When an electronic device (1000) comes into contact with or is positioned on an area of ​​a card, reference information may be acquired as described in steps S1100 to S1300. Here, the reference information may include medium type, medium ID, and location information (coordinate information and / or area identification information). For example, referring to FIG. 20, the reference information may be {card (medium type), CARD000 (medium ID), (1308.59, 393.417) (coordinate information)}.

[0461] The main server (2000) can load prompts corresponding to reference information from the database (2150). The prompts can be pre-stored in the database (2150) and input into the large language model of the artificial intelligence server (3000) without additional modification or processing.

[0462] Meanwhile, depending on the area that the electronic device (1000) comes into contact with, the location information among the reference information may change, and the prompt to be loaded may vary depending on whether the location information is included in the question area or the description area (if it is coordinate information) or whether the location information is included in the question area or the description area (if it is area identification information).

[0463] For example, if location information is included in the question area or indicates the question area, the main server (2000) may load the second prompt illustrated in FIG. 21. The second prompt may include information about the content printed on the card, a question generation guide for starting a conversation, and the like. The main server (2000) may transmit the second prompt to the artificial intelligence server (3000), and the artificial intelligence server (3000) may use the second prompt to generate response data. The response data may be, for example, "Have you heard of flying squirrels?"

[0464] Thereafter, a voice corresponding to the response data may be output to the user according to steps S1700 and S1800.

[0465] After a voice corresponding to the response data is output to the user, the user can obtain user voice data by activating the recording function of the electronic device (1000) to respond to the output voice. Thereafter, step S2100 is performed to obtain voice text.

[0466] The main server (2000) can generate subsequent prompts using response text and spoken text. Here, response text refers to response data generated using a previously stored prompt.

[0467] For example, referring to FIG. 22, the main server (2000) can load a second follow-up prompt form from the database (2150). The second follow-up prompt form can be text that includes response information and a portion where user voice information is inserted. The main server (2000) can modify the second follow-up prompt form using the previously obtained response text, "Have you heard of flying squirrels?" and the voice text, "Squirrels know." The main server (2000) can modify the second follow-up prompt form to generate the second follow-up prompt.

[0468] As another example, the main server (2000) may generate a follow-up prompt by modifying or adding content to the spoken text. As another example, the main server (2000) may generate a follow-up prompt that only includes the spoken text.

[0469] The subsequent prompt may include descriptive text describing the spoken text. Additionally, the prompt may further include guidance text describing considerations for the AI ​​server (3000) when generating a response.

[0470] Thereafter, steps S2300 to S2500 may be performed so that subsequent response voice data corresponding to the subsequent prompt may be output to the user.

[0471] After the follow-up response voice data is output, the follow-up prompt generation method can be performed again. Specifically, if the user additionally records voice using the electronic device (1000), a follow-up prompt can be generated using the acquired voice text, and follow-up voice response data for the follow-up prompt can be generated and output to the user. This means that the conversation between the interaction system (100) and the user continues. At this time, the follow-up prompt format may vary.

[0472] Meanwhile, if the medium ID changes (e.g., the user touches a different card using an electronic device (1000), the session may end and a new session may begin. In this case, a session refers to a conversation conducted on a single medium, and the beginning of a new session means that the method of assisting the user's use of the medium using the previously-saved prompts is re-executed.

[0473] The main server (2000) can store the prompt and response data generated in the session as session information. The session information can be used to generate analysis information as described above. For example, the main server (2000) or the artificial intelligence server (3000) can use the session information to generate at least one of a conversation log representing a conversation between a user and an electronic device (1000), a conversation summary summarizing a conversation between a user and an electronic device (1000), a pronunciation evaluation representing an evaluation of the user's pronunciation, keywords frequently used by the user, the user's medium usage time, the user's medium usage time by type, and the user's areas of interest.

[0474]

[0475] 5. Interaction using medium context information

[0476] Hereinafter, with reference to FIGS. 23 to 27, a method for an interaction system (100) to assist a user in using a medium by using medium context information is described.

[0477] Users (especially children) may express questions or thoughts that arise while reading a book. In this case, the user's voice can be input into a large language model to provide response data. However, the user's questions or thoughts arise while reading the book. To generate more appropriate response data, the large language model needs to understand the context in which the user spoke, including the intent and purpose of the words.

[0478] As a way to understand the user's context in a large-scale language model, the large-scale language model can be provided with content information printed on the medium. More specifically, if the user's voice is input into the large-scale language model along with information such as the title or content of the book the user is reading, or their interests, the resulting response data can be expected to be generated within the user's context.

[0479] FIG. 23 is a flowchart illustrating a method for assisting a user's use of a medium by using medium context information according to the seventh embodiment.

[0480] Figure 24 is a diagram showing a process for obtaining medium context information according to the seventh embodiment.

[0481] FIG. 25 is a diagram showing a prompt being generated using medium context information according to the seventh embodiment.

[0482] Referring to FIG. 23, the medium use assistance method may include a step of obtaining a code image (S3100), a step of analyzing the code image to obtain code data (S3200), a step of obtaining reference information from the code data (S3300), a step of obtaining medium context information based on the reference information (S3400), a step of obtaining voice text corresponding to the user's voice (S3500), a step of generating a prompt using the voice text and medium context information (S3600), a step of generating response data using the prompt (S3700), a step of obtaining response voice data corresponding to the response data (S3800), and a step of outputting the response voice data (S3900).

[0483] Each step is described in detail below. However, step S3100 is identical to step S1100, step S3200 is identical to step S1200, step S3800 is identical to step S1700, and step S3900 is identical to step S1800, so the contents described in FIG. 15 can be applied in the same manner.

[0484] After steps S3100 and S3200 are performed, reference information can be obtained from the code data (S3300). At this time, the code image obtained in step S3100 can be obtained by the user using an electronic device (1000) to touch a text area on a specific page of the medium where text is printed.

[0485] Referring to FIG. 24, the main server (2000) can decode code data to obtain at least one of medium type, medium ID, page information, and location information (coordinate information and / or area identification information). For example, the reference information obtained by the main server (2000) in FIG. 24 may include {Book (medium type), BOOK000 (medium ID), 14 (page information), A2 (area identification information)}.

[0486] Medium context information can be acquired based on reference information (S3400). Specifically, referring to FIG. 24, the main server (2000) can retrieve information corresponding to the reference information among the information about the medium from the database (2150) to acquire the medium title and page-by-page text. As described above, the page-by-page text is information about the flow of the story printed on the medium and can be understood as medium context information. For example, the medium title acquired in FIG. 24 is 'Snow White and the Seven Dwarfs', and the page-by-page text is 'From one day...even the magic mirror's answers changed. "Snow White is the fairest in the world." "What? Snow White is prettier than me?" No matter how many times I asked, the magic mirror only repeated that Snow White was the fairest. The queen's eyes burned with jealousy.' (hereinafter, 'From one day...burned with jealousy').

[0487] Meanwhile, while the medium context information is described above as page-specific text, the technical concept of the present disclosure is not limited thereto. The medium context information may be the entire text or image printed on the medium, or a portion of the text or image that the user is interested in. Specifically, the medium context information may include text related to the portion of the medium that the user touched with the electronic device (1000), or a summary or modified text thereof. The medium context information may be a word, sentence, or paragraph, and may also include an image.

[0488] A voice text corresponding to a user's voice can be obtained (S3500). Specifically, the electronic device (1000) can record the user's voice by activating the microphone of the electronic device input unit (1300) through the user's button operation. The electronic device (1000) transmits the recorded user voice data to the main server (2000), and the main server (2000) can convert the user voice data into voice text using the STT model (2110). Meanwhile, the main server (2000) may not include the STT model (2110), in which case the main server (2000) can obtain voice text corresponding to the user voice data using an external voice-to-text conversion service.

[0489] A prompt can be generated using spoken text and medium context information (S3600). Specifically, the main server (2000) can generate a prompt using at least spoken text and medium context information. For example, the main server (2000) can load a prompt form from the database (2150) and modify the prompt form using spoken text and medium context information to generate a prompt.

[0490] The main server (2000) can load a prompt form corresponding to reference information (e.g., medium type, medium ID, page information, area identification information) or contextual information (e.g., medium title) from the database (2150). Meanwhile, the main server (2000) can utilize other information, such as contextual information or user information other than medium contextual information, when modifying the prompt form.

[0491] For example, referring to FIG. 25, the main server (2000) can load a third prompt form from the database (2150). The third prompt form is text that includes a section where a medium title, medium context information, and user voice information are inserted. The main server (2000) can modify the third prompt form using the previously acquired medium title, “Snow White and the Seven Dwarfs,” the page-specific text, “From which day on… I was burning with jealousy,” and the voice text, “Where does the princess live?” The main server (2000) can modify the third prompt form to generate the third prompt.

[0492] As another example, the main server (2000) may generate a prompt by modifying or adding content to the medium context information and spoken text. As another example, the main server (2000) may separately generate a prompt containing only the medium context information and a prompt containing only the spoken text, and sequentially transmit these to the artificial intelligence server (3000).

[0493] The prompt may include descriptive text describing the medium's contextual information. Additionally, the prompt may further include guidance text describing considerations for the AI ​​server (3000) to consider when generating a response.

[0494] Response data can be generated using a prompt (S3700). Specifically, the main server (2000) transmits the prompt generated in step S3600 to the artificial intelligence server (3000), and the artificial intelligence server (3000) can input the received prompt into a large language model to generate response data. For example, the artificial intelligence server (3000) can receive the third prompt illustrated in FIG. 25 and generate response data such as "The princess lives in the castle with the queen" or "The princess lives in the castle with the queen. What else are you curious about?"

[0495] The response data can be understood as data generated by the artificial intelligence server (2000) considering the contextual information of the medium. That is, when a user uses the medium, the interaction system (100) can induce a conversation with the user, and the topic of the conversation can be related to the part of the story printed on the medium that the user is reading (or is interested in).

[0496] The response data may be text data or voice data. Specifically, the large language model included in the artificial intelligence server (3000) may be implemented in a text-based or voice-based manner, and if the input prompt is in text format, it may output text-based response data, and if the input prompt is in the form of a voice signal, it may output voice response data.

[0497] Thereafter, steps S3800 and S3900 are performed so that response voice data corresponding to the response data can be output to the user.

[0498] After a voice corresponding to the response data is output to the user, the user can obtain user voice data by activating the recording function of the electronic device (1000) to respond to the output voice.

[0499] In this case, the main server (2000) generates a follow-up prompt and transmits it to the artificial intelligence server (3000), the artificial intelligence server (3000) outputs follow-up response data for the follow-up prompt, and the follow-up response data is transmitted to the user, thereby allowing a follow-up conversation to continue.

[0500] Below, we describe how a follow-up prompt is generated with reference to Figure 26.

[0501] Figure 26 is a flowchart illustrating a method for generating a follow-up prompt according to the seventh embodiment.

[0502] Referring to FIG. 26, a method for generating a follow-up prompt may include a step of obtaining a voice text corresponding to a user's voice (S4100), a step of generating a follow-up prompt using medium context information and the voice text (S4200), a step of generating follow-up response data using the follow-up prompt (S4300), a step of obtaining follow-up response voice data corresponding to the follow-up response data (S4400), and a step of outputting the follow-up response voice data (S4500).

[0503] Here, step S4100 is identical to step S2100, step S4300 is identical to step S2300, step S4400 is identical to step S2400, and step S4500 is identical to step S2500, so that the contents described in FIG. 18 can be applied identically.

[0504] Accordingly, only step S4200 will be described in detail.

[0505] After step S4100 is performed and the spoken text is acquired, a subsequent prompt can be generated using the response data and the spoken text (S4200). Specifically, the main server (2000) can generate the subsequent prompt using at least the response data and the spoken text. Here, the response data refers to the data generated in step S3700.

[0506] For example, the main server (2000) can generate a prompt by loading a follow-up prompt form from the database (2150) and modifying the follow-up prompt form using response data and spoken text. Here, the main server (2000) can load a follow-up prompt form corresponding to reference information or contextual information from the database (2150). Meanwhile, the main server (2000) can utilize other information, such as medium contextual information or user information, in addition to response data and spoken text, when modifying the follow-up prompt form.

[0507] For another example, the main server (2000) may generate a subsequent prompt that includes previously acquired response data and spoken text. At this time, the response data and spoken data may each be partially modified or have additional content added to them.

[0508] As another example, the main server (2000) may generate a follow-up prompt by modifying or adding content to the spoken text. As another example, the main server (2000) may generate a follow-up prompt that only includes the spoken text.

[0509] Thereafter, steps S4300 to S4500 are performed so that subsequent response voice data can be output to the user through the electronic device (1000).

[0510]

[0511] 6. Interaction using artificial context information and medium context information

[0512] Hereinafter, with reference to FIGS. 27 to 29, a method for an interaction system (100) to assist a user in using a medium by using artificial context information and medium context information is described.

[0513] FIG. 27 is a flowchart illustrating a method for assisting a user's use of a medium by using artificial context information and medium context information according to the eighth embodiment.

[0514] Figure 28 is a diagram showing a process for obtaining artificial context information and medium context information according to the eighth embodiment.

[0515] FIG. 29 is a diagram showing a prompt being generated using artificial context information and medium context information according to the eighth embodiment.

[0516] Referring to FIG. 27, the medium use assistance method may include a step of obtaining a code image (S5100), a step of analyzing the code image to obtain code data (S5200), a step of obtaining reference information from the code data (S5300), a step of obtaining artificial context information and medium context information based on the reference information (S5400), a step of generating a prompt using the artificial context information and the medium context information (S5500), a step of generating response data using the prompt (S5600), a step of obtaining response voice data corresponding to the response data (S5700), and a step of outputting the response voice data (S5800).

[0517] Each step is described in detail below. However, step S5100 is identical to step S1100, step S5200 is identical to step S1200, step S5700 is identical to step S1700, and step S5800 is identical to step S1800, so the contents described in FIG. 15 can be applied identically.

[0518] After steps S5100 and S5200 are performed, reference information can be obtained from the code data (S5300). At this time, the code image obtained in step S5100 can be obtained by a user touching a question area on a specific page of the medium using an electronic device (1000).

[0519] Referring to FIG. 28, the main server (2000) can decode code data to obtain at least one of medium type, medium ID, page information, and location information (coordinate information and / or area identification information). For example, the reference information obtained by the main server (2000) in FIG. 28 can include {Book (medium type), BOOK000 (medium ID), 14 (page information), A1 (area identification information)}.

[0520] Based on reference information, artificial context information and medium context information can be acquired (S5400). Specifically, referring to FIG. 28, the main server (2000) can retrieve information corresponding to the reference information among information about the medium from the database (2150) to acquire the medium title, page-by-page text, and questions. As described above, the page-by-page text is information about the flow of the story printed on the medium and can be understood as medium context information. In addition, the questions have little relevance to the flow of the story printed on the medium and are created by the content creator to enhance the user's interest and comprehension, and can be understood as artificial context information.

[0521] For example, the medium title obtained in FIG. 28 is 'Snow White and the Seven Dwarfs', the page-specific text is 'From which day... I was filled with jealousy', and the question is 'Why is the queen jealous of the princess?' Meanwhile, although the medium context information is described as page-specific text in the above, the technical idea of ​​the present disclosure is not limited thereto. The medium context information may be printed text or picture information not only in the page-specific text but also throughout the medium. Alternatively, the medium context information may be a specific sentence within a specific page. Furthermore, although the artificial context information is described as a question in the above, the technical idea of ​​the present disclosure is not limited thereto, and contents that help understanding of the medium without affecting the flow of the story, such as instructions for specific learning, may be obtained as artificial context information.

[0522] Prompts can be generated using artificial context information and medium context information (S5500). Specifically, the main server (2000) can generate prompts using at least artificial context information and medium context information. For example, the main server (2000) can load a prompt form from the database (2150) and modify the prompt form using artificial context information and medium context information to generate a prompt.

[0523] The main server (2000) can load a prompt form corresponding to reference information (e.g., medium type, medium ID, page information, area identification information) or contextual information (e.g., medium title) from the database (2150). Meanwhile, the main server (2000) can further utilize other information, such as artificial contextual information and contextual information other than medium contextual information or user information, when modifying the prompt form.

[0524] For example, referring to FIG. 29, the main server (2000) can load a fourth prompt form from the database (2150). The fourth prompt form is text that includes a portion where a medium title, medium context information, and artificial context information are inserted. The main server (2000) can modify the fourth prompt form using the previously acquired medium title 'Snow White and the Seven Dwarfs', the page-specific text 'From which day... I was consumed by jealousy', and the question information 'Why is the queen jealous of the princess?'. The main server (2000) can modify the fourth prompt form to generate the fourth prompt.

[0525] As another example, the main server (2000) may generate a prompt by partially modifying or adding artificial context information and medium context information. As another example, the main server (2000) may separately generate a prompt containing only artificial context information or a prompt containing only medium context information and sequentially transmit the prompt to the artificial intelligence server (3000).

[0526] The prompt may include descriptive text describing the artificial contextual information. Additionally, the prompt may further include guidance text describing what the AI ​​server (3000) should consider when generating a response.

[0527] Response data can be generated using a prompt (S5600). Specifically, the main server (2000) transmits the prompt generated in step S5500 to the artificial intelligence server (3000), and the artificial intelligence server (3000) inputs the received prompt into a large language model to generate response data. For example, the artificial intelligence server (3000) can receive the fourth prompt illustrated in FIG. 29 and generate response data such as "Why is the queen jealous of the princess?" or "Why is the queen jealous of the princess? What did the mirror say?"

[0528] The response data can be understood as data generated by the artificial intelligence server (2000) by considering artificial contextual information and medium contextual information. That is, when a user uses the medium, the interaction system (100) can induce a conversation with the user. The topic of the conversation is related to the part of the story printed on the medium that the user is reading (or is interested in), and the conversation can be understood as developing by considering the deep learning (to enhance understanding of the medium) induced by the content creator.

[0529] The response data may be text data or voice data. Specifically, the large language model included in the artificial intelligence server (3000) may be implemented in a text-based or voice-based manner, and if the input prompt is in text format, it may output text-based response data, and if the input prompt is in the form of a voice signal, it may output voice response data.

[0530] Thereafter, steps S5700 and S5800 are performed so that response voice data corresponding to the response data can be output to the user.

[0531] After a voice corresponding to the response data is output to the user, the user can obtain user voice data by activating the recording function of the electronic device (1000) to respond to the output voice.

[0532] In this case, the follow-up prompt generation method described in Figure 18 can be performed, and follow-up response data for the follow-up prompt can be generated. Accordingly, the follow-up response data can be delivered to the user, allowing a follow-up conversation to continue. In generating the follow-up prompt, it goes without saying that medium context information can be utilized in addition to artificial context information.

[0533]

[0534] 7. When the data types processed in the generative model are diverse

[0535] Below, referring to FIG. 30, a method for assisting a user's use of a medium by using contextual information when a generative model stored in an artificial intelligence server (3000) processes data in a format other than text data is described.

[0536] Fig. 30 is a flowchart illustrating a method for assisting a user's use of a medium using contextual information according to the ninth embodiment.

[0537] Referring to FIG. 30, the medium use assistance method may include a step of obtaining a code image (S6100), a step of analyzing the code image to obtain code data (S6200), a step of obtaining reference information from the code data (S6300), a step of obtaining context information based on the reference information (S6400), a step of converting the context information according to a specific data type (S6500), a step of obtaining voice data in which the user's voice is recorded (S6600), a step of converting the voice data according to a specific data type (S6700), a step of generating a prompt using the converted context information and the converted voice data (S6800), a step of generating response data using the prompt (S6900), and a step of outputting the response data (S7000).

[0538] Each step is described in detail below. However, step S6100 is identical to step S1100, and step S6200 is identical to step S1200, so the contents described in FIG. 15 can be applied equally.

[0539] After steps S6100 and S6200 are performed, reference information can be obtained from the code data (S6300). At this time, the code image obtained in step S6100 can be obtained by the user using an electronic device (1000) to touch an area (e.g., a text area, a picture area, or a question area) of a specific page of the medium.

[0540] Reference information may include at least one of medium type, medium ID, page information, and location information (coordinate information and / or area identification information).

[0541] Contextual information can be acquired based on reference information (S6400). Here, the contextual information may include artificial contextual information and / or medium contextual information. For example, the method described in step S1400 may be applied to the artificial contextual information. For example, the method described in step S3400 may be applied to the medium contextual information. For example, the method described in step S5500 may be applied to the artificial contextual information and medium contextual information.

[0542] Meanwhile, contextual information may include at least one of text, sound, image, or video. For example, contextual information may include text printed on a specific page of the medium, a voice reading the printed text, a sound related to the printed text, an image related to the printed text, or a video related to the printed text. For another example, contextual information may include an image printed on a specific page of the medium, text describing the printed image, a voice describing the printed image, a sound related to the printed image, or a video related to the printed image.

[0543] Contextual information can be converted based on a specific data type (S6500). For example, the main server (2000) can convert contextual information based on the data type that the generative model stored in the artificial intelligence server (3000) can process. Here, the data type that the generative model can process refers to the data type of input data supported by the generative model (or the data type that can receive input and output an appropriate response), and can include at least one of text, sound, image, or video.

[0544] For example, if contextual information is composed of text and the type of data that the generative model can input is sound, the main server (2000) can convert the text of the contextual information into sound (e.g., voice). In this case, the aforementioned TTS model (2130) may be used, or an external server capable of performing the function of converting text into voice may be used.

[0545] As another example, if the contextual information consists of an image or video and the type of data that the generative model can input is text, the main server (2000) can convert the image or video of the contextual information into text. Here, the text means a description of the image or video.

[0546] Meanwhile, if the generative model can process multiple data types, and the data type of the contextual information is included in the data types that the generative model can process, the main server (2000) may not convert the contextual information. For example, if the contextual information consists of sound and the data types that the generative model can input are text and sound, the main server (2000) may not convert the contextual information. That is, step S6500 may be omitted.

[0547] The main server (2000) can obtain voice data in which the user's voice is recorded (S6600). Specifically, the electronic device (1000) can record the user's voice by activating the microphone of the electronic device input unit (1300) through the user's button operation. The electronic device (1000) can transmit the recorded user voice data to the main server (2000). Here, the voice data can be understood as data containing voice information or a voice signal.

[0548] Voice data can be converted according to a specific data type (S6700). For example, the main server (2000) can convert voice data based on the data type that the generative model stored in the artificial intelligence server (3000) can process. Here, the data type that the generative model can process refers to the data type of input data supported by the generative model (or the data type that can receive input and output an appropriate response), and can include at least one of text, sound, image, or video.

[0549] For example, if the type of data that the generative model can receive is text, the main server (2000) can convert voice data into text. In this case, the aforementioned STT model (2110) may be used, or an external server capable of converting voice into text may be used.

[0550] Meanwhile, the data types that the generative model can input may include sound. In this case, the main server (2000) may not convert the sound data. That is, step S6700 may be omitted.

[0551] A prompt can be generated using the converted context information and converted voice data (S6800). Specifically, the main server (2000) can generate a prompt using the converted context information obtained in step S6500 and the converted voice data obtained in step S6700.

[0552] For example, the main server (2000) can generate a prompt by concatenating converted context information and converted speech data.

[0553] As another example, the main server (2000) can generate a prompt by using the aforementioned prompt form, but processing the prompt form using converted context information and converted voice data.

[0554] Meanwhile, the data type of the converted context information and the data type of the converted speech data may be the same or different. The main server (2000) may generate a prompt after unifying the data types of the converted context information and the converted speech data into one.

[0555] Response data can be generated using a prompt (S6900). For example, the main server (2000) can transmit the generated prompt to the artificial intelligence server (3000), and the artificial intelligence server (3000) can generate response data from the prompt using a generative model.

[0556] The data type of the generated response data may include at least one of text, sound, image, and video. The data type of the response data may vary depending on the process by which the generative model was trained. The generative model may be trained to input data including at least one of the data types of text, sound, image, and video and output data including at least one of the data types of text, sound, image, and video.

[0557] Response data can be output to the user (S7000). Specifically, the main server (2000) can transmit the response data to the electronic device (1000) with or without conversion, and at this time, the form in which the response data is output from the electronic device (1000) can vary depending on the data type of the response data.

[0558] For example, if the response data is text, the main server (2000) converts the response data into response voice data and provides it to the electronic device (1000), and the electronic device (1000) can output the response voice data through the speaker of the electronic device output unit (1400). At this time, the TTS model (2130) of the main server (2000) can be used.

[0559] As another example, if the response data is voice (or sound), the main server (2000) provides the response data to the electronic device (1000) without converting it, and the electronic device (1000) can output the response data through the speaker of the electronic device output unit (1400).

[0560] As another example, if the response data is an image or video, the main server (2000) converts the response data into response voice data and provides it to the electronic device (1000), and the electronic device (1000) can output the response data through the speaker of the electronic device output unit (1400).

[0561] As another example, if the response data is an image or video, the main server (2000) provides the response data to the electronic device (1000) as is without converting it, and the electronic device (1000) can output the response data through the display of the electronic device output unit (1400).

[0562] Since the generative model stored in the artificial intelligence server (3000) can process various data types, the response data output to the user can be understood as having been generated with greater consideration to the user's context. For example, when a user asks a question by pointing to an image in a medium using an electronic device (1000), if the generative model can only process text, a process of converting the image into text must be followed, and since the converted text is related to the image but is not the image itself, the response data that the generative model can generate cannot be considered to have been generated with clear recognition of the user's interest or question intent. On the other hand, if the generative model can process both text and images, the generative model can receive the original image as input, more clearly recognize the user's interest or question intent, and generate response data.

[0563] Accordingly, a more appropriate response can be provided to the user, that is, a response more appropriate to the intent of the user's question.

[0564]

[0565] The features, structures, effects, etc. described in the embodiments above are included in at least one embodiment of the present specification, and are not necessarily limited to just one embodiment. Furthermore, the features, structures, effects, etc. exemplified in each embodiment can be combined or modified in other embodiments by a person skilled in the art to which the embodiments pertain. Therefore, the contents related to such combinations and modifications should be construed as being included within the scope of the present specification.

[0566] In addition, although the above description focuses on the embodiments, these are merely examples and do not limit the technical idea of ​​this specification. Those with ordinary skill in the art to which this specification pertains will recognize that various modifications and applications not exemplified above are possible without departing from the essential characteristics of this embodiment. In other words, each component specifically shown in the embodiments can be modified and implemented. In addition, differences related to such modifications and applications should be interpreted as being included within the scope of this specification as defined in the appended claims.

[0567] -

Claims

1. A server that communicates with electronic devices and artificial intelligence servers. memory; Department of Communications; and It includes a control unit that controls the above memory and the above communication unit; The above control unit, [1] Obtain reference information encoded in the code data obtained by the above electronic device; - At this time, the reference information is received from the electronic device through the communication unit or the control unit receives the code data from the electronic device through the communication unit and then obtains the code data by decoding the code data. The above reference information relates to the medium through which the user interacts using the electronic device; The above reference information includes ID information of the medium, page information indicated by the electronic device, and location information within the page indicated by the electronic device. [2] Based on the above reference information, the content printed on the medium and / or the content of interest about a part of the content printed on the medium that the user is interested in is acquired from the memory to generate medium context information, [3] Obtaining a voice text corresponding to the user's voice obtained by the electronic device; -At this time, the above voice text is obtained by converting the above voice into STT (Speech-To-Text)-, [4] Based on the acquired medium context information and the spoken text, a prompt is generated. -At this time, the prompt includes the medium context information and the spoken text-, [5] Transmit the above prompt to the artificial intelligence server through the above communication unit, [6] Receive a response from the artificial intelligence server to the transmitted prompt through the communication unit, [7] Obtain a response voice corresponding to the received response, [8] Transmitting the above response voice to the electronic device through the communication unit. A server that communicates with electronic devices and artificial intelligence servers.

2. In paragraph 1, The above memory stores a plurality of medium texts, each medium text corresponding to at least a portion of a text printed on a particular medium, The above control unit, Loading a medium text corresponding to the medium ID information of the reference information among the plurality of medium texts from the memory and obtaining the medium context information. A server that communicates with electronic devices and artificial intelligence servers.

3. In paragraph 1, The above memory stores multiple page-specific texts, each page-specific text corresponding to at least a portion of the text printed on a particular page of a particular medium, The above control unit, Loading the page-specific text corresponding to the medium ID information and the page information of the reference information from the plurality of page-specific texts from the memory and obtaining it as the medium context information. A server that communicates with electronic devices and artificial intelligence servers.

4. In paragraph 1, The above memory stores a plurality of position-specific texts, each position-specific text corresponding to at least a portion of text printed at a particular position within a particular page of a particular medium, The above control unit, Loading the medium ID information, the page information, and the position-specific text corresponding to the position information among the plurality of position-specific texts from the memory to obtain the medium context information. A server that communicates with electronic devices and artificial intelligence servers.

5. In paragraph 1, The above prompt is generated by concatenating the spoken text with the medium context information. A server that communicates with electronic devices and artificial intelligence servers.

6. In paragraph 1, The above prompt further includes descriptive text about the medium context information. A server that communicates with electronic devices and artificial intelligence servers.

7. In paragraph 1, The above prompt further includes pre-stored user information, The user information includes at least one of age and gender for the user. A server that communicates with electronic devices and artificial intelligence servers.

8. In paragraph 1, The above prompt further includes pre-saved guide text, The above guide text includes matters to be considered by the artificial intelligence server when generating the response. A server that communicates with electronic devices and artificial intelligence servers.

9. In paragraph 1, The above control unit, Obtain the prompt form from the above memory, Generating the prompt by modifying the prompt form using the medium context information and the spoken text. A server that communicates with electronic devices and artificial intelligence servers.

10. In paragraph 1, wherein said prompt comprises at least a portion of a past prompt transmitted to said artificial intelligence server prior to obtaining said code data from said electronic device; A server that communicates with electronic devices and artificial intelligence servers.

11. In paragraph 10, The control unit transmits the past prompt to the artificial intelligence server through the communication unit and receives the past response to the past prompt from the artificial intelligence server, The above prompt includes at least a portion of the above past response, A server that communicates with electronic devices and artificial intelligence servers.

12. In paragraph 1, The above reference information further includes area identification information regarding a question or instruction printed on a medium with which the user is interacting using the electronic device, wherein the area identification information identifies a preset area within the medium; The control unit obtains artificial context information corresponding to the question or instruction printed in the preset area from the memory based on the reference information, The above prompt further includes the artificial context information, A server that communicates with electronic devices and artificial intelligence servers.

13. In paragraph 1, The above control unit, Obtaining a subsequent voice text corresponding to the subsequent voice of the user obtained by the electronic device; -At this time, the above subsequent voice text is obtained by converting the above subsequent voice into STT-, Based on the above follow-up spoken text, generate a follow-up prompt. Transmitting the above follow-up prompt to the artificial intelligence server via the above communication unit, Receive a follow-up response from the artificial intelligence server to the follow-up prompt transmitted through the communication unit; Obtain a follow-up response voice corresponding to the received follow-up response, Transmitting the above follow-up response voice to the electronic device through the communication unit, A server that communicates with electronic devices and artificial intelligence servers.

14. In paragraph 13, The above follow-up prompt further includes the medium context information, A server that communicates with electronic devices and artificial intelligence servers.

15. In paragraph 1, The above artificial intelligence server includes an LLM (Large Language Model). A server that communicates with electronic devices and artificial intelligence servers.

16. In paragraph 1, The above electronic device, Includes an image sensor, Configured to photograph an area of ​​the medium by the operation of the user and to obtain the code data from the photographed image, A server that communicates with electronic devices and artificial intelligence servers.

17. In paragraph 1, The above electronic device has a pen-shape having one end and the other end, When the user points one end of the electronic device at an area of ​​the medium, the electronic device is configured to photograph the medium using the image sensor arranged on the end. A server that communicates with electronic devices and artificial intelligence servers.

18. A server that communicates with electronic devices and artificial intelligence servers. memory; Department of Communications; and It includes a control unit that controls the above memory and the above communication unit; The above control unit, [1] Obtain reference information encoded in the code data obtained by the above electronic device; - At this time, the reference information is received from the electronic device through the communication unit or the control unit receives the code data from the electronic device through the communication unit and then obtains the code data by decoding the code data. The above reference information relates to the medium through which the user interacts using the electronic device; The above reference information includes ID information of the medium, page information indicated by the electronic device, and location information within the page indicated by the electronic device. [2] Based on the above reference information, the content printed on the medium and / or the content of interest about a part of the content printed on the medium that the user is interested in is acquired from the memory to generate medium context information, [3] Obtain voice data corresponding to the user's voice obtained by the electronic device, [4] Generate a prompt using the acquired medium context information and the voice data. [5] Transmit the above prompt to the artificial intelligence server through the above communication unit, [6] Receive a response from the artificial intelligence server to the transmitted prompt through the communication unit, [7] Transmitting the received response or the response voice corresponding to the response to the electronic device through the communication unit. A server that communicates with electronic devices and artificial intelligence servers.

19. In paragraph 18, The data type of the above medium context information is any of text, sound, image or video. A server that communicates with electronic devices and artificial intelligence servers.

20. In paragraph 18, If the data type of the above medium context information is text and the data type supported by the above artificial intelligence server is sound, The control unit converts the medium context information so that the data type of the medium context information becomes sound before generating the prompt. A server that communicates with electronic devices and artificial intelligence servers.

Citation Information

Patent Citations

  • Electronic pen and program to be used for the same

    JP2009003531A

  • Using language models to generate common sense explanations

    JP2022522712A

  • Electronic Pen based on Artificial Intelligence, System for Playing Contents using Electronic Pen based on artificial intelligence and Method thereof

    KR102082181B1

  • Generative-discriminative language modeling for controllable text generation

    US20210374341A1

  • Methods and apparatus for natural language interface for constructing complex database queries

    US20230315722A1