Apparatus, method, or program for making characters speak in digital signage
The system allows digital signage characters to interact with viewers by generating speech content based on input, addressing the issue of unawareness and information insufficiency, thereby enhancing engagement and reliability.
Patent Information
- Application Number
- JP2025033313
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-12-11
- Estimated Expiration
- 2045-01-07
AI Technical Summary
Digital signage systems often fail to attract the attention of passersby and provide insufficient information due to unawareness of their existence and inability to interact effectively.
A method and system that enables a character on digital signage to speak by receiving viewer input, searching for relevant knowledge data, generating speech content using a generative AI model, and transmitting it through a control device, along with associated metadata, to engage viewers.
Enhances viewer interaction and provides highly reliable information by presenting accurate speech content and metadata, thereby attracting attention and improving engagement.
Smart Images

Figure 0007784217000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an apparatus, method, or program for making a character in digital signage speak. [Background technology]
[0002] Digital signage is a system that uses displays connected to a communications network to convey various information to passersby, and is installed indoors and outdoors in public transportation, public facilities, and various other locations.
[0003] This system can capture the attention of viewers by displaying visually appealing content on a high-performance display and providing content using data collected in real time via a communications network. Summary of the Invention [Problem to be solved by the invention]
[0004] However, in order for digital signage to be fully effective, passersby need to be aware of its existence and actively use it. However, currently, passersby are often unaware of its existence, or even if they are aware of it, are unable to extract sufficient information from it.
[0005] The present invention has been made in view of the above points, and an object of the present invention is to promote the provision of useful information to passersby in digital signage. [Means for solving the problem]
[0006] In order to achieve this object, a first aspect of the present invention is a method for causing a character displayed on a digital signage system comprising a control device connected to a display and a server capable of communicating with the control device to make the character speak, the method comprising the steps of: receiving input from a viewer of the display; searching for one or more pieces of knowledge data related to the input; acquiring the searched one or more pieces of knowledge data; making a request to a generative AI model including an instruction to generate a speech content for the character corresponding to the input, the instruction including all or part of the one or more pieces of knowledge data; acquiring a response to the request including the speech content; and transmitting the speech content and at least a part of the knowledge data or metadata associated therewith to the control device.
[0007] A second aspect of the present invention is a method of the first aspect, wherein the response includes one or more identifiers of one or more pieces of knowledge used to generate the utterance content, and the at least part is at least a portion of metadata associated with knowledge data identified by at least any of the one or more identifiers.
[0008] Furthermore, a third aspect of the present invention is the method of the first aspect, wherein the step of acquiring one or more pieces of knowledge data is a step of acquiring the one or more pieces of knowledge data and metadata associated with each of them.
[0009] Furthermore, a fourth aspect of the present invention is the method of the second aspect, further comprising, after obtaining the response, obtaining metadata associated with knowledge data identified by at least one of the one or more identifiers included in the response.
[0010] A fifth aspect of the present invention is the method according to the third or fourth aspect, wherein each piece of metadata includes a video or a URL relating to the knowledge data with which it is associated.
[0011] A sixth aspect of the present invention is the method according to any one of the first to fifth aspects, wherein the instructions include one or more characteristics of the character.
[0012] A seventh aspect of the present invention is the method of the sixth aspect, wherein the one or more characteristics include at least one of the character's name, age, occupation, likes, personality, and appearance.
[0013] Furthermore, an eighth aspect of the present invention is a method of any one of the first to fifth aspects, wherein the instructions include a description of outputting one or more parameters representing the generated speech content or the emotion of the character uttering it.
[0014] A ninth aspect of the present invention is the method of any one of the first to eighth aspects, wherein the input is a voice input.
[0015] A tenth aspect of the present invention is the method according to any one of the first to eighth aspects, wherein the input is a selection of a button displayed on the display.
[0016] An eleventh aspect of the present invention is a method of any one of the first to tenth aspects, further comprising a step of transmitting to the control device the content of a speech made to a passerby in front of the display, and the step of receiving input from the viewer is a step of receiving input from the passerby in response to the speech.
[0017] A twelfth aspect of the present invention is a method of the eleventh aspect, further comprising the steps of receiving video captured by a camera connected to the control device, analyzing the video to detect the presence of a passerby, and extracting characteristics of the detected passerby, wherein the content of the speech spoken in the call is in accordance with the characteristics.
[0018] A thirteenth aspect of the present invention is the method of the eleventh aspect, further comprising the step of sending a notification to the control device according to the characteristic.
[0019] A fourteenth aspect of the present invention is a program for causing a server constituting digital signage, the server comprising a control device connected to a display and a server capable of communicating with the control device, to execute a method for causing a character displayed on the display to speak, the method including the steps of receiving input from a viewer of the display, searching for one or more knowledge data related to the input, acquiring the searched one or more knowledge data, making a request to a generative AI model including an instruction to generate speech content for the character corresponding to the input, the instruction including all or part of the one or more knowledge data, acquiring a response to the request including the speech content, and transmitting the speech content and at least a part of the knowledge data or metadata associated therewith to the control device.
[0020] A fifteenth aspect of the present invention is a server constituting a digital signage system comprising a control device connected to a display and a server capable of communicating with the control device, the server receiving an input from a viewer of the display on which a character is displayed, searching for one or more knowledge data related to the input, obtaining the one or more knowledge data, and generating a speech content of the character corresponding to the input, the instruction being configured to make a request including all or part of the one or more knowledge data to a generative AI model, obtaining a response including the speech content, and transmitting the speech content and at least a part of the knowledge data or metadata associated therewith used to generate the speech content to the control device. [Effects of the Invention]
[0021] According to one aspect of the present invention, in addition to the speech by the character in response to input from the viewer, by presenting the knowledge data used to generate the speech content or the metadata previously associated with it as reliable information, it is possible to provide highly reliable information while strongly attracting the viewer's interest through interaction with the character. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a diagram illustrating a digital signage system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a diagram showing a flow of a method for using digital signage according to the first embodiment of the present invention. [Figure 3] FIG. 2 is a diagram showing an example of a format of knowledge data according to the first embodiment of the present invention. [Figure 4] FIG. 3 is a diagram showing an example of a prompt according to the first embodiment of the present invention. [Figure 5] FIG. 4 is a diagram showing an example of a response of a character according to the first embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing a call flow according to the second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0024] (First embodiment) FIG. 1 shows a system according to a first embodiment of the present invention. The device 100 communicates with a control device 110 connected to a display 111 and a platform 120 that provides an AI model via an IP network such as the Internet. The device 100 can also communicate with a database 104 in which customer data collected from customers who use the service provided by the device 100 (hereinafter also referred to as a "digital signage service") is registered, and acquire required data. The control device 110 is further connected to sensors 112 such as a camera and a microphone. The sensors 112 connected to the control device 110 may be integrated in whole or in part into the display 111. While the AI model is described as being provided by a platform 120 that can communicate with the device 100, an application for providing the AI model may be executed on the device 100, allowing the device 100 to provide the AI model.
[0025] The device 100 includes a communication unit 101 such as a communication interface, a processing unit 102 such as a processor or CPU, and a storage unit 103 including a storage device or storage medium such as a memory or hard disk, and can be configured by executing a program for performing each process or operation in the processing unit 102. The device 100 may include one or more devices, computers, or servers. The program may include one or more programs, and may be recorded on a computer-readable storage medium to form a non-transitory program product. The program may be stored in a storage device or storage medium such as the storage unit 103 or a database 104 accessible from the device 100 via an IP network, and instructions included in the program may be executed by at least one processor of the processing unit 102. Data described below as being stored in a storage device or storage medium such as the database 104 may also be stored in the storage unit 103, and vice versa.
[0026] We will explain the flow of information provision that is triggered when a passerby watches the display 111 and makes an input. First, we will explain the registration of customer data in the database 104 (S201), and then we will explain the flow of information provision in this embodiment.
[0027] A customer of a digital signage service provides the operator of the service with customer data that the customer wishes to reflect in communications with passersby. This provision may be made by any method, such as by entering the customer data through a webpage provided by the device 100 or by the customer's representative electronically providing the data to the operator's representative. As an example, the database 104 may be a vector database, and the device 100 may register the customer data in the database 104 by requesting the database 104 to create an index of the provided customer data. In this case, the customer data is divided into units called "chunks" and stored in the database 104. However, the registration of customer data in the database 104 is not limited to indexing in a vector database. Ultimately, the database 104 may register multiple pieces of "knowledge data," which represent knowledge about the customer or their products or services. For example, the provided customer data itself may consist of multiple pieces of knowledge data. Each piece of knowledge data may be text data, image data, or a URL.
[0028] Each piece of knowledge data registered in the database 104 can be associated with metadata. Each piece of metadata includes a description of the associated piece of knowledge data and at least one of a video related to the knowledge data or a URL for the video. When indexing customer data, the database 104 acquires the customer data and the associated metadata, and can associate the metadata with each of multiple pieces of knowledge data based on the customer data. The video can be an image, a video, etc.
[0029] Thereafter, the device 100 receives input from a viewer of the display 111 on which the character is displayed (S202). The input may be an input to the display 111, which is a touch panel, such as selecting a button displayed on the display 111, a voice input to the microphone 112, a gesture to the camera 112, or the like. As an example, the reception of voice by the microphone 112 may be started by selecting an "voice input" button displayed on the display 111. The input is often a question, but may not be in the form of a question or may not be a question in content.
[0030] Next, the device 100 searches for one or more pieces of knowledge data related to the input (S203). The device 100 may use the input directly to search the database 104, or may convert the input into data suitable for the search. As an example, the input may be translated into a language such as English or Japanese. The search can be performed by creating a query depending on the type of database 104 and inputting the query into the database 104. If the database 104 is a vector database, the search extracts one or more chunks of knowledge data that are relatively highly related to the input from among the registered pieces of knowledge data.
[0031] The device 100 acquires one or more pieces of searched knowledge data, and if they are registered, may further acquire the respective metadata (S204). The knowledge data may be in a format as shown in FIG. 3, including associated metadata. In the example of FIG. 3, an identifier such as a knowledge number is set in the variable {index}, the title of the knowledge is set in the variable {title}, a description of the knowledge is set in the variable {description}, a URL of an image, video, or other video corresponding to the knowledge is set in the variable {url}, and the content of the knowledge is set in the variable {knowledge}. The knowledge data includes the identifier and content, and may further include at least one of a title, a description, and a URL. The title, description, and URL may be included in metadata in a format separate from the knowledge data, as necessary. For example, the knowledge identifier can be included in the metadata to be associated with the knowledge data.
[0032] Then, device 100 makes a request to the generative AI model (hereinafter also referred to as "first generative AI model") including an instruction to generate utterance content of the character displayed on display 111 corresponding to the input (S205). More specifically, device 100 can create an instruction to generate utterance content of the character corresponding to the input, and make a request (hereinafter also referred to as "first request") including the instruction to the first generative AI model.
[0033] As used herein, an "AI model" refers to a machine learning model trained to predict an output for a given input, and a "generative AI model" refers to a large-scale language model (LLM) trained to generate an output for a given input that is not included in the input. While an LLM employing the Transformer architecture is particularly preferred as a generative AI model, the name of the architecture is expected to change as technology advances. Therefore, as used herein, the term "Transformer architecture" encompasses architectures that utilize one or more features of the Transformer architecture or improvements thereof. While the example in Figure 2 illustrates a single API call to fulfill a desired request, it is also possible to split the request into multiple requests and fulfill them with multiple API calls. In this case, the multiple API calls may be directed to the same or different generative AI models. As used herein, whether "generative AI models" are identical is determined by whether the type of generative AI model specified by the user is the same. In the case of the OpenAI API, if the value of the variable "model" is the same, the generative AI models are considered to be identical. If multiple generative AI models are not identical, they may be provided on the same platform 120. If multiple requests are directed to a generative AI model provided on the same platform 120, the requests may be fulfilled by the execution of a single code.
[0034] FIG. 4 shows an example of an instruction according to the first embodiment of the present invention. In the example of FIG. 4, a viewer input acquired by the device 100 is set in the variable {input}, and the input is described in the instruction. Data obtained by processing the viewer input, i.e., data based on the input, may be set as the value of the variable. In addition, one or more pieces of knowledge data acquired from the database 104 may be described in the instruction as many times as necessary, for example, in the format described with reference to FIG. 3. Some, but not all, of the one or more pieces of knowledge data acquired from the database 104 may be included in the instruction. FIG. 4 shows an example in which two pieces of knowledge are provided, and specific knowledge content is set in the variables {knowledge_1} and {knowledge_2}. Furthermore, if there is knowledge that is to be provided regardless of the viewer input, knowledge data representing the knowledge may be included in the instruction in advance. The instructions may include one or more characteristics of the character to be displayed on the display, such as the name, age, occupation, likes, personality, appearance, etc.
[0035] The instructions generated by device 100 describe the steps of: considering a response by a character in response to a user input; and, if one or more pieces of knowledge are used in the response, outputting the identifiers of the one or more pieces of knowledge along with the response. More generally, the instructions generated by device 100 may include a description of outputting the character's utterance corresponding to the user input and one or more identifiers of the one or more pieces of knowledge used to generate the utterance. The instructions can be generated by storing code for generating instructions for the first generative AI model in memory 103 and executing the code by device 100. Then, code for a first request to the first generative AI model can be stored in memory 103 and executed by device 100 to call the OpenAI API and issue a request including the generated instructions to an AI model provided on platform 120. The code for generating instructions for the first generative AI model may be included as part of the code for the first request to the first generative AI model. The Open API described above is an example, and other APIs may be used.
[0036] The instructions included in the first request may include a description of generating and outputting one or more parameters that represent the generated speech content or the emotions, such as joy, anger, sadness, or happiness, of the character uttering the speech. By outputting such one or more parameters, it is possible to adjust the speed, pitch, volume, and other tones of the character's final speech displayed on display 111.
[0037] Additionally, the instructions may include a statement that if input by a viewer of display 111 includes a desired language, the speech content is to be output in that language.
[0038] Next, the device 100 obtains a response to the first request from the platform 120 (S206), and for at least one of one or more identifiers included in the response, identifies metadata associated with knowledge data identified by the identifier (S207).The device 100 then transmits the utterance content and at least a portion of the identified metadata to the control device 110 (S208).
[0039] If the first request contains only one piece of knowledge, the knowledge used to generate the utterance content is uniquely determined even if the response from the first generative AI model does not contain an identifier. Also, if all of one or more pieces of knowledge included in the first request are used, the knowledge used to generate the utterance content is uniquely determined even if the response from the first generative AI model does not contain an identifier. As such, since it may be possible to identify metadata even if an identifier is not explicitly included in the response, the response from the first generative AI model does not necessarily have to include an identifier.
[0040] In the above explanation, an example was given in which when one or more pieces of knowledge data are searched for in the database 104, metadata associated with each piece of knowledge data is also obtained. However, after a response is obtained from the first generative AI model, metadata associated with knowledge data identified by at least one of one or more identifiers included in the response may also be obtained.
[0041] Then, based on the received speech content, control device 110 causes a speaker (not shown) connected to control device 110 to output the speech of the character displayed on display 111, and also displays on display 111 the description included in the received metadata and the image included in the metadata or an image based on the URL included in the metadata (S209). The synthesis and playback of voice based on the speech content, and the movement of the character in response to the playback of voice, can be performed appropriately using known techniques. In addition to playing back voice based on the speech content, the speech content may also be displayed on display 111. The speaker may be directly or indirectly connected to control device 110, and may be provided in display 111 or camera 112, for example.
[0042] 5 shows an example of a display screen during speech according to the first embodiment of the present invention. In this example, a display screen 500 displays, in addition to a character 510, speech content 520, a description 531, and an image 532.
[0043] Thus, according to one embodiment of the present invention, in addition to the character's speech in response to input from the viewer, by presenting metadata that has been previously associated with the knowledge data used to generate the speech content as highly reliable and accurate information, it is possible to provide highly reliable information while strongly attracting the viewer's interest through interaction with the character.
[0044] Furthermore, since knowledge data based on data collected from customers can be said to be highly reliable in itself, even if the device 100 transmits at least a portion of the generated utterance content and one or more pieces of knowledge used to generate the utterance content to the control device 110, it is possible to provide highly reliable information while achieving the effect of improving engagement through interaction with characters.
[0045] The device 100 may measure the duration of a viewer's stay on the display 111 as a measure of the effectiveness of an interaction with a character. The measurement may start when the viewer makes an input, when the viewer enters the camera 112's range of capture, when the viewer stops in front of the display 111, etc. The measurement may end when the viewer leaves the camera 112's range of capture, when the viewer starts to move away from in front of the display 111, etc.
[0046] (Second embodiment) The digital signage according to the first embodiment can be introduced in various environments, such as airports, train stations and other public transportation facilities, outdoors in restaurants, retail stores and other stores, tourist spots, etc. If passersby in front of the display 111 are not aware of its presence, the use of the digital signage will not be promoted, so the system according to this embodiment calls out to passersby.
[0047] First, the device 100 receives video captured by the camera 112 from the control device 110 (S601). Then, the device 100 analyzes the video to detect the presence of a passerby (S602). This analysis can be performed using, for example, an AI model, and may be executed on the device 100 or on a server (not shown) that can communicate with the device 100. As an example, this detection may be performed by detecting a person's face.
[0048] When the device 100 detects the presence of a passerby, it extracts the characteristics of the passerby as necessary (S603). The extraction can be performed using, for example, an AI model, and may be executed on the device 100 or on a server (not shown) that can communicate with the device 100. The server may be the same as the server for video analysis described above. The extracted characteristics are Examples of characteristics that can be used include age, sex, clothing, facial expression, and emotion.
[0049] Then, the device 100 generates an utterance content to call out to the passerby using, for example, a generative AI model, and transmits the generated utterance content to the control device 110, which outputs the utterance content to a speaker as the speech of the character displayed on the display 111 (S604). The utterance content may correspond to the extracted characteristics. The display 111 may display a notification that corresponds to the extracted characteristics and that the control device 110 received from the device 100.
[0050] When a passerby makes an input in response to the utterance, the device 100 can start the processing or operation according to the first embodiment in response to the input.
[0051] Thus, according to one embodiment of the present invention, the character actively speaks to passersby to attract their attention, turning them into viewers and making more people aware of the existence of the digital signage.
[0052] It should be noted that in the above embodiments, unless the word "only" is used, such as "based only on," "only in accordance with," "only in the case of," or "with reference only," this specification assumes that additional information may be taken into consideration. Also, as an example, it should be noted that the phrase "do b when a" does not necessarily mean "always do b when a" or "do b immediately after a" unless explicitly stated otherwise. Furthermore, the phrase "each a constituting A" does not necessarily mean that A is composed of multiple components, but includes the case where the component is singular.
[0053] It should be noted that the disclosure of this specification includes any combination of the above-described embodiments of the present invention within the scope of not contradicting each other.
[0054] Also, just to be clear, even if there is an aspect of a method, program, terminal, device, server, or system (hereinafter referred to as a "method, etc.") that performs an operation different from that described in this specification, each aspect of the present invention is directed to an operation that is identical to one of the operations described in this specification, and the existence of an operation different from that described in this specification does not make the method, etc. outside the scope of each aspect of the present invention.
[0055] Furthermore, the "start" and "end" shown in Figure 6 are merely examples, and do not necessarily mean that the method according to this embodiment necessarily starts or ends in the illustrated procedure. [Explanation of symbols]
[0056] 100 devices 101 Communications Department 102 Processing section 103 Storage section 104 Database 110 Control equipment 111 Display 112 Sensors 120 Platform 500 display screen 510 characters 520 Speech content 531 Description 532 images
Claims
1. A method for causing a character displayed on a display to speak, the method comprising: a control device connected to a display; and a server capable of communicating with the control device, the server configuring digital signage, the method comprising: receiving input by a viewer of said display; retrieving one or more pieces of knowledge data related to said input; acquiring one or more pieces of searched knowledge data; making a request to a generative AI model, the request including an instruction to generate an utterance content of the character corresponding to the input, the instruction including one or more characteristics of the character and all or part of the one or more knowledge data; obtaining a response to the request, the response including the utterance content; transmitting the utterance content and at least a portion of the knowledge data used to generate the utterance content, or the utterance content and at least a portion of the metadata associated with the knowledge data used to generate the utterance content, to the control device; Includes:
2. 10. The method of claim 1, The response includes one or more identifiers of one or more pieces of knowledge used to generate the utterance.
3. 10. The method of claim 1, The step of acquiring one or more pieces of knowledge data is a step of acquiring the one or more pieces of knowledge data and metadata associated with each of the pieces of knowledge data.
4. 4. A method according to any one of claims 1 to 3, comprising: Each piece of metadata contains the image or its URL for the knowledge data with which it is associated.
5. 4. A method according to any one of claims 1 to 3, comprising: The instructions include a description of outputting one or more parameters that represent the generated speech content or the emotion of the character that utters it.
6. 4. A method according to any one of claims 1 to 3, comprising: The method further includes a step of transmitting to the control device a speech content to be spoken to a passerby in front of the display, The step of receiving the input from the viewer is a step of receiving an input from the passerby in response to the call.
7. A program for causing a server constituting digital signage, the server comprising a control device connected to a display and capable of communicating with the control device, to execute a method for making a character displayed on the display speak, the method comprising: receiving input by a viewer of said display; retrieving one or more pieces of knowledge data related to said input; acquiring one or more pieces of searched knowledge data; making a request to a generative AI model, the request including an instruction to generate an utterance content of the character corresponding to the input, the instruction including one or more characteristics of the character and all or part of the one or more knowledge data; obtaining a response to the request, the response including the utterance content; transmitting the utterance content and at least a portion of the knowledge data used to generate the utterance content, or the utterance content and at least a portion of the metadata associated with the knowledge data used to generate the utterance content, to the control device; Includes:
8. The server constituting the digital signage includes a control device connected to a display and a server capable of communicating with the control device, receiving input by a viewer of said display on which the character is displayed; searching for one or more pieces of knowledge data related to the input to obtain the one or more pieces of knowledge data; a request including an instruction to generate an utterance content of the character corresponding to the input, the instruction including one or more characteristics of the character and all or part of the one or more knowledge data, to a generative AI model, and obtaining a response including the utterance content; The utterance content and at least a portion of the knowledge data used to generate the utterance content, or the utterance content and at least a portion of the metadata associated with the knowledge data used to generate the utterance content, are configured to be transmitted to the control device.
Citation Information
Patent Citations
Information output device, method, and program
JP2019128910A
Advertisement output device, learning device, advertisement method, and program
JP2021033603A
Information processing system, information processing device, information processing method, and program
JP2023007309A
JPP7529204B