Information processing apparatus, information processing system, and information processing method
The information processing device automates the generation of inquiries for display device malfunctions by capturing images and using a machine learning model, thereby simplifying user interaction and reducing documentation burden.
Patent Information
- Application Number
- JP2024119443
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2026-02-05
- Estimated Expiration
- 2044-07-25
AI Technical Summary
Users are burdened with the task of documenting display device malfunctions, which complicates the resolution process.
An information processing device that captures a display image, generates an image description, and uses a machine learning model to provide candidate inquiries based on the image and user context, reducing the need for user documentation.
Reduces user burden by automatically generating and presenting potential inquiries related to display device malfunctions, facilitating easier troubleshooting.
Smart Images

Figure 2026018234000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing system, and an information processing method. [Background technology]
[0002] Conventionally, answers to inquiries from users have been automatically presented. For example, Patent Document 1 discloses a device that, when a new question sentence is input by a user, presents past answers to the past question sentence that are associated with past questions similar to the new question sentence. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2024-9724 Summary of the Invention [Problem to be solved by the invention]
[0004] When a user inquires about a malfunction of a display device and resolves the malfunction, the user is required to write a document describing the situation at the time of the malfunction, which places a burden on the user.
[0005] The present invention has been made in consideration of these points, and aims to reduce the burden on users when an event such as a malfunction occurs. [Means for solving the problem]
[0006] An information processing device according to a first aspect of the present invention includes a first acquisition unit that acquires a capture image that captures a display image displayed by a display device; a generation unit that generates an image description that explains the display content of the capture image based on the capture image acquired by the first acquisition unit; a second acquisition unit that inputs the image description generated by the generation unit as explanatory information from a user to a machine learning model that outputs an answer to explanatory information regarding the status of the display device input from the user, and acquires an answer corresponding to the image description from the machine learning model; and an output unit that outputs the answer acquired by the second acquisition unit.
[0007] The second acquisition unit may input the image description generated by the generation unit to the machine learning model, which outputs one or more candidate inquiries from the user corresponding to the explanatory information as the answer to the input of the explanatory information, and acquire one or more candidate inquiries corresponding to the explanatory information from the machine learning model, and the output unit may output the one or more candidate inquiries acquired by the second acquisition unit.
[0008] The second acquisition unit may input the image description generated by the generation unit into a machine learning model that outputs, in response to the input of the explanatory information, one or more candidate inquiries from the user corresponding to the explanatory information and inquiry answers that are answers to the inquiries corresponding to each of the one or more candidate inquiries, and acquire from the machine learning model one or more inquiries corresponding to the image description and inquiry answers corresponding to each of the one or more inquiries.
[0009] The first acquisition unit may acquire contract information of a user of the display device for displaying the display image on the display device, and the second acquisition unit may input the image description generated by the generation unit and the contract information acquired by the first acquisition unit to a machine learning model that outputs the answer in response to input of the description information from the user and the contract information corresponding to the user, and acquire an answer corresponding to the image description from the machine learning model.
[0010] The first acquisition unit may acquire the contract information related to use of the display device of the user or a display control device that causes the display device to display the display image.
[0011] The first acquisition unit may acquire device identification information for identifying the display device or a display control device that displays the display image on the display device, and the second acquisition unit may input the image description generated by the generation unit and the device identification information acquired by the first acquisition unit to a machine learning model that outputs the answer in response to input of the description information from a user and the device identification information corresponding to the user, and acquire an answer corresponding to the image description from the machine learning model.
[0012] The first acquisition unit may acquire operation log information indicating an operation log until the display image is displayed on the display device, or an operation log of a display control device that displays the display image on the display device until the display image is displayed on the display device, and the second acquisition unit may input the image description generated by the generation unit and the operation log information acquired by the first acquisition unit to a machine learning model that outputs the answer in response to input of the explanation information from a user and the operation log information corresponding to the user, and acquire an answer corresponding to the image description from the machine learning model.
[0013] The second acquisition unit may acquire a partial description from the image description, which is a description regarding the display device or a display control device that displays the display image on the display device, input the partial description acquired from the image description generated by the generation unit to the machine learning model, and acquire an answer corresponding to the partial description from the machine learning model.
[0014] The second acquisition unit may identify a position in the captured image corresponding to each of a plurality of partial descriptions included in the image description, and acquire a partial description from the image description that is a description regarding the display device or a display control device that displays the display image on the display device based on the position of each of the plurality of partial descriptions.
[0015] The first acquisition unit may acquire, from the display device or the display control device, a captured image that captures the display image displayed by the display device at the time the malfunction is detected, in response to the display device or the display control device detecting a malfunction in itself, the generation unit may generate the image description in response to the first acquisition unit acquiring the captured image, the second acquisition unit may input the image description to the machine learning model in response to the generation unit generating the image description and acquire an answer corresponding to the image description from the machine learning model, and the output unit may output the answer to the display device in response to the second acquisition unit acquiring the answer.
[0016] The second acquisition unit may input the captured image acquired by the first acquisition unit and the image description generated by the generation unit into a machine learning model that outputs the answer in response to the input of the captured image and the image description, and acquire an answer corresponding to the captured image and the image description from the machine learning model.
[0017] An information processing system according to a second aspect of the present invention is an information processing system having a display device and an information processing device, wherein the display device has a display unit, an image generation unit that generates a capture image by capturing a display image displayed on the display unit, and a transmission unit that transmits the capture image generated by the image generation unit to the information processing device, and the information processing device has a first acquisition unit that acquires the capture image, and a generation unit that generates an image caption that explains the display content of the capture image based on the capture image acquired by the first acquisition unit, a second acquisition unit that inputs the image caption generated by the generation unit as explanatory information from the user to a machine learning model that outputs an answer to the explanatory information input from the user regarding the status of the display device, and acquires an answer corresponding to the image caption from the machine learning model, and an output unit that outputs the answer acquired by the second acquisition unit to the display device, and the display device may further have a display control unit that displays the answer output from the information processing device on the display unit.
[0018] The image generating section may generate the capture image in response to a predetermined operation being performed by an operation section that operates the display device.
[0019] An information processing method according to a third aspect of the present invention is executed by a computer and includes the steps of acquiring a capture image by capturing a display image displayed by a display device, generating an image description explaining the display content of the captured image based on the acquired capture image, inputting the generated image description as explanatory information from a user to a machine learning model that outputs an answer to explanatory information regarding the status of the display device input by a user, acquiring an answer corresponding to the image description from the machine learning model, and outputting the acquired answer. [Effects of the Invention]
[0020] The present invention has the effect of reducing the burden on the user when an event such as a malfunction occurs. [Brief explanation of the drawings]
[0021] [Figure 1] FIG. 1 is a diagram illustrating an overview of an information processing system. [Figure 2] FIG. 2 is a diagram illustrating a functional configuration of an information processing device. [Figure 3] FIG. 10 is a diagram illustrating an example of generating an image description. [Figure 4] FIG. 10 is a diagram showing query candidates acquired from a second generation AI (Artificial Intelligence). [Figure 5] FIG. 10 is a diagram showing an example in which candidate information is superimposed on a display image displayed on a display device. [Figure 6] FIG. 2 is a diagram illustrating a functional configuration of a set-top box. [Figure 7] FIG. 2 is a sequence diagram showing a processing flow in the information processing system. DETAILED DESCRIPTION OF THE INVENTION
[0022] [Outline of Information Processing System S] 1 is a diagram showing an overview of an information processing system S. The information processing system S has an information processing device 1, a set-top box (hereinafter referred to as "STB") 2, and a display device 3, and is a system that, when a malfunction or the like occurs in the display device 3 or the STB 2, outputs to the user suggestions for inquiries to be made to customer support for the STB 2 regarding the malfunction or the like.
[0023] When a malfunction or the like occurs in the STB 2 or the display device 3, the user performs a predetermined operation, such as pressing a support button, on the remote controller 24 of the STB 2. When the STB 2 receives the predetermined operation, it generates a capture image by capturing the display image displayed on the display device 3 ((1) in FIG. 1), and transmits the generated capture image to the information processing device 1 ((2) in FIG. 1).
[0024] When the information processing device 1 acquires a capture image from the STB 2, it generates an image caption that explains the display content of the capture image based on the capture image ((3) in FIG. 1). The information processing device 1 stores a machine learning model that has been trained in advance using training data and that outputs one or more query candidates as a response to an input of explanatory information from a user regarding the status of the display device. The information processing device 1 inputs the generated image caption into the machine learning model as explanatory information from the user, acquires one or more query candidates as a response corresponding to the image caption from the machine learning model ((4) in FIG. 1), and outputs candidate information indicating the acquired query candidates to the STB 2 ((5) in FIG. 1).
[0025] When the STB 2 acquires the candidate information, it displays the inquiry candidates indicated by the acquired candidate information on the display device 3 ((6) in FIG. 1). This allows the user to select the content of their inquiry from the inquiry candidates indicated by the candidate information displayed on the display device 3 and make an inquiry when a malfunction or the like occurs with the display device 3 or the STB 2, without having to create a message indicating the content of the inquiry to be made to customer support or explain the malfunction or other event to customer support. Therefore, the information processing system S can reduce the burden on the user when a malfunction or the like occurs with the display device 3 or the STB 2.
[0026] [Functional configuration of information processing device 1] Next, the functions of the information processing device 1 and STB 2 of the information processing system S will be described. First, the functional configuration of the information processing device 1 will be described. FIG. 2 is a diagram showing the functional configuration of the information processing device 1. The information processing device 1 has a communication unit 11, a storage unit 12, and a control unit 13.
[0027] The communication unit 11 is a communication interface for transmitting and receiving data to and from an external device such as the STB 2 via a communication network such as the Internet or a mobile phone line. The storage unit 12 is a storage medium that stores various types of data, and includes a ROM (Read Only Memory), a RAM (Random Access Memory), a hard disk, etc. The storage unit 12 stores programs that are executed by the control unit 13. The storage unit 12 stores programs that cause the control unit 13 to function as a
[0028] The control unit 13 is, for example, a CPU (Central Processing Unit). The control unit 13 executes a program stored in the storage unit 12, thereby functioning as a first acquisition unit 131, a generation unit 132, a second acquisition unit 133, an output unit 134, and a third acquisition unit 135.
[0029] The first acquisition unit 131 acquires a capture image obtained by capturing a display image displayed by the display device 3. For example, the first acquisition unit 131 acquires, from the STB 2, the capture image captured by the STB 2 and an STB ID (Identification) for identifying the STB 2. For example, the storage unit 12 stores STB user information that associates the STB ID with contract information of a user regarding use of the STB 2 as contract information for displaying a display image on the display device 3, and the first acquisition unit 131 references the STB user information and acquires contract information of the user of the display device 3 that corresponds to the acquired STB ID.
[0030] The user's contract information regarding the use of the STB2 includes contract information regarding the rental of the STB2 to the user from the service provider, contract information regarding the video distribution service provided by the STB2, and contract information for displaying a channel specified by the user from among multiple pay channels distributed via the STB2 on the display device 3. Note that the contract information for displaying a display image on the display device 3 may also include contract information regarding the use of an Internet line.
[0031] The generation unit 132 generates an image caption that explains the display content of a capture image based on the capture image acquired by the first acquisition unit 131. The generation unit 132 generates an image caption corresponding to the capture image in response to the capture image being acquired by the first acquisition unit 131. For example, the generation unit 132 inputs the capture image acquired by the first acquisition unit 131 to a first generation AI that accepts input of a capture image and outputs an image caption corresponding to the capture image, and acquires an image caption from the first generation AI, thereby generating an image caption corresponding to the capture image.
[0032] The first generation AI is realized by executing a program related to the first generation AI stored in the storage unit 12, but is not limited to this. The first generation AI may be provided by an external device different from the information processing device 1. In this case, the generation unit 132 accesses the external device and acquires an image caption corresponding to the captured image, thereby generating an image caption corresponding to the captured image.
[0033] Figure 3 shows an example of generated image captions. Figure 3(a) shows an example of a captured image, and Figure 3(b) shows the image caption output from the first generation AI when the captured image shown in Figure 3(a) is input to the first generation AI. As shown in Figure 3(b), it can be seen that an image caption corresponding to the captured image shown in Figure 3(a) has been generated.
[0034] The second acquisition unit 133 inputs the image caption generated by the generation unit 132 as explanatory information about the status of the display device 3 from the user to a machine learning model that outputs one or more candidate inquiries from the user corresponding to the explanatory information as an answer to the explanatory information in response to input of explanatory information about the status of the display device from the user, and acquires one or more candidate inquiries as an answer corresponding to the image caption from the machine learning model. In response to the generation unit 132 generating the image caption, the second acquisition unit 133 inputs the image caption to the machine learning model and acquires one or more candidate inquiries corresponding to the image caption from the machine learning model.
[0035] The machine learning model is a second generation AI that outputs one or more candidate inquiries in response to input of an image caption corresponding to a captured image and user contract information for displaying a display image corresponding to the captured image on the display device 3. For example, the second generation AI is assumed to have previously trained using a data set in which questions corresponding to FAQs regarding the display control device 2 are input data and at least some of the answers to those questions are output data as training data. The second generation AI is not limited to this, and may also be trained using training data in which explanatory information regarding the status of the display device previously received from a user is input data and candidate inquiries corresponding to that explanatory information are output data. Furthermore, the second generation AI may additionally train using an image caption corresponding to a captured image as input data and inquiry information selected by the user from one or more pieces of inquiry information output in response to the image caption as output data.
[0036] In addition, the second generation AI may learn from a dataset that uses explanatory information and contract information regarding the status of the display device received from the user as input data and inquiries from the user regarding the display device corresponding to the captured image as output data as training data.
[0037] Here, the second generation AI is realized by executing a program related to the second generation AI stored in the storage unit 12, but is not limited to this. The second generation AI may be provided by an external device different from the information processing device 1. Furthermore, when the second generation AI is provided by an external device, the second generation AI may be a general-purpose natural language processing model such as a large-scale language model. In this case, the second acquisition unit 133 accesses the external device and acquires one or more query candidates corresponding to the image caption, thereby acquiring one or more query candidates corresponding to the image caption.
[0038] To acquire one or more query candidates corresponding to the image caption, the second acquisition unit 133 first generates input information to be input to the second generation AI, including the contract information and the image caption. For example, the second acquisition unit 133 generates the input information shown below.
[0039] A customer who has been subscribing to the unlimited viewing plan since 2019 is experiencing problems with the following screen when using the video content service. Please list possible issues in the FAQ. This is the content selection screen for a video streaming service. In the center, information about the currently focused anime film "Sunny Child" is displayed prominently, along with details such as the rating, synopsis, and voice actors.
[0040] Of the information entered above, "Subscribed to unlimited viewing since 2019" and "Use of video content service" are based on the user's contract information. Furthermore, of the information entered above, the following text is included in the image description: "This is the content selection screen for the video streaming service. Information about the currently focused animated film, "Sunny Child," is displayed prominently in the center, along with details such as ratings, plot summary, and voice actors."
[0041] Then, the second acquisition unit 133 inputs input information including the image description generated by the generation unit 132 and the contract information acquired by the first acquisition unit 131 to the second generation AI, and acquires one or more query candidates corresponding to the image description and contract information from the second generation AI. Figure 4 is a diagram showing the query candidates acquired from the second generation AI. In Figure 4, it can be seen that multiple query candidates are displayed.
[0042] The content of the inquiry includes inquiries such as "I can't view the STB-related service even though I should have subscribed to it," or "I can't subscribe to the service." In response to this, the information processing device 1 inputs contract information related to the use of the STB 2 into the second generation AI and can acquire inquiry candidates that take into account the contract information, so that it is possible to acquire inquiry candidates that are likely to be answered by the user in accordance with the contract status of the STB 2 user.
[0043] Note that the second generation AI is configured to output one or more candidate inquiries from users regarding the display device corresponding to the captured image in response to input information including an image description corresponding to the captured image and contract information, but this is not limited to this.
[0044] The second generation AI may output one or more candidate queries in response to input of explanatory information about the status of the display device received from the user and the model number of the STB as device identification information for identifying the STB. In this case, the second generation AI is assumed to have previously learned from a dataset using as training data a question corresponding to an FAQ or explanatory information about the status of the display device previously received from a user and the model number of the STB as input data, and an inquiry from the user about the display device corresponding to the question or explanatory information as output data.
[0045] For example, the first acquisition unit 131 acquires from the STB 2 a capture image and the model number of the STB 2 for identifying the STB 2. Then, the second acquisition unit 133 inputs the image description generated by the generation unit 132 and the model number of the STB 2 acquired by the first acquisition unit 131 to the second generation AI, and acquires one or more query candidates corresponding to the image description and the model number of the STB 2 from the second generation AI.
[0046] Since malfunctions and other issues vary depending on the model number of the STB, the information processing device 1 can input the model number of the STB2 into the second generation AI and obtain candidate inquiries about malfunctions and other issues that correspond to the STB2 of that model number, thereby obtaining candidate inquiries that are likely to correspond to the STB2 used by the user.
[0047] Furthermore, the second generation AI may output one or more candidate queries in response to input of explanatory information from the user regarding the status of the display device and operation log information indicating the operation log in the STB 2 until the display image corresponding to the captured image is displayed on the display device 3. In this case, the second generation AI is assumed to have previously learned a dataset using as training data a question corresponding to an FAQ or explanatory information regarding the status of the display device previously received from the user, and operation log information indicating the operation log of a predetermined number of STB operations performed until the display image corresponding to the captured image is displayed on the display device, as input data, and an inquiry from the user regarding the display device corresponding to the question or explanatory information as output data.
[0048] For example, the first acquisition unit 131 acquires from the STB 2 a captured image and operation log information indicating a log of operations in the STB 2 performed a predetermined number of times before the display image corresponding to the captured image is displayed on the display device 3, as operation log information indicating a log of operations in the STB 2 until the display image corresponding to the captured image is displayed on the display device 3. Then, the second acquisition unit 133 inputs the image description generated by the generation unit 132 and the operation log information acquired by the first acquisition unit 131 to the second generation AI, and acquires one or more query candidates corresponding to the image description and operation log information from the second generation AI.
[0049] As a result of operations being performed until a display image corresponding to the captured image is displayed, there is a high probability that the display image is displayed, and the operation log indicating the operations is highly likely to indicate the cause of the occurrence of a malfunction or the like at the time the display image was displayed. In response to this, the information processing device 1 inputs the operation log of the STB 2 until the display image corresponding to the captured image is displayed to the second generation AI, and can acquire candidate inquiries for the user, thereby making it possible to acquire candidate inquiries that are highly likely to correspond to the user using the STB 2.
[0050] In addition, the second generation AI may output one or more query candidates in response to input of a partial description that is a description regarding a display device or STB contained in an image description corresponding to a captured image.
[0051] The second acquisition unit 133 acquires partial descriptions, which are descriptions relating to the display device 3 or the STB 2, from the image descriptions generated by the generation unit 132. The second acquisition unit 133 identifies positions in the captured image corresponding to each of the multiple partial descriptions included in the image description, and acquires the partial description corresponding to the display device 3 or the STB 2 from the image description based on the positions of each of the multiple partial descriptions.
[0052] In this case, the generation unit 132 generates an image description including a plurality of partial descriptions. Each partial description is associated with position information indicating a position or area in the captured image corresponding to the partial description. For example, the second acquisition unit 133 acquires, from among the plurality of partial descriptions, a partial description corresponding to a predetermined position or area in the captured image as a partial description corresponding to the display device 3 or the STB 2. The second acquisition unit 133 then inputs the partial description to the second generation AI and acquires one or more inquiry candidates corresponding to the partial description from the second generation AI. Because the partial description is highly likely to indicate a malfunction or the like of the display device 3 or the STB 2, the second acquisition unit 133 can accurately acquire candidates for user inquiry information.
[0053] The second generation AI may also output one or more query candidates in response to input of a capture image and explanatory information about the status of the display device corresponding to the capture image. In this case, the second generation AI is assumed to have previously learned from a dataset in which the capture image and explanatory information about the status of the display device corresponding to the capture image are used as input data and user inquiries about the display device corresponding to the capture image and the explanatory information are used as output data. The second acquisition unit 133 then inputs the capture image acquired by the first acquisition unit 131 and the image caption generated by the generation unit 132 to the second generation AI, and acquires one or more query candidates corresponding to the capture image and the image caption from the second generation AI. In this way, the second acquisition unit 133 can acquire candidate query information for the user while taking into account elements of the capture image that are not included in the image caption.
[0054] Furthermore, the second generation AI only outputs one or more query candidates corresponding to the image description when an image description is input, but this is not limited to this. The second generation AI may also output one or more query candidates corresponding to the image description when an image description is input, and query answers that are answers corresponding to each of the one or more query candidates.
[0055] In this case, the second generation AI is assumed to have previously learned a data set using as training data a dataset in which explanatory information from a user regarding the status of the display device is input data, and inquiries from the user regarding the display device corresponding to the explanatory information and answers to the inquiries are output data. The second acquisition unit 133 then inputs the image caption generated by the generation unit 132 to the second generation AI, and acquires from the second generation AI one or more candidate inquiries corresponding to the image caption and answers to the inquiries corresponding to each of the one or more candidate inquiries.
[0056] In this way, the information processing device 1 can display the inquiry candidates and the inquiry answers corresponding to the inquiry candidates at the same time on the display device 3, allowing the user to understand the answers without having to select the inquiry corresponding to the user from the inquiry candidates.
[0057] The output unit 134 outputs one or more query candidates as answers acquired by the second acquisition unit 133. In response to the second acquisition unit 133 acquiring one or more query candidates, the output unit 134 outputs the one or more query candidates to the display device 3. For example, the output unit 134 outputs candidate information indicating the one or more query candidates acquired by the second acquisition unit 133 to the STB 2. This allows the STB 2 to display the candidate information by superimposing it on a display image corresponding to the captured image displayed on the display device 3.
[0058] 5 is a diagram showing an example in which candidate information is superimposed on a display image displayed on the display device 3. This allows the user to select an inquiry item from the candidate information that corresponds to the status of the display device 3 and STB 2, without the user having to explain the status of the display device 3 and STB 2.
[0059] The third acquisition unit 135 acquires selection information indicating an inquiry item selected by a user from one or more inquiry candidates indicated by the candidate information, and acquires inquiry answer information indicating an answer corresponding to the inquiry item indicated by the acquired selection information.
[0060] For example, the storage unit 12 is provided with a program that functions as a third generation AI, which is a machine learning model that outputs inquiry answer information indicating an inquiry answer corresponding to an input inquiry content corresponding to an inquiry item indicated by the selection information. Here, the third generation AI is assumed to have previously learned from a dataset that uses as training data the inquiry content as input data and the inquiry answer information corresponding to the inquiry content as output data.
[0061] The third acquisition unit 135 executes a program that functions as the third generation AI, and causes the control unit 13 to function as the third generation AI. The third acquisition unit 135 inputs the inquiry content corresponding to the inquiry item indicated by the selection information to the third generation AI, and acquires inquiry response information corresponding to the inquiry content from the third generation AI.
[0062] In response to the third acquisition unit 135 acquiring the inquiry answer information, the output unit 134 outputs the inquiry answer information to the display device 3. For example, the output unit 134 outputs the inquiry answer information acquired by the third acquisition unit 135 to the STB 2. This allows the STB 2 to display the inquiry answer information superimposed on a display image corresponding to the capture image displayed on the display device 3.
[0063] [STB2 function configuration] Next, the functional configuration of the STB 2 will be explained. Here, a description of the general functions of the STB 2 as a set-top box will be omitted, and the functions of the STB 2 according to the present invention will be explained. Fig. 6 is a diagram showing the functional configuration of the STB 2. The STB 2 has a communication unit 21, a storage unit 22, and a control unit 23.
[0064] The communication unit 21 is a communication interface for transmitting and receiving data to and from an external device such as the information processing device 1 via a communication network such as the Internet or a mobile phone line. The storage unit 22 is a storage medium that stores various types of data, and includes a ROM, a RAM, etc. The storage unit 22 stores programs executed by the control unit 23. The storage unit 22 stores programs that cause the control unit 23 to function as an image generation unit 231, a transmission unit 232, a reception unit 233, a display control unit 234, and a selection reception unit 235.
[0065] The control unit 23 is, for example, a CPU. The control unit 23 executes a program stored in the storage unit 22, thereby functioning as an image generation unit 231, a transmission unit 232, a reception unit 233, a display control unit 234, and a selection reception unit 235.
[0066] The image generation unit 231 generates a capture image by capturing a display image displayed on the display device 3 by the STB 2. For example, the remote controller 24 serving as the operation unit of the STB 2 is provided with a support button that allows customer support for the STB 2 to provide support to the user when a problem or the like occurs in the STB 2. The image generation unit 231 detects that the support button on the remote controller 24 has been pressed. In response to the support button being pressed, the image generation unit 231 generates a capture image by capturing a display image displayed on the display device 3.
[0067] In response to the image generation unit 231 generating a capture image, the transmission unit 232 transmits the capture image to the information processing device 1. The transmission unit 232 also transmits an STB ID for identifying the STB2 to the information processing device 1. If the model number of the device is used as input information for the second generation AI, the transmission unit 232 may transmit the model number as device identification information of the STB2 to the information processing device. If operation log information is used as input information for the second generation AI, the transmission unit 232 may transmit operation log information indicating the operation log of the STB2 to the information processing device 1.
[0068] The receiving unit 233 receives candidate information indicating one or more inquiry candidates output from the information processing device 1. The display control unit 234 causes the display device 3 to display one or more inquiry candidates indicated by the candidate information received by the receiving unit 233.
[0069] The selection receiving unit 235 receives, via the remote controller 24, a selection from the inquiry candidates that the display control unit 234 has displayed on the display device 3. In response to the selection receiving unit 235 receiving the selection from the inquiry candidates, the transmission unit 232 transmits selection information indicating the selected inquiry item to the information processing device 1.
[0070] The receiving unit 233 receives inquiry answer information corresponding to the inquiry item indicated by the selection information output from the information processing device 1. The display control unit 234 causes the display device 3 to display the inquiry answer information received by the receiving unit 233.
[0071] [Processing flow in information processing system S] Next, a description will be given of the flow of processing in the information processing system S. FIG. First, in response to pressing the support button on the remote controller 24, the image generation unit 231 generates a capture image by capturing the display image displayed on the display device 3 (S1). Next, the transmission unit 232 transmits the capture image generated by the image generation unit 231 and the STB ID for identifying the STB 2 to the information processing device 1 (S2). The first acquisition unit 131 of the information processing device 1 acquires the capture image and the STB ID from the STB 2.
[0072] Next, the first acquisition unit 131 refers to the management server that stores STBIDs and user contract information in association with each other, and acquires the user contract information corresponding to the acquired STBID (S3). Next, the generation unit 132 generates an image description that explains the display content of the capture image based on the capture image acquired by the first acquisition unit 131 (S4).
[0073] Next, the second acquisition unit 133 inputs input information including an image description corresponding to the captured image and the user's contract information to the second generation AI, and acquires one or more query candidates from the second generation AI (S5). Next, the output unit 134 transmits candidate information indicating the candidate queries acquired by the second acquisition unit 133 to the STB 2 (S6).
[0074] The receiving unit 233 of the STB 2 receives candidate information indicating one or more query candidates output from the information processing device 1. The display control unit 234 causes the display device 3 to display one or more query candidates indicated by the candidate information received by the receiving unit 233 (S7).
[0075] The selection receiving unit 235 receives, via the remote controller 24, a selection from the inquiry candidates that the display control unit 234 has displayed on the display device 3 (S8). In response to the selection receiving unit 235 receiving the selection from the inquiry candidates, the transmission unit 232 transmits selection information indicating the selected inquiry item to the information processing device 1 (S9).
[0076] When the third acquisition unit 135 of the information processing device 1 acquires selection information indicating an inquiry item selected by the user from one or more inquiry candidates indicated by the candidate information, the third acquisition unit 135 acquires inquiry answer information indicating an inquiry answer corresponding to the inquiry item indicated by the acquired selection information (S10). In response to the third acquisition unit 135 acquiring the inquiry answer information, the output unit 134 outputs the answer information to the STB 2 (S11).
[0077] The receiving unit 233 of the STB 2 receives the inquiry response information output from the information processing device 1. The display control unit 234 causes the display device 3 to display the inquiry response information received by the receiving unit 233 (S12).
[0078] [Variation 1] In the above-described embodiment, the STB 2 and the display device 3 are different devices, but this is not limited to this. The display device 3 may have functions related to the STB 2. In this case, the first acquisition unit 131 may acquire contract information related to the user's use of the display device 3 instead of the contract information of the STB 2.
[0079] Furthermore, when the model number of the device is used as input information for the second generation AI, the first acquisition unit 131 may acquire the model number of the display device 3 instead of acquiring the model number of STB2 as the device identification information of STB2, or may acquire two model numbers, the model number of STB2 and the model number of the display device 3. Furthermore, when operation log information is used as input information for the second generation AI, the first acquisition unit 131 may acquire operation log information indicating the operation log up to the display image related to the captured image on the display device 3 instead of acquiring operation log information indicating the operation log of STB2.
[0080] [Variation 2] Furthermore, in the above-described embodiment, the image generation unit 231 of the STB 2 generates a capture image in response to the user pressing the support button on the remote controller 24, and the first acquisition unit 131 of the information processing device 1 acquires the capture image, but this is not limited to this. The image generation unit 231 may generate a capture image in response to detecting a malfunction or the like of the STB 2. Here, malfunctions or the like of the STB 2 may include the signal level of a broadcast signal received by the STB 2 falling below a first threshold, the remaining memory capacity of the STB 2 falling below a second threshold, the signal level of a Wifi (registered trademark) or BLE (registered trademark) signal falling below a third threshold, or the temperature of the CPU of the STB 2 exceeding a fourth threshold.
[0081] Then, the transmission unit 232 transmits the capture image generated in response to the detection of a malfunction or the like of the STB 2 to the information processing device 1. The first acquisition unit 131 of the information processing device 1 acquires the capture image. In this way, the information processing device 1 can automatically display query candidates on the display device 3 when a malfunction or the like of the STB 2 occurs.
[0082] [Effects of information processing device 1] As described above, the information processing device 1 according to this embodiment acquires a captured image obtained by capturing a display image displayed by the display device 3, and generates an image caption that explains the display content of the captured image. The information processing device 1 then inputs the generated image caption to a machine learning model as explanatory information from the user regarding the status of the display device 3, acquires one or more query candidates from the machine learning model as answers corresponding to the image caption, and outputs the acquired one or more query candidates. In this way, the information processing device 1 can reduce the burden on the user when an event such as a malfunction occurs.
[0083] Furthermore, this invention will make it possible to contribute to Goal 9 of the United Nations' Sustainable Development Goals (SDGs), which is "Build resilient infrastructure, promote inclusive and sustainable industrialization, and promote innovation and resilience."
[0084] The present invention has been described above using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, all or part of the device can be configured by functionally or physically distributing or integrating any unit. Furthermore, new embodiments resulting from any combination of multiple embodiments are also included in the embodiments of the present invention. The effects of the new embodiments resulting from the combination also have the effects of the original embodiments. [Explanation of symbols]
[0085] 1. Information processing equipment 2. Set-top boxes 3 Display device 11 Communications Department 12 Storage section 13 Control Unit 21 Communications Department 22 Memory section 23 Control Unit 24 Remote Controller 131 First acquisition part 132 Generation part 133 Second Acquisition Department 134 Output section 135 Third acquisition part 231 Image Generation Unit 232 Transmitter 233 Receiving Unit 234 Display control unit 235 Selection Reception Department
Claims
1. a first acquisition unit that acquires a captured image obtained by capturing a display image displayed by the display device; a generation unit that generates an image caption that explains the display content of the capture image based on the capture image acquired by the first acquisition unit; a second acquisition unit that inputs an image description generated by the generation unit as explanatory information from a user to a machine learning model that outputs an answer to explanatory information input from a user regarding a status of the display device, and acquires an answer corresponding to the image description from the machine learning model; an output unit that outputs the answer acquired by the second acquisition unit; An information processing device having the above.
2. the second acquisition unit inputs the image description generated by the generation unit to the machine learning model, which outputs, as the answer to the input of the description information, one or more candidate inquiries from the user corresponding to the description information, and acquires, from the machine learning model, the one or more candidate inquiries corresponding to the description information; the output unit outputs the one or more query candidates acquired by the second acquisition unit. The information processing device according to claim 1 .
3. the second acquisition unit inputs the image description generated by the generation unit into a machine learning model that outputs, in response to the input of the explanatory information, one or more candidate inquiries from the user corresponding to the explanatory information and inquiry answers that are answers to the inquiries corresponding to each of the one or more candidate inquiries, and acquires, from the machine learning model, the one or more inquiries corresponding to the image description and the inquiry answers corresponding to each of the one or more inquiries; The information processing device according to claim 2 .
4. the first acquisition unit acquires contract information of a user of the display device for displaying the display image on the display device; the second acquisition unit inputs the image description generated by the generation unit and the contract information acquired by the first acquisition unit into a machine learning model that outputs the answer in response to input of the description information from a user and the contract information corresponding to the user, and acquires an answer corresponding to the image description from the machine learning model. The information processing device according to claim 1 .
5. the first acquisition unit acquires the contract information related to use of the display device of the user or a display control device that causes the display image to be displayed on the display device; The information processing device according to claim 4 .
6. the first acquisition unit acquires device identification information for identifying the display device or a display control device that causes the display device to display the display image; the second acquisition unit inputs the image description generated by the generation unit and the device identification information acquired by the first acquisition unit into a machine learning model that outputs the answer in response to input of the description information from a user and the device identification information corresponding to the user, and acquires an answer corresponding to the image description from the machine learning model. The information processing device according to claim 1 .
7. the first acquisition unit acquires operation log information indicating an operation log until the display image is displayed on the display device, or an operation log of a display control device that causes the display image to be displayed on the display device, the operation log indicating an operation log until the display image is displayed on the display device; the second acquisition unit inputs the image description generated by the generation unit and the operation log information acquired by the first acquisition unit into a machine learning model that outputs the answer in response to input of the explanation information from a user and the operation log information corresponding to the user, and acquires an answer corresponding to the image description from the machine learning model; The information processing device according to claim 1 .
8. The second acquisition unit acquires, from the image caption, a partial caption which is a caption relating to the display device or a display control device which causes the display device to display the display image, inputs the partial caption acquired from the image caption generated by the generation unit to the machine learning model, and acquires an answer corresponding to the partial caption from the machine learning model. The information processing device according to claim 1 .
9. The second acquisition unit identifies positions in the captured image corresponding to each of a plurality of partial descriptions included in the image description, and acquires partial descriptions from the image description, which are descriptions related to the display device or a display control device that causes the display device to display the display image, based on the positions of each of the plurality of partial descriptions. The information processing device according to claim 8 .
10. the first acquisition unit acquires, in response to detection of a malfunction of the display device or a display control device that causes the display device to display the display image, from the display device or the display control device, a captured image that captures a display image displayed by the display device at a time when the malfunction is detected; the generation unit generates the image caption in response to the first acquisition unit acquiring the capture image; The second acquisition unit inputs the image description into the machine learning model in response to the generation unit generating the image description, and acquires an answer corresponding to the image description from the machine learning model; the output unit outputs the answer to the display device in response to the second acquisition unit acquiring the answer. The information processing device according to claim 1 .
11. the second acquisition unit inputs the captured image acquired by the first acquisition unit and the image caption generated by the generation unit into a machine learning model that outputs the answer in response to the input of the captured image and the image caption, and acquires an answer corresponding to the captured image and the image caption from the machine learning model; The information processing device according to claim 1 .
12. An information processing system having a display device and an information processing device, The display device includes: A display unit; an image generating unit that generates a captured image by capturing a display image displayed on the display unit; a transmission unit that transmits the capture image generated by the image generation unit to the information processing device; The information processing device includes: a first acquisition unit that acquires the capture image; a generation unit that generates an image caption that explains the display content of the capture image based on the capture image acquired by the first acquisition unit; a second acquisition unit that inputs an image description generated by the generation unit as explanatory information from a user to a machine learning model that outputs an answer to explanatory information input from a user regarding a status of the display device, and acquires an answer corresponding to the image description from the machine learning model; an output unit that outputs the answer acquired by the second acquisition unit to the display device; and The display device includes: a display control unit that displays the answer output from the information processing device on the display unit; Information processing system.
13. the image generation unit generates the capture image in response to a predetermined operation being performed by an operation unit that operates the display device. The information processing system according to claim 12.
14. The computer executes acquiring a captured image by capturing a display image displayed by the display device; generating an image description that explains the display content of the captured image based on the acquired captured image; a step of inputting the generated image description as explanatory information from a user to a machine learning model that outputs an answer to explanatory information input from a user regarding the status of the display device, and acquiring an answer corresponding to the image description from the machine learning model; outputting the obtained answer; An information processing method comprising:
Citation Information
Patent Citations
Dialogue reply candidate proposal system and dialogue reply candidate proposal method
JP2024009724A