Information processing apparatus, information processing system, information processing method, and program

The information processing device addresses the limitation of existing technologies by generating images based on unique participant data, enhancing meeting visualization and understanding.

JP2025141758APending Publication Date: 2025-09-29RICOH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024141960
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-13
Filing Date
2024-08-23
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Existing technologies only search for and output illustrations or images corresponding to linguistic information recognized from speech or text, failing to account for the unique information of each participant in a meeting.

Method used

An information processing device that acquires unique data from participant behavior, generates images using a machine learning model trained on unique information, and displays these images to visualize each participant's contributions.

Benefits of technology

Enables visualization of information specific to each participant, facilitating a deeper understanding of the meeting content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025141758000001_ABST
    Figure 2025141758000001_ABST
Patent Text Reader

Abstract

To enable visualization according to unique information of each of participants in a meeting.SOLUTION: An information processing apparatus includes: an acquisition unit which acquires unique data indicating unique information based on actions at a meeting of each of participants in the meeting; and an image generation unit which generates, for each of the participants, corresponding image corresponding to the unique data acquired by the acquisition unit, based on the unique data acquired by the acquisition unit and a machine learning model trained using learning data including images and the unique data indicating the unique information.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing system, an information processing method, and a program. [Background technology]

[0002] In situations where multiple people participate in a dialogue, such as a meeting or other gathering, graphic recording can be used to represent the content of the discussion in diagrams, thereby increasing common understanding of the content of the gathering.

[0003] Patent Document 1 discloses a communication system, a display device, a display control method, and a display control program that can present appropriate visual information related to information transmission as information transmission progresses through voice or text input, etc., and discloses a language information input means that accepts input of language information, a recognition means that recognizes the input language information, and an image display means that displays an image corresponding to the input language information on a display means that displays images based on the recognition result by the recognition means. Summary of the Invention [Problem to be solved by the invention]

[0004] However, the technology of Patent Document 1 only searches for and outputs illustrations or images that correspond to linguistic information recognized from speech or text.

[0005] The present invention has been made in view of the above points, and has as its object to enable visualization according to the unique information of each of the participants in a meeting. [Means for solving the problem]

[0006] In order to solve the above problem, the information processing device has an acquisition unit that acquires unique data indicating unique information of each participant of a gathering based on the behavior of the participant related to the gathering, and an image generation unit that generates a corresponding image for each participant that corresponds to the unique data acquired by the acquisition unit based on the unique data acquired by the acquisition unit and a machine learning model trained using learning data that includes the unique data indicating the unique information and an image. [Effects of the Invention]

[0007] It is possible to visualize information specific to each participant in a meeting. [Brief explanation of the drawings]

[0008] [Figure 1] 1 is a diagram illustrating an example of a configuration of an information processing system 1 according to a first embodiment. [Figure 2] 1 is a diagram illustrating an example of a hardware configuration of a terminal device 10 according to a first embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a functional configuration of an information processing system 1 according to a first embodiment. [Figure 4] FIG. 2 is a sequence diagram illustrating an example of a processing procedure executed by the information processing system 1 according to the first embodiment. [Figure 5] FIG. 3 is a diagram showing an example of a screen displayed in the first embodiment. [Figure 6] FIG. 10 is a diagram showing a first specific example of a screen 530. [Figure 7] FIG. 10 is a diagram showing a second specific example of the screen 530. [Figure 8] FIG. 10 is a diagram showing a third specific example of a screen 530. [Figure 9] FIG. 10 is a diagram illustrating an example of a processing procedure executed by the information processing system 1 according to the second embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of a screen of the cooperative conference system. [Figure 11] FIG. 10 is a diagram showing a first display example of a screen of a cooperative conference system including a corresponding image. [Figure 12] FIG. 10 is a diagram showing a second display example of the screen of the cooperative conference system including the corresponding image. [Figure 13] FIG. 10 is a diagram showing an example of display of minutes data. [Figure 14] FIG. 10 is a diagram illustrating an example of a functional configuration of an information processing system 1 according to a third embodiment. [Figure 15] FIG. 13 is a diagram illustrating an example of a hardware configuration of a terminal device 10 according to a fourth embodiment. [Figure 16] FIG. 10 is a diagram illustrating an example of a functional configuration of an information processing system 1 according to a fourth embodiment. [Figure 17] FIG. 13 is a diagram illustrating an example of a functional configuration of an information processing system 1 according to a fifth embodiment. [Figure 18] 13A and 13B are diagrams showing examples of displaying corresponding images and the like in the fifth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In this embodiment, a conference will be described as an example of a gathering, but this embodiment is also effective for gatherings of other types than a conference, such as training or classes. FIG. 1 is a diagram showing an example of the configuration of an information processing system 1 in a first embodiment. FIG. 1 shows two configuration examples, (A) and (B), of the information processing system 1.

[0010] (A) is a configuration corresponding to an online conference such as a web conference or a video conference. In (A), a terminal device 10 is provided for each participant of the conference. Each terminal device 10 is connected to a server device 40 via a communication network 100 such as the Internet.

[0011] The server device 40 is one or more computers that generate images based on the audio data collected by each terminal device 10.

[0012] (B) is a configuration corresponding to a face-to-face conference. In (B), one terminal device 10 is provided for each of the multiple participants participating in the conference. A microphone 109b is connected to or built into the terminal device 10. The number of microphones 109b may be one or more. When multiple microphones 109b are provided, a highly directional microphone may be provided for each participant, making it possible to identify the participant who spoke according to the microphone 109b that collected the audio. In the case of (B), the terminal device 10 generates an image based on the audio data collected by the microphone 109b.

[0013] In the following, the configuration of (A) will be referred to as "Configuration A" and the configuration of (B) will be referred to as "Configuration B."

[0014] In the present embodiment, the term "conference" refers to a discussion such as a conversation or dialogue that is held among two or more people, and the format of the discussion is not limited to a specific one.

[0015] Fig. 2 is a diagram showing an example of the hardware configuration of the terminal device 10 in the first embodiment. In Fig. 2, the hardware configuration of the terminal device 10 will be described, but the hardware configuration of the server device 40 is similar and therefore will not be described. In Fig. 2, the reference numbers of the components of the terminal device 10 start with "1", and the reference numbers of the components of the server device 40 start with "4".

[0016] The terminal device 10 is constructed by a computer and, as shown in FIG. 2, includes a CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a HD (Hard Disk) 104, a HDD (Hard Disk Drive) controller 105, a display I / F 106, and a communication I / F 107.

[0017] Of these, the CPU 101 controls the overall operation of the terminal device 10. The ROM 102 stores programs such as an IPL (Initial Program Loader) used to drive the CPU 101. The RAM 103 is used as a work area for the CPU 101.

[0018] The HD 104 stores various data such as programs, etc. The HDD controller 105 controls reading and writing of various data from and to the HD 10 under the control of the CPU 101.

[0019] The display I / F 106 is a circuit that displays images on a display 106a. The display 106a is a type of display unit, such as a liquid crystal display or organic electroluminescence (EL) display, that displays various information such as a cursor, menu, window, text, or image. The communication I / F 107 is an interface used for communication with other devices.

[0020] The communication I / F 107 is, for example, a network interface card (NIC) that supports transmission control protocol (TCP) / internet protocol (IP).

[0021] The terminal device 10 also includes a sensor I / F 108 , a sound input / output I / F 109 , an input I / F 110 , a media I / F 111 , and a DVD-RW (Digital Versatile Disk Rewritable) drive 112 .

[0022] The sensor I / F 108 is an interface that receives detection information from various sensors. The sound input / output I / F 109 is a circuit that processes input and output of sound signals between the speaker 109a and the microphone 109b under the control of the CPU 101. The input I / F 110 is an interface that connects a predetermined input means to the terminal device 10.

[0023] The keyboard 110a is a type of input means having multiple keys for inputting characters, numbers, various instructions, etc. The mouse 110b is a type of input means for selecting and executing various instructions, selecting a processing target, moving the cursor, and performing operations on the display screen, etc.

[0024] The media I / F 111 controls reading and writing (storing) of data from and to a recording medium 111a such as a flash memory. The DVD-RW drive 112 controls reading and writing of various data from and to a DVD-RW 112a, which is an example of a removable recording medium. Note that the medium is not limited to a DVD-RW, and a DVD-R or the like may also be used. The DVD-RW drive 112 may also be a Blu-ray drive that controls reading and writing of various data from and to a Blu-ray Disc (registered trademark).

[0025] The terminal device 10 also includes a bus line 113. The bus line 113 is an address bus, a data bus, or the like for electrically connecting the components such as the CPU 101.

[0026] The recording media, such as HDs and CD-ROMs, on which the above programs are stored can be provided domestically or internationally as program products. The terminal device 10, for example, realizes the information processing method according to the present invention by executing the program according to the present invention.

[0027] Fig. 3 is a diagram showing an example of the functional configuration of the information processing system 1 in the first embodiment. In Fig. 3, the terminal device 10 has an input unit 11, an input data transmission unit 12, a display control unit 13, etc. Each of these units is realized by a process in which one or more programs installed in the terminal device 10 are executed by a CPU 101.

[0028] The input unit 11 accepts input from users (conference participants) of the terminal device 10. The input unit 11 accepts, for example, information indicating the agenda (topic) of the conference, the number of participants, etc. (hereinafter referred to as "conference information"), and inputs audio data of speech of the conference participants collected by the microphone 109b during the conference. The input unit 11 may also accept, for example, input of an image (hereinafter referred to as "input image") related to the agenda of the conference from the participants.

[0029] The input data transmission unit 12 transmits the conference information or data received by the input unit 11 to the server device 40 .

[0030] The display control unit 13 controls the display of a screen or the like based on the display information transmitted from the server device 40.

[0031] On the other hand, the server device 40 has an input data acquisition unit 41, an image generation unit 42, a relevance calculation unit 43, a display information generation unit 44, a display information transmission unit 45, etc. Each of these units is realized by a process executed by a CPU 401 of one or more programs installed in the server device 40. The server device 40 also uses an input data storage unit 451 and a model storage unit 452. Each of these storage units can be realized using, for example, the HD 404 or a storage device connectable to the server device 40 via a network.

[0032] The input data acquisition unit 41 acquires (receives) the conference information or data transmitted from the input data transmission unit 12, and records the acquired conference information or data in the input data storage unit 451. Note that when the server device 40 has the input data acquisition unit 41 as in FIG. 3, the input data acquisition unit 41 has a receiving function, but in the case of configuration B, that is, when the terminal device 10 has the input data acquisition unit 41, the input data acquisition unit 41 has an input interface function.

[0033] The image generation unit 42 generates text data (i.e., text data indicating the content of speech of the conference participants) for each participant from the voice data acquired by the input data acquisition unit 41, and generates an image corresponding to the text data (hereinafter referred to as a "corresponding image") for each participant based on the text data and the image generation model m1 stored in the model storage unit 452. Specifically, the image generation unit 42 inputs the text data related to each participant to the image generation model m1, and acquires, as the corresponding image, an image generated by the image generation model m1 based on the text data.

[0034] Here, the text indicating the content of speeches made by the meeting participants is an example of unique information based on the actions of each of the meeting (gathering) participants at the meeting (gathering), and the text data indicating the content of speeches made by the meeting participants is an example of unique data indicating the unique information of each of the meeting participants.

[0035] The unique information of each conference participant may include, in addition to or instead of text indicating the content of the speech of the conference participant, at least one of characters or operations entered by the conference participant, image information of each conference participant captured by a photographing device, and biometric information of each conference participant detected by a biometric information detection unit.

[0036] The image generation model m1 is an example of a machine learning model (generative AI) that generates images from text data, and is an example of a machine learning model (generative AI) that generates images from unique data indicating unique information about each conference participant. The image generation model m1 is trained using training data including, for example, text data and images. The training data includes, for example, training text data as input and an image as a correct answer for output. For example, training may be performed so that an image generated by the image generation model m1 to which text data included in the training data is input approaches the correct image included in the training data. A commercially available model such as Stable Diffusion or DALL·E or an independently developed model may be used as the image generation model m1.

[0037] The image generation unit 42 also generates an integrated corresponding image corresponding to a set of multiple pieces of text data based on the text data of each of the multiple participants and the image generation model m1. Specifically, the image generation unit 42 inputs the set of multiple pieces of text data to the image generation model m1, and obtains an image generated by the image generation model m1 based on the set as the integrated corresponding image. That is, while a corresponding image is generated for each participant, one integrated corresponding image is generated for all participants.

[0038] When an input image (an image input in relation to a conference) is input, the image generation unit 42 may generate a corresponding image for each participant based on the text data for each participant, the input image, and the image generation model m1, and may also generate an integrated corresponding image corresponding to the set of text data for each of the multiple participants and the input image. In this case, the image generation model m1 inputs the text data and input image for each participant to the image generation model m1, and obtains, as the corresponding image, an image generated by the image generation model m1 based on the text data and input image. Furthermore, the image generation model m1 inputs a set of multiple pieces of text data to the image generation model m1, and obtains, as the integrated corresponding image, an image generated by the image generation model m1 based on the set. In this case, the corresponding image or the integrated corresponding image may be, for example, an image obtained by adding changes to the input image in accordance with the content of the participant's speech.

[0039] The relevance calculation unit 43 calculates the relevance of each corresponding image generated by the image generation unit 42 with other images. The relevance is an index showing the degree of relevance. For a certain corresponding image, "other images" are, for example, other corresponding images. In this case, the relevance between the corresponding images is calculated. Furthermore, the "other images" may be integrated corresponding images. In this case, the relevance between each corresponding image and the integrated corresponding image is calculated. Furthermore, the "other images" may be input images or images generated in advance based on predetermined text data, etc.

[0040] The display information generating unit 44 generates display information that displays each corresponding image (for example, display information of a screen including the corresponding image). In this case, if the display information generating unit 44 can grasp the correspondence between each participant and the corresponding image, the display information generating unit 44 may generate display information that displays participant identification information that identifies the participant in association with the corresponding image. Furthermore, the display information generating unit 44 may generate the display information that displays the degree of association in association with the participant identification information and the corresponding image.

[0041] The display information transmitting unit 45 transmits the display information generated by the display information generating unit 44 to the terminal device 10.

[0042] When the information processing system 1 has the configuration B, the terminal device 10 may have the same components as the server device 40 in FIG.

[0043] The following describes the processing procedures executed by the information processing system 1. Fig. 4 is a sequence diagram for explaining an example of the processing procedures executed by the information processing system 1 in the first embodiment. Fig. 4 shows an example of the processing procedures when the information processing system 1 has configuration A, but when the information processing system 1 has configuration B, the terminal device 10 may also execute the processing procedures of the server device 40.

[0044] In step S101, the terminal device 10 receives input of conference information from one of the conference participants (for example, the organizer) via the displayed screen. The conference information includes, for example, the agenda and the number of participants.

[0045] 5 is a diagram showing an example of a screen displayed in the first embodiment. In step S101, a screen 510 is displayed on the terminal device 10. The screen 510 may be displayed based on display information received by the terminal device 10 from the server device 40 before execution of step S101. The screen 510 includes an input area 511, an input area 512, a start button 513, etc. The input area 511 is an area for accepting input of an agenda. The input area 512 is an area for accepting input of the number of participants. The start button 513 is a button for accepting a request to start a conference.

[0046] When an agenda item is input in input area 511, the number of participants is input in input area 512, and start button 513 is pressed, input data transmission unit 12 transmits conference information including the input agenda item and number of participants to server device 40 (S102). Note that an input image may also be input via screen 510. For example, a file for storing an input image may be selected via screen 510, and the input image stored in that file may be transmitted to server device 40 together with the conference information.

[0047] When the input data acquisition unit 41 of the server device 40 receives the conference information, etc., it records the conference information, etc. in the input data storage unit 451 (S103).

[0048] Next, the display information sending unit 45 transmits display information, for example, displaying an instruction to start the conference, to each terminal device 10 (S104). The display control unit 13 of each terminal device 10 displays the screen 520 of FIG. 5 based on the display information. The screen 520 includes a message instructing the participants to start a discussion on the agenda. After viewing the message, each participant starts the conference (hereinafter referred to as the "target conference") and speaks about their understanding and thoughts on the agenda. When one microphone 109b is provided for multiple participants in the target conference, known techniques for separating the speech of each participant can be applied. However, in consideration of the ease of separating the speech of each participant, the conference may be conducted so that participants speak one by one. When a microphone 109b is provided for each participant in the target conference (including in the case of an online conference), multiple people may be allowed to speak at the same time.

[0049] Note that the conference may be started without executing steps S101 to S104.

[0050] During the target conference, the input unit 11 of the terminal device 10 accepts input of voice data from the microphone 109b (S105). The input data transmission unit 12 transmits the voice data to the server device 40 (S106). The input data acquisition unit 41 of the server device 40 records the received voice data in the input data storage unit 451 (S107). Steps S105 to S107 are repeatedly executed until the target conference ends. Therefore, the input data storage unit 451 stores voice data until the target conference ends. When one microphone 109b is provided for multiple participants, one voice data is recorded for multiple participants. On the other hand, when a microphone 109b is provided for each participant (including the case of an online conference), voice data is recorded for each participant. In this case, the input data acquisition unit 41 may record voice data from the terminal device 10 or microphone 109b related to the participant in the input data storage unit 451 in association with the participant's identification information (hereinafter referred to as "participant identification information"). The correspondence between each terminal device 10 or each microphone 109b and the participant identification information may be set in advance.

[0051] The duration of the target conference may be set in advance in the server device 40, or may be a variable time until the microphone 109b or the terminal device 10 is turned off.

[0052] When the conference ends, the image generation unit 42 uses natural language processing to perform transcription, text shaping, and the like on the voice data to generate text data (S108). At this time, if the voice data is distinguished for each participant, the image generation unit 42 generates text data for each voice data, thereby generating text data for each participant. If the voice data is not distinguished for each participant, the image generation unit 42 also performs speaker separation on the voice data to generate text data for each participant. At this time, speaker separation may be performed using the number of participants included in the conference information. Note that text shaping includes, for example, inserting punctuation marks and deleting unnecessary character data and spaces.

[0053] Next, the image generation unit 42 generates prompts to be input to the image generation model m1 based on the agenda and the text data for each participant (S109). At this time, the image generation unit 42 generates a prompt for each text data (hereinafter referred to as a "corresponding prompt") and a prompt based on a collection of all text data (hereinafter referred to as an "integrated corresponding prompt"). The corresponding prompt based on certain text data is text data for causing the image generation model m1 to generate a corresponding image based on the agenda and the text data. The integrated corresponding prompt is text data for causing the image generation model m1 to generate an integrated corresponding image based on the agenda and the utterances of all participants. Note that for both the corresponding prompt and the integrated corresponding prompt, the image generation unit 42 may generate a prompt using words with relatively high importance (hereinafter referred to as "important words") and audio data by performing morphological analysis and key word extraction on the original text data. Furthermore, if the image generation model m1 used supports prompts in only a specific language (e.g., English), the image generation unit 42 may perform language conversion (translation) on each prompt to convert the prompt into a language supported by the image generation model m1. Furthermore, the image generation unit 42 may weight each important word in the prompt based on the importance of each important word obtained in the important word extraction. This increases the influence of the important words on the generated image data.

[0054] For example, a prompt that combines a topic and a key word is generated as follows:

[0055] "<Agenda> inspired by ('work': 2.692), ('companion': 1.923), 'new': 1.538), ('future': 1.538), ('ai': 1.154), ('robot': 1.154)" Here, <agenda> is a string indicating the agenda recorded as meeting information. "Work", "companion", "new", "future", "ai", and "robot" are important words. The numbers paired with important words indicate the importance of the important words. Note that the format of the above prompt follows the grammar of the text that image generation model m1 expects to input as a prompt. In other words, it is known to image generation model m1 that the words and importance are contained within parentheses. Furthermore, this format differs depending on the image generation model used.

[0056] Important words can be extracted using various known techniques, such as TF-IDF (Term Frequency-Inverse Document Frequency), TextRank, RAKE (Rapid Automatic Keyword Extraction), and analysis using Transformer.

[0057] Next, the image generation unit 42 inputs the corresponding prompt to the image generation model m1 for each corresponding prompt, and acquires an image (corresponding image) generated by the image generation model m1 based on the corresponding prompt (S110). Thus, a corresponding image is generated for each corresponding prompt (for each participant).

[0058] Next, the image generating unit 42 inputs the integration-compatible prompt to the image generation model m1, and acquires an image (integration-compatible image) generated by the image generation model m1 based on the integration-compatible prompt (S111).

[0059] The process of generating a corresponding image for each participant (S109, S110) and the process of generating an integrated corresponding image (S109, S111) may be executed in parallel for each image.

[0060] Alternatively, only one of the corresponding image and the integrated corresponding image may be generated. In this case, the user may be able to set which one to generate. When generating a corresponding image for each participant, the seed value of the image generation model m1 is fixed so that when exactly the same prompt is input, exactly the same image data is generated.

[0061] When corresponding images are generated, the relevance calculation unit 43 calculates the relevance of each corresponding image with other images (S112). Here, the other images are, as described above, other corresponding images, integrated corresponding images, input images, images generated in advance based on predetermined text data, etc. For example, if there are three corresponding images, images 1 to 3, and the relevance between the corresponding images is calculated, three relevances are calculated: between image 1 and image 2, between image 1 and image 3, and between image 2 and image 3.

[0062] For example, a correlation coefficient may be used as the degree of association. There are various methods for calculating the correlation coefficient, such as MSE (Mean Squared Error), SSIM (Structural Similarity Index), histogram comparison, feature vector comparison using a neural network, and comparison of objects detected in an image, and any of these may be used, or a value obtained by linearly combining the results of calculations by a plurality of methods may be used.

[0063] Furthermore, in consideration of ease of understanding for users, the value of the relevance may be normalized to a range of 0 to 100. In this case, the larger the value, the higher the correlation (the stronger the association).

[0064] The relevance calculation unit 43 also calculates the average of the relevance degrees as the common understanding degree. The common understanding degree is an index showing the degree of common understanding of the participants regarding the agenda.

[0065] It should be noted that if no corresponding image is generated (that is, if only one integrated corresponding image is generated), step S112 does not need to be executed.

[0066] Next, the display information generation unit 44 generates display information (e.g., a screen) for displaying the generated images (each corresponding image or the integrated corresponding image, or each corresponding image and the integrated corresponding image) (S113). At this time, if a degree of association has been generated for each corresponding image, the display information generation unit 44 generates display information that displays each corresponding image in association with the degree of association and also displays the common understanding level. Furthermore, if the correspondence between each corresponding image and each participant (participant identification information) is clear, the display information generation unit 44 may generate display information that further displays participant identification information in association with each corresponding image.

[0067] Next, the display information transmitting unit 45 transmits the generated display information to each terminal device 10 (S114). When the display control unit 13 of each terminal device 10 receives the display information, it displays screen 530-1 or screen 530-2 (hereinafter referred to as "screen 530" when the screens are not distinguished) in Fig. 5 based on the display information (S115). Screen 530 may be displayed on display 106a or on a projector connected to the terminal device 10.

[0068] Screen 530-1 is an example of a screen displayed when the number of participants in the target conference is three and a corresponding image is generated for each participant. In addition to the corresponding images, screen 530-1 includes, as display elements, the relevance between the corresponding images (i.e., the relevance of each of the three pairs of corresponding images) and the common understanding level. If the correspondence between each corresponding image and the participant can be grasped, participant identification information may be associated with each corresponding image and displayed.

[0069] Screen 530-2 is an example of a screen that is displayed when one integrated corresponding image is generated for all participants.

[0070] When both the corresponding image and the integrated corresponding image are generated, a screen including the display contents of screen 530-1 and screen 530-2 may be displayed.

[0071] Here, a specific example of a screen 530 that displays a corresponding image or an integrated corresponding image generated when three people hold a meeting on the topic "About the office of the future" will be shown below.

[0072] Fig. 6 is a diagram showing a first specific example of screen 530. Screen 530-1 shown in Fig. 6 is a screen having the same configuration as screen 530-1 shown in Fig. 5. Fig. 6 shows an example of how screen 530-1 would appear if three people made exactly the same statement. In this case, the contents of the three corresponding images would be the same, and the relevance between the corresponding images would all be 100%. In addition, the common understanding level, which is the average of the relevance levels, would be 100%.

[0073] FIG. 7 is a diagram showing a second specific example of the screen 530. Like FIG. 6, the screen 530-1 shown in FIG. 7 has the same configuration as the screen 530-1 shown in FIG. 5. However, FIG. 7 shows an example of the display of the screen 530-1 when there is a discrepancy in the content of the statements made by the three participants. In this case, the content of the three corresponding images differs depending on the statements made by each participant. Therefore, the relevance between the corresponding images is not 100%, and the common understanding level is 76%. Note that in this example, the common understanding level is calculated by normalizing the linear sum of the correlation coefficient of the feature vectors and the correlation coefficient when the grayscale images are converted into one dimension.

[0074] From the content and relevance (85%) of the corresponding images in Figure 7, it can be seen that image 2 (center) and image 3 (right) are corresponding images based on similar statements. Also, although the content of image 1 (left) and images 2 and 3 appear to be very different at first glance, there are commonalities in the brightness patterns when converted to grayscale, which shows that there are two different orientations regarding the "office of the future": a bird's-eye view (the exterior of the office) and a subjective view (the interior of the office), which can be useful for subsequent discussions.

[0075] Fig. 8 is a diagram showing a third specific example of screen 530. Screen 530-2 shown in Fig. 8 is a screen having the same configuration as screen 530-2 shown in Fig. 5. That is, one image included in screen 530-2 in Fig. 8 is an integrated corresponding image based on the content of statements made by three people. Because this image is generated with the statements of the three people mixed together, it is possible to stimulate and deepen the discussion based on the visualized information, for example, by saying, "This part is different from what I imagined!" or "I can relate to this part!"

[0076] Participants who view screen 530 such as that shown in Figures 6 to 8 can quantitatively grasp the level of common understanding as a number (in the case of Figure 6 or Figure 7), and can also share the direction of understanding from each perspective in a visible form, that is, an image.

[0077] As a modification of this embodiment, the corresponding image may be displayed without displaying the degree of common understanding as a numerical value. Even in this case, by visually comparing the corresponding image based on one's own utterance with the corresponding image or integrated corresponding image based on the utterances of others, one can visually recognize the difference between one's impression of one's own utterance and the impression of the utterances of others or the impression of the utterances as a whole.

[0078] As described above, according to the first embodiment, an image (such as a corresponding image or an integrated image) is generated using text data indicating the content of speech by each participant in a conference and the image generation model m1, and the image is displayed. That is, an image is dynamically generated according to the content of speech. Therefore, visualization according to the unique information of each participant in a meeting or other gathering can be achieved. For example, visualization of qualitative events contained in speech at a meeting or other gathering can be achieved. A qualitative event is a qualitative, invisible, and ambiguous event, such as a person's thoughts or psychology contained in speech.

[0079] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Therefore, unless otherwise specified, the second embodiment may be the same as the first embodiment. Furthermore, the functions and configurations described in the second embodiment may be combined with the functions and configurations described in the first embodiment.

[0080] In the second embodiment, an example will be described in which the information processing system 1 cooperates with an online conference system such as Zoom (registered trademark) or Teams (registered trademark), generates prompts based on text data of each speaker for a certain time interval generated by a transcription function provided by the cooperated online conference system (hereinafter referred to as the "cooperative conference system"), and generates corresponding images for each speaker.

[0081] Fig. 9 is a diagram for explaining an example of a processing procedure executed by the information processing system 1 in the second embodiment. In Fig. 9, the step numbers in Fig. 4 are written in parentheses for steps that are the same as or correspond to those in Fig. 4.

[0082] In step S201, the input data acquiring unit 41 waits for a certain period of time (for example, 20 seconds). Note that the screen of the cooperative conference system is displayed on the terminal device 10 of each participant of the online conference.

[0083] Fig. 10 is a diagram showing an example of a screen of a collaborative conference system. In Fig. 10, screen 610 is an example of a screen of the collaborative conference system. Screen 610 includes a window for displaying an image of each participant, and a text area 611 that contains the speech content of each participant in chronological order, which has been converted into text using a transcription function.

[0084] After a certain period of time has elapsed, the input data acquisition unit 41 acquires from the linked conference system text data generated for each participant of the online conference during the most recent certain period of time, in association with the participant identification information of the participant who spoke the content of the text data (S202). Note that the acquired text data corresponds to a character string obtained by dividing the text data obtained from the start to the end of the online conference in chronological order.

[0085] In subsequent steps S203 to S206, the same processing as steps S109 to S112 in Fig. 4 is executed for each acquired text data. As a result, an integrated corresponding image for a certain period of time immediately preceding the acquisition and corresponding images for each participant of the online conference are generated. Note that either one or both of steps S205 and S206 may not be executed. In other words, an integrated corresponding image may not be generated. Furthermore, the degree of association may not be calculated.

[0086] Next, the display information generating unit 44 includes the corresponding image on the screen 610 of the cooperative conference system (S207).

[0087] FIG. 11 is a diagram showing a first display example of a screen of a cooperative conference system including a corresponding image.

[0088] 11, a window for each participant is overlaid with a corresponding image of the participant corresponding to that window. Each window contains participant identification information for that participant. Therefore, the corresponding image and participant identification information are displayed in association with each other on screen 610.

[0089] 12 is a diagram showing a second display example of a screen of the linked conference system including the corresponding image. A pop-up screen 612 including the corresponding image of each participant and the participant identification information in association with each other is superimposed on the screen 610 shown in FIG.

[0090] When an integrated corresponding image is generated or when a degree of association is generated, the integrated corresponding image or the degree of association may also be displayed in Fig. 11 or 12. The degree of association may be displayed so that the correspondence between the corresponding image related to the degree of association and the degree of association can be understood.

[0091] Steps S201 to S207 are repeated until the end of the online conference, so that each participant can check the corresponding image of each participant at regular intervals as the conference progresses, and can hold a discussion. As a result, it is expected that the discussion will become more lively and new ideas will be generated.

[0092] When the online conference ends (Yes in S208), the display information generation unit 44 adds (pastes) thumbnails of the corresponding images generated at the times corresponding to the timings at which each corresponding image was generated in the minutes data, which includes text data indicating the speech content of each participant (i.e., character strings obtained by dividing the text data of the entire conference in chronological order) generated by the minutes system of the linked conference system and which is associated with the participant identification information of the participants corresponding to the text data, at locations corresponding to the timings at which each corresponding image was generated (S209). The minutes data is an example of display information including character strings obtained by dividing the text data obtained from the start to the end of the online conference in chronological order. Note that the display information generation unit 44 may generate the minutes data based on the text data acquired in step S202 without using the minutes system.

[0093] The minutes data is transmitted to the terminal device 10 in response to a predetermined operation on the terminal device 10, and the display control unit 13 of the terminal device 10 controls the display of the minutes data.

[0094] Fig. 13 is a diagram showing an example of display of minutes data. As shown in Fig. 13, the minutes data 620 includes, for each time interval of a certain time, text data (character strings separated in chronological order) corresponding to the content of speech during that time interval. The minutes data 620 also includes, for each participant, thumbnails of corresponding images based on the text data corresponding to the content of speech by each participant during that time interval, at a location corresponding to the end of the time interval (i.e., the timing at which the corresponding image is generated). Clicking on the thumbnail image may allow the corresponding image to be enlarged or saved.

[0095] Such minutes data 620 allows a user to visually review the meeting and can serve as a trigger for memory.

[0096] In the first embodiment, a corresponding image and an integrated corresponding image may also be generated at regular intervals, and the corresponding image and the integrated image may be displayed. Whether or not to generate a corresponding image at regular intervals may be set by the user. In other words, the portion of the text data from the start to the end of the conference that image generation unit 42 uses to generate a corresponding image may be changeable based on the user's designation of the "start" and "end" points, or on predetermined conditions set in the system.

[0097] When corresponding images or the like are generated at regular intervals, if the relevance or common understanding remains significantly low, a message may be output urging the user to reconsider the agenda.

[0098] Next, a third embodiment will be described. In the third embodiment, differences from the first embodiment will be described. Therefore, unless otherwise specified, the third embodiment may be the same as the first embodiment. Furthermore, the functions and configurations described in the third embodiment may be combined with the functions and configurations described in the first embodiment.

[0099] Fig. 14 is a diagram showing an example of the functional configuration of the information processing system 1 according to the third embodiment. In Fig. 14, the same components as those in Fig. 3 are given the same reference numerals, and their description will be omitted.

[0100] 14, the server device 40 further includes a virtual data generation unit 46. A language model m2 is further stored in the model storage unit 452. The virtual data generation unit 46 is realized by a process in which one or more programs installed in the server device 40 are executed by the CPU 401.

[0101] The language model m2 is a machine learning model that inputs text and outputs a sentence in text format with content corresponding to the text. For example, an existing large-scale language model (LLM (Large Language Model)) can be used as the language model m2.

[0102] The virtual data generation unit 46 performs virtual (pseudo) speech in the target conference as a virtual participant (hereinafter referred to as a "virtual participant"), like a chatbot. The virtual data generation unit 46 generates unique data (text data indicating the speech content) indicating virtual unique information (text indicating the speech content) of the virtual participant based on the unique data (text data indicating the speech content) of the other participants, for example, at timing (hereinafter referred to as a "virtual speech interval") every certain time during the target conference (hereinafter referred to as a "virtual speech interval"). The virtual data generation unit 46 uses a language model m2 when generating the unique data.

[0103] For example, each time a virtual utterance timing arrives, the virtual data generation unit 46 generates text data by performing the same process as in step S108 for each piece of voice data recorded in the input data storage unit 451 in step S107 executed during the immediately preceding virtual utterance interval (the entire time period of the virtual utterance interval or the last 1 to 3 minutes of the virtual utterance interval). The virtual data generation unit 46 inputs a prompt including the text data and including an instruction such as "As a conference participant, speak based on the agenda and the utterances of the other participants" to the language model m2. The virtual data generation unit 46 generates voice data having the text data generated by the language model m2 based on the prompt as the utterance content, outputs the voice data to each terminal device 10, and records the text data in the input data storage unit 451 in association with the identification information of the virtual participant. Note that the generation of voice data having the text data as the utterance content can use a known voice synthesis technique.

[0104] In the third embodiment, from step S109 onwards, the text data (utterance content) relating to the virtual participant is treated in the same way as other participants. That is, a corresponding prompt is also generated for the text data relating to the virtual participant (S109), and a corresponding image is generated based on the corresponding prompt. The subsequent steps may be the same as in the first embodiment.

[0105] The virtual utterance timing does not necessarily have to be set at regular intervals. For example, the virtual utterance timing may be set at a timing when the number of utterances in the target conference decreases (for example, when the number of utterances in the most recent regular period is equal to or less than a threshold).

[0106] As described above, according to the third embodiment, by adding virtual participants, it is possible to expect livelier discussions in the target conference.

[0107] Next, a fourth embodiment will be described. In the fourth embodiment, differences from the first embodiment will be described. Therefore, unless otherwise specified, the fourth embodiment may be the same as the first embodiment. Furthermore, the functions and configurations described in the fourth embodiment may be combined with the functions and configurations described in the first embodiment.

[0108] 15 is a diagram showing an example of the hardware configuration of the terminal device 10 in the fourth embodiment. In FIG. 15, an image capturing device 114 and a biometric information detecting unit 115 are connected to the terminal device 10.

[0109] The image capturing device 114 is a camera that captures images including the face or posture of the user (participant in the target conference) of the terminal device 10. The images captured by the image capturing device 114 may be moving images or still images captured at regular time intervals.

[0110] The biological information detection unit 115 is a sensor or the like that measures biological information (blood pressure, pulse rate, etc.) of the user of the terminal device 10 (participant in the target conference).

[0111] In the fourth embodiment, the unique data includes first type unique data including first type unique information and second type unique data including second type unique information. The first type unique information is text indicating at least one of characters entered by a participant during a conference and the content of the participant's speech. The second type unique information refers to unique information based on a reaction of any of the other participants to the first type unique information. The second type unique information includes at least one of operation information (e.g., entered characters) indicating an operation entered by a participant who reacted to the first type unique information of any of the other participants, captured image information of the participant captured by the imaging device 114, and biometric information of the participant detected by the biometric information detection unit 115. The captured image information can be changed to be displayed on the display unit or not.

[0112] That is, the image capturing device 114 and the biological information detecting unit 115 are used to detect reactions to speeches and the like made by other participants.

[0113] Fig. 16 is a diagram showing an example of the functional configuration of the information processing system 1 in the fourth embodiment. In Fig. 16, the same components as those in Fig. 3 are given the same reference numerals, and the description thereof will be omitted.

[0114] 16, the model storage unit 452 further stores an image generation model m3. The image generation model m3 is a machine learning model that receives a first type of unique data of a first participant and a second type of unique data of a second participant as input, and outputs a corresponding image corresponding to the first type of unique data and the second type of unique data (i.e., a corresponding image related to a reaction to an utterance of another participant). Therefore, the corresponding image is an image related to the reaction of the second participant to an utterance of the first participant.

[0115] Specifically, the image generation model m3 includes, as input data, text (the text of the utterances of other participants), operation information (operation information on the terminal device 10) of a participant who responded to another participant (hereinafter referred to as "participant R"), captured image information of participant R (face, posture, etc.), and information including biometric information of participant R (blood pressure, pulse, etc.) (hereinafter, the operation information, captured image information, and biometric information of participant R are referred to as "reaction information"). The image generation model m3 has been trained using training data including a correct image (hereinafter referred to as a "correct image"). That is, the image generation model m3 is trained so that the image output by the image generation model m3, to which input data included in the training data is input, approaches the correct image in the training data. Here, if the reaction information related to the input data of the training data indicates a positive reaction to the utterance of the other participant, an image that has a relatively high relevance to the corresponding image generated by the image generation model m1 for the utterance of the other participant, becomes the correct image in the training data. On the other hand, if the reaction information related to the input data of the learning data indicates a negative reaction to the utterances of other participants, an image that has a relatively low relevance to the corresponding image generated by the image generation model m1 for the utterances of the other participants will be the correct image in the learning data.

[0116] Therefore, when the input reaction information (participant R's reaction information) is positive toward the utterances of other participants, the image generation model m3 generates a corresponding image that is relatively highly related to the corresponding image of the text of the utterance content of the other participants that is input, and when the input reaction information (participant R's reaction information) is negative toward the utterances of other participants, the image generation model m3 generates a corresponding image that is relatively less related to the corresponding image of the text of the utterance content of the other participants that is input.

[0117] The processing procedure in the fourth embodiment will be described with reference to Fig. 4. The input data transmission unit 12 of each terminal device 10 transmits reaction information of the user of the terminal device 10 to the server device 40 in response to utterances by the participant related to the user of the other terminal device 10. The input data acquisition unit 41 of the server device 40 records the received reaction information in the input data storage unit 451 for each participant. Therefore, for the period until the target conference is ended, reaction information of the other participants is stored in the input data storage unit 451 for each piece of audio data related to the utterances of each participant.

[0118] In step S110, the image generation unit 42 further inputs, for each participant, the participant's reaction information to the other participant and the text data of the other participant's utterance into the image generation model m3, and obtains the image output by the image generation model m3 as a corresponding image related to the participant's reaction.

[0119] For example, if person B smiles (shows a positive reaction) in response to person A's speech, this means that person B agrees with person A's speech (is positive about person A's speech), and a corresponding image for person B that is highly related (highly relevant) to the corresponding image related to person A's speech will be generated.

[0120] On the other hand, if person B shows a dissatisfied expression (negative reaction) in response to person A's speech, since person B does not agree with person A's speech (is negative about person A's speech), a corresponding image for person B that is less relevant to the corresponding image related to person A's speech will be generated.

[0121] Here, we have shown an example in which facial expressions are used as reaction information, but by labeling positive and negative reactions to the participant's posture, pulse rate, etc. and having the image generation model m3 learn these, it is possible to generate corresponding images in the same way as facial expressions.

[0122] The rest may be the same as in the first embodiment.

[0123] As a result of the above, in the fourth embodiment, corresponding images relating to the reactions to the utterances of other participants are also displayed.

[0124] The fourth embodiment may be combined with any of the above embodiments.

[0125] As described above, according to the fourth embodiment, the reaction of one participant to another participant's utterance can also be expressed by an image.

[0126] Next, a fifth embodiment will be described. In the fifth embodiment, differences from the third embodiment will be described. Therefore, unless otherwise specified, the fifth embodiment may be the same as the third embodiment. Furthermore, the functions and configurations described in the fifth embodiment may be combined with the functions and configurations described in the third embodiment.

[0127] Fig. 17 is a diagram showing an example of the functional configuration of the information processing system 1 in the fifth embodiment. In Fig. 17, the same components as those in Fig. 14 are given the same reference numerals, and the description thereof will be omitted.

[0128] 17, the server device 40 further includes a layout determination unit 47. The layout determination unit 47 is realized by processing that one or more programs installed in the server device 40 cause the CPU 401 to execute.

[0129] The layout determination unit 47 determines at least one of the arrangement, size of each corresponding image, and decoration when displaying multiple corresponding images based on at least one of the unique information (text indicating the content of the utterance or audio of the utterance) and the corresponding images.

[0130] Specifically, the layout determination unit 47 determines the layout when displaying a plurality of corresponding images based on the degree of association calculated by the degree-of-association calculation unit 43 between each corresponding image and other images.

[0131] The layout determination unit 47 also determines the decoration or size of the corresponding image based on at least one of the amount of speech of the participant, the volume of speech, image information showing the facial expressions of the other participants photographed by the photographing device 114, and changes in the biometric information of the other participants detected by the biometric information detection unit 115, which are included in the unique information. In this case, the terminal device 10 may have the hardware configuration shown in Fig. 15, and reaction information may be recorded in the input data storage unit 451, as described in the fourth embodiment.

[0132] In the fifth embodiment, the display information generating unit 44 generates display information that displays a plurality of corresponding images in at least one of the arrangement, size, and decoration determined by the layout determining unit 47.

[0133] Specifically, before the display information generation unit 44 generates display information in step S113 of Fig. 4, the layout determination unit 47 determines the layout position (arrangement order) of each corresponding image based on the relevance (relevance with other images) calculated in step S112 for each corresponding image for each participant (including virtual participants). For example, the layout determination unit 47 determines the layout position of each corresponding image so that a corresponding image with a relatively high relevance to a certain image is relatively close to the corresponding image. The certain image may be any one of the corresponding images, an integrated corresponding image, or an agenda image. An agenda image is an image generated by inputting text data indicating an agenda into the image generation model m1.

[0134] The layout determination unit 47 also increases the size of each corresponding image and adds decoration (highlight display, etc.) to make the corresponding image more conspicuous, the greater the amount of speech by the participant related to the corresponding dialogue image or the louder the volume of the speech. The layout determination unit 47 also increases the size of each corresponding image and adds decoration (highlight display, etc.) to make the corresponding image more conspicuous, the greater the reaction indicated by the reaction information to the speech related to the corresponding image for each corresponding image (i.e., the greater the impact on other participants).

[0135] The amount of speech of a participant related to a certain dialogue image can be evaluated based on the number of characters in the text data corresponding to the corresponding image (or the participant), or the length of the audio data corresponding to the text data. The volume of speech of a participant related to a certain dialogue image can be evaluated based on the audio data corresponding to the corresponding image (or the participant). For example, the average or maximum volume of the audio data may be used as the volume of speech related to the audio data. The magnitude of the reaction indicated by the reaction information can be evaluated using publicly known technology (for example, the technology disclosed at https: / / aicam.jp / tech / optical_flow).

[0136] The layout determination unit 47 may structure the conversation (using the order of speech and natural language processing) from the flow of the conversation based on the text data (content of speech) corresponding to each corresponding image, and determine the layout of the corresponding images based on the structure. The structuring refers to, for example, the following (1) and (2).

[0137] (1) A time-series conversation structure, that is, an order of speech. In this case, the layout determination unit 47 determines the layout according to the time-series flow of speech corresponding to each corresponding image.

[0138] (2) This refers to converting conversations into a network-like structure using a graph generated by natural language processing (knowledge graph, co-occurrence network). In this graph, each utterance is a node, and related utterances are connected by edges. In this case, the layout determination unit 47 determines the layout of corresponding images based on the relevance of the utterances.

[0139] In step S113, the display information generation unit 44 generates display information (e.g., a screen) for displaying the generated images (each corresponding image or the integrated corresponding image, or each corresponding image and the integrated corresponding image) based on the display mode determined by the layout determination unit 47.

[0140] As a result, in the fifth embodiment, each corresponding image etc. is displayed on each terminal device 10 in the display mode determined by the layout determination unit 47.

[0141] Fig. 18 is a diagram showing a display example of a corresponding image etc. in the fifth embodiment. In Fig. 18, the vertical direction corresponds to the time axis (the passage of time), and the horizontal direction corresponds to the degree of relevance with the agenda image.

[0142] In this case, the position of each corresponding image is determined based on the chronological order and the degree of relevance to the agenda image. The size of each corresponding image is also determined by the method described above.

[0143] In the first embodiment, an example in which step S108 and subsequent steps are executed when the conference ends has been described, but step S108 and subsequent steps may be repeated at regular intervals during the conference. Fig. 18 shows an example in which corresponding images, etc., generated at regular intervals are arranged on the time axis in such a case. Fig. 18 also shows an example in which agenda images are arranged at regular intervals, but the agendas corresponding to the agenda images may be the same or different. In other words, the agenda may change during the conference.

[0144] The fifth embodiment may be combined with any of the above embodiments. In Fig. 18, the corresponding image to which the character string "(unspoken)" is attached indicates the corresponding image generated based on the reaction information in the fourth embodiment.

[0145] As described above, according to the fifth embodiment, the display mode of each corresponding image can be changed depending on the degree of influence of the corresponding image in the conference, thereby visually expressing the magnitude of the influence of each corresponding image.

[0146] Each function of this embodiment can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and a conventional circuit module designed to execute each function described above.

[0147] Additionally, the devices described herein represent only one of several computing environments for implementing the embodiments disclosed herein.

[0148] In one embodiment, server apparatus 40 includes multiple computing devices, such as a server cluster, configured to communicate with each other over any type of communications link, including a network, shared memory, etc., and to perform the processes disclosed herein.

[0149] In the present embodiment, the terminal device 10 and the server device 40 are examples of information processing devices.

[0150] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to such specific embodiments, and various modifications and variations are possible within the scope of the gist of the present invention as described in the claims.

[0151] For example, aspects of the present invention are as follows.

[0152] <1> an acquisition unit that acquires unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation unit that generates, for each participant, an image corresponding to the unique data acquired by the acquisition unit based on the unique data acquired by the acquisition unit and a machine learning model trained using learning data including the unique data indicating the unique information and an image; An information processing device comprising:

[0153] <2> a virtual data generation unit that generates virtual unique data of a virtual participant based on the unique data of the other participants; The image generation unit generating a corresponding image for the virtual participant based on the virtual unique data and the machine learning model, the corresponding image corresponding to the virtual unique data; <1> The information processing device described.

[0154] <3> the unique data includes first type unique data including first type unique information, and second type unique data including second type unique information based on a reaction to any of the first type unique information; The image generation unit generating the corresponding image for the second participant, the corresponding image corresponding to the first type of unique data and the second type of unique data, based on the first type of unique data for the first participant, the second type of unique data for the second participant, and a machine learning model; <1> or <2> The information processing device described.

[0155] <4> the first type of unique information is text indicating at least one of characters input by the participant and speech content of the participant; The second type of unique information includes at least one of operation information indicating an operation input by the participant, photographed image information of the participant photographed by the photographing device, and biometric information of each participant detected by the biometric information detection unit. <3> The information processing device described.

[0156] <5> The image generation unit further generates an integrated image corresponding to the set based on the set of unique data acquired by the acquisition unit for each of the plurality of participants and the machine learning model. Characterized by <1> ~ <4> Any of the information processing devices described above.

[0157] <6> the machine learning model is trained using training data including text and images; The image generation unit generates an image corresponding to the text data for each participant based on the text data indicating the content of each participant's speech and the machine learning model. <1> ~ <5> Any of the information processing devices described above.

[0158] <7> the image generation unit generates the corresponding image based on the text data, image data input in relation to the meeting, and the machine learning model. Characterized by <6> The information processing device described.

[0159] <8> a display information generating unit that generates display information that displays participant identification information that identifies the participant and the corresponding image generated for the participant in association with the participant identification information; characterized in that it has <1> ~ <7> Any of the information processing devices described above.

[0160] <9> a correlation calculation unit that calculates a correlation between the corresponding image and another image; the display information generation unit generates the display information that displays the degree of association in association with the participant identification information and the corresponding image. Characterized by <8> The information processing device described.

[0161] <10> the association degree calculation unit generates the association degree between the corresponding images; Characterized by <9> The information processing device described.

[0162] <11> a correlation calculation unit that generates a correlation indicating a correlation between the corresponding image and the integrated corresponding image; a display information generating unit that generates display information that displays participant identification information that identifies the participant, the corresponding image generated for the participant, and the degree of association generated for the corresponding image in association with each other; characterized in that it has <5> The information processing device described.

[0163] <12> a layout determination unit that determines at least one of an arrangement of the corresponding images when the corresponding images are displayed, a size of each of the corresponding images, and a decoration based on at least one of the unique information and the corresponding images; a display information generating unit that generates display information for displaying the plurality of corresponding images in at least one of the arrangement, size, and decoration determined by the layout determining unit; characterized in that it has <1> ~ <7> Any of the information processing devices described above.

[0164] <13> a correlation calculation unit that calculates a correlation between the corresponding image and another image; The layout determination unit determines the layout based on the degree of association. <12> The information processing device described.

[0165] <14> The layout determination unit determines the decoration or size of the corresponding image based on at least one of the amount of speech of the participant, the volume of the speech, image information showing the facial expressions of the other participants photographed by a photographing device, and changes in the biometric information of the other participants detected by a biometric information detection unit, which are included in the unique information. <12> or <13> The information processing device described.

[0166] <15> the display information generation unit generates the display information that displays the participant identification information, the corresponding image generated for the participant, and the unique data related to the participant in association with each other. Characterized by <8> The information processing device described.

[0167] <16> the image generation unit generates, for each participant, the corresponding image corresponding to the character string based on the character string obtained by dividing the text data included in the unique data acquired by the acquisition unit in time series and the machine learning model; the display information generation unit generates the display information that displays the character strings in chronological order. Characterized by <15> The information processing device described.

[0168] <17> A portion of the unique data that the image generation unit uses to generate the corresponding image is changeable. Characterized by <1> ~ <16> Any of the information processing devices described above.

[0169] <18> an acquisition unit that acquires unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation unit that generates, for each participant, an image corresponding to the unique data acquired by the acquisition unit based on the unique data acquired by the acquisition unit and a machine learning model trained using training data including the unique data indicating the unique information and an image; An information processing system comprising:

[0170] <19> an acquisition step of acquiring unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation step of generating, for each participant, an image corresponding to the unique data acquired in the acquisition step, based on the unique data acquired in the acquisition step and a machine learning model trained using training data including the unique data indicating the unique information and an image; An information processing method characterized by being executed by a computer.

[0171] <20> an acquisition step of acquiring unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation step of generating, for each participant, an image corresponding to the unique data acquired in the acquisition step, based on the unique data acquired in the acquisition step and a machine learning model trained using training data including the unique data indicating the unique information and an image; A program characterized by causing a computer to execute the above. [Explanation of symbols]

[0172] 10 Terminal Equipment 11 Input section 12 Input data transmission unit 13 Display control unit 40 Server device 41 Input data acquisition unit 42 Image generation unit 43 Relevance calculation unit 44 Display information generation section 45 Display information transmission unit 46 Virtual Data Generation Unit 47 Layout determination section 451 Input data storage unit 452 Model Memory Unit m1 image generation model m2 language model m3 image generation model [Prior art documents] [Patent documents]

[0173] [Patent Document 1] Japanese Patent Publication No. 2022-64301

Claims

1. an acquisition unit that acquires unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation unit that generates, for each participant, an image corresponding to the unique data acquired by the acquisition unit based on the unique data acquired by the acquisition unit and a machine learning model trained using training data including the unique data indicating the unique information and an image; An information processing device comprising:

2. a virtual data generation unit that generates virtual unique data of a virtual participant based on the unique data of the other participants; The image generation unit The information processing device according to claim 1 , further comprising: generating, for the virtual participant, the corresponding image corresponding to the virtual unique data based on the virtual unique data and the machine learning model.

3. the unique data includes first type unique data including first type unique information, and second type unique data including second type unique information based on a reaction to any of the first type unique information; The image generation unit 2. The information processing device of claim 1, wherein the corresponding image corresponding to the first type of unique data and the second type of unique data is generated for the second participant based on the first type of unique data of the first participant, the second type of unique data of the second participant, and a machine learning model.

4. the first type of unique information is text indicating at least one of characters input by the participant and speech content of the participant; 4. The information processing device according to claim 3, wherein the second type of unique information includes at least one of operation information indicating operations input by the participant, photographed image information of the participant photographed by a photographing device, and biometric information of each participant detected by a biometric information detection unit.

5. The image generation unit further generates an integrated image corresponding to the set based on the set of unique data acquired by the acquisition unit for each of the plurality of participants and the machine learning model.

2. The information processing apparatus according to claim 1, wherein:

6. the machine learning model is trained using training data including text and images; The information processing device according to claim 1 or 5, characterized in that the image generation unit generates a corresponding image corresponding to the text data for each participant of the meeting based on text data indicating the speech content of each participant and the machine learning model.

7. the image generation unit generates the corresponding image based on the text data, image data input in relation to the meeting, and the machine learning model.

7. The information processing apparatus according to claim 6,

8. a display information generating unit that generates display information that displays participant identification information that identifies the participant and the corresponding image generated for the participant in association with the participant identification information; 6. The information processing apparatus according to claim 1, further comprising:

9. a correlation calculation unit that calculates a correlation between the corresponding image and another image; the display information generation unit generates the display information that displays the degree of association in association with the participant identification information and the corresponding image.

9. The information processing apparatus according to claim 8,

10. the association degree calculation unit generates the association degree between the corresponding images; 10. The information processing apparatus according to claim 9,

11. a correlation calculation unit that generates a correlation indicating a correlation between the corresponding image and the integrated corresponding image; a display information generating unit that generates display information that displays participant identification information that identifies the participant, the corresponding image generated for the participant, and the degree of association generated for the corresponding image in association with each other; 6. The information processing apparatus according to claim 5, further comprising:

12. a layout determination unit that determines at least one of an arrangement of the corresponding images when the corresponding images are displayed, a size of each of the corresponding images, and a decoration based on at least one of the unique information and the corresponding images; a display information generating unit that generates display information for displaying the plurality of corresponding images in at least one of the arrangement, size, and decoration determined by the layout determining unit; 2. The information processing apparatus according to claim 1, further comprising:

13. a correlation calculation unit that calculates a correlation between the corresponding image and another image; The information processing apparatus according to claim 12 , wherein the layout determination unit determines the arrangement based on the degree of association.

14. The information processing device of claim 12, wherein the layout determination unit determines the decoration or size of the corresponding image based on at least one of the amount of speech of the participant, the volume of the speech, image information showing the facial expressions of other participants photographed by a photographing device, and changes in the biometric information of other participants detected by a biometric information detection unit, which are included in the unique information.

15. the display information generation unit generates the display information that displays the participant identification information, the corresponding image generated for the participant, and the unique data related to the participant in association with each other.

9. The information processing apparatus according to claim 8,

16. the image generation unit generates, for each participant, the corresponding image corresponding to the character string based on the character string obtained by dividing the text data included in the unique data acquired by the acquisition unit in time series and the machine learning model; the display information generation unit generates the display information that displays the character strings in chronological order.

16. The information processing apparatus according to claim 15,

17. A portion of the unique data that the image generation unit uses to generate the corresponding image is changeable.

2. The information processing apparatus according to claim 1, wherein:

18. an acquisition unit that acquires unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation unit that generates, for each participant, an image corresponding to the unique data acquired by the acquisition unit based on the unique data acquired by the acquisition unit and a machine learning model trained using learning data including the unique data indicating the unique information and an image; An information processing system comprising:

19. an acquisition step of acquiring unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation step of generating, for each participant, an image corresponding to the unique data acquired in the acquisition step, based on the unique data acquired in the acquisition step and a machine learning model trained using training data including the unique data indicating the unique information and an image; An information processing method characterized by being executed by a computer.

20. an acquisition step of acquiring unique data indicating unique information of each participant of a gathering based on the participant's behavior at the gathering; an image generation step of generating, for each participant, an image corresponding to the unique data acquired in the acquisition step, based on the unique data acquired in the acquisition step and a machine learning model trained using training data including the unique data indicating the unique information and an image; A program characterized by causing a computer to execute the above.

Citation Information

Patent Citations

  • Communication system, display device, display control method, and display control program

    JP2022064301A