Information processing device, information processing system, and information processing method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2022-09-09
- Publication Date
- 2026-05-15
Smart Images

Figure 0007859448000001 
Figure 0007859448000002 
Figure 0007859448000003
Abstract
Description
[Technical Field]
[0001] This disclosure relates to an information processing device, an information processing system, and an information processing method. [Background technology]
[0002] For example, in online meetings, online seminars, and online classes, presentation materials and the presenter's face may be displayed simultaneously. Similarly, in online live commentary of games and sports, the video of the subject being commentated on and the commentator's face may be displayed simultaneously. For example, as a technology related to online meetings, a technology has been developed that allows two or more terminals to mutually send and receive information about each other's meeting rooms (see, for example, Patent Document 1). [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2008-258779 [Overview of the project] [Problems that the invention aims to solve]
[0004] However, in typical online meetings, online seminars, and online classes, the on-screen positions of materials (presentations) and the presenter's face are predetermined, or one of them is fixed, while the other can be arbitrarily changed by the user via the user interface. Similarly, in online live broadcasts, the on-screen positions of the video being broadcast (presentations) and the presenter's face are usually predetermined and fixed.
[0005] When the document and the presenter's face are fixed on the screen in this way, the face is displayed next to the document to prevent it from overlapping, resulting in inefficient use of screen space. Furthermore, if either the document or the face is changed, and the face's position on the screen remains unchanged, important parts of the document may be obscured by the face on certain pages. To avoid this, the presenter must manually adjust the face's position on the screen according to the page, which is cumbersome. These issues also occur in online live broadcasts.
[0006] Therefore, this disclosure proposes an information processing device, an information processing system, and an information processing method that can appropriately display the presented material and the presenter while making effective use of the screen area. [Means for solving the problem]
[0007] The information processing apparatus according to the embodiment of the present disclosure includes: a selection unit that identifies a region to be composited in the second image, including a face image, while avoiding change regions where the amount of information in the first image and the second image has changed; and a compositing unit that composites the face image into the region to be composited in the second image, thereby generating a composite image including the second image and the face image.
[0008] An information processing system according to an embodiment of the present disclosure is an information processing system having a plurality of information processing devices, each of which includes: a selection unit that identifies a region to be composited in the second image, including a face image, while avoiding change regions where the amount of information in the first image and the second image has changed; and a compositing unit that composites the face image in the region to be composited in the second image, thereby generating a composite image including the second image and the face image.
[0009] An information processing method according to an embodiment of the present disclosure includes: identifying a region in the second image that is to be composited, including the face image, while avoiding change regions where the amount of information in the first image and the second image has changed; and compositing the face image onto the region to be composited in the second image to generate a composite image including the second image and the face image. [Brief explanation of the drawing]
[0010] [Figure 1] This figure shows an example of a schematic configuration of an information processing system according to an embodiment of this disclosure. [Figure 2] This figure shows an example of a schematic configuration of a server device according to an embodiment of this disclosure. [Figure 3] This figure shows an example of a schematic configuration of a terminal device for viewers according to an embodiment of this disclosure. [Figure 4] This figure shows an example of a schematic configuration of a terminal device for presenters according to an embodiment of this disclosure. [Figure 5] This flowchart shows the flow of a first processing example of a terminal device for presenters according to an embodiment of this disclosure. [Figure 6] This diagram illustrates the flow of a first processing example of a terminal device for presenters according to an embodiment of this disclosure. [Figure 7] This figure illustrates a modified example of the information processing system according to the present disclosure. [Figure 8] This flowchart shows the flow of a second processing example of a terminal device for presenters according to an embodiment of this disclosure. [Figure 9] This diagram illustrates the flow of a second processing example of a terminal device for presenters according to the present disclosure. [Figure 10] This diagram illustrates the flow of a third processing example of a terminal device for presenters according to the present disclosure. [Figure 11] This figure shows an example of a schematic configuration of the hardware according to the embodiment of this disclosure. [Modes for carrying out the invention]
[0011] Embodiments of this disclosure will be described in detail below with reference to the drawings. Note that these embodiments do not limit the apparatus, systems, methods, etc., related to this disclosure. Furthermore, in each of the following embodiments, the same reference numerals are used for essentially the same parts to avoid redundant explanations.
[0012] The one or more embodiments (including examples and modifications) described below can each be implemented independently. On the other hand, at least some of the embodiments described below may be implemented in appropriate combination with at least some of the other embodiments. These embodiments may contain novel features that differ from each other. Therefore, these embodiments may contribute to solving different objectives or problems and may produce different effects.
[0013] This disclosure will be explained in the order of the items shown below. 1. Embodiment 1-1. Example of an Information Processing System Configuration 1-2. Example of Server Device Configuration 1-3. Example configuration of a terminal device for viewers 1-4. Example of terminal device configuration for presenters 1-5. First processing example of the terminal device for the presenter 1-6. Modifications of Information Processing Systems 1-7. Second processing example of the terminal device for the presenter 1-8. Third processing example of the terminal device for the presenter 1-9. Action and Effects 2. Other Embodiments 3. Example Hardware Configuration 4. Addendum
[0014] <1. Embodiments> <1-1. Example of an information processing system configuration> An example of the configuration of the information processing system 1 according to this embodiment will be described with reference to Figure 1. Figure 1 is a diagram showing an example of the schematic configuration of the information processing system 1 according to this embodiment.
[0015] As shown in Figure 1, the information processing system 1 comprises a server device 10 and a plurality of terminal devices 20 and 30. The server device 10 and each terminal device 20 and 30 are configured to transmit and receive (communicate) various types of information via a network N that is either wireless or wired, or both. The server device 10 and each terminal device 20 and 30 each function as an information processing device. Note that Figure 1 is an example, and the number of server devices 10, each terminal device 20 and 30, and the network N are not limited.
[0016] Network N is a communication network such as a LAN (Local Area Network), WAN (Wide Area Network), cellular network, fixed telephone network, regional IP (Internet Protocol) network, or the Internet. Network N may include wired networks or wireless networks. Network N may also include a core network. A core network is, for example, an EPC (Evolved Packet Core) or a 5GC (5G Core network). Network N may also include data networks other than the core network. For example, a data network may be a service network of a telecommunications carrier, such as an IMS (IP Multimedia Subsystem) network. Alternatively, a data network may be a private network, such as an internal corporate network.
[0017] For example, a communication device such as a terminal device 20 may be configured to connect to the network N using radio access technologies (RATs) such as LTE (Long Term Evolution), NR (New Radio), Wi-Fi (registered trademark), and Bluetooth (registered trademark). In this case, the communication device may be configured to use different radio access technologies. For example, the communication device may be configured to use NR and Wi-Fi. Also, the communication device may be configured to use different cellular communication technologies (e.g., LTE and NR). LTE and NR are types of cellular communication technologies that enable mobile communication of communication devices such as the terminal device 20 by arranging multiple base stations in a cell-like structure over the area they cover.
[0018] Server device 10 is a server that manages and relays various types of information. For example, this server device 10 manages various types of information such as identification information related to each terminal device 20, 30, and also relays various types of information including image information such as images of presentation materials and images of the presenter's face (for example, images of presentation materials and the presenter's face). For example, server device 10 may be a cloud server, a PC server, a midrange server, or a mainframe server.
[0019] Terminal device 20 is a device owned by the user, the viewer. This terminal device 20 exchanges various information with server device 10, including image information such as images of the presented object and facial images of the presenter. For example, terminal device 20 can be a personal computer (e.g., a laptop or desktop computer), a smart device (e.g., a smartphone or tablet), a PDA (Personal Digital Assistant), or a mobile phone. Terminal device 20 may also be an xR device such as an AR (Augmented Reality) device, a VR (Virtual Reality) device, or an MR (Mixed Reality) device. Here, the xR device may be a glasses-type device (e.g., AR / MR / VR glasses), or a head-mounted or goggle-type device (e.g., an AR / MR / VR headset, or AR / MR / VR goggles). These xR devices may display an image to only one eye, or they may display an image to both eyes.
[0020] Terminal device 30 is a device owned by the presenter, who is the user. This terminal device 30 exchanges various types of information with the server device 10, including image information such as images of the presented object and the presenter's face. For example, terminal device 30 can be a personal computer (e.g., a laptop or desktop computer), a smart device (e.g., a smartphone or tablet), or a PDA (Personal Digital Assistant).
[0021] <1-2. Example of Server Device Configuration> An example of the configuration of the server device 10 according to this embodiment will be described with reference to Figure 2. Figure 2 is a diagram showing an example of the schematic configuration of the server device 10 according to this embodiment.
[0022] As shown in Figure 2, the server device 10 comprises a communication unit 11, a storage unit 12, and a control unit 13. Note that the configuration shown in Figure 2 is a functional configuration, and the hardware configuration may differ. Furthermore, the functions of the server device 10 may be distributed and implemented across multiple physically separated configurations. For example, the server device 10 may be composed of multiple server devices.
[0023] The communication unit 11 is a communication interface for communicating with other devices. For example, the communication unit 11 is a LAN (Local Area Network) interface such as a NIC (Network Interface Card). The communication unit 11 may be a wired interface or a wireless interface. The communication unit 11 communicates with each terminal device 20, 30, etc., according to the control of the control unit 13.
[0024] The memory unit 12 is a data read / write storage device such as DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), flash memory, or hard disk. This memory unit 12 stores various information as needed, according to the control of the control unit 13. For example, the memory unit 12 stores identification information related to each terminal device 20, 30, as well as various other information including image information such as images of the presented object and images of the presenter's face.
[0025] The control unit 13 is a controller that controls various parts of the server device 10. This control unit 13 is implemented by a processor such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit). For example, the control unit 13 is implemented by the processor executing various programs stored in the internal storage device of the server device 10 using RAM (Random Access Memory) or the like as a working area. The control unit 13 may also be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). CPU, MPU, ASIC, and FPGA can all be considered controllers. In addition to or instead of a CPU, the control unit 13 may also be implemented by a GPU (Graphics Processing Unit).
[0026] <1-3. Example of a terminal device configuration for viewers> An example of the configuration of the viewer terminal device 20 according to this embodiment will be described with reference to Figure 3. Figure 3 is a diagram showing an example of the schematic configuration of the viewer terminal device 20 according to this embodiment.
[0027] As shown in Figure 3, the terminal device 20 includes a communication unit 21, a storage unit 22, an input unit 23, an output unit 24, and a control unit 25. Note that the configuration shown in Figure 3 is a functional configuration, and the hardware configuration may differ. Furthermore, the functions of the terminal device 20 may be distributed and implemented across multiple physically separated configurations.
[0028] The communication unit 21 is a communication interface for communicating with other devices. For example, the communication unit 21 is a LAN interface such as a NIC. The communication unit 21 may be a wired interface or a wireless interface. The communication unit 21 communicates with the server device 10 and other devices according to the control of the control unit 25.
[0029] The memory unit 22 is a data read / write storage device such as DRAM, SRAM, flash memory, or hard disk. This memory unit 22 stores various information as needed, according to the control of the control unit 25.
[0030] The input unit 23 is an input device that receives various inputs from the outside. This input unit 23 is equipped with an operating device that receives input operations. The operating device is a device for the user to perform various operations, such as a keyboard, mouse, or operation keys. If the terminal device 20 employs a touch panel, the touch panel is also included as an operating device. In this case, the user performs various operations by touching the screen with their finger or stylus. Alternatively, the operating device may be a voice input device (e.g., a microphone) that receives input operations by the operator's voice.
[0031] The output unit 24 is a device that outputs various types of information to the outside, such as sound, light, vibration, and images. The output unit 24 is equipped with a display device that displays various types of information. The display device is, for example, a liquid crystal display or an organic EL (Electro-Luminescence) display. If a touch panel is used in the terminal device 20, the display device may be an integrated device with the operating device of the input unit 23. The output unit 24 outputs various types of information to the user according to the control of the control unit 25.
[0032] The control unit 25 is a controller that controls various parts of the terminal device 20. The control unit 25 is implemented by a processor such as a CPU or MPU. For example, the control unit 25 is implemented by the processor executing various programs stored in the memory unit 22 using RAM or the like as a working area. The control unit 25 may also be implemented by an integrated circuit such as an ASIC or FPGA. A CPU, MPU, ASIC, and FPGA can all be considered as controllers. In addition, the control unit 25 may be implemented by a GPU in addition to, or instead of, the CPU.
[0033] <1-4. Example of terminal device configuration for presenters> An example of the configuration of the presenter's terminal device 30 according to this embodiment will be described with reference to Figure 4. Figure 4 is a diagram showing an example of the schematic configuration of the presenter's terminal device 30 according to this embodiment.
[0034] As shown in Figure 4, the terminal device 30 includes a communication unit 31, a storage unit 32, an input unit 33, an output unit 34, and a control unit 35. Note that the configuration shown in Figure 4 is a functional configuration, and the hardware configuration may differ. Furthermore, the functions of the terminal device 30 may be implemented in a distributed manner across multiple physically separated configurations.
[0035] The communication unit 31 is a communication interface for communicating with other devices. For example, the communication unit 31 is a LAN interface such as a NIC. The communication unit 31 may be a wired interface or a wireless interface. The communication unit 31 communicates with the server device 10 and other devices according to the control of the control unit 35.
[0036] The memory unit 32 is a data read / write storage device such as DRAM, SRAM, flash memory, or hard disk. This memory unit 32 stores various information as needed, according to the control of the control unit 35.
[0037] The input unit 33 is an input device that receives various inputs from the outside. This input unit 33 includes an imaging device that acquires images and an operating device that accepts input operations. The imaging device is, for example, a camera having an image sensor. The operating device is, for example, a device for the user to perform various operations, such as a keyboard, mouse, or operation keys. If the terminal device 30 employs a touch panel, the touch panel is also included in the operating device. In this case, the user performs various operations by touching the screen with their finger or stylus. The operating device may also be a voice input device (for example, a microphone) that accepts input operations by the operator's voice.
[0038] The output unit 34 is a device that outputs various types of information to the outside, such as sound, light, vibration, and images. The output unit 34 is equipped with a display device that displays various types of information. The display device is, for example, a liquid crystal display or an organic EL display. If a touch panel is used in the terminal device 30, the display device may be an integrated device with the operating device of the input unit 33. The output unit 34 outputs various types of information to the user according to the control of the control unit 35.
[0039] The control unit 35 is a controller that controls various parts of the terminal device 20. The control unit 35 is implemented by a processor such as a CPU or MPU. For example, the control unit 35 is implemented by the processor executing various programs stored in the memory unit 32 using RAM or the like as a working area. The control unit 35 may also be implemented by an integrated circuit such as an ASIC or FPGA. A CPU, MPU, ASIC, and FPGA can all be considered controllers. In addition, the control unit 35 may be implemented by a GPU in addition to, or instead of, the CPU.
[0040] The control unit 35 comprises a creation unit 35a, a specification unit 35b, and a synthesis unit 35c. Each block constituting the control unit 35 (creation unit 35a, specification unit 35b, and synthesis unit 35c) is a functional block that represents the function of the control unit 35. These functional blocks may be software blocks or hardware blocks. For example, each block may be a single software module implemented in software (including a microprogram), or a single circuit block on a semiconductor chip (die). Of course, each block may also be a single processor or a single integrated circuit. The control unit 35 may be composed of functional units different from the above-mentioned blocks. The configuration method of each block is arbitrary. Furthermore, some or all of the operations of each block may be performed by other devices. The operation (processing examples) of each block constituting such a control unit 35 will be described later.
[0041] <1-5. First example of processing on the terminal device for the presenter> A first processing example of the terminal device 30 for a presenter according to this embodiment will be described with reference to FIGS. 5 and 6. FIG. 5 is a flowchart showing the flow of the first processing example of the terminal device 30 for a presenter according to this embodiment. FIG. 6 is a diagram for explaining the flow of the first processing example of the terminal device 30 for a presenter according to this embodiment.
[0042] In the examples of FIGS. 5 and 6, the information processing system 1 is a system that simultaneously presents presentation materials (presentations) and the face of a presenter (presenter) to viewers in an online meeting, online seminar, online class, etc.
[0043] As shown in FIGS. 5 and 6, in step S11, the creation unit 35a creates a difference image B N-1 from the previous screen (previous image A N ) and the current screen (current image A N ). In step S12, the specifying unit 35b specifies a synthesis target region R1 while avoiding a change region R2 in which the amount of information has changed in the difference image B N . In step S13, the synthesis unit 35c synthesizes a face image C N into the synthesis target region R1 of the current screen (current image A N ). Note that the previous image A N-1 and the current image A N are, for example, consecutive images. The previous image A N-1 corresponds to the first image, and the current image A N corresponds to the second image. Also, the synthesis target region R1 is, for example, the region with the least amount of information among the difference image B N in the size of the face image C N .
[0044] Specifically, as shown in FIG. 6, the creation unit 35a calculates the difference in the amount of information by comparing the previous image A N-1 (previous page) of the material with the current image A N (current page), and creates a difference image B N regarding the difference in the amount of information. Next, the specifying unit 35b is based on the difference image B N created by the creation unit 35a, and for that difference image B NIn this case, avoiding the change region R2, the face image C including the presenter's face is displayed. N The composite region R1 of the size is identified. The composite region 35c is the current image A N In this process, the region to be synthesized R1 identified by the specific unit 35b is used to create the facial image C N By combining them, the current image A N and facial image C N Composite image D including N This generates the current image A. N In this context, the face image C is positioned in a location that does not interfere with obtaining information from the document. N They are superimposed.
[0045] Here, for example, the materials (multiple images) are stored in the storage unit 32 of the terminal device 30 for the presenter, and are read out from the storage unit 32 for use. Also, face image C N This is acquired by an imaging device, which is part of the input unit 33 of the terminal device 30 for the presenter. The imaging device captures the upper body, including the presenter's face, etc., and therefore captures a face image C N This includes parts other than the face. Therefore, face image C N This is an image that includes at least a face.
[0046] Composite image D N The composite image D is displayed by a display device which is part of the output unit 34 of the presenter's terminal device 30, and is also transmitted to each viewer's terminal device 20 via the server device 10, where it is displayed by a display device which is part of the output unit 24 of each viewer's terminal device 20. The server device 10 receives the composite image D from the presenter's terminal device 30. N Received the composite image D. N This is transmitted via network N to terminal devices 20 for viewers participating in online meetings, online seminars, online classes, etc.
[0047] According to this process, the specific unit 35b generates the difference image B N In this case, avoid the change region R2 and capture the face image C N Identify the region R1 to be composited, which is of a certain size. This region R1 to be composited is, for example, the previous image A N-1 and current image AN In this case, facial image C N The region with the least change in information content at a given size (for example, face image C) N This is the region where the amount of information does not change at this size (current image A). N The identified synthesis region R1 contains the face image C N Since these are combined, face image C N This is the current image A N It is composited in an appropriate position, that is, in a position that does not interfere with obtaining information from the material as much as possible. For example, in the example in Figure 6, composite image D N In the current image A N The above information includes face image C N Although they are superimposed, the information for that area is in the previous image A. N-1 However, it was displayed and was known to the viewers. Therefore, the face image C is placed in that location. N Overlapping images poses no problem. In this way, it becomes possible to overlay the materials (presentations) and the face (presenter) while minimizing the loss of information, thus enabling effective use of the screen area. Furthermore, by tracking changes in the materials (presentations), effective use of the screen area can be achieved automatically without requiring any effort from the presenter or viewers.
[0048] Here, facial image C N The size is predetermined, but is not limited to this, and may be changed (enlarged or reduced) by the user, for example. For example, the user who is presenting the image may operate the input unit 33 to input a face image C N The size may be changed, or the user, who is a viewer, may operate the input unit 23 to change the size of the face image C N You may change the size of the face image C. N The size may be changed automatically. For example, the specific part 35b may change the face image C modified by the user. N The size of, or the automatically modified face image C N The region R1 to be synthesized is identified according to its size.
[0049] Furthermore, the specific part 35b is the difference image B NBased on this, the change region R2 is determined, and the region to be combined R1 is identified while avoiding the determined change region R2, but this is not limited to this. For example, the identification unit 35b identifies the difference image B N Without using the previous image A N-1 and current image A N Alternatively, changes in entropy (e.g., randomness) and blank areas (e.g., white areas) can be determined, and based on the obtained information on changes in entropy and white areas, the change region R2 can be determined, and the region to be combined R1 can be identified while avoiding the determined change region R2.
[0050] Furthermore, the composite section 35c is shown in the previous image A. N-1 From current image A N At the moment of switching, composite image D N Alternatively, you may generate the previous image A N-1 A part of it (for example, the animation) changes to the current image A N At this timing, composite image D N These may be generated. While these timings are pre-set, they are not limited to these and may be changed by the user, for example, as described above. Note that face image C N This is updated at predetermined intervals (for example, at regular intervals). The predetermined interval is set in advance, but is not limited to this, and may be changed by the user, for example, as described above.
[0051] <1-6. Variations of Information Processing Systems> A modified example of the information processing system 1 according to this embodiment will be described with reference to Figure 7. Figure 7 is a diagram illustrating a modified example of the information processing system 1 according to this embodiment.
[0052] As shown in Figure 7, the information processing system 1, as described above, uses a terminal device 30 for the presenter to create a composite image D N The system generates (A: synthesized on the transmitting side), but is not limited to this. For example, the information processing system 1 generates a synthesized image D on the server device 10. N Alternatively, the composite image D may be generated (B: synthesized on the server side), and the viewer terminal device 20 will display the composite image D. NYou may generate this (C: synthesized on the receiving end).
[0053] A: When synthesis is performed on the transmitting side, the presenter's terminal device 30 has a creation unit 35a, a specification unit 35b, and a synthesis unit 35c, as described above. This presenter's terminal device 30 produces a synthesized image D through the above process. N Generate and create the composite image D N The composite image D transmitted from the terminal device 30 is sent to the server device 10. N Received the composite image D. N The composite image D transmitted from the server device 10 is sent to each viewer's terminal device 20. Each viewer's terminal device 20 receives the composite image D transmitted from the server device 10. N The signal is received and displayed by the display device, which is part of the output unit 24. The presenter then uses the input unit 33 of the terminal device 30 (for example, a mouse or touch panel) to input the face image C. N It is also possible to change its position.
[0054] B: When the synthesis is performed on the server side, the presenter's terminal device 30 displays the current image A of the document. N and face image C N The data is transmitted directly to the server device 10. This server device 10 has a creation unit 35a, a specification unit 35b, and a synthesis unit 35c. The server device 10 receives the original image A of the data transmitted from the terminal device 30. N and face image C N Upon receiving the image, the above process is performed to create a composite image D. N The system generates and transmits the composite image D to each viewer's terminal device 20. Each viewer's terminal device 20 receives the composite image D transmitted from the server device 10. N The data is received and displayed by a display device, which is part of the output unit 24. In this type of processing, a series of synthesis processes are performed on the server device 10, which reduces the processing load on the other terminal devices 20 and 30.
[0055] C: When synthesis is performed on the receiving end, the presenter's terminal device 30 displays the original image A of the document. N and face image C NThe data is sent directly to the server device 10. The server device 10 receives the current image A of the material sent from the presenter's terminal device 30. N and face image C N Received, current image A of the received material N and facial image C N This is transmitted directly to each viewer's terminal device 20. Each of these viewer terminal devices 20 has a creation unit 35a, a specification unit 35b, and a synthesis unit 35c. Each viewer terminal device 20 receives the original image A of the material transmitted from the server device 10. N and facial image C N Upon receiving the image, the above process is performed to create a composite image D. N It generates and displays it on the display device, which is part of the output unit 24. The viewer can input the face image C by using the input unit 23 of the terminal device 20 (for example, a mouse or touch panel). N It is also possible to change its position.
[0056] Although the presenter's terminal device 30 and the viewer's terminal device 20 communicate via the server device 10, this is not the only option. For example, direct communication (P2P) without going through the server device 10 may also be used. In this case, A: the sender may perform the synthesis, or C: the receiver may perform the synthesis.
[0057] <1-7. Second example of processing on the presenter's terminal device> A second processing example of the presenter's terminal device 30 according to this embodiment will be described with reference to Figures 8 and 9. Figure 8 is a flowchart showing the flow of the second processing example of the presenter's terminal device 30 according to this embodiment. Figure 9 is a diagram illustrating the flow of the second processing example of the presenter's terminal device 30 according to this embodiment.
[0058] In the examples of Figures 8 and 9, similar to the examples of Figures 5 and 6, Information Processing System 1 is a system that simultaneously displays presentation materials (presentations) and the presenter's face to viewers during online meetings, online seminars, online classes, etc.
[0059] As shown in FIGS. 8 and 9, in step S11, the creation unit 35a creates a difference image B from the previous screen (previous image A N-1 ) and the current screen (current image A N ). In step S21, the specifying unit 35b specifies a composite region R1 from predetermined candidate regions Ra, Rb, and Rc while avoiding a change region R2 where the amount of information has changed in the difference image B N . In step S13, the composite unit 35c composites the face image C N into the composite region R1 of the current screen (current image A N ).
[0060] Each of the candidate regions Ra, Rb, and Rc is a region that is a candidate for the position (composite region R1) where the face image C N is overlaid, and is determined in advance. The individual sizes of these candidate regions Ra, Rb, and Rc are, for example, the same as the size of the face image C N . In the example of FIG. 9, three predetermined candidate regions Ra, Rb, and Rc are determined in advance on the right side of the screen. The specifying unit 35b selects the composite region R1 from among those candidate regions Ra, Rb, and Rc. As a result, the face image C N is overlaid at the best position among the candidate regions Ra, Rb, and Rc.
[0061] Note that a priority (priority order) may be assigned in advance to each of the candidate regions Ra, Rb, and Rc. In the example of FIG. 9, among the three candidate regions Ra, Rb, and Rc, the upper candidate region Ra has the first priority, the lower candidate region Rc has the second priority, and the middle candidate region Rb has the third priority. In the example of FIG. 9, among the candidate regions Ra, Rb, and Rc, the candidate region Rc overlaps with the change region R2, so the composite region R1 is selected based on the priority from the two candidate regions Ra and Rb. As a result, the candidate region Ra is specified as the composite region R1. In this way, by setting a candidate region that is preferentially selected, that candidate region is preferentially determined as the composite region R1, and the face image C N is preferentially displayed in the composite region R1, so that the viewer can easily predict the position of the presenter's face.
[0062] Here, each candidate region Ra, Rb, Rc and its predetermined priority are set in advance, but are not limited to these and may be changed by the user, for example. For example, the presenter user may change each candidate region Ra, Rb, Rc and its predetermined priority by operating the input unit 33, or the viewer user may change each candidate region Ra, Rb, Rc and its predetermined priority by operating the input unit 23.
[0063] Furthermore, the individual sizes of each candidate region Ra, Rb, and Rc are as follows: N The size is the same as, but not limited to, the size of each candidate region Ra, Rb, and Rc relative to the face image C. N The sizes may be larger than the specified size, and they may also be different from each other. Furthermore, the area to the right or below the center of the screen may be set as a predetermined candidate area. Note that the number of each candidate area, Ra, Rb, and Rc, is not limited.
[0064] <1-8. Third Processing Example of the Presenter's Terminal Device> A third processing example of the presenter's terminal device 30 according to this embodiment will be described with reference to Figure 10. Figure 10 is a diagram illustrating the flow of the third processing example of the presenter's terminal device 30 according to this embodiment.
[0065] In the example shown in Figure 10, similar to the examples in Figures 5 and 6, the information processing system 1 is a system that simultaneously presents images (presentations) of the subject of the live commentary and the face of the commentator (presenter) to viewers during online live commentary of games, sports, etc. The processing flow in the example shown in Figure 10 is the same as in the examples in Figures 5 and 6.
[0066] As shown in Figure 10, the creation unit 35a, the identification unit 35b, and the synthesis unit 35c basically perform the same processing as steps S11 to S13 in Figure 5. Specifically, for example, the creation unit 35a captures the screen at predetermined intervals (for example, at regular intervals) and the previous screen of the game screen (previous image A N-1 ) and the current screen (current image A) N) are compared to calculate the difference in information, and difference image B is created regarding the difference in information. N Next, the specific unit 35b creates the difference image B created by the creation unit 35a. N Within this, avoid the change region R2 (for example, the region with movement) where the amount of information changes, and select the face image C. N The size of the composite region R1 (for example, face image C) N The composite unit 35c identifies a region (without movement) of a certain size. The composite unit 35c then places the face image C onto the identified composite region R1. N By combining them, the current image A N and facial image C N Composite image D including N Generates.
[0067] According to this process, the specific unit 35b generates the difference image B N In this case, avoid the change region R2 and capture the face image C N Identify the region R1 to be composited, which is of a certain size. This region R1 to be composited is, for example, the previous image A N-1 and current image A N In this case, facial image C N The region with the least change in information content at a given size (for example, face image C) N The region where the amount of information does not change at the same size, i.e., the facial image C N This is a stationary area of this size (Image A). N The identified synthesis region R1 contains the face image C N Since these are combined, face image C N This is the current image A N It is composited in an appropriate position, that is, a position that does not interfere with obtaining information from the screen as much as possible. For example, while displaying a game image (game video) that fills the entire screen, a face image C is composited in a position that does not interfere with the expression of movement, at a timing that does not interfere with the expression of movement. N The images can be superimposed. The composite unit 35c superimposes the facial image C in the moving region. N Because it does not synthesize, the face image C is not visible in the continuous images as long as there is movement. N Sometimes it may not be displayed.
[0068] Here, the predetermined interval for the screen capture (for example, a fixed interval) is set in advance, but is not limited to this, and may be changed by the user, for example. For example, the user who is presenting may change the predetermined interval by operating the input unit 33, or the user who is viewing may change the predetermined interval by operating the input unit 23. The predetermined interval may also be changed automatically.
[0069] Furthermore, the third processing example can also be applied to the first and second processing examples. For example, in the first or second processing example, the specific unit 35b generates a face image C based on the movement of a display object such as a mouse pointer that is displayed in response to a mouse which is part of the input unit 33. N The size of the composite region R1 may be specified. For example, a region with movement is a region where the amount of information changes, and a region without movement is a region where the amount of information does not change.
[0070] Furthermore, the composite section 35c is in the current image A N Depending on whether or not there is a moving area within the face image C N The size may be changed. For example, the composite part 35c is the current image A N If there is no area of movement within the face image C N You may enlarge this face image C from its original size. N After enlargement, current image A N If there is a moving area within the face image C N You can shrink it and then return it to its original size.
[0071] Furthermore, the composite section 35c is in the current image A N Depending on whether or not there are areas of movement within the current image A N Face image C relative to the composite region R1 N The synthesis process may be switched between being performed and not being performed. For example, the synthesis unit 35c processes the current image A N If there is no area of movement within the current image A N Face image C relative to the composite region R1 N The synthesis was performed, resulting in the current image A NIf there is a moving region within the current image A N Face image C relative to the composite region R1 N The synthesis will not be performed.
[0072] Furthermore, the composite section 35c is in the current image A N Depending on the situation, facial image C N The size may be changed. For example, if the scene is a live commentary of a game or sport, i.e., if the presenter is providing commentary, the face image C may appear larger compared to when there is no commentary. N You may enlarge this face image C. N After the enlargement, when the live commentary scene ends, face image C N You may reduce the size and return it to its original size. Also, when the scene is in the middle of a question and answer session, the facial image C is different compared to when it is not in the middle of a question and answer session. N You may enlarge this face image C. N After the enlargement, when the Q&A session ends, face image C N You can shrink it and then return it to its original size.
[0073] <1-9. Action and Effects> As described above, according to this embodiment, the terminal device 30 for the presenter (an example of an information processing device) displays the first image (for example, previous image A N-1 ) and a second image (for example, current image A) N Avoiding the change region R2 where the amount of information changes in ), the face image C including the face is generated. N A identifying unit 35b identifies a composite region R1 of a certain size, and a face image C is placed in the composite region R1 of the second image. N The two images are combined to form the second image and the face image C. N Composite image D including N It includes a synthesis unit 35c that generates a face image C including the presenter's face. N However, in the second image, which is the presented object, it is overlaid in a position that does not interfere with obtaining information from the presented object. Therefore, the presented object and the presenter can be displayed appropriately while making effective use of the screen area.
[0074] Furthermore, the presenter's terminal device 30 generates a difference image B from the first image and the second image.N The creation unit 35a further comprises a creation unit 35a that creates a difference image B, and the identification unit 35b creates a difference image B N The change region R2 may be determined based on this, and the region to be combined R1 may be identified while avoiding the determined change region R2. This ensures that the region to be combined R1 of the second image is reliably identified.
[0075] Furthermore, the identification unit 35b may determine the change region R2 based on the change in entropy or blank area between the first image and the second image, and identify the composite region R1 while avoiding the determined change region R2. This ensures that the composite region R1 of the second image is reliably identified.
[0076] Furthermore, the specific part 35b is the difference image B N The region to be synthesized R1 may be identified from multiple predetermined candidate regions within (for example, candidate regions Ra, Rb, Rc). This allows each candidate region to be set, so that only one of those candidate regions is used for the face image C. N Because this is displayed, viewers can more easily predict the position of the face.
[0077] Furthermore, the designated candidate area may be pre-set by the user. This allows the user, such as the presenter or viewer, to place the face image C in the desired position on the screen. N It can be displayed.
[0078] Furthermore, the identification unit 35b may identify the region to be synthesized R1 from a plurality of predetermined candidate regions based on a predetermined priority. As a result, a candidate region with a high priority is set, and the face image C is synthesized in that candidate region with a high priority. N Because it is displayed preferentially, viewers can more easily predict the position of the face.
[0079] Furthermore, predetermined priorities may be set in advance by the user. This allows the user, such as the presenter or viewer, to place the face image C in the desired area within each candidate area. N It can be displayed.
[0080] Furthermore, the specific part 35b is the face image CN As the image is enlarged, the enlarged facial image C N The size of the composite region R1 may be specified. This allows for the identification of the face image C N The size of the face image C will increase, N It can be made easier to see.
[0081] Furthermore, the specific part 35b is the face image C N In accordance with the reduction, the reduced face image C N The size of the composite region R1 may be specified. This allows for the identification of the face image C N Since the size will be reduced, the face image C will be placed in the appropriate position in the second image. N It can be reliably synthesized.
[0082] Furthermore, the composite unit 35c generates the composite image D at the timing when switching from the first image to the second image. N This may generate the composite image D at the appropriate time. N You can obtain this.
[0083] Furthermore, the composite unit 35c generates the composite image D at the timing when a part of the first image changes to become the second image. N This may generate the composite image D at the appropriate time. N You can obtain this.
[0084] Furthermore, the synthesis unit 35c adjusts the facial image C depending on whether or not there is a moving region within the second image. N The size may be changed. This will change the face image C depending on whether there is movement from the first image to the second image. N Enlarge the image, or also, face image C N It can be reduced in size.
[0085] Furthermore, if there are no moving regions in the second image, the composite unit 35c will process the face image C N When the second image is enlarged and there is a moving region, face image C N It may be reduced in size. This means that if there is no movement from the first image to the second image, the face image C NSince it is enlarged, the viewer can be drawn to the presenter. On the other hand, if there is movement from the first image to the second image, face image C N Because the first image is reduced in size, the viewer's attention can be drawn to the second image, which is the presentation.
[0086] Furthermore, the synthesis unit 35c adjusts the face image C for the region to be synthesized R1 depending on whether or not there is a moving region within the second image. N The synthesis process may be repeated, with or without it being performed. This allows the face image C to be generated depending on whether or not there is movement from the first image to the second image. N This involves performing synthesis, and also, facial image C N It is possible to choose not to perform the synthesis.
[0087] Furthermore, if there are no moving regions in the second image, the composite unit 35c will process the face image C for the region to be composited R1. N When the synthesis is performed and there is a moving region in the second image, the face image C is compared to the region R1 to be synthesized. N The synthesis may be stopped. In this case, if there is no movement from the first image to the second image, the face image C will be added to the second image. N Since these are combined, the viewer's attention can be drawn to the presenter. On the other hand, if there is movement from the first image to the second image, the face image C is added to the second image. N Since the two images are not combined, the viewer's attention can be drawn to the second image, which is the presented object.
[0088] Furthermore, the synthesis unit 35c processes the face image C according to the scene relating to the second image. N The size may be changed. This allows viewers to focus their attention on the presentation or the presenter, depending on the context.
[0089] Furthermore, if the scene is being broadcast live, the composite unit 35c will capture the face image C N The image may be reduced in size. This allows the viewer to focus on the second image, which is the presentation.
[0090] Furthermore, if the scene is in the middle of a question-and-answer session, the composite unit 35c will process the face image C NYou may enlarge it. This will draw the viewer's attention to the presenter.
[0091] <2. Other Embodiments> The processes described in the above-described embodiments (or variations) may be carried out in various other forms (variations) besides those described in the above embodiments. For example, all or part of the processes described as being performed automatically in the above embodiments may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.
[0092] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.
[0093] Furthermore, the embodiments (or modifications) described above can be combined as appropriate, provided that the processing content is not contradictory. Also, the effects described herein are merely examples and not limiting, and other effects may exist.
[0094] <3. Hardware Configuration Examples> Specific hardware configuration examples of information devices such as the server device 10 and each terminal device 20, 30 according to the above-described embodiment (or modified version) will now be explained. The information devices such as the server device 10 and each terminal device 20, 30 according to the embodiment (or modified version) may be implemented by a computer 500 having a configuration such as that shown in Figure 11. Figure 11 is a diagram showing an example of hardware configuration that realizes the functions of the information devices such as the server device 10 and each terminal device 20, 30 according to the embodiment (or modified version).
[0095] As shown in Figure 11, the computer 500 has a CPU 510, RAM 520, ROM (Read Only Memory) 530, HDD (Hard Disk Drive) 540, a communication interface 550, and an input / output interface 560. The various parts of the computer 500 are connected by a bus 570.
[0096] The CPU 510 operates based on programs stored in the ROM 530 or HDD 540, and controls various parts. For example, the CPU 510 loads the programs stored in the ROM 530 or HDD 540 into the RAM 520 and executes processing corresponding to the various programs.
[0097] ROM530 stores boot programs such as the BIOS (Basic Input Output System) that are executed by the CPU510 when the computer 500 starts up, as well as programs that depend on the computer 500's hardware.
[0098] HDD540 is a computer-readable recording medium that non-temporarily records programs executed by CPU510 and data used by such programs. Specifically, HDD540 is a recording medium that records an information processing program related to this disclosure, which is an example of program data 541.
[0099] The communication interface 550 is an interface for the computer 500 to connect to an external network 580 (for example, the internet). For example, the CPU 510 receives data from other devices and transmits data it generates to other devices via the communication interface 550.
[0100] The input / output interface 560 is an interface for connecting the input / output device 590 and the computer 500. For example, the CPU 510 receives data from input devices such as a keyboard or mouse via the input / output interface 560. The CPU 510 also transmits data to output devices such as a display, speaker, or printer via the input / output interface 560.
[0101] The input / output interface 560 may also function as a media interface for reading programs and other data recorded on a predetermined recording medium (media). Examples of media include optical recording media such as DVDs (Digital Versatile Discs) and PDs (Phase Change Rewritable Disks), magneto-optical recording media such as MOs (Magneto-Optical disks), tape media, magnetic recording media, or semiconductor memory.
[0102] Here, for example, if the computer 500 functions as an information device such as a server device 10 or terminal devices 20, 30 according to the embodiment (or modified version), the CPU 510 of the computer 500 executes an information processing program loaded onto the RAM 520 to realize all or part of the functions of each part of the server device 10 or terminal devices 20, 30 according to the embodiment (or modified version). The HDD 540 stores the information processing program and data according to this disclosure (various information including image information such as images of the presented object and facial images of the presenter). The CPU 510 reads and executes the program data 541 from the HDD 540, but as another example, these programs may be obtained from other devices via an external network 580.
[0103] <4. Addendum> Furthermore, this technology can also be configured as follows. (1) A selection unit that identifies the area to be composited in the second image, including the face image, by avoiding the change region where the amount of information in the first and second images changes, A synthesis unit that synthesizes the face image into the area to be synthesized in the second image and generates a composite image including the second image and the face image, An information processing device equipped with the following features. (2) The system further comprises a creation unit that creates a difference image from the first image and the second image, The identifying unit determines the change region based on the difference image, and identifies the region to be combined while avoiding the determined change region. The information processing device described in (1) above. (3) The identifying unit determines the change region based on the change in entropy or blank area between the first image and the second image, and identifies the area to be combined while avoiding the determined change region. The information processing device described in (1) above. (4) The identification unit identifies the region to be combined from a plurality of predetermined candidate regions in the difference image. The information processing device described in (2) above. (5) The aforementioned predetermined candidate area is set in advance by the user. The information processing device described in (4) above. (6) The identifying unit identifies the region to be synthesized from a plurality of predetermined candidate regions based on a predetermined priority. The information processing device described in (4) above. (7) The aforementioned predetermined priority is set in advance by the user. The information processing device described in (6) above. (8) The specified unit identifies the area to be composited in the size of the enlarged face image in accordance with the enlargement of the face image. An information processing device as described in any one of the above (1) to (7). (9) The identifying unit identifies the area to be composited in the reduced size of the face image in accordance with the reduction of the face image. An information processing device as described in any one of the above (1) to (7). (10) The synthesis unit generates the synthesized image at the timing when switching from the first image to the second image. An information processing device as described in any one of the above (1) through (9). (11) The synthesis unit generates the composite image at the timing when a part of the first image changes to become the second image. An information processing device as described in any one of the above (1) through (9). (12) The synthesis unit changes the size of the face image depending on whether or not there is a moving region in the second image. An information processing device as described in any one of the above (1) through (11). (13) The synthesis unit enlarges the face image if the moving region does not exist in the second image, and reduces the face image if the moving region exists in the second image. The information processing device described in (12) above. (14) The synthesis unit repeatedly performs and does not perform the synthesis of the face image on the region to be synthesized, depending on whether or not there is a region with movement in the second image. An information processing device as described in any one of the above (1) through (11). (15) The synthesis unit performs the synthesis of the face image to the region to be synthesized if the region with movement does not exist in the second image, and stops the synthesis of the face image to the region to be synthesized if the region with movement exists in the second image. The information processing device described in (14) above. (16) The synthesis unit changes the size of the face image according to the scene relating to the second image. An information processing device as described in any one of the above (1) through (15). (17) The aforementioned synthesis unit reduces the size of the face image when the scene is being broadcast live. The information processing device described in (16) above. (18) The synthesis unit, when the scene is in the middle of a question-and-answer session, enlarges the facial image. The information processing device described in (16) above. (19) An information processing system having multiple information processing devices, Any one of the aforementioned information processing devices is A selection unit that identifies the area to be composited in the second image, including the face image, by avoiding the change region where the amount of information in the first and second images changes, A synthesis unit that synthesizes the face image into the area to be synthesized in the second image and generates a composite image including the second image and the face image, An information processing system equipped with the following features. (20) To identify the area of the face image, including the face, that is to be composited in the second image, while avoiding the change region where the amount of information changes in the first and second images, The process involves compositing the face image onto the area to be composited in the second image, thereby generating a composite image that includes the second image and the face image. Information processing methods, including those mentioned above. (twenty one) An information processing system comprising an information processing device described in any one of the above (1) to (18). (twenty two) An information processing method using an information processing device described in any one of the above (1) to (18). [Explanation of Symbols]
[0104] 1. Information Processing System 10 Server devices 11 Communications Department 12 Storage section 13 Control Unit 20 Terminal devices 21 Communications Department 22 Memory section 23 Input section 24 Output section 25 Control Unit 30 Terminal devices 31 Communications Department 32 Storage section 33 Input section 34 Output section 35 Control Unit 35a Creation Section 35b Specific part 35c Synthesis Department 500 Computers 541 Program Data 550 Communication Interfaces 560 Input / Output Interfaces 570 Bus 580 External Network 590 Input / Output Devices A N Current image A N-1 Previous B N Difference image C N Face image D N Composite image N Network R1 Area to be synthesized R2 Change Area Ra candidate area Rb candidate area Rc candidate area
Claims
1. A selection unit that identifies the area to be composited in the second image, including the face image, by avoiding the change region where the amount of information in the first image and the second image has changed, A synthesis unit that synthesizes the face image into the region to be synthesized in the second image and generates a composite image including the second image and the face image, An information processing device equipped with the following features.
2. The system further comprises a creation unit that creates a difference image from the first image and the second image, The identifying unit determines the change region based on the difference image, and identifies the region to be combined while avoiding the determined change region. The information processing apparatus according to claim 1.
3. The identifying unit determines the change region based on the change in entropy or blank area between the first image and the second image, and identifies the area to be combined while avoiding the determined change region. The information processing apparatus according to claim 1.
4. The identification unit identifies the region to be combined from a plurality of predetermined candidate regions in the difference image. The information processing apparatus according to claim 2.
5. The aforementioned predetermined candidate area is set in advance by the user. The information processing apparatus according to claim 4.
6. The identifying unit identifies the region to be synthesized from a plurality of predetermined candidate regions based on a predetermined priority. The information processing apparatus according to claim 4.
7. The aforementioned predetermined priority is set in advance by the user. The information processing apparatus according to claim 6.
8. The specified unit identifies the area to be composited in the size of the enlarged face image in accordance with the enlargement of the face image. The information processing apparatus according to claim 1.
9. The identifying unit identifies the area to be composited in the reduced size of the face image in accordance with the reduction of the face image. The information processing apparatus according to claim 1.
10. The synthesis unit generates the synthesized image at the timing when switching from the first image to the second image. The information processing apparatus according to claim 1.
11. The synthesis unit generates the composite image at the timing when a part of the first image changes to become the second image. The information processing apparatus according to claim 1.
12. The synthesis unit changes the size of the face image depending on whether or not there is a moving region in the second image. The information processing apparatus according to claim 1.
13. The synthesis unit enlarges the face image if the moving region does not exist in the second image, and reduces the face image if the moving region exists in the second image. The information processing apparatus according to claim 12.
14. The synthesis unit repeatedly performs and does not perform the synthesis of the face image on the region to be synthesized, depending on whether or not there is a region with movement in the second image. The information processing apparatus according to claim 1.
15. The synthesis unit performs the synthesis of the face image to the region to be synthesized if the region with movement does not exist in the second image, and stops the synthesis of the face image to the region to be synthesized if the region with movement exists in the second image. The information processing apparatus according to claim 14.
16. The synthesis unit changes the size of the face image according to the scene relating to the second image. The information processing apparatus according to claim 1.
17. The aforementioned synthesis unit reduces the size of the face image when the scene is being broadcast live. The information processing apparatus according to claim 16.
18. The synthesis unit, when the scene is in the middle of a question-and-answer session, enlarges the facial image. The information processing apparatus according to claim 16.
19. An information processing system having multiple information processing devices, Any one of the aforementioned information processing devices is A selection unit that identifies the area to be composited in the second image, including the face image, by avoiding the change region where the amount of information in the first image and the second image has changed, A synthesis unit that synthesizes the face image into the region to be synthesized in the second image and generates a composite image including the second image and the face image, An information processing system equipped with the following features.
20. To identify the area to be composited in the second image, including the face image, while avoiding the change region where the amount of information in the first and second images changes, The process involves compositing the face image onto the composite region of the second image to generate a composite image including the second image and the face image. Information processing methods, including those mentioned above.