Image processing method, image processing system and program

The image processing method dynamically adjusts the framing of both the speaker and virtual background to address the unnatural display issue in teleconferences, providing a more realistic and comfortable viewing experience.

JP2025142653APending Publication Date: 2025-10-01YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024042128
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-10-01

AI Technical Summary

Technical Problem

In teleconferences and similar events, the use of composite images combining a speaker's image with a virtual background results in an unnatural display when the virtual background does not change with the speaker's position, causing discomfort to viewers.

Method used

An image processing method that adjusts the framing of both the speaker and virtual background images in real-time based on changes in the speaker's position, ensuring a natural and dynamic composite image display.

Benefits of technology

The method achieves a natural and comfortable viewing experience by dynamically adjusting the virtual background to match the speaker's movements, enhancing the realism and smoothness of remote conferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025142653000001_ABST
    Figure 2025142653000001_ABST
Patent Text Reader

Abstract

To provide an image processing method that provides a no-discomfort display using a composite image composed of an image of a speaker which changes as the speaker is traced and an image of a virtual background after framing processing based upon the change of the speaker in a remote conference, etc.SOLUTION: An image processing method includes: determining second framing information representing the position and size of a virtual background in an image of the virtual background based upon first framing information having changed when the first framing information has changes; performing second framing processing on the image of the virtual background based upon the determined second frame information; and composing an image of an image including a speaker after the first framing processing and an image of the virtual background after the second framing.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] An embodiment of the present invention relates to an image processing method, an image processing device, and a program. [Background technology]

[0002] Patent Document 1 describes an image generating device that generates an image of a virtual conference space for a video conference. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 11-282479 Summary of the Invention [Problem to be solved by the invention]

[0004] In teleconferences and other similar events, a composite image is sometimes used, combining an image of a speaker captured by a camera with an image of a virtual background. Conventionally, when a framing process is performed to track the speaker with a camera, the virtual background does not change even if the image of the speaker changes, which can cause users to feel uncomfortable when viewing the composite image.

[0005] One embodiment of the present invention aims to provide an image processing method that realizes a natural display in remote conferences, etc., by using a composite image that combines an image of a speaker that changes to track the speaker with an image of a virtual background. [Means for solving the problem]

[0006] An image processing method according to an embodiment of the present invention includes: Acquire an image containing the speaker, performing a first framing process on the image including the speaker based on first framing information indicating a position and a size of the speaker within the image including the speaker; Acquires an image of a virtual background stored in a storage unit; synthesizing the image including the speaker that has undergone the first framing process with the image of the virtual background; 1. An image processing method, comprising: When the first framing information has changed, second framing information indicating a position and a size of the virtual background within the image of the virtual background is determined based on the changed first framing information; performing a second framing process on the image of the virtual background based on the determined second framing information; The image including the speaker that has undergone the first framing process is combined with the image of the virtual background that has undergone the second framing process. [Effects of the Invention]

[0007] According to an image processing method of one embodiment of the present invention, in a remote conference or the like, a natural display can be achieved by using a composite image that combines an image of the speaker that changes by tracking the speaker with an image of a virtual background that has been framed based on the changes in the speaker's image. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram of an image processing system according to the first embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of the configuration of the PC 1. [Figure 3] FIG. 3 is a flowchart showing an example of processing executed by the processor 15. [Figure 4] FIG. 4 is a diagram showing an example of an image M1 captured by the camera 6 and an example of an image M2 cut out from the image M1. [Figure 5] FIG. 5 is a diagram showing an example of an image M3 of speaker U1 for synthesis. [Figure 6] FIG. 6 is a diagram showing an example of a virtual background image M4 and an example of a virtual background image M5 cut out from the virtual background image M4. [Figure 7]FIG. 7 shows a composite image M6 obtained by combining an image M3 of the speaker U1 with an image M5 of a virtual background. [Figure 8] FIG. 8 is a diagram showing an example of a virtual background image M7 for synthesis when the zoom magnification is increased. [Figure 9] FIG. 9 shows a synthesized image M8 obtained by synthesizing an image M3 of the speaker U1 for synthesis with an image M7 of a virtual background. [Figure 10] FIG. 10 is a diagram showing an example of a portion cut out from the image M1 of the speaker U1 when the amount of horizontal movement is changed. [Figure 11] FIG. 11 is a diagram showing an example of a portion cut out from the virtual background image M4 when the horizontal movement amount is changed. DETAILED DESCRIPTION OF THE INVENTION

[0009] [First embodiment] 1 is a block diagram of an image processing system according to a first embodiment. The image processing system includes a PC (Personal Computer) 1 and a PC 2 connected via the Internet 3. In the following description of the embodiment, all sound signals will be described as digital signals unless otherwise specified.

[0010] PC1 and PC2 are each used to hold a remote conference with a far-end PC connected via the Internet 3. PC1 establishes communication with far-end PC2 via the Internet 3. PC1 communicates with PC2 wirelessly or via a wired connection. A user of PC1 (hereinafter referred to as speaker U1) holds a remote conference with a user of far-end PC2 by using PC1 that has established communication with PC2. PC1 is an example of an image processing device in this application. Note that the number of PCs connected to PC1 via the Internet 3 is not limited to the number shown in the first embodiment.

[0011] It should be noted that PC1 does not necessarily have to be connected to PC2 via the Internet 3. If PC1 and PC2 are located in the same building, PC1 and PC2 may be connected to each other via a LAN (Local Area Network).

[0012] The configuration of PC2 is the same as that of PC1, and therefore, a description of the configuration of PC2 will be omitted below.

[0013] Fig. 2 is a block diagram showing an example of the configuration of the PC 1. As shown in Fig. 2, the PC 1 includes a user interface 10, a communication interface 11, an external device connection interface 12, a flash memory 13, a RAM (Random Access Memory) 14, a processor 15, and a display 16. The processor 15 is, for example, a CPU (Central Processing Unit).

[0014] The user interface 10 includes a mouse, a keyboard, a touch panel, etc. The user interface 10 accepts operations by the user of the PC 1.

[0015] The communication interface 11 is an interface conforming to standards such as LAN or HDMI (registered trademark). The communication interface 11 transmits an audio signal related to the voice of the speaker of PC1 to the far-end PC2. The communication interface 11 receives an audio signal related to the voice of the user of the far-end PC2 from PC2.

[0016] As an example, the external device connection interface 12 is an interface based on the USB (Universal Serial Bus) standard. Of course, the external device connection interface 12 may be an interface other than an interface based on the USB standard. As shown in FIG. 2, a microphone 4, a speaker 5, and a camera 6 are connected to the external device connection interface 12.

[0017] The microphone 4 acquires a sound signal related to the voice of the speaker U1 of the PC 1. The microphone 4 transmits the acquired sound signal to the external device connection interface 12.

[0018] The microphone 4 does not necessarily have to be an external microphone, but may be a microphone built into the PC 1.

[0019] The speaker 5 receives an audio signal relating to the voice of the user of the far-end PC 2 received by the communication interface 11 via the external device connection interface 12. The speaker 5 outputs a sound based on the audio signal.

[0020] The speaker 5 does not necessarily have to be an external speaker, but may be a speaker built into the PC 1.

[0021] It should be noted that the external device connection interface 12 may not be connected to the microphone 4 and the speaker 5 separately, but may be connected to a device such as a headset that integrates a microphone and a speaker.

[0022] The camera 6 captures an image of the speaker U1 and an image of the background behind the speaker U1 on the PC 1. The camera 6 transmits the captured image of the speaker U1 and the captured image of the background as image data to the external device connection interface 12. The external device connection interface 12 outputs the received image data to the processor 15. The camera 6 is an example of an acquisition unit in this application.

[0023] The camera 6 does not necessarily have to be an external camera, but may be a camera built into the PC 1.

[0024] The flash memory 13 stores various programs. The various programs include programs that operate the PC 1. The flash memory 13 is an example of a storage unit in the present application. Note that the flash memory 13 does not necessarily have to store the various programs. The various programs may be stored in another device, such as a server. In this case, the PC 1 receives the various programs from the other device, such as the server.

[0025] The processor 15 executes various operations by reading programs stored in the flash memory 13 into the RAM 14. For example, the processor 15 executes processing related to a remote conference with the PC 2 by reading a program related to the remote conference with the PC 2 into the RAM 14. The processor 15 receives image data acquired by the camera 6 of the far-end PC 2 via the Internet 3. The processor 15 decodes the received image data to acquire a signal related to the image data. The processor 15 outputs the signal related to the image data to the display 16.

[0026] The display 16 is a liquid crystal display, an organic EL display, or the like. The display 16 outputs an image (an image of the user of PC2) based on a signal related to the image data input from the processor 15. The speaker U1 of PC1 can see the image of the user of PC2 displayed on the display 16.

[0027] The following describes in detail the process relating to the remote conference executed by the processor 15 with reference to the drawings.

[0028] For example, when an application program related to a remote conference is executed, the processor 15 starts the process (FIG. 3: START).

[0029] After starting, the processor 15 acquires an image of the speaker U1 and the background (an image including the speaker U1) captured by the camera 6 (FIG. 3: step S10), as shown in FIG. 4. FIG. 4 is a diagram showing an example of an image M1 acquired by the camera 6 and an example of an image M2 cut out from the image M1.

[0030] After step S10, the processor 15 performs a framing process by zooming to cut out and enlarge the image of the speaker U1 in the image M1 (FIG. 3: step S11).

[0031] For example, the processor 15 recognizes the face of the speaker U1 in the image M1 using an image analysis process such as a deep neural network (DNN). For example, the flash memory 13 stores a trained model that has been trained to determine the relationship between an input image and an object (such as a human face) captured in the input image. The processor 15 inputs the image M1 into the trained model. The trained model recognizes an image of a human face included in the object in the image M1. The processor 15 sets a bounding box at the position of the recognized human face.

[0032] Furthermore, the processor 15 determines a zoom magnification based on the size of the bounding box. The flash memory 13 stores information indicating a target value for the size of the bounding box. The processor 15 reads the information indicating the target value from the flash memory 13. The processor 15 determines the zoom magnification based on the information indicating the target value and information indicating the size of the bounding box of the speaker U1. The processor 15 determines the zoom magnification so as to bring the size of the bounding box of the speaker U1 closer to the target value. As shown in FIG. 4, the processor 15 acquires an image M2 by zooming the image M1 based on the determined zoom magnification. The information indicating the zoom magnification is an example of "first framing information" in this application.

[0033] For example, if speaker U1 moves away from camera 6, processor 15 zooms image M1 at a magnification higher than the current zoom magnification in order to keep the size of speaker U1 in the image constant. On the other hand, if speaker U1 moves closer to camera 6, processor 15 zooms image M1 at a magnification lower than the current zoom magnification in order to keep the size of speaker U1 in the image constant. This zooming is an example of the "first framing process for the speaker's image" in this application.

[0034] After step S11, processor 15 acquires an image of the speaker to be synthesized (FIG. 3: step S12). FIG. 5 is a diagram showing an example of image M3 of speaker U1 to be synthesized. Processor 15 acquires image M3 of speaker U1 to be synthesized by making all parts of image M2 transparent except for the image of speaker U1. For example, processor 15 identifies the face of speaker U1 recognized by image analysis processing such as DNN and the part (body) connected to the face of speaker U1, and makes all parts transparent except for the identified area.

[0035] After step S12, processor 15 acquires an image of the virtual background (FIG. 3: step S13). FIG. 6 shows an example of virtual background image M4 and an example of virtual background image M5 cut out from virtual background image M4.

[0036] The virtual background image M5 is a background image (for example, an image of a natural object such as a mountain as shown in FIG. 6) that is different from the background image acquired by the camera 6. For example, the flash memory 13 stores the virtual background image M4. The processor 15 acquires the image M4 by reading the virtual background image M4 from the flash memory 13.

[0037] Processor 15 acquires virtual background image M5 by zooming in on a predetermined portion of virtual background image M4. Processor 15 acquires virtual background image M5 by cutting out a portion of virtual background image M4 with a predetermined size (for example, a range half that of virtual background image M4) and enlarging the cut-out image.

[0038] After step S13, processor 15 combines image M3 of speaker U1 with image M5 of the virtual background (FIG. 3: step S14). Specifically, processor 15 superimposes image M3 of speaker U1 on top of image M5 of the virtual background. FIG. 7 shows combined image M6 obtained by combining image M3 of speaker U1 with image M5 of the virtual background.

[0039] After step S14, processor 15 determines whether the size of speaker U1 in the frame of image M1 of speaker U1 has changed (FIG. 3: step S15). Specifically, if the zoom magnification of the zoom on image M1 has been changed in step S11, processor 15 determines that the size of speaker U1 has changed.

[0040] When the processor 15 determines that the size of the speaker U1 in the image frame has changed (FIG. 3: step S15: Yes), the processor 15 redetermines the zoom magnification for the virtual background image M4 (FIG. 3: step S16). For example, the processor 15 determines the zoom magnification for the virtual background image M4 based on the zoom magnification for the image M1 of the speaker U1 in step S11.

[0041] For example, when speaker U1 moves away from camera 6, processor 15 increases the zoom magnification of the virtual background image M4 from the current zoom magnification, similar to the zoom for image M1. On the other hand, when speaker U1 moves closer to camera 6, processor 15 decreases the zoom magnification of the virtual background image M4 from the current zoom magnification, similar to the zoom for image M1. The zoom magnification of the virtual background image M4 is an example of "second framing information indicating the size of the virtual background within the frame of the image" in this application.

[0042] After step S16, processor 15 zooms virtual background image M4 at the re-determined zoom magnification (FIG. 3: step S17). FIG. 8 is a diagram showing an example of virtual background image M7 for synthesis when the zoom magnification is increased. Processor 15 acquires virtual background image M7 shown in FIG. 8 by zooming image M4 at a magnification higher than the zoom magnification in step S12. The process of zooming virtual background image M4 at the re-determined zoom magnification is an example of "second framing process for a virtual background image based on the determined second framing information" in this application.

[0043] After the processing of step S17, the processor 15 combines the image M3 of the speaker U1 for synthesis with the image M7 of the virtual background for synthesis (FIG. 3: step S18). FIG. 9 shows a combined image M8 obtained by combining the image M3 of the speaker U1 for synthesis with the image M7 of the virtual background. The processor 15 transmits the combined image M8 to the PC 2 on the far-end side as image data.

[0044] In step S15, if the processor 15 determines that the size of the speaker U1 in the image frame has not changed (FIG. 3: step S15 No), it does not perform the processes from step S16 to step S18. In this case, the processor 15 transmits the composite image M6 acquired in step S14 to the PC 2 as image data.

[0045] PC2 receives the image data and outputs it to display 16 of PC2. Display 16 of PC2 displays an image (composite image M6 or composite image M8) based on the image data. The user of PC2 views composite image M6 or composite image M8 displayed on display 16 of PC2.

[0046] As described above, processor 15 performs a zoom process on the image of the virtual background so that the virtual background moves in accordance with the movement of speaker U1. Processor 15 acquires a composite image M8 by combining speaker image M3, which changes as speaker U1 is tracked, with virtual background image M7, which has undergone a second framing process based on the changes in speaker image M3. Therefore, during a remote conference or the like, the user of PC2 can view the virtual background that is framed in accordance with the first framing process that tracks speaker U1. This reduces the likelihood that the user of PC2 will perceive the background as stationary even though the camera 6 is tracking speaker U1. Therefore, the user of PC2 viewing composite image M8 can hold a remote conference with speaker U1 on PC1 without feeling any discomfort. As a result, the user of PC2 feels as if they are participating in a real conference or the like, and can enjoy a customer experience of smooth communication with speaker U1.

[0047] The processor 15 ends the process when, for example, an application program related to the remote conference is terminated (FIG. 3: END).

[0048] It should be noted that the virtual background image in this application does not include an image (for example, an image painted in a single color) in which the user on the far end side (user of PC2) cannot recognize the movement of the background.

[0049] In step S16, the processor 15 may change the zoom magnification of the virtual background image M4 less than the zoom magnification of the speaker image M1. This allows the processor 15 to reproduce a sense of distance from the real background, whereby changes in the background farther from the speaker U1 are smaller than changes in the speaker. This makes it even less likely that the user of PC2 viewing the composite image M8 will feel uncomfortable. As a result, the user of PC2 can feel as if they are participating in a more realistic meeting or the like, and can enjoy a customer experience in which they can communicate more smoothly with the speaker U1.

[0050] Processor 15 may zoom the image of the virtual background at a different zoom factor depending on the type of virtual background. For example, flash memory 13 stores information indicating the zoom factor associated with each virtual background image as data. Processor 15 references the data from flash memory 13 to zoom the image of the virtual background at a different zoom factor depending on the type of virtual background.

[0051] For example, the zoom factor for a virtual background, such as a mountain, that is perceived as being far from the speaker is smaller than the zoom factor for a virtual background, such as a chair, that is perceived as being close to the speaker. In this case, when zoomed by processor 15, changes in the virtual background, such as a mountain, are smaller than changes in the virtual background, such as a chair. This allows for a more realistic sense of distance in the background, where changes in a background (such as a mountain) that is far from the speaker are smaller than changes in a background (such as a chair) that is close to the speaker. This makes it even less likely that the user of PC2 viewing composite image M8 will feel uncomfortable. As a result, the user of PC2 feels as if they are participating in a more realistic meeting, etc., and can enjoy a customer experience that allows for smoother communication with speaker U1.

[0052] It should be noted that the processor 15 does not necessarily need to acquire an image of the speaker using the camera 6. For example, the processor 15 may acquire an image of the speaker U1 (an image of a character expressed in a CG image) by sensing the movement of the speaker U1 using LiDAR (Light Detecting and Ranging) or the like and performing a process of representing the movement of the speaker U1 in a CG image of a predetermined character.

[0053] The zoom processing in this application may be digital zoom processing or analog zoom processing. Specifically, digital zoom processing is processing to enlarge or reduce an image by cutting out a part of the image and processing (enlarging or reducing) the cut-out part of the image. Specifically, analog zoom processing is processing to optically enlarge or reduce an image by changing the magnification of the lens of camera 6.

[0054] [Variation 1] The PC1 according to Modification 1 will be described below with reference to the drawings. The configuration of the PC1 according to Modification 1 is the same as the configuration of the PC1 according to the first embodiment shown in Fig. 3, and therefore Fig. 3 will be applied mutatis mutandis in the description.

[0055] The processor 15 according to the first modification tracks the speaker U1 by performing image analysis processing (recognizing the face of the speaker U1) such as DNN on the image M1. The processor 15 calculates the horizontal and vertical movement amounts as the range for cutting out the image so that the image M1 of the speaker U1 is positioned at the center with a predetermined size.

[0056] FIG. 10 is a diagram showing an example of a portion cut out from image M1 of speaker U1 when the amount of horizontal movement is changed. In FIG. 10, speaker U1 has moved to the left from the initial position (the position of speaker U1 in FIG. 4) as seen from the user of PC2. Processor 15 calculates the amount of horizontal movement after the change (hereinafter referred to as the first amount of movement) so that speaker U1 after the movement is positioned at the center of a predetermined size. Processor 15 obtains image M3 of the speaker to be synthesized by cutting out a part of image M1 of speaker U1 based on the first amount of movement.

[0057] 11 is a diagram showing an example of a portion cut out from virtual background image M4 when the horizontal movement amount is changed. Based on the first movement amount, processor 15 calculates the horizontal movement amount after the change (hereinafter referred to as the second movement amount) as the range from which to cut out the image so that virtual background image M4 is positioned at the center of a predetermined size. As shown in FIG. 11, processor 15 acquires virtual background image M7 for compositing by cutting out a portion of virtual background image M4 based on the second movement amount.

[0058] For example, processor 15 may obtain a second movement amount that is greater than the first movement amount. As an example, processor 15 obtains the second movement amount by multiplying the first movement amount by a predetermined value (e.g., 1.02 times). Processor 15 obtains a virtual background image M7 for compositing by cutting out a part of virtual background image M4 using the second movement amount that is greater than the first movement amount.

[0059] The processor 15 obtains a composite image M8 by combining an image M3 of the speaker U1 for synthesis obtained based on the first movement amount and an image M7 of the virtual background for synthesis obtained based on a second movement amount that is larger than the first movement amount. In this case, the change in the image of the virtual background included in the composite image M8 is greater than the change in the image of the speaker U1 included in the composite image M8.

[0060] Through the above processing, processor 15 can reproduce the sense of distance of the actual background, in which changes in the background farther away from speaker U1 are greater than changes in speaker U1. This makes it even less likely that the user of PC2 viewing composite image M8 will feel uncomfortable. As a result, the user of PC2 can feel as if they are participating in a more realistic meeting or the like, and can enjoy a customer experience in which they can communicate more smoothly with speaker U1.

[0061] The second movement amount does not necessarily have to be greater than the first movement amount. The second movement amount may be the same as the first movement amount. In this case, the change in the image of the virtual background included in the composite image M8 will be the same as the change in the image of the speaker U1 included in the composite image M8.

[0062] Processor 15 may calculate different second movement amounts depending on the type of virtual background. For example, flash memory 13 stores, as data, information indicating a value by which the first movement amount is multiplied (hereinafter referred to as a multiplication value) for each virtual background image. Processor 15 refers to the data and multiplies the first movement amount by a different multiplication value for each type of virtual background, thereby calculating different second movement amounts for each type of virtual background.

[0063] For example, a multiplication value associated with an image of a virtual background, such as a mountain, that is perceived as being far from the speaker is greater than a multiplication value associated with an image of a virtual background, such as a chair, that is perceived as being close to the speaker. In this case, changes in the virtual background, such as a mountain, are greater than changes in the virtual background, such as a chair. This allows for a more realistic sense of distance in the background, where changes in a background (such as a mountain) that is farther from the speaker are greater than changes in a background (such as a chair) that is closer to the speaker. This makes it even less likely that a user of PC2 viewing composite image M8 will feel uncomfortable. As a result, the user of PC2 feels as if they are participating in a more realistic meeting, etc., and can enjoy a customer experience that allows for smoother communication with speaker U1.

[0064] In the above description, the processor 15 only calculates the horizontal movement amount as the range from which the image is cut out. However, it goes without saying that the processor 15 may calculate the vertical movement amount as the range from which the image is cut out.

[0065] The processing of the processor 15 according to the first modification is applicable to the processing of the processor 15 according to the first embodiment.

[0066] [Variation 2] The PC1 according to Modification 2 will be described below with reference to the drawings. The configuration of the PC1 according to Modification 2 is the same as the configuration of the PC1 according to the first embodiment, and therefore will be described with reference to FIG.

[0067] In the second modification, the camera 6 further has a physical PTZ function. Specifically, the camera 6 has a function to rotate horizontally (Pan), a function to rotate vertically (Tilt), and a function to enlarge / reduce the camera's angle of view (Zoom). The camera 6 automatically tracks the speaker U1 and rotates horizontally or vertically. Alternatively, the camera 6 automatically tracks the speaker U1 and changes (enlarges / reduces) the lens magnification of the camera 6. The process of rotating the camera 6 or the process of changing the lens magnification is an example of the "first framing process" in this application.

[0068] The camera 6 transmits information indicating the horizontal rotation angle, vertical rotation angle, or lens magnification of the camera 6 to the processor 15. The information indicating the horizontal rotation angle, vertical rotation angle, or lens magnification of the camera 6 is an example of "camera control information" in this application.

[0069] The information indicating the horizontal rotation angle of camera 6 is an example of information used to calculate the "horizontal movement amount" in Modification 1. The information indicating the vertical rotation angle of camera 6 is an example of information used to calculate the "vertical movement amount" in Modification 1. The information indicating the lens magnification of camera 6 corresponds to the "information indicating the zoom magnification" in the first embodiment.

[0070] The processor 15 receives information from the camera 6 indicating the horizontal rotation angle, vertical rotation angle or lens magnification of the camera 6 .

[0071] Based on information received from camera 6 indicating the horizontal rotation angle or vertical rotation angle (hereinafter referred to as the first angle) of camera 6, processor 15 determines a horizontal angle or vertical angle (hereinafter referred to as the second angle) as a range for cutting out the image so that virtual background image M4 is centered and of a predetermined size. For example, processor 15 determines an angle greater than the first angle as the second angle. As an example, processor 15 determines an angle obtained by multiplying the first angle by a predetermined value (e.g., 1.02 times) as the second angle. Of course, the second angle does not necessarily have to be greater than the first angle and may be the same as the first angle.

[0072] Processor 15 redetermines the zoom magnification for zooming in on virtual background image M4 based on the information indicating the lens magnification received from camera 6. Specifically, if processor 15 determines that the lens magnification has increased, it increases the zoom magnification for zooming in on virtual background image M4, and if it determines that the lens magnification has decreased, it decreases the zoom magnification for zooming in on virtual background image M4. The process for determining the zoom magnification for zooming in on image M4 in Modification 2 is the same as the process for determining the zoom magnification for zooming in on image M4 in the first embodiment, and therefore description thereof will be omitted.

[0073] Through the processing by the processor 15, the user of PC2 can see a virtual background that moves in accordance with the movement of the speaker U1, like a background in real space, in the same way as in the first embodiment or modification 1. Therefore, the user of PC2 who sees the composite image M8 can hold a remote conference with the speaker U1 of PC1 without feeling any discomfort. As a result, the user of PC2 can feel as if he or she is actually holding a conference or the like, and can enjoy the customer experience of being able to communicate smoothly with the speaker U1.

[0074] [Variation 3] In Modification 3, processor 15 performs further image processing on image M7 of the virtual background to be synthesized. For example, processor 15 may perform image processing to adjust the brightness of image M7 of the virtual background to be synthesized based on the rotation direction of camera 6. Image M7 of the virtual background to be synthesized is an example of the "image of the virtual background to which second framing processing has been performed" in this application.

[0075] For example, when the camera 6 rotates vertically upward, the processor 15 increases the brightness of the virtual background image M7 for synthesis compared to its current brightness. On the other hand, when the camera 6 rotates vertically downward, the processor 15 decreases the brightness of the virtual background image M7 for synthesis compared to its current brightness. This allows the processor 15 to reproduce a bright image, such as one captured by pointing the camera 6 toward a light source such as a lamp or the sun (vertically upward), and a dark image, such as one captured by pointing the camera 6 in the opposite direction from the light source (vertically downward). This makes it even less likely that the user of PC2 viewing the composite image M8 will feel uncomfortable. As a result, the user of PC2 can feel as if they are participating in a more realistic meeting or the like, and can enjoy a customer experience in which they can communicate more smoothly with the speaker U1.

[0076] The processing of the processor 15 according to the third modification is applicable to the processing of the processor 15 according to either the first embodiment or the first or second modification.

[0077] [Variation 4] In the fourth modification, the processor 15 performs processing to increase or decrease the resolution of the image M7 of the virtual background based on the zoom magnification of the zoom on the virtual image.

[0078] For example, when image M1 is zoomed in at a magnification lower than the current zoom magnification, processor 15 reduces the resolution of image M7 of the virtual background to be lower than the current resolution. On the other hand, when image M1 is zoomed in at a magnification higher than the current zoom magnification, processor 15 increases the resolution of image M7 of the virtual background to be higher than the current resolution.

[0079] Through the above processing, the processor 15 can express the depth of field in the virtual background, and can more accurately reproduce the sense of distance in the real background. This makes it even less likely that the user of PC2 viewing the composite image M8 will feel uncomfortable. As a result, the user of PC2 can feel as if they are participating in a more realistic meeting, etc., and can enjoy a customer experience in which they can communicate more smoothly with the speaker U1. The processing of the processor 15 according to the fourth modification is applicable to the processing of the processor 15 according to the first embodiment or any of the first to third modifications.

[0080] [Variation 5] In the fifth modification, after the process of step S17 shown in FIG. 3, the processor 15 further detects the movement of the speaker U1 from the image M1 of the speaker U1, and further performs image processing on the image M7 of the virtual background based on the detected movement.

[0081] For example, processor 15 performs image analysis processing such as DNN on image M1 to recognize the image of speaker U1's mouth and set a bounding box. Processor 15 identifies the size of the mouth based on the size of the bounding box. Processor 15 performs image processing to adjust the brightness of virtual background image M7 in accordance with changes in the size of speaker U1's mouth.

[0082] For example, processor 15 may calculate the change in speaker U1's mouth size by calculating the difference in area or vertical length between the bounding box size at its minimum and maximum. Processor 15 determines whether the change in speaker U1's mouth size is equal to or greater than a predetermined change value. If processor 15 determines that the change in speaker U1's mouth size is equal to or greater than the predetermined change value, it increases the brightness of virtual background image M7 from its current brightness. This allows the user of PC2 viewing composite image M8 to enjoy a customer experience in which they can visually determine whether speaker U1's mouth is moving significantly, i.e., whether speaker U1 is speaking loudly. The movement of the speaker's mouth is an example of the "speaker's behavior" in this application.

[0083] If the processor 15 determines that the change in the size of the mouth of the speaker U1 is less than a predetermined change value, the processor 15 may lower the brightness of the virtual background image M7 below the current brightness.

[0084] Note that processor 15 does not necessarily have to perform image processing to adjust the brightness of virtual background image M7 in accordance with changes in the size of speaker U1's mouth. Processor 15 may determine whether speaker U1 is facing the direction in which camera 6 is located, and, if it is determined that speaker U1 is facing the direction in which camera 6 is located, may increase the brightness of virtual background image M7 from the current brightness.

[0085] Note that processor 15 does not necessarily have to perform image processing on virtual background image M7 based on the movement of speaker U1 detected from image M1 of speaker U1. For example, processor 15 may perform image processing on virtual background image M7 based on the level of the voice of speaker U1 acquired from microphone 4. For example, when the level of the voice is equal to or greater than a predetermined threshold, processor 15 increases the brightness of virtual background image M7 from the current brightness. This allows the user of PC 2 to enjoy the customer experience of being able to visually judge the voice level of speaker U1.

[0086] The processing of the processor 15 according to the fifth modification is applicable to the processing of the processor 15 according to the first embodiment or any of the first to fourth modifications.

[0087] [Variation 6] In Modification 6, the processor 15 of the PC 1 performs image processing on the virtual background image M7 based on the time that has elapsed since the start of the remote conference. The time that has elapsed since the start of the remote conference is an example of time information in the present application.

[0088] For example, processor 15 counts the time that has elapsed since the start of the remote conference. When the elapsed time exceeds a predetermined time, processor 15 performs image processing on virtual background image M7 to cause virtual background image M7 to blink. As a result, the virtual background image included in composite image M8 blinks. Therefore, a user of PC2 viewing composite image M8 can enjoy the customer experience of visually knowing the time that has elapsed since the start of the remote conference by seeing the blinking of the virtual background image included in composite image M8.

[0089] The processing of the processor 15 according to the sixth modification is applicable to the processing of the processor 15 according to the first embodiment or any of the first to fifth modifications.

[0090] The above-described embodiments and modifications should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments or modifications, but by the claims. Furthermore, the scope of the present invention is intended to include all modifications that are equivalent to and within the scope of the claims. [Explanation of symbols]

[0091] 1, 2: PC, 3: Internet, 4: Microphone, 5: Speaker, 6: Camera, 10: User interface, 11: Communication interface, 12: External device connection interface, 13: Flash memory, 14: RAM, 15: Processor, 16: Display

Claims

1. Acquire an image containing the speaker, performing a first framing process on the image including the speaker based on first framing information indicating a position and a size of the speaker within the image including the speaker; Acquires an image of a virtual background stored in a storage unit; synthesizing the image including the speaker that has undergone the first framing process with the image of the virtual background; 1. An image processing method, comprising: When the first framing information has changed, second framing information indicating a position and a size of the virtual background within the image of the virtual background is determined based on the changed first framing information; performing a second framing process on the image of the virtual background based on the determined second framing information; synthesizing the image including the speaker subjected to the first framing process with the image of the virtual background subjected to the second framing process; Image processing methods.

2. acquiring an image including the speaker using a camera; The image processing method according to claim 1 .

3. the first framing information includes control information for the camera; performing the first framing process by controlling the camera based on the control information; determining the second framing information based on the control information; The image processing method according to claim 2 .

4. a change in the image of the virtual background before and after the second framing process is different from a change in the image including the speaker before and after the first framing process; 4. The image processing method according to claim 1.

5. further image processing is performed on the image of the virtual background that has been subjected to the second framing processing; 4. The image processing method according to claim 1.

6. Detecting a movement of the speaker from an image including the speaker; performing the image processing on the virtual background based on the detected speaker's movement; The image processing method according to claim 5 .

7. obtaining a level of the speaker's voice; performing the image processing on the virtual background based on the level of the speaker's voice; The image processing method according to claim 5 .

8. Get the time information, performing the image processing on the virtual background in accordance with the time information; The image processing method according to claim 5 .

9. a storage unit that stores an image of a virtual background; an acquisition unit that acquires an image including a speaker; a processor that acquires an image of the virtual background stored in the storage unit, performs a first framing process on the image including the speaker based on first framing information that indicates a position and a size of the speaker within the image including the speaker, and synthesizes the image including the speaker that has undergone the first framing process with the image of the virtual background; An image processing device comprising: The processor: When the first framing information has changed, second framing information indicating a position and a size of the virtual background within the image of the virtual background is determined based on the changed first framing information; performing a second framing process on the image of the virtual background based on the determined second framing information; synthesizing the image including the speaker subjected to the first framing process with the image of the virtual background subjected to the second framing process; Image processing device.

10. the acquisition unit is a camera, the camera acquires an image including the speaker; The image processing device according to claim 9 .

11. the first framing information includes control information for the camera; The processor: performing the first framing process by controlling the camera based on the control information; determining the second framing information based on the control information; The image processing device according to claim 10.

12. a change in the image of the virtual background before and after the second framing process is different from a change in the image including the speaker before and after the first framing process; 12. The image processing device according to claim 9.

13. the processor further performs image processing on the image of the virtual background that has been subjected to the second framing processing.

12. The image processing device according to claim 9.

14. The processor: Detecting a movement of the speaker from an image including the speaker; performing the image processing on the virtual background based on the detected speaker's movement; The image processing device according to claim 13 .

15. The processor: obtaining a level of the speaker's voice; performing the image processing on the virtual background based on the level of the speaker's voice; The image processing device according to claim 13 .

16. The processor: Get the time information, performing the image processing on the virtual background in accordance with the time information; The image processing device according to claim 13 .

17. a process of performing a first framing process on an image including a speaker acquired by an acquisition unit based on first framing information indicating a position and a size of the speaker in the image including the speaker; a process of synthesizing an image including the speaker that has been subjected to the first framing process with an image of a virtual background stored in a storage unit; A program for causing a processor to execute the following: When the first framing information has changed, a process of determining second framing information indicating a position and a size of the virtual background within the image of the virtual background based on the changed first framing information; performing a second framing process on the image of the virtual background based on the determined second framing information; a process of synthesizing an image including the speaker that has been subjected to the first framing process and an image of the virtual background that has been subjected to the second framing process; A program that causes the processor to execute the above.

Citation Information

Patent Citations

  • Karaoke device with addition of function for photographing and recording singing figure by video camera

    JP1999282479A