Image processing methods, apparatus, computer equipment, storage media and program products

By using transparent screens and embedded cameras in video conferencing systems to remove identical colors and adjust image colors, the problem of not being able to look directly at each other caused by the camera's top-down view was solved, thus improving the communication effect and quality of video conferencing.

CN119520723BActive Publication Date: 2025-10-31CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411679773.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-10-31
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In existing video conferencing systems, cameras are usually located above the display screen, resulting in a downward viewing angle. Users cannot see each other's eyes clearly, thus failing to achieve a "eye contact" effect, which affects communication effectiveness and meeting quality.

Method used

By using a transparent screen and an embedded camera, the RGB values ​​of pixels on the transparent screen are obtained, and image processing algorithms are used to remove the same color. Combined with a reference image from a second camera, the image color is adjusted to achieve a visual interaction effect.

Benefits of technology

It enables both parties in a video conference to freely make eye contact, improving communication effectiveness and meeting quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119520723B_ABST
    Figure CN119520723B_ABST
Patent Text Reader

Abstract

This application relates to an image processing method, apparatus, computer device, storage medium, and program product. The method includes: during the display of an image on a transparent screen, if it is determined that each pixel on the transparent screen within the field of view range has the same color based on the RGB values ​​corresponding to a first camera, then using the image captured by the first camera as an initial image; removing the same color from the initial image to obtain a target image; and outputting the target image to another display device. This method enables one user to clearly see the other user's eyes during video conferencing, creating a sense of "eye contact," improving communication effectiveness and enhancing the quality of remote meetings. This method can also update any image captured by a first camera based on an image captured by a second camera to avoid situations where the image captured by the first camera is unavailable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to an image processing method, apparatus, computer equipment, storage medium, and program product. Background Technology

[0002] A video conferencing system is a conferencing system that enables people in two or more locations to have face-to-face conversations via communication devices and networks, and is typically used for remote meetings.

[0003] Most video conferencing systems currently in use come with built-in cameras. One approach is to integrate the camera into an all-in-one conferencing unit within the top bezel of the display screen. While this is convenient for deployment, the integrated camera's configuration is relatively poor. Another approach is to separate the camera from the display screen, deploying the camera independently, typically positioned directly above the screen. Both approaches share a common problem: the camera is positioned above the screen, resulting in a top-down view when capturing images. This top-down view makes it difficult for the other party in the video conference to see the other person's eyes, preventing eye contact and reducing communication effectiveness, ultimately lowering the quality of the remote meeting.

[0004] Therefore, how to ensure that one user can clearly see the other user's eyes during a video conference, creating a sense of "eye contact," can improve the communication effectiveness between the two parties and enhance the quality of remote meetings. Summary of the Invention

[0005] Therefore, it is necessary to address the aforementioned technical problems by providing an image processing method, apparatus, computer equipment, storage medium, and program product that enables one user to clearly see the other user's eyes during a video conference, creating a sense of "eye contact," thereby improving the communication effect between the two parties and enhancing the quality of remote meetings.

[0006] In a first aspect, an image processing method is provided, applied to a display device, the display device including a body, a transparent screen, a plurality of first cameras, and a processor; wherein the transparent screen and the plurality of first cameras are all disposed on the body, the transparent screen covering the lenses of the first cameras; the transparent screen and the plurality of first cameras are all signal-connected to the processor;

[0007] The method includes:

[0008] During the process of displaying images on the transparent screen, the RGB values ​​of pixels on the transparent screen within the field of view of each of the first cameras are obtained;

[0009] If it is determined that each pixel on the transparent screen within the field of view has the same color based on the RGB value corresponding to the first camera, the image captured by the first camera is used as the initial image.

[0010] The same color is removed from the initial image to obtain the target image;

[0011] The target image is output to another display device.

[0012] In one embodiment, the display device further includes a second camera disposed on the body, and the lens of the second camera is not covered by the transparent screen;

[0013] The method further includes:

[0014] When it is determined, based on the RGB values ​​of each of the first cameras, that there are multiple pixels with different colors on the transparent screen within the corresponding field of view, the image captured by the second camera is used as a reference image.

[0015] Using an image captured by any of the first cameras as the image to be processed, for the same user captured between the image to be processed and the reference image, determine the first image region where the same user is located in the image to be processed and the second image region where the same user is located in the reference image.

[0016] The color of the pixels in the first image region of the image to be processed is replaced with the color of the pixels in the second image region of the reference image to obtain the target image.

[0017] In one embodiment, the method further includes:

[0018] Extract a first image feature from the image to be processed and a second image feature from the reference image; wherein the first image feature includes at least the grayscale of the image to be processed, and the second image feature includes at least the grayscale of the reference image;

[0019] The first image feature and the second image feature are input into an image recognition classifier, and the same user collected between the image to be processed and the reference image is determined based on the output information of the image recognition classifier.

[0020] In one embodiment, before playing the image on the transparent screen, the method further includes:

[0021] For each of the first cameras, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera;

[0022] The step of obtaining the RGB values ​​of pixels on the transparent screen within the field of view of each of the first cameras includes:

[0023] For each of the first cameras, the RGB values ​​of the pixels on the transparent screen within the field of view of the first camera are determined based on the RGB values ​​of each pixel in the image displayed on the transparent screen and the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera.

[0024] In one embodiment, removing the same color from the initial image includes:

[0025] For the first camera corresponding to the initial image, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera;

[0026] Based on the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera, determine the sub-region in the initial image to be processed to remove the same color;

[0027] Remove the same color from the sub-region to be processed.

[0028] In one embodiment, before obtaining the pixel coordinates of the corresponding pixels on the transparent screen within the field of view for each of the first cameras, the method further includes:

[0029] The coordinates of the first camera in the two-dimensional coordinate system are calibrated.

[0030] In a second aspect, a display device is provided, comprising: a body, a transparent screen, a plurality of first cameras, and a processor;

[0031] The transparent screen and the plurality of first cameras are all disposed on the main body, and the transparent screen covers the lens of the first camera;

[0032] The transparent screen and the multiple first cameras are all signal-connected to the processor;

[0033] The transparent screen is configured to display images;

[0034] The processor is configured as follows:

[0035] During the process of displaying images on the transparent screen, the RGB values ​​of pixels on the transparent screen within the field of view of each of the first cameras are obtained;

[0036] If it is determined that each pixel on the transparent screen within the field of view has the same color based on the RGB value corresponding to the first camera, the image captured by the first camera is used as the initial image.

[0037] The same color is removed from the initial image to obtain the target image;

[0038] The target image is output to another display device.

[0039] In one embodiment, the display device further includes a second camera disposed on the body, and the lens of the second camera is not covered by the transparent screen;

[0040] The processor is also configured to:

[0041] When it is determined, based on the RGB values ​​of each of the first cameras, that there are multiple pixels with different colors on the transparent screen within the corresponding field of view, the image captured by the second camera is used as a reference image.

[0042] The image captured by any of the first cameras is used as the image to be processed;

[0043] For the same user captured between the image to be processed and the reference image, determine the first image region where the same user is located in the image to be processed and the second image region where the same user is located in the reference image;

[0044] The color of the pixels in the first image region of the image to be processed is replaced with the color of the pixels in the second image region of the reference image to obtain the target image.

[0045] Thirdly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0046] Fourthly, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0047] The aforementioned image processing methods, devices, computer equipment, storage media, and program products utilize the "transparent" feature of a transparent screen, combined with a first camera embedded under the transparent screen. Through flexible camera scheduling and algorithm optimization, the system solves the common problem in current video conferencing systems where the two parties cannot make eye contact, allowing them to freely exchange glances, thus improving communication effectiveness and video conferencing quality. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is an application environment diagram of an image processing method in one embodiment;

[0050] Figure 2 This is a flowchart illustrating an image processing method in one embodiment;

[0051] Figure 3 This is a schematic diagram of a display device in one embodiment;

[0052] Figure 4 This is a partial flowchart of an image processing method in one embodiment;

[0053] Figure 5 This is a schematic diagram illustrating communication between a display device and another display device in one embodiment;

[0054] Figure 6 This is a partial flowchart of an image processing method in one embodiment;

[0055] Figure 7 This is a flowchart illustrating the image processing method in yet another embodiment;

[0056] Figure 8 This is a schematic diagram of the structure of a display device in one embodiment;

[0057] Figure 9 This is a structural block diagram of an image processing device in one embodiment;

[0058] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0060] The image processing method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, tablets, smart TVs, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0061] It should be noted that the image processing method provided in any of the following embodiments is not only applicable to video conferencing scenarios, but also to other two-party video scenarios such as video conferencing via an application.

[0062] In one exemplary embodiment, such as Figure 2 As shown, an image processing method is provided, which is applied to... Figure 1 The terminal 102 in the example is a display device.

[0063] Please see Figure 3 As shown, the display device includes a main body, a transparent screen, multiple first cameras, and a processor. The transparent screen and the multiple first cameras are all mounted on the main body, with the transparent screen covering the lenses of the first cameras. The transparent screen and the multiple first cameras are all signal-connected to the processor. Figure 3 The multiple first cameras shown are camera 1, camera 2, camera 3, and camera 4. The number and type of these first cameras can be set according to actual needs, and this embodiment does not limit them.

[0064] The image processing method includes steps 201 to 204. Wherein:

[0065] Step 201: During the process of displaying images on the transparent screen, obtain the RGB values ​​of each pixel on the transparent screen within the field of view of the first camera.

[0066] Transparent screens, also known as transparent display screens, are a new type of display method that combines the characteristics of a screen with transparency. They can be used as a screen and also replace transparent flat glass. Static images and dynamic images (such as video images) can be displayed on transparent screens. Users can see the content displayed on the transparent screen while also seeing objects behind it. Similarly, a camera behind the transparent screen can capture images of the user and other images (such as the backdrop of a conference room).

[0067] The main types of transparent screens include transparent organic light-emitting diode (OLED) displays, transparent light-emitting diode (LED) displays, and transparent liquid crystal displays (LCDs). The type of transparent screen used in this embodiment can be selected according to actual needs, and this embodiment does not impose any limitations.

[0068] The first camera, also known as the first sensor, is used to capture images. The lens of this first camera is covered by a transparent screen; therefore, the image captured by the first camera consists of two parts: one part is the image displayed on the transparent screen, and the other part is the image captured through the transparent screen. For example, in a conference system, the image captured by the first camera through the transparent screen can include the user, the conference room setting, etc. Preferably, the first camera is a color camera (RGB camera, R for Red, G for Green, B for Blue), which can simultaneously capture the three main color channels of the image—red, green, and blue—to generate a complete color image.

[0069] The field of view (FOV) refers to the angular range within which a camera can receive images in an imaging scene; it is also commonly referred to as the field of view. The FOVs of different first cameras can be the same or different; this embodiment does not impose any limitations.

[0070] Since the coordinates and RGB values ​​of each pixel on the transparent screen are known during image display, the coordinates of the pixels on the transparent screen within the field of view of each first camera can be determined. Therefore, for each first camera, based on the coordinates of the pixels on the transparent screen within the field of view of that first camera, the coordinates of each pixel on the transparent screen, and the RGB values ​​of each pixel, the RGB values ​​of the pixels on the transparent screen within the field of view of that first camera can be determined.

[0071] RGB values ​​can also be understood as color codes. A color code is a numerical representation of a color, typically used to describe its specific components and intensity. In computer science and design, color codes are usually represented using the RGB (Red, Green, Blue) color model. This model defines a color using three values ​​(representing the intensity of the red, green, and blue color channels, respectively). For example, the RGB value for red is (255, 0, 0), meaning the red channel has the maximum intensity (255), while the green and blue channels have an intensity of 0.

[0072] Step 202: If it is determined that each pixel on the transparent screen within the field of view has the same color based on the RGB value corresponding to the first camera, the image captured by the first camera is used as the initial image.

[0073] The RGB value corresponding to the first camera refers to the RGB value of the pixel on the transparent screen within the field of view of the first camera obtained in step 202.

[0074] An RGB value represents a color; therefore, the color of each pixel on the transparent screen within the field of view can be determined using the RGB value corresponding to the first camera. For example, red can be the same color.

[0075] When pixels on a transparent screen within the field of view in front of the first camera lens have color, the colors captured by the first camera through these pixels are no longer accurate. This is especially true when pixels embedded behind the transparent screen within the camera's field of view contain two or more colors; even with algorithms, it's difficult to restore the original colors. Image processing algorithms, however, can handle images with only one pure color superimposed on the original color image, removing the superimposed pure color to restore the original color image.

[0076] Therefore, preferably, when each pixel on the transparent screen within the defined field of view has the same color, the image captured by the first camera is used as the initial image in order to more easily and quickly restore the original color image (i.e., the target image).

[0077] It should be noted that the display device is equipped with multiple first cameras. If multiple pixels on the transparent screen corresponding to one first camera have different colors, it can be determined whether each pixel on the transparent screen corresponding to other first cameras has the same color. This process continues until it is determined that each pixel on the transparent screen corresponding to a first camera has the same color, and then the image captured by that first camera is used as the initial image. For example, if multiple pixels on the transparent screen corresponding to camera 1 and camera 2 have different colors, but each pixel on the transparent screen corresponding to camera 3 has the same color, then the image captured by camera 3 is used as the initial image.

[0078] Step 203: Remove the same color from the initial image to obtain the target image.

[0079] After removing the same color, the resulting target image can be understood as the real image captured by the first camera without being obstructed by the transparent screen.

[0080] Alternatively, an image processing algorithm can be used to remove the same color from the initial image.

[0081] Image processing algorithms can process images where only one pure color is superimposed on the original color image. By removing this superimposed pure color, the original color image (i.e., the target image) can be restored. Therefore, by using an image processing algorithm to remove the same color from the initial image, the target image is obtained.

[0082] Image processing algorithms, such as machine learning-based segmentation and clustering algorithms, can segment and classify colors in an image and then remove regions of a specified color. See also... Figure 4 The processing diagram shown illustrates that after determining the initial image, the pixel coordinates of each pixel in the initial image are obtained, and the RGB values ​​(color codes) of the same color in the initial image are obtained. The image processing algorithm then removes the same color based on the pixel coordinates and RGB values ​​to restore the original color image or video, thus obtaining the target image.

[0083] Optionally, noise reduction, smoothing, etc., can be performed on the initial image before processing it.

[0084] Step 204: Output the target image to another display device.

[0085] The display device is one communication terminal, and the other display device is another communication terminal. These two display devices are two communication terminals in a video conferencing system. Please refer to [link / reference]. Figure 5 The display device and the other display device can communicate with each other via the Internet.

[0086] After obtaining the target image, in order for the user on another display device to see the user on this display device, the target image is output from this display device to the other display device, where it is displayed. Thus, when displaying images or videos through a transparent screen, both parties watching the video can see the user through the transparent screen, achieving the effect of simultaneous viewing and visual interaction.

[0087] The image processing method described above utilizes the "transparent" feature of the transparent screen, combined with the first camera embedded under the transparent screen. Through flexible camera scheduling and algorithm optimization, it solves the problem that the two parties in a video conferencing conference cannot "make eye contact" in the current market video conferencing systems. This allows the two parties in the video conferencing conference to freely make eye contact, improving the communication effect between the two parties and enhancing the quality of the video conferencing.

[0088] In one exemplary embodiment, such as Figure 3 As shown, the display device also includes a second camera. This second camera is mounted on the main body, and its lens is not covered by the transparent screen. For example... Figure 3 The second camera shown is camera 5.

[0089] Please see Figure 6 The image processing method further includes steps 601 to 603. Step 601 can be performed after step 201.

[0090] Step 601: If it is determined that there are multiple pixels with different colors on the transparent screen within the corresponding field of view based on the RGB value of each first camera, the image captured by the second camera is used as a reference image.

[0091] The presence of multiple pixels with different colors on the transparent screen corresponding to the first camera, such as red, green, and blue, proves that the colors of all pixels on the transparent screen within the field of view of the first camera are not the same.

[0092] Since the lens of the second camera is not obstructed by the transparent screen, the pixel colors captured by the second camera are the colors of the unobstructed pixels. Therefore, if it is determined that the colors of all pixels on the transparent screen within the field of view of each first camera are not the same, the image captured by the second camera is selected as the reference image.

[0093] Step 602: Using any image captured by the first camera as the image to be processed, for the same user captured between the image to be processed and the reference image, determine the first image region where the same user is located in the image to be processed and the second image region where the same user is located in the reference image.

[0094] Each image captured by the first camera can be used as the image to be processed. For example, the image captured by camera 1 can be selected as the image to be processed.

[0095] Content matching is performed between the image to be processed and the reference image. Identified elements are marked based on the comparison. Identified elements include, for example, the same user, the same object, or the same building.

[0096] To achieve a two-way eye-to-eye effect in the video, only the same user captured between the image to be processed and the reference image can be marked. A first image region containing the same user in the image to be processed and a second image region containing the same user in the reference image are determined. There may be more than one such user; for example, if there are two users, the first image region includes the first image region corresponding to one user and the first image region corresponding to the other user, and the second image region includes the second image region corresponding to both users.

[0097] Step 603: Replace the color of the pixels in the first image region of the image to be processed with the color of the pixels in the second image region of the reference image to obtain the target image.

[0098] In other words, by replacing the color of pixels in the first image region with the color of pixels in the second image region, the true image of the same user can be restored. For example, the color of pixels in the first image region of camera 1 can be replaced with the color of pixels in the second image region of camera 5.

[0099] After the color of the image to be processed is replaced, the target image can be obtained. Since the image to be processed is captured through a transparent screen, the target image appears to the user on the other side of the video as if they are looking directly at each other.

[0100] In this embodiment, when it is determined that multiple pixels on the transparent screen within the corresponding field of view have different colors based on the RGB values ​​corresponding to each first camera, the image captured by any of the first cameras is color-replaced based on the image captured by the second camera. This solves the problem that existing image processing algorithms cannot remove multiple colors. Furthermore, it avoids the problem of unusable images captured by the first cameras. Through flexible camera scheduling and algorithm optimization, it solves the common problem in current video conferencing systems where the two parties cannot make eye contact, allowing them to freely exchange glances, improving communication effectiveness and enhancing the quality of the video conference.

[0101] In one exemplary embodiment, prior to step 602, when determining the same user captured between the image to be processed and the reference image, the process includes:

[0102] Step 1: Extract the first image features of the image to be processed and the second image features of the reference image.

[0103] The first image feature includes at least the grayscale of the image to be processed, and may also include the texture information of the image to be processed, the RGB values ​​of the pixels, etc.

[0104] The second image feature includes at least the grayscale of the reference image, and may also include the texture information of the reference image, the RGB values ​​of the pixels, etc.

[0105] Step two: Input the first image feature and the second image feature into the image recognition classifier, and determine the same user collected between the image to be processed and the reference image based on the output information of the image recognition classifier.

[0106] This image recognition classifier is a machine learning or deep learning model used to identify objects in an image and classify them into predefined categories. The classifier uses first image features to identify elements in the image to be processed and second image features to identify elements in a reference image. It then matches the elements in the image to be processed and the elements in the reference image to identify the same users.

[0107] In one exemplary embodiment, before playing the image on the transparent screen in step 201, the image processing method further includes:

[0108] Step 1: For each of the first cameras, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera.

[0109] For each of the first cameras, the pixel coordinates of all pixels on the transparent screen within the field of view of the first camera are marked. For example, for camera 1 as described above, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 1 are marked. For camera 2, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 2 are marked. For camera 3, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 4 are marked.

[0110] Step 201, obtaining the RGB values ​​of each pixel on the transparent screen within the field of view of the first camera, includes:

[0111] Step 2: For each of the first cameras, determine the RGB values ​​of the pixels on the transparent screen within the field of view of the first camera, based on the RGB values ​​of each pixel in the image displayed on the transparent screen and the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera.

[0112] Understandably, before displaying an image on a transparent screen, it's necessary to know the RGB values ​​and pixel coordinates of each pixel in the image to be displayed, so that the transparent screen can be controlled to display the image based on the known RGB values. The displayed image is a complete image.

[0113] Based on the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera, a sub-image region corresponding to the field of view of the first camera can be determined from the displayed image. The RGB value of the sub-image region is the RGB value of the pixels on the transparent screen within the field of view of the first camera.

[0114] In one exemplary embodiment, step 203 includes:

[0115] Step 1: For the first camera corresponding to the initial image, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera.

[0116] For each of the first cameras, the pixel coordinates of all pixels on the transparent screen within the field of view of the first camera are marked. For example, for camera 1 as described above, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 1 are marked. For camera 2, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 2 are marked. For camera 3, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 3 are marked. For camera 4, the pixel coordinates of all pixels on the transparent screen within the field of view of camera 4 are marked.

[0117] Optionally, before step one, the coordinates of the first camera in the two-dimensional coordinate system are calibrated to improve the accuracy of the pixel coordinates.

[0118] Step 2: Based on the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera, determine the sub-region in the initial image where the same color needs to be removed.

[0119] The coordinates of each pixel on the transparent screen are pre-calibrated. Based on the coordinates of each pixel on the transparent screen and the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera, the coordinates of each pixel in the sub-region to be processed can be determined.

[0120] Step 3: Remove the same color from the sub-area to be processed.

[0121] Image processing algorithms can be used to remove the same color from the initial image.

[0122] Image processing algorithms can process images where only one pure color is superimposed on the original color image. By removing this superimposed pure color, the original color image (i.e., the target image) can be restored. Therefore, by using an image processing algorithm to remove the same color from the initial image, the target image is obtained.

[0123] Image processing algorithms, such as machine learning-based segmentation and clustering algorithms, can segment and classify colors in an image and then remove regions of a specified color.

[0124] The method provided in this embodiment utilizes the "transparent" feature of a transparent screen, combined with a first camera embedded under the transparent screen. Through flexible camera scheduling and algorithm optimization, it solves the problem that the two parties in a video conferencing conference cannot make eye contact, which is common in current video conferencing systems on the market. This allows the two parties participating in the video conference to freely exchange glances, improving the communication effect and enhancing the quality of the video conference.

[0125] For ease of understanding, the image processing method will be explained in more detail by referring to the multiple first cameras included in the display device as camera 1, camera 2, camera 3, and camera 4, and the second camera as camera 5. Please also refer to [link to relevant documentation]. Figure 7 Another flowchart illustrating this image processing method is shown. (See diagram below.) Figure 7 As shown, the image processing method includes:

[0126] The first step is coordinate system initialization. This includes: marking the coordinates of cameras 1, 2, 3, 4, and 5; marking the pixel coordinates of all pixels on the transparent screen within the field of view of camera 1; marking the pixel coordinates of all pixels on the transparent screen within the field of view of camera 2; marking the pixel coordinates of all pixels on the transparent screen within the field of view of camera 3; and marking the pixel coordinates of all pixels on the transparent screen within the field of view of camera 4.

[0127] The second step is for the transparent screen to begin playing images.

[0128] The third step involves acquiring the transparent screen pixels within the field of view of each of the four cameras (camera 1, camera 2, camera 3, and camera 4) during the image playback process. This is to determine which camera's field of view contains all pixels of the same color and color code on the transparent screen. The color code of each pixel's displayed color is also acquired to prepare the algorithm for image processing.

[0129] Fourth, the system first determines whether the colors of all pixels on the transparent screen within the field of view of camera 1 are pure colors. If they are all the same color, the image captured by camera 1 is used as the initial image. Then, the system extracts the color codes (RGB values) of all pixels on the transparent screen within the field of view of camera 1, and uses an algorithm to remove any superimposed single color (i.e., the same color) from the initial image to obtain the target image. After obtaining the target image, the system transmits it as the final output image to the other party in the video conference.

[0130] Fifth, if it is determined that the colors of all pixels on the transparent screen within the field of view of camera 1 are not the same color, the system will continue to determine whether the colors of all pixels on the transparent screen within the field of view of camera 2 are the same color. If they are the same color, the image captured by camera 2 is used as the initial image. Then, the color code (RGB value) of all pixels on the transparent screen within the field of view of camera 2 is extracted, and the single color superimposed in the initial image is processed by an algorithm to obtain the target image. After obtaining the target image, the system will transmit the target image as the final output image to the other party in the video conference.

[0131] Step 6: If it is determined that the colors of all pixels on the transparent screen within the field of view of camera 2 are not the same, the system will continue to determine whether the colors of all pixels on the transparent screen within the field of view of camera 3 are the same. If they are the same color, the image captured by camera 3 is used as the initial image. Then, the color code (RGB value) of all pixels on the transparent screen within the field of view of camera 2 is extracted, and the single color superimposed in the initial image is processed by an algorithm to obtain the target image. After obtaining the target image, the system will transmit the target image as the final output image to the other party in the video conference.

[0132] Step 7: If it is determined that the colors of all pixels on the transparent screen within the field of view of camera 3 are not the same color, the system will continue to determine whether the colors of all pixels on the transparent screen within the field of view of camera 4 are the same color. If they are the same color, the image captured by camera 4 is used as the initial image. Then, the color code (RGB value) of all pixels on the transparent screen within the field of view of camera 2 is extracted, and the single color superimposed in the initial image is processed by an algorithm to obtain the target image. After obtaining the target image, the system will transmit the target image as the final output image to the other party in the video conference.

[0133] Step 8: If the system determines that the colors of all pixels on the transparent screen within the field of view of camera 4 are not the same, the system will select the image captured by camera 1 as the image to be processed. At the same time, the system will perform content matching between the image captured by camera 1 and the image captured by camera 5 (i.e., the second camera). The two images will be compared and matched to find the same elements and mark them. Then, the colors of the elements in the image captured by camera 1 (the image to be processed) will be replaced with the colors of the same elements in the image captured by camera 5 (the reference image). The system will then use the image after replacing the colors of all elements in the image captured by camera 1 with the colors of the same elements captured by camera 5 as the target image (i.e., the final output image of the system) and transmit it to the other party in the video conference.

[0134] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0135] Please see Figure 8 One embodiment of this application also provides a display device 800. The display device 800 includes a body 801, a transparent screen 802, a plurality of first cameras 803, and a processor 804.

[0136] The transparent screen 802 and the multiple first cameras 803 are all disposed on the main body 801. The transparent screen 802 covers the lens of the first camera 803. The placement of the first camera 803 under the transparent screen 802 can be selected according to actual needs, and is not limited in this embodiment. The number of the first cameras 803 can be selected according to actual needs, and is not limited in this embodiment.

[0137] The transparent screen 802 and the multiple first cameras 803 are all connected to the processor 804 via signal transmission. The processor 804 can be integrated onto the transparent screen 802 or located in other positions; this embodiment does not impose any limitations on this.

[0138] The transparent screen 802 is configured to display images.

[0139] The transparent screen 802, or transparent display screen, is a new type of display method that combines the characteristics of a screen and transparency. It can be used as a screen and also replace transparent flat glass. The transparent screen 802 can display static images and dynamic images (such as video images). While users see the content displayed on the transparent screen 802, they can also see objects behind it. Similarly, the first camera 803 behind the transparent screen 802 can also capture images of the user and other images (such as the conference room setting).

[0140] The main types of transparent screens 802 include transparent organic light-emitting diode (OLED) displays, transparent light-emitting diode (LED) displays, and transparent liquid crystal displays (LCDs). The type of transparent screen 802 used in this embodiment can be selected according to actual needs, and this embodiment does not limit it.

[0141] The first camera 803, also known as the first camera, is used to capture images. The lens of the first camera 803 is covered by a transparent screen 802; therefore, the image captured by the first camera 803 includes two parts: one part is the image displayed on the transparent screen 802, and the other part is the image captured through the transparent screen 802. For example, in a conference system, the image captured by the first camera 803 through the transparent screen 802 may include users, conference room scenery, etc. Preferably, the first camera 803 is a color camera (RGB camera, R for Red, G for Green, B for Blue), which can simultaneously capture the three main color channels of the image: red, green, and blue, thereby generating a complete color image.

[0142] The processor 804 is configured to perform the image processing method provided in any of the preceding embodiments.

[0143] Specifically, the processor 804 is configured to execute steps 201 to 204. A detailed description of steps 201 to 204 can be found in the relevant descriptions above, and will not be repeated here.

[0144] The aforementioned display device 800 utilizes the "transparent" feature of the transparent screen 802, combined with the first camera 803 embedded under the transparent screen 802, and through flexible camera scheduling and algorithm optimization, solves the problem that the two parties in a video conferencing conference cannot "make eye contact" in a common video conferencing system on the market. This allows the two parties participating in the video conferencing conference to freely make eye contact, improving the communication effect between the two parties and enhancing the quality of the video conferencing.

[0145] In one exemplary embodiment, please refer to Figure 8 The display device also includes a second camera 805, which is disposed on the main body 801, and the lens of the second camera 805 is not covered by the transparent screen 802. The placement of the second camera 805 on the main body 801 can be selected according to actual needs. Preferably, the second camera 805 is disposed on the upper side of the transparent screen 802.

[0146] The processor 804 is also configured to execute steps 601 to 603. For a detailed description of steps 601 to 603, please refer to the relevant descriptions above; they will not be repeated here.

[0147] To achieve a two-way eye-to-eye effect in the video, only the same user captured between the image to be processed and the reference image can be marked. A first image region containing the same user in the image to be processed and a second image region containing the same user in the reference image are determined. There may be more than one such user; for example, if there are two users, the first image region includes the first image region corresponding to one user and the first image region corresponding to the other user, and the second image region includes the second image region corresponding to both users.

[0148] Based on the same inventive concept, this application also provides an image processing apparatus for implementing the image processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more image processing apparatus embodiments provided below can be found in the limitations of the image processing method described above, and will not be repeated here.

[0149] In one exemplary embodiment, such as Figure 9 As shown, an image processing apparatus 900 is provided, including: an acquisition module 901, a processing module 902, and a communication module 903, wherein:

[0150] The acquisition module 901 is used to acquire the RGB values ​​of each pixel on the transparent screen within the field of view of the first camera during the process of displaying images on the transparent screen.

[0151] The processing module 902 is used to take the image captured by the first camera as the initial image when it is determined from the RGB value corresponding to the first camera that each pixel on the transparent screen within the field of view has the same color.

[0152] The processing module 902 removes the same color from the initial image to obtain the target image.

[0153] The communication module 903 is used to output the target image to another display device.

[0154] Optionally, the display device further includes a second camera, which is disposed on the main body and the lens of the second camera is not covered by the transparent screen; the processing module 902 is further configured to, when determining that there are multiple pixels with different colors on the transparent screen within the corresponding field of view range based on the RGB values ​​corresponding to each of the first cameras, use the image captured by the second camera as a reference image; use any image captured by the first camera as the image to be processed, for the same user captured between the image to be processed and the reference image, determine the first image region where the same user is located in the image to be processed and the second image region where the same user is located in the reference image; replace the color of the pixels in the first image region of the image to be processed with the color of the pixels in the second image region of the reference image to obtain the target image.

[0155] Optionally, the processing module 902 is further configured to extract a first image feature of the image to be processed and a second image feature of the reference image; wherein the first image feature includes at least the grayscale of the image to be processed and the second image feature includes at least the grayscale of the reference image; the first image feature and the second image feature are input into an image recognition classifier, and the same user collected between the image to be processed and the reference image is determined according to the output information of the image recognition classifier.

[0156] Optionally, the processing module 902 is further configured to, for each of the first cameras, obtain the pixel coordinates of pixels on the transparent screen within the field of view of the first camera. The processing module 902 is configured to...

[0157] For each of the first cameras, the RGB values ​​of the pixels on the transparent screen within the field of view of the first camera are determined based on the RGB values ​​of each pixel in the image displayed on the transparent screen and the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera.

[0158] Optionally, the processing module 902 is further configured to, for the first camera corresponding to the initial image, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera; determine the processing sub-region in the initial image from which the same color is to be removed based on the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera; and remove the same color from the processing sub-region.

[0159] Optionally, the processing module 902 is also used to calibrate the coordinates of the first camera in a two-dimensional coordinate system.

[0160] Each module in the aforementioned image processing apparatus 900 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0161] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the image processing method provided in any of the preceding embodiments.

[0162] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the image processing method provided in any of the preceding embodiments.

[0163] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0164] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0165] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image processing method, characterized in that, The invention is applied to a display device, which includes a body, a transparent screen, multiple first cameras, and a processor; wherein the transparent screen and the multiple first cameras are all disposed on the body, and the transparent screen covers the lenses of the first cameras; the transparent screen and the multiple first cameras are all signal-connected to the processor. The method includes: During the process of displaying images on the transparent screen, the RGB values ​​of pixels on the transparent screen within the field of view of each of the first cameras are obtained; If it is determined that each pixel on the transparent screen within the field of view has the same color based on the RGB value corresponding to the first camera, the image captured by the first camera is used as the initial image. The same color is removed from the initial image to obtain the target image; The target image is output to another display device.

2. The method according to claim 1, characterized in that, The display device further includes a second camera, which is disposed on the main body and the lens of the second camera is not covered by the transparent screen; The method further includes: When it is determined, based on the RGB values ​​of each of the first cameras, that there are multiple pixels with different colors on the transparent screen within the corresponding field of view, the image captured by the second camera is used as a reference image. Using an image captured by any of the first cameras as the image to be processed, for the same user captured between the image to be processed and the reference image, determine the first image region where the same user is located in the image to be processed and the second image region where the same user is located in the reference image. The color of the pixels in the first image region of the image to be processed is replaced with the color of the pixels in the second image region of the reference image to obtain the target image.

3. The method according to claim 2, characterized in that, The method further includes: Extract a first image feature from the image to be processed and a second image feature from the reference image; wherein the first image feature includes at least the grayscale of the image to be processed, and the second image feature includes at least the grayscale of the reference image; The first image feature and the second image feature are input into an image recognition classifier, and the same user collected between the image to be processed and the reference image is determined based on the output information of the image recognition classifier.

4. The method according to claim 2, characterized in that, Before playing the image on the transparent screen, the method further includes: For each of the first cameras, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera; The step of obtaining the RGB values ​​of pixels on the transparent screen within the field of view of each of the first cameras includes: For each of the first cameras, the RGB values ​​of the pixels on the transparent screen within the field of view of the first camera are determined based on the RGB values ​​of each pixel in the image displayed on the transparent screen and the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera.

5. The method according to any one of claims 1-4, characterized in that, Removing the same color from the initial image includes: For the first camera corresponding to the initial image, obtain the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera; Based on the pixel coordinates of the pixels on the transparent screen within the field of view of the first camera, determine the sub-region in the initial image to be processed to remove the same color; Remove the same color from the sub-region to be processed.

6. The method according to claim 4, characterized in that, Before obtaining the pixel coordinates of the corresponding pixels on the transparent screen within the field of view for each of the first cameras, the method further includes: The coordinates of the first camera in the two-dimensional coordinate system are calibrated.

7. A display device, characterized in that, include: The main body, transparent screen, multiple front cameras, and processor; The transparent screen and the plurality of first cameras are all disposed on the main body, and the transparent screen covers the lens of the first camera; The transparent screen and the multiple first cameras are all signal-connected to the processor; The transparent screen is configured to display images; The processor is configured as follows: During the process of displaying images on the transparent screen, the RGB values ​​of pixels on the transparent screen within the field of view of each of the first cameras are obtained; If it is determined that each pixel on the transparent screen within the field of view has the same color based on the RGB value corresponding to the first camera, the image captured by the first camera is used as the initial image. The same color is removed from the initial image to obtain the target image; The target image is output to another display device.

8. The display device according to claim 7, characterized in that, The display device further includes a second camera, which is disposed on the main body and the lens of the second camera is not covered by the transparent screen; The processor is also configured to: When it is determined, based on the RGB values ​​of each of the first cameras, that there are multiple pixels with different colors on the transparent screen within the corresponding field of view, the image captured by the second camera is used as a reference image. The image captured by any of the first cameras is used as the image to be processed; For the same user captured between the image to be processed and the reference image, determine the first image region where the same user is located in the image to be processed and the second image region where the same user is located in the reference image; The color of the pixels in the first image region of the image to be processed is replaced with the color of the pixels in the second image region of the reference image to obtain the target image.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Mobile device with display overlaid with at least a light sensor

    CN108369471A

  • Page display method and device, equipment and storage medium

    CN118433340A