Video Call Image Optimization Method, Device, Medium, and Computing Device

By extracting the characteristics of video call images and determining the image type, and combining historical optimization information for clarity optimization, the problem of poor clarity of video call images is solved and the user experience is improved.

CN118714251BActive Publication Date: 2025-06-17SHENZHEN YOUYISHUOYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410991203.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2025-06-17
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

The current video call image has poor clarity, which affects the user experience.

Method used

By detecting that the video call status is turned on, the target video image optimized based on the pre-stored historical optimization information is obtained, the image features are extracted, the image type is determined, and the target video image is optimized based on the clarity optimization information.

Benefits of technology

It improves the clarity of video call images, improves the user experience, and accurately determines the image type of the target video image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118714251B_ABST
    Figure CN118714251B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, medium and computing device for optimizing video call images, including: when detecting that the video call status is turned on, obtaining an optimized target video image based on pre-stored historical optimization information; extracting multiple image features of the target video image; determining an image type corresponding to the target video image according to the multiple image features; wherein the image type at least includes a portrait type, a landscape type, a text type and a food type; determining clarity optimization information corresponding to the image type; if the clarity optimization information is different from the historical optimization information, performing clarity optimization on the target video image according to the clarity optimization information, and outputting the optimized target video image to improve the clarity of the optimized target video image. The present invention can specifically improve the clarity of the target video image, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video calls, and in particular, to a method, device, medium, and computing device for optimizing video call images. Background Art

[0002] Currently, with the improvement of the speed of wireless networks, many portable devices are equipped with cameras and can conduct video calls. To ensure the stability of video calls, the quality of video call images is usually sacrificed. Therefore, video call images are usually of poor clarity, thus reducing the user experience. Summary of the Invention

[0003] Embodiments of the present invention provide a method, device, medium, and computing device for optimizing video call images to improve the clarity of video call images and enhance the user experience.

[0004] According to one aspect of the embodiments of the present invention, a method for optimizing video call images is provided, including:

[0005] When it is detected that the video call state is turned on, obtain a target video image optimized based on pre-stored historical optimization information;

[0006] Extract multiple image features of the target video image;

[0007] According to the multiple image features, determine the image type corresponding to the target video image; wherein the image type at least includes a portrait type, a landscape type, a text type, and a food type;

[0008] Determine clarity optimization information corresponding to the image type;

[0009] If the clarity optimization information is different from the historical optimization information, perform clarity optimization on the target video image according to the clarity optimization information, and output the optimized target video image.

[0010] As an optional implementation, the determining the image type corresponding to the target video image according to the multiple image features includes:

[0011] Determine the feature area of each image feature and determine the image area of the target video image; wherein one image feature corresponds to one feature area;

[0012] According to the respective feature areas and the image area, determine the first ratio of each image feature in the target video image; wherein one image feature corresponds to one first ratio;

[0013] Delete the image features corresponding to the first ratio that is less than the first preset ratio threshold, and obtain multiple image features to be screened after deletion;

[0014] Determine the position information of each image feature to be screened in the target video image; wherein, one image feature to be screened corresponds to one position information;

[0015] If there is target position information located in the central region of the target video image in the position information, then determine the image features to be screened corresponding to each target position information as target image features;

[0016] Determine the image type corresponding to the target video image according to the target image feature with the largest feature area.

[0017] As an optional implementation manner, if there is no target position information located in the central region of the target video image in the position information, the method further includes:

[0018] Perform recognition on the target video image through an optical character recognition technology to obtain an optical character recognition result;

[0019] If the optical character recognition result indicates that the target video image contains text information, then determine the image type corresponding to the target video image as a text type;

[0020] If the optical character recognition result indicates that the target video image does not contain text information, then determine the image type corresponding to the target video image as a landscape type.

[0021] As an optional implementation manner, when the image type corresponding to the target video image is a portrait type, the determining the clarity optimization information corresponding to the image type includes:

[0022] Determine the area of the face image of the face image from the target video image corresponding to the portrait type;

[0023] Determine the second ratio of the face image in the target video image according to the area of the face image and the area of the image;

[0024] If the second ratio is greater than or equal to the second preset ratio threshold, then perform gender recognition on the face image to obtain a gender recognition result;

[0025] Determine the face beautification parameters matching the gender recognition result;

[0026] Determine the optimization region matching the face image, wherein the optimization region includes the face image;

[0027] Determine the face beautification parameters and the optimization area together as the clarity optimization information corresponding to the portrait type.

[0028] As an alternative implementation, outputting the optimized target video image includes:

[0029] Collect the light parameters of the environment where the terminal device for video call is located through a light sensor;

[0030] Determine the exposure level corresponding to the target video image according to the light parameters;

[0031] Determine the preset brightness interval corresponding to the exposure level;

[0032] Output the optimized target video image based on the preset brightness interval, so that the brightness of the output optimized target video image is within the preset brightness interval.

[0033] According to another aspect of the embodiments of the present invention, there is also provided a video call image optimization device, including:

[0034] An acquisition unit, configured to acquire the optimized target video image based on pre-stored historical optimization information when detecting that the video call state is turned on;

[0035] An extraction unit, configured to extract multiple image features of the target video image;

[0036] A type determination unit, configured to determine the image type corresponding to the target video image according to the multiple image features; wherein, the image type at least includes portrait type, landscape type, text type, and food type;

[0037] An information determination unit, configured to determine the clarity optimization information corresponding to the image type;

[0038] An optimization unit, configured to, if the clarity optimization information is different from the historical optimization information, perform clarity optimization on the target video image according to the clarity optimization information, and output the optimized target video image.

[0039] According to yet another aspect of the embodiments of the present invention, there is also provided a computing device, which includes: at least one processor, a memory, and an input / output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the above video call image optimization method.

[0040] According to yet another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the above video call image optimization method.

[0041] In an embodiment of the present invention, when the video call state is turned on, the video image can be optimized first through pre-stored historical optimization information, and feature extraction can be performed on the optimized target video image to obtain a plurality of image features. Furthermore, by analyzing the plurality of image features, the image type corresponding to the target video image can be determined, and the clarity optimization information matching the image type can be determined. By optimizing the target video image through the clarity optimization information, the clarity of the target video image can be improved specifically, thereby enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0043] Figure 1 is a schematic flowchart of an optional video call image optimization method according to an embodiment of the present invention;

[0044] Figure 2 is a schematic flowchart of a method for determining an image type according to an embodiment of the present invention;

[0045] Figure 3 is a schematic structural diagram of an optional video call image optimization device according to an embodiment of the present invention;

[0046] Figure 4 schematically shows a structural diagram of a medium according to an embodiment of the present invention;

[0047] Figure 5 schematically shows a structural diagram of a computing device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0049] It should be noted that the terms "first", "second", etc. in the description, claims and the above drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0050] The following refers to Figure 1 , Figure 1 which is a schematic flow chart of a video call image optimization method provided by an embodiment of the present invention. It should be noted that the embodiments of the present invention can be applied to any applicable scenario.

[0051] Figure 1 The flow of the video call image optimization method provided by an embodiment of the present invention shown in the figure includes:

[0052] Step S101, when it is detected that the video call state is turned on, obtain a target video image optimized based on pre-stored historical optimization information.

[0053] In the embodiments of the present invention, the terminal device for video calls can be a smart phone, a tablet computer, a laptop computer, a smart home, a smart watch, etc. The embodiments of the present invention do not make any limitations in this regard.

[0054] In the embodiments of the present invention, the pre-stored historical optimization information can be the image optimization information used in the previous video call, or the image optimization information set in advance by the user of the terminal device, or the image optimization information used the most times in history. The historical optimization information may include an optimization area and optimization parameters; among them, the optimization area can be the area in the video image of the video call that needs to be optimized, and the area of this area can be less than or equal to the total area of the video image. The number of optimization areas can be one or more, and multiple optimization areas in the same video image can be in a separated and / or adjacent state. The optimization parameters can include, but are not limited to, image brightness, contrast, and image resolution, etc. The embodiments of the present invention do not make any limitations in this regard.

[0055] Step S102, extract multiple image features of the target video image.

[0056] Step S103, determine the image type corresponding to the target video image according to the multiple image features.

[0057] In an embodiment of the present invention, the image types at least include portrait type, landscape type, text type, and food type. For example, when there are portrait features in the target video image, the image type of the target video image can be determined as the portrait type; if there are no portrait features or the portrait features are small and the landscape features are large in the target video image, the image type of the target video image can be determined as the landscape type; if there is text information in the target video image, the image type of the target video image can be determined as the text type.

[0058] In another embodiment of the present invention, in order to accurately determine the image type corresponding to the target video image, the image features with a relatively small ratio of the feature area to the image area of the target video image can be deleted, so that the obtained multiple image features to be screened are all features with a large area, thereby being able to more accurately determine the relatively important image features in the target video image, such as Figure 2 As shown, the above step S103 is replaced by the following steps S201 to S206:

[0059] Step S201, determine the feature area of each image feature, and determine the image area of the target video image.

[0060] In an embodiment of the present invention, one image feature corresponds to one feature area; when an image feature is recognized, the contour information of the image feature in the target video image can be determined, and then the feature area of the image feature in the target video image can be determined based on the obtained contour information of the image feature.

[0061] Step S202, according to the respective feature areas and the image area, determine the first ratio of each image feature in the target video image.

[0062] In an embodiment of the present invention, one image feature corresponds to one first ratio.

[0063] Step S203, delete the image features corresponding to the first ratios less than the first preset ratio threshold, and obtain multiple image features to be screened after deletion.

[0064] In an embodiment of the present invention, if the feature area is small, it can be considered that the image feature corresponding to the feature area is not important enough. Therefore, the image features with a small feature area can be deleted to obtain multiple image features to be screened after deletion.

[0065] Step S204, determine the position information of each image feature to be screened in the target video image.

[0066] In an embodiment of the present invention, a to-be-screened image feature corresponds to a location information. The center point of each to-be-screened image feature can be determined according to the contour information corresponding to each to-be-screened image feature, that is, one to-be-screened image feature corresponds to one center point; the position coordinates of the center point in the target video image are determined as the location information of the to-be-screened image feature in the target video image.

[0067] Step S205, if there is target location information located in the central area of the target video image in the location information, then determine the to-be-screened image features respectively corresponding to each target location information as target image features.

[0068] In an embodiment of the present invention, the central area of the target video image can be a preset area; it can also be a circular area with the center point of the target video image as the center and a preset length as the radius; in addition, the target video image can also be divided into 9 sub-areas of 3×3, and then the sub-area located in the central position in the target video image can be determined as the central area.

[0069] Step S206, determine the image type corresponding to the target video image according to the target image feature with the largest feature area.

[0070] By implementing the above steps S201 to S206, the image features with a relatively small ratio of the feature area to the image area of the target video image can be deleted, so that the obtained multiple to-be-screened image features are all features with a relatively large area, thereby being able to more accurately determine the relatively important image features in the target video image, and further accurately determining the image type corresponding to the target video image.

[0071] As an alternative embodiment, if there is no target location information located in the central area of the target video image in the location information, the following steps are further included:

[0072] Identify the target video image through an optical character recognition technology to obtain an optical character recognition result;

[0073] If the optical character recognition result indicates that the target video image contains text information, then determine the image type corresponding to the target video image as the text type;

[0074] If the optical character recognition result indicates that the target video image does not contain text information, then determine the image type corresponding to the target video image as the landscape type.

[0075] Among them, when implementing this implementation manner, when it is determined that there are no image features at the center position of the target video image, the target video image can be recognized through an optical character recognition technology to determine whether there is text information in the target video image; if so, the image type corresponding to the target video image can be determined as the text type; otherwise, the image type corresponding to the target video image can be determined as the landscape type, thereby improving the accuracy of recognizing the text type and landscape type of the target video image.

[0076] In an embodiment of the present invention, the optical character recognition technology can be an OCR (Optical Character Recognition) technology. After it is recognized that the target video image contains text information, the number of characters included in the recognized text information can be determined; if the number of characters is less than a preset character threshold, it can be considered that the number of characters included in the target video image is small, and the characters are not the key features of the target video image, so the image type of the target video image cannot be determined as the text type; if the number of characters is greater than or equal to the preset character threshold, it can be considered that the number of characters included in the target video image is large, and the characters are the key features of the target video image, so the image type of the target video image can be determined as the text type.

[0077] Step S104, determine the clarity optimization information corresponding to the image type.

[0078] In an embodiment of the present invention, different image types can correspond to different clarity optimization information. For example, the optimized area in the clarity optimization information of the portrait type can be the area corresponding to the face image, and the optimization parameters in the clarity optimization information can include, but are not limited to, parameters such as brightness, hue, image resolution, and beauty information. Among them, the brightness can adjust the brightness of the face image in the optimized area; the hue can adjust the hue of the face image in the optimized area so that the hue of the adjusted face image conforms to the hue of the real face; the image resolution can improve the resolution of the face image in the optimized area to make the face image clearer; the beauty information can beautify the face image in the optimized area to improve the user experience.

[0079] In the embodiments of the present invention, the optimization area in the clarity optimization information of the landscape type can be the entire target video image, and the optimization parameters in the clarity optimization information can be adjusted according to different landscape features included in the target video image. For example, if all the landscape image features recognized in the target video image of the landscape type are fixed-type landscape image features (such as mountains, trees, buildings, etc.), the clarity of the target video image can be optimized using unified clarity optimization information; if the landscape image features recognized in the target video image of the landscape type include both fixed-type landscape image features and dynamic-type landscape image features (such as clouds, rivers, pedestrians, animals, etc.), different clarity optimization information can be formulated for different types of landscape image features. The fixed-type landscape image features can correspond to fixed-type clarity optimization information, and the dynamic-type landscape image features can correspond to dynamic-type clarity optimization information; among them, the dynamic-type clarity optimization information can eliminate the jitter of the landscape image corresponding to the dynamic-type landscape image features, making the landscape image corresponding to the dynamic-type landscape image features clearer.

[0080] In the embodiments of the present invention, the optimization area in the clarity optimization information of the text type can be the area containing all the text, and the optimization parameters in the clarity optimization information can include but are not limited to parameters such as contrast, brightness, sharpness, etc. Optimizing all the text in the optimization area based on the optimization parameters can make the text information in the optimization area clearer. In addition, the text information can be extracted through optical character recognition technology to obtain all the characters, and all the extracted characters can be output to the text output area of the terminal device display, so that the user of the terminal device can more conveniently copy, edit, and other operations on the output characters. Among them, the text output area can be superimposed on the target video image or set at a position adjacent to the target video image.

[0081] As an alternative embodiment, when the image type corresponding to the target video image is the portrait type, the manner of determining the clarity optimization information corresponding to the image type in step S104 can specifically be:

[0082] Determine the face image area of the face image from the target video image corresponding to the portrait type;

[0083] According to the face image area and the image area, determine the second ratio of the face image in the target video image;

[0084] If the second ratio is greater than or equal to the second preset ratio threshold, perform gender recognition on the face image to obtain a gender recognition result;

[0085] Determine the face beautification parameters matching the gender recognition result;

[0086] Determine an optimized region that matches the face image, where the optimized region includes the face image;

[0087] jointly determine the face beautification parameter and the optimized region as the clarity optimization information corresponding to the portrait type.

[0088] Among them, by implementing this implementation manner, when it is determined that the proportion of the face in the target video image is relatively large, the face beautification parameter can be determined according to the gender recognition result of the face, and then the face image can be optimized accordingly, improving the optimization effect of the target video image of the portrait type.

[0089] In the embodiment of the present invention, the face beautification parameter may include but is not limited to parameters such as skin color, eyebrow shape, pupil color, lip color, hair color, etc., and the embodiment of the present invention does not make any limitations in this regard.

[0090] In another embodiment, if the second ratio is less than the second preset ratio threshold, it can be considered that the proportion of the face image in the target video image is relatively small, and there is no need to perform face beautification operations. Therefore, the entire portrait image corresponding to the determined portrait type can be directly optimized to make the portrait image clearer.

[0091] Step S105, if the clarity optimization information is different from the historical optimization information, perform clarity optimization on the target video image according to the clarity optimization information, and output the optimized target video image.

[0092] As an optional implementation manner, the specific way to output the optimized target video image in step S105 may be:

[0093] Collect the light parameters of the environment where the terminal device for video call is located through a light sensor;

[0094] Determine the exposure level corresponding to the target video image according to the light parameters;

[0095] Determine the preset brightness interval corresponding to the exposure level;

[0096] Output the optimized target video image based on the preset brightness interval, so that the brightness of the output optimized target video image is within the preset brightness interval.

[0097] Among them, when implementing this implementation manner, the light sensor can be used to collect the light parameters of the environment where the terminal device is located. Furthermore, the brightness of the output target video image can be adjusted based on the obtained light parameters, so that the brightness of the target video image falls within the preset brightness interval corresponding to the light parameters, making the brightness of the output target video image more matched with the environment of the terminal device, thereby protecting the eyesight of the user of the terminal device.

[0098] In the embodiments of the present invention, the light parameters can be parameters such as light intensity, luminous flux, color temperature, etc. The corresponding exposure level can be determined according to the light parameters. The exposure level can be multiple preset levels. Different exposure levels correspond to different preset brightness intervals, and there is no overlapping area between any two preset brightness intervals. The brightness of the optimized target video image output is within the preset brightness interval, which can ensure that the light emitted by the display of the terminal device causes the least harm to the eyesight of the user of the terminal device.

[0099] When the video call state is turned on, the present invention can first optimize the video image through the pre-stored historical optimization information, and can extract multiple image features from the optimized target video image; furthermore, through the analysis of multiple image features, the image type corresponding to the target video image can be determined; and the clarity optimization information matching the image type can be determined, and the target video image can be optimized through the clarity optimization information, which can specifically improve the clarity of the target video image, thereby improving the user experience. In addition, the present invention can more accurately determine the relatively important image features in the target video image, and then can accurately determine the image type corresponding to the target video image. In addition, the present invention can also improve the accuracy of recognizing the text type and landscape type of the target video image. In addition, the present invention can also improve the optimization effect of the target video image of the portrait type. In addition, the present invention can make the brightness of the output target video image more matched with the environment of the terminal device, thereby protecting the eyesight of the user of the terminal device.

[0100] After introducing the method of the exemplary embodiment of the present invention, next, refer to Figure 3 A video call image optimization device according to an exemplary embodiment of the present invention will be described. The device includes:

[0101] An acquisition unit 301, configured to acquire an optimized target video image based on pre-stored historical optimization information when detecting that the video call state is turned on;

[0102] An extraction unit 302, configured to extract multiple image features of the target video image acquired by the acquisition unit 301;

[0103] A type determination unit 303, configured to determine an image type corresponding to the target video image according to the multiple image features obtained by the extraction unit 302; wherein, the image type at least includes a portrait type, a landscape type, a text type, and a food type;

[0104] An information determination unit 304, configured to determine clarity optimization information corresponding to the image type determined by the type determination unit 303;

[0105] An optimization unit 305, configured to, if the clarity optimization information determined by the information determination unit 304 is different from the historical optimization information, perform clarity optimization on the target video image according to the clarity optimization information, and output the optimized target video image.

[0106] As an optional implementation manner, the manner in which the type determination unit 303 determines the image type corresponding to the target video image according to the multiple image features is specifically as follows:

[0107] Determine the feature area of each image feature, and determine the image area of the target video image; wherein, one image feature corresponds to one feature area;

[0108] According to the respective feature areas and the image area, determine a first ratio of each image feature in the target video image; wherein, one image feature corresponds to one first ratio;

[0109] Delete the image features corresponding to the first ratios that are less than the first preset ratio threshold, and obtain multiple image features to be screened after deletion;

[0110] Determine the position information of each image feature to be screened in the target video image; wherein, one image feature to be screened corresponds to one position information;

[0111] If there is target position information located in the central area of the target video image in the position information, determine the image features to be screened corresponding to the respective target position information as target image features;

[0112] According to the target image feature with the largest feature area, determine the image type corresponding to the target video image.

[0113] Among them, by implementing this implementation manner, the image features with a relatively small ratio of the feature area to the image area of the target video image can be deleted, so that the multiple image features to be screened obtained are all features with a relatively large area, thereby being able to more accurately determine the relatively important image features in the target video image, and further accurately determining the image type corresponding to the target video image.

[0114] As an alternative implementation, the type determination unit 303 is further configured to:

[0115] If there is no target position information located in the central region of the target video image in the position information, the target video image is recognized by text recognition technology to obtain a text recognition result;

[0116] If the text recognition result indicates that the target video image contains text information, the image type corresponding to the target video image is determined as the text type;

[0117] If the text recognition result indicates that the target video image does not contain text information, the image type corresponding to the target video image is determined as the landscape type.

[0118] Among them, by implementing this implementation, when it is determined that there is no image feature at the center position of the target video image, the target video image can be recognized by text recognition technology to determine whether there is text information in the target video image; if there is, the image type corresponding to the target video image can be determined as the text type; otherwise, the image type corresponding to the target video image can be determined as the landscape type, thereby improving the accuracy of recognizing the text type and landscape type of the target video image.

[0119] As an alternative implementation, when the image type corresponding to the target video image is the portrait type, the manner in which the information determination unit 304 determines the clarity optimization information corresponding to the image type is specifically as follows:

[0120] Determine the face image area of the face image from the target video image corresponding to the portrait type;

[0121] According to the face image area and the image area, determine the second ratio of the face image in the target video image;

[0122] If the second ratio is greater than or equal to the second preset ratio threshold, perform gender recognition on the face image to obtain a gender recognition result;

[0123] Determine the face beautification parameters matching the gender recognition result;

[0124] Determine the optimization region matching the face image, where the optimization region includes the face image;

[0125] Determine the face beautification parameters and the optimization region together as the clarity optimization information corresponding to the portrait type.

[0126] Among them, when implementing this implementation manner, in the case where it is determined that the proportion of the human face in the target video image is relatively large, the face beautification parameters can be determined according to the gender recognition result of the human face, and then the face image can be optimized accordingly, improving the optimization effect of the target video image of the portrait type.

[0127] As an optional implementation manner, the specific way for the optimization unit 305 to output the optimized target video image is as follows:

[0128] Collect the light parameters of the environment where the terminal device for video call is located through a light sensor;

[0129] Determine the exposure level corresponding to the target video image according to the light parameters;

[0130] Determine the preset brightness interval corresponding to the exposure level;

[0131] Output the optimized target video image based on the preset brightness interval, so that the brightness of the output optimized target video image is within the preset brightness interval.

[0132] Among them, when implementing this implementation manner, the light parameters of the environment where the terminal device is located can be collected through a light sensor, and then the brightness of the output target video image can be adjusted based on the obtained light parameters, so that the brightness of the target video image falls into the preset brightness interval corresponding to the light parameters, making the brightness of the output target video image more matched with the environment of the terminal device, thereby protecting the eyesight of the user of the terminal device.

[0133] After introducing the methods and devices of the exemplary embodiments of the present invention, next, refer to Figure 4 to describe the computer-readable storage medium of the exemplary embodiments of the present invention. Please refer to Figure 4 , the shown computer-readable storage medium is an optical disc 40, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will implement the steps recorded in the above method embodiments. For example, when it is detected that the video call state is turned on, obtain the optimized target video image based on the pre-stored historical optimization information; extract multiple image features of the target video image; according to the multiple image features, determine the image type corresponding to the target video image; wherein, the image type at least includes portrait type, landscape type, text type, and food type; determine the clarity optimization information corresponding to the image type; if the clarity optimization information is different from the historical optimization information, then perform clarity optimization on the target video image according to the clarity optimization information, and output the optimized target video image; the specific implementation manners of each step will not be repeated here.

[0134] It should be noted that examples of the computer-readable storage medium may further include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0135] After introducing the methods, media, and devices of the exemplary embodiments of the present invention, next, reference is made to Figure 5 a computing device for optimizing video call images according to the exemplary embodiments of the present invention.

[0136] Figure 5 The block diagram of an exemplary computing device 50 suitable for implementing the embodiments of the present invention is shown. The computing device 50 may be a computer system or a server. Figure 5 The shown computing device 50 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0137] As Figure 5 shown, the components of the computing device 50 may include, but are not limited to: one or more processors or processing units 501, a system memory 502, and a bus 503 connecting different system components (including the system memory 502 and the processing unit 501).

[0138] The computing device 50 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computing device 50, including volatile and non-volatile media, removable and non-removable media.

[0139] The system memory 502 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 5021 and / or cache memory 5022. The computing device 50 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 5023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 5 not shown in the figure, usually referred to as a "hard disk drive"). Although not shown in Figure 5As shown, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical medium) can be provided. In these cases, each drive can be connected to the bus 503 through one or more data medium interfaces. The system memory 502 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present invention.

[0140] A program / utility 5025 having a set (at least one) of program modules 5024 can be stored, for example, in the system memory 502, and such program modules 5024 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 5024 generally perform the functions and / or methods in the embodiments described in the present invention.

[0141] The computing device 50 can also communicate with one or more external devices 504 (such as a keyboard, pointing device, display, etc.). Such communication can be carried out through the input / output (I / O) interface 605. Also, the computing device 50 can further communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 506. As Figure 5 shown, the network adapter 506 communicates with other modules (such as the processing unit 501, etc.) of the computing device 50 through the bus 503. It should be understood that although Figure 5 not shown, other hardware and / or software modules can be used in combination with the computing device 50.

[0142] The processing unit 501 executes various functional applications and data processing by running the programs stored in the system memory 502. For example, when the video call status is detected to be turned on, it obtains the target video image optimized based on the pre-stored historical optimization information; extracts multiple image features of the target video image; determines the image type corresponding to the target video image according to the multiple image features, where the image type at least includes portrait type, landscape type, text type, and food type; determines the clarity optimization information corresponding to the image type; if the clarity optimization information is different from the historical optimization information, it optimizes the clarity of the target video image according to the clarity optimization information and outputs the optimized target video image. The specific implementation manners of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the video call image optimization device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0143] In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0144] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0145] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. Also, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0146] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0147] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit.

[0148] If the described functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0149] Finally, it should be noted that: the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0150] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

Claims

1. A video call image optimization method, comprising: When it is detected that the video call state is turned on, a target video image optimized based on pre-stored historical optimization information is obtained; wherein the pre-stored historical optimization information is the image optimization information used during the last video call, or the pre-stored historical optimization information is the image optimization information pre-set by the user of the terminal device, or the pre-stored historical optimization information is the image optimization information that has been used most times in history; the historical optimization information includes an optimization area and optimization parameters; the optimization area is an area in the video image of the video call that needs to be optimized, the area of ​​the area is less than or equal to the total area of ​​the video image, and the number of the optimization areas is one or more, and multiple optimization areas in the same video image are separated and / or adjacent; Extracting a plurality of image features of the target video image; Determine the image type corresponding to the target video image according to the multiple image features; wherein the image types include at least portrait type, landscape type, text type and food type; the target video image of the portrait type has portrait features; the target video image of the landscape type has no portrait features or the existing landscape features are greater than the portrait features; There is text information in the target video image of the text type; Determining clarity optimization information corresponding to the image type; If the definition optimization information is different from the historical optimization information, optimizing the definition of the target video image according to the definition optimization information, and outputting the optimized target video image; Wherein, when the image type corresponding to the target video image is a portrait type, determining the definition optimization information corresponding to the image type includes: Determining a face image area of ​​a face image from a target video image corresponding to the portrait type; Determining a second proportion of the face image in the target video image according to the face image area and the image area; If the second ratio is greater than or equal to a second preset ratio threshold, gender recognition is performed on the face image to obtain a gender recognition result; Determine face beautification parameters that match the gender recognition result; wherein the face beautification parameters include at least skin color, eyebrow shape, pupil color, lip color, and hair color; Determining an optimized region matching the facial image, wherein the facial image is included in the optimized region; Determine the face beautification parameter and the optimization area together as clarity optimization information corresponding to the portrait type; Furthermore, the optimization area in the clarity optimization information of the scenery type is the entire target video image, and the optimization parameters in the clarity optimization information of the scenery type are adjusted according to different scenery features contained in the target video image; if the scenery features identified in the target video image of the scenery type are all scenery features of a fixed type, then the optimization parameters are unified clarity optimization information; if the scenery features identified in the target video image of the scenery type include both scenery features of a fixed type and scenery features of a dynamic type, then the optimization parameters include the clarity optimization information corresponding to the scenery features of the fixed type and the clarity optimization information corresponding to the scenery features of the dynamic type; And, the optimized area in the clarity optimization information of the text type is an area containing all texts, and the optimization parameters in the clarity optimization information of the text type include but are not limited to contrast, brightness and sharpness; And, after determining the definition optimization information corresponding to the image type, the method further includes: The text information in the target video image is extracted by text recognition technology to obtain all characters, and all the extracted characters are output to the text output area of ​​the terminal device display so that the user of the terminal device can copy or edit the output characters; wherein the text output area is superimposed on the target video image or is set at a position adjacent to the target video image.

2. The video call image optimization method according to claim 1, wherein determining the image type corresponding to the target video image according to the plurality of image features comprises: Determine the characteristic area of ​​each image feature, and determine the image area of ​​the target video image; wherein one image feature corresponds to one characteristic area; Determine a first proportion of each image feature in the target video image according to each feature area and the image area; wherein one image feature corresponds to one first proportion; Deleting image features corresponding to a first ratio that is smaller than a first preset ratio threshold, to obtain a plurality of image features to be screened after deletion; Determine the position information of each image feature to be screened in the target video image; wherein one image feature to be screened corresponds to one position information; If there is target position information located in the central area of ​​the target video image in the position information, the image features to be screened corresponding to each piece of target position information are determined as target image features; The image type corresponding to the target video image is determined according to the target image feature with the largest feature area.

3. The video call image optimization method according to claim 2, if the location information does not contain target location information located in the center area of ​​the target video image, the method further comprises: Recognize the target video image by using a text recognition technology to obtain a text recognition result; If the text recognition result indicates that the target video image contains text information, determining the image type corresponding to the target video image as a text type; If the text recognition result indicates that the target video image does not contain text information, the image type corresponding to the target video image is determined to be a landscape type.

4. The video call image optimization method according to claim 2 or 3, wherein the outputting of the optimized target video image comprises: The light sensor is used to collect light parameters of the environment in which the terminal device used for the video call is located; Determining an exposure level corresponding to the target video image according to the light parameters; Determine a preset brightness range corresponding to the exposure level; The optimized target video image is output based on the preset brightness range, so that the brightness of the output optimized target video image is within the preset brightness range.

5. A video call image optimization device, comprising: An acquisition unit is used to acquire a target video image optimized based on pre-stored historical optimization information when it is detected that a video call is turned on; wherein the pre-stored historical optimization information is the image optimization information used during the last video call, or the pre-stored historical optimization information is the image optimization information pre-set by a user of a terminal device, or the pre-stored historical optimization information is the image optimization information that has been used most times in history; the historical optimization information includes an optimization area and optimization parameters; the optimization area is an area in the video image of the video call that needs to be optimized, the area of ​​the area is less than or equal to the total area of ​​the video image, and the number of the optimization areas is one or more, and multiple optimization areas in the same video image are separated and / or adjacent; An extraction unit, used for extracting a plurality of image features of the target video image; A type determination unit is used to determine the image type corresponding to the target video image according to the multiple image features; wherein the image types include at least portrait type, landscape type, text type and food type; the target video image of the portrait type has portrait features; the target video image of the landscape type has no portrait features or the existing landscape features are greater than the portrait features; the target video image of the text type has text information; an information determination unit, configured to determine definition optimization information corresponding to the image type; an optimization unit, configured to optimize the definition of the target video image according to the definition optimization information if the definition optimization information is different from the historical optimization information, and output the optimized target video image; Wherein, when the image type corresponding to the target video image is a portrait type, the information determination unit determines the definition optimization information corresponding to the image type in the following manner: Determining a face image area of ​​a face image from a target video image corresponding to the portrait type; Determining a second proportion of the face image in the target video image according to the face image area and the image area; If the second ratio is greater than or equal to a second preset ratio threshold, gender recognition is performed on the face image to obtain a gender recognition result; Determine face beautification parameters that match the gender recognition result; wherein the face beautification parameters include at least skin color, eyebrow shape, pupil color, lip color, and hair color; Determining an optimized region matching the facial image, wherein the facial image is included in the optimized region; Determine the face beautification parameter and the optimization area together as clarity optimization information corresponding to the portrait type; Furthermore, the optimization area in the clarity optimization information of the scenery type is the entire target video image, and the optimization parameters in the clarity optimization information of the scenery type are adjusted according to different scenery features contained in the target video image; if the scenery features identified in the target video image of the scenery type are all scenery features of a fixed type, then the optimization parameters are unified clarity optimization information; if the scenery features identified in the target video image of the scenery type include both scenery features of a fixed type and scenery features of a dynamic type, then the optimization parameters include the clarity optimization information corresponding to the scenery features of the fixed type and the clarity optimization information corresponding to the scenery features of the dynamic type; And, the optimized area in the clarity optimization information of the text type is an area containing all texts, and the optimization parameters in the clarity optimization information of the text type include but are not limited to contrast, brightness and sharpness; And, the information determination unit is further used for: The text information in the target video image is extracted by text recognition technology to obtain all characters, and all the extracted characters are output to the text output area of ​​the terminal device display so that the user of the terminal device can copy or edit the output characters; wherein the text output area is superimposed on the target video image or is set at a position adjacent to the target video image.

6. The video call image optimization device according to claim 5, wherein the type determination unit determines the image type corresponding to the target video image according to the plurality of image features in the following manner: Determine the feature area of ​​each image feature, and determine the image area of ​​the target video image; wherein, One image feature corresponds to one feature area; Determine a first proportion of each image feature in the target video image according to each feature area and the image area; wherein one image feature corresponds to one first proportion; Deleting image features corresponding to a first ratio that is smaller than a first preset ratio threshold, to obtain a plurality of image features to be screened after deletion; Determine the position information of each image feature to be screened in the target video image; wherein one image feature to be screened corresponds to one position information; If there is target position information located in the central area of ​​the target video image in the position information, the image features to be screened corresponding to each piece of target position information are determined as target image features; The image type corresponding to the target video image is determined according to the target image feature with the largest feature area.

7. The video call image optimization device according to claim 6, wherein the type determination unit is further configured to: If the target position information located in the central area of ​​the target video image does not exist in the position information, the target video image is recognized by using a text recognition technology to obtain a text recognition result; If the text recognition result indicates that the target video image contains text information, determining the image type corresponding to the target video image as a text type; If the text recognition result indicates that the target video image does not contain text information, the image type corresponding to the target video image is determined to be a landscape type.

8. A computing device, comprising: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method according to any one of claims 1 to 4.

9. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Image processing method and device and terminal

    CN110163810A

  • Definition enhancement method, terminal and storage medium

    CN113989136A