Video picture display method and device, equipment, program product and storage medium
By utilizing a reference facial feature information database in video footage, the problem of low facial image clarity due to network bandwidth limitations is solved, achieving high-quality display of facial images and improving the security and effectiveness of video conferencing.
Patent Information
- Application Number
- CN202410620055.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-18
AI Technical Summary
Due to network bandwidth limitations, the clarity of facial images in video feeds is low, and existing technologies lack specificity, making it difficult to recognize facial expressions and details in video conferencing, thus affecting security and effectiveness.
By utilizing a database of stored reference facial feature information in the edge device, the image quality of facial images is matched and improved. High-quality facial feature information is obtained by using methods such as similarity calculation and weighted calculation. Combined with style management and component processing of facial images, high-definition facial images are generated.
It improves the clarity of facial images in video footage, enhances the ability to recognize expressions and details, ensures the security and effectiveness of video conferencing, and reduces the computing power requirements of edge devices.
Smart Images

Figure CN120980185A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to methods, apparatus, devices, systems and storage media for displaying video images. Background Technology
[0002] In video conferencing, clear facial images convey richer emotions and details, enabling users to better understand each other's expressions and intentions. Furthermore, clear facial images help users more accurately identify themselves, ensuring the security and effectiveness of video conferences. However, due to network bandwidth limitations, video images may be transmitted at reduced resolution, resulting in lower clarity of the video received by the end device, and consequently, lower clarity of facial images within the video frame. Therefore, there is an urgent need for a video display method to improve the clarity of facial images in video feeds. Summary of the Invention
[0003] This application provides a method, apparatus, device, program product, and storage medium for displaying video images, used to improve the clarity of facial images in video images.
[0004] In a first aspect, this application provides a method for displaying a video frame, the method comprising: receiving a first video frame and acquiring a first facial image in the first video frame; acquiring first facial feature information matching the first facial image from a database, the database including multiple reference facial feature information, each of the multiple reference facial feature information corresponding to multiple reference facial images; processing the first facial image based on the first facial feature information to obtain a second facial image, the second facial image having a higher image quality than the first facial image; acquiring a second video frame based on the second facial image and the first video frame, and displaying the second video frame.
[0005] This method obtains first facial feature information that matches a first facial image by using reference facial feature information stored in a database. This first facial feature information is then used to improve the image quality of the first facial image, thereby improving the image quality of the facial image in the video frame, i.e., increasing the clarity of the facial image in the video frame. Since the reference facial feature information is stored in the database beforehand, the computational power required to obtain the first facial feature information used to improve the clarity of the first facial image is relatively small, thus reducing the computational power requirements on the video display device (i.e., the end-side device). Furthermore, the obtained first facial feature information matches the first facial image, enabling a more accurate improvement in the clarity of the first facial image.
[0006] In one possible implementation, the process of obtaining first facial feature information matching a first facial image from a database includes: obtaining multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image; and obtaining the first facial feature information from the multiple reference facial feature information based on the multiple first similarities. In this approach, the accuracy of the first facial feature information is improved by confirming the first facial feature information matching the first facial image based on similarity.
[0007] In one possible implementation, the process of obtaining first facial feature information matching a first facial image from a database includes: sending a first facial image to a server storing the database; the first facial image being used by the server to obtain multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image; obtaining the first facial feature information from the multiple reference facial feature information based on the multiple first similarities; and receiving the first facial feature information sent by the server. In this method, the first facial feature information is determined by the server, reducing the processing pressure on the edge device and saving the resources of the edge device.
[0008] In one possible implementation, the process of obtaining first facial feature information from multiple reference facial feature information based on multiple first similarities includes: if the maximum similarity among the multiple first similarities is greater than a similarity threshold, obtaining first facial feature information based on the reference facial feature information corresponding to the maximum similarity; or, if the maximum similarity is less than or equal to the similarity threshold, obtaining first facial feature information based on the reference facial feature information corresponding to a second similarity, wherein the second similarity is greater than any of the multiple first similarities except for the second similarity.
[0009] In this approach, different methods for determining the first facial feature information are proposed for different situations. This allows the first facial feature information to be determined by using multiple larger similarities instead of the maximum similarity when the maximum similarity is less than or equal to the similarity threshold. This makes the acquisition of the first facial feature information more flexible.
[0010] In one possible implementation, the process of obtaining first facial feature information based on reference facial feature information corresponding to a second similarity includes: when there are multiple second similarities, performing a weighted calculation on multiple reference facial feature information corresponding to multiple second similarities based on multiple second similarities, and obtaining the first facial feature information based on the weighted calculation result. Thus, by using weighted calculation, the different weights corresponding to different similarities are fully considered, improving the accuracy of the first facial feature information.
[0011] In one possible implementation, the database further includes style information corresponding to multiple reference facial feature information; the process of obtaining first facial feature information from multiple reference facial feature information based on multiple first similarities includes: obtaining at least one second facial feature information from multiple reference facial feature information based on multiple first similarities; if the style information corresponding to at least one second facial feature information includes the required style of the first facial image, obtaining third facial feature information whose style information in at least one second facial feature information is the required style; if the style information corresponding to at least one second facial feature information does not include the required style of the first facial image, obtaining third facial feature information after converting at least one facial feature information to the required style; and obtaining first facial feature information based on the third facial feature information.
[0012] In this approach, reference facial images may have different styles. Users can determine the first facial feature information based on the desired style to ensure that the processed second facial image conforms to the desired style. Furthermore, if the desired style is not found in the database, it can be derived from existing styles. This enables style management of facial images in video footage.
[0013] In one possible implementation, the first facial feature information includes multiple component facial feature information. The process of processing the first facial image based on the first facial feature information to obtain a second facial image includes: acquiring at least one component facial image from the first facial image, each component facial image corresponding to one of the multiple component facial feature information; processing the multiple component facial images based on the multiple component facial feature information to obtain processed multiple component facial images; and obtaining the second facial image based on the processed multiple component facial images. In this approach, the facial image can be segmented into different components during processing, and by processing each component facial image, the facial image processing becomes more accurate. Optionally, processing can be performed on only at least one component; different components can achieve different levels of processing as needed while improving image quality.
[0014] In one possible implementation, the process of acquiring a first facial image from a first video frame includes: identifying a face in the first video frame; determining image positioning information based on the identified face, the image positioning information including at least one of the position of the identified face in the first video frame or the size of an image frame; and acquiring the first facial image from the first video frame based on the image positioning information. In this method, determining the first facial image based on the image positioning information corresponding to the face improves the accuracy of the first facial image.
[0015] In one possible implementation, the process of obtaining the second video frame based on the second facial image and the first video frame includes: updating the first facial image in the first video frame to a second facial image based on image location information, thereby obtaining the second video frame. In this method, the first facial image can be processed without affecting other parts of the first video frame, thus obtaining the second video frame.
[0016] Secondly, a video display device is provided, the device comprising:
[0017] The receiving module is used to receive the first video frame;
[0018] The acquisition module is used to acquire the first facial image in the first video frame;
[0019] The acquisition module is also used to acquire first facial feature information that matches the first facial image in the database. The database includes multiple reference facial feature information, and the multiple reference facial feature information corresponds to multiple reference facial images respectively.
[0020] The processing module is used to process the first facial image based on the first facial feature information to obtain a second facial image, wherein the image quality of the second facial image is higher than that of the first facial image.
[0021] The acquisition module is also used to acquire a second video frame based on the second facial image and the first video frame;
[0022] The display module is used to display the second video frame.
[0023] In one possible implementation, the acquisition module is configured to acquire multiple first similarities between multiple reference facial feature information and facial feature information of a first facial image; and to acquire the first facial feature information from the multiple reference facial feature information based on the multiple first similarities.
[0024] In one possible implementation, the acquisition module includes a sending submodule and a receiving submodule. The sending submodule is used to send a first facial image to a server storing a database. The first facial image is used by the server to obtain multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image, and to obtain the first facial feature information from the multiple reference facial feature information based on the multiple first similarities. The receiving submodule is used to receive the first facial feature information sent by the server.
[0025] In one possible implementation, the acquisition module is configured to acquire first facial feature information based on reference facial feature information corresponding to the maximum similarity when the maximum similarity among a plurality of first similarities is greater than a similarity threshold; or, when the maximum similarity is less than or equal to the similarity threshold, acquire first facial feature information based on reference facial feature information corresponding to a second similarity, wherein the second similarity is greater than any of the first similarities other than the second similarity among a plurality of first similarities.
[0026] In one possible implementation, the acquisition module is used to perform weighted calculation on multiple reference facial feature information corresponding to multiple second similarities based on multiple second similarities when there are multiple second similarities, and to acquire first facial feature information based on the weighted calculation result.
[0027] In one possible implementation, the database further includes style information corresponding to multiple reference facial feature information; an acquisition module is configured to acquire at least one second facial feature information from the multiple reference facial feature information based on multiple first similarities; if the style information corresponding to at least one second facial feature information includes the required style of the first facial image, acquire third facial feature information whose style information in the at least one second facial feature information is the required style; if the style information corresponding to at least one second facial feature information does not include the required style of the first facial image, acquire third facial feature information after converting at least one facial feature information to the required style; and acquire first facial feature information based on the third facial feature information.
[0028] In one possible implementation, the first facial feature information includes multiple component facial feature information; the processing module is used to acquire at least one component facial image in the first facial image, the at least one component facial image corresponding to the multiple component facial feature information respectively; process the multiple component facial images based on the multiple component facial feature information to obtain the processed multiple component facial images, and acquire a second facial image based on the processed multiple component facial images.
[0029] In one possible implementation, the acquisition module is used to identify faces in a first video frame, determine image positioning information based on the identified faces, the image positioning information including at least one of the position of the identified faces in the first video frame or the size of an image frame; and acquire a first facial image in the first video frame based on the image positioning information.
[0030] In one possible implementation, the acquisition module is used to update the first facial image in the first video frame to a second facial image based on image location information, thereby obtaining the second video frame.
[0031] Thirdly, a video display device is provided, the device including a memory and a processor; the memory stores at least one computer instruction, and the at least one computer instruction is loaded and executed by the processor to enable the video display device to implement the video display method of the first aspect or any possible embodiment of the first aspect.
[0032] Fourthly, a computer-readable storage medium is provided, which stores at least one instruction that is loaded and executed by a processor to enable the computer to implement the video display method described in the above aspects.
[0033] Fifthly, a computer program (product) is provided, which, when executed by a computer, enables the processor or computer to execute the video display methods described in the above aspects.
[0034] In a sixth aspect, a chip is provided, including a processor for retrieving and executing instructions stored in a memory, causing a computer equipped with the chip to execute the video display methods described in the above aspects.
[0035] In a seventh aspect, another chip is provided, comprising: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected via an internal connection path. The processor is used to execute code in the memory. When the code is executed, the computer with the chip installed executes the video display method described in the above aspects.
[0036] It should be understood that the beneficial effects of the technical solutions of the second to seventh aspects of this application and the corresponding possible implementations can be referred to the above-described technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description
[0037] Figure 1 A flowchart illustrating a video image processing method provided in related technologies;
[0038] Figure 2 This is a schematic diagram of a facial image feature region provided in related technologies.
[0039] Figure 3 This is a schematic diagram of a facial image processing process provided in related technologies;
[0040] Figure 4 A schematic diagram illustrating the implementation environment of a video display method provided in this application embodiment;
[0041] Figure 5 A flowchart illustrating a method for displaying video frames provided in an embodiment of this application;
[0042] Figure 6 A schematic diagram illustrating a facial image processing procedure provided in an embodiment of this application;
[0043] Figure 7 This application provides a schematic diagram of a process for processing a first facial image.
[0044] Figure 8 This is a schematic diagram illustrating another process for processing a first facial image provided in an embodiment of this application;
[0045] Figure 9 A schematic diagram of a video display scene provided in an embodiment of this application;
[0046] Figure 10 A schematic diagram of the structure of a video display device provided in an embodiment of this application;
[0047] Figure 11 This application provides a schematic diagram of the structure of a server according to an embodiment of the present application.
[0048] Figure 12 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0050] In video conferencing, improved facial clarity can enhance the transmission of emotions and the capture of details. Higher-resolution facial images allow users to more accurately observe facial expressions and subtle movements, leading to a deeper understanding of the other person's emotions and intentions. Users can also more accurately perceive the other person's emotional state, allowing them to adjust their communication style to better adapt to changes in the other person's mood. Furthermore, higher-resolution facial images enable users to more accurately identify participants, ensuring the security and effectiveness of the meeting. However, due to network bandwidth limitations, video feeds may be transmitted at reduced resolution to minimize latency. This reduced resolution results in lower clarity of facial images received by the end device. Therefore, end devices need to improve the clarity of facial images received in the video feed before displaying them.
[0051] In related technology one, the end-side device enhances the received video image by remapping pixel values using interpolated values generated from the enhanced and filtered pixel values. This reduces facial color variations in digital video caused by media processing or aging characteristics. See also Figure 1The flowchart illustrates a video image processing method. It processes the pixel values of a face in a frame (video image) to generate enhanced pixel values. Methods for generating enhanced pixel values include, but are not limited to, increasing contrast, brightness, and color saturation. The enhanced pixel values are then processed to generate filtered values, which are used to smooth noise in the image, sharpen image edges, etc. Finally, interpolated values associated with the enhanced and filtered pixel values are used to remap the facial pixel values. The interpolated values can be values between the enhanced and filtered pixel values, and are used to remap (or replace) the original facial pixel values, thereby improving the display quality of the face in the video image. However, in related technology one, the processing method is consistent for any facial image, lacking specificity and resulting in poor processing effects.
[0052] In related technology two, the end-side device determines multiple feature regions on the facial image based on feature points in the facial image, see [link to related technology]. Figure 2 The diagram shown illustrates a facial image feature region, in which... Figure 2 The facial image shown includes feature regions such as eyebrows, eyes, nose, and lips. Each feature region is assigned an initial enhancement weight coefficient (i.e., initial enhancement weight coefficient), which indicates the importance of the corresponding image feature region during image enhancement. By correcting the initial enhancement weight coefficients of the feature regions, the importance of the feature regions is adjusted as needed, for example, increasing the importance of the eyes. A discrete enhancement weight map of the face is obtained. The corrected enhancement weight coefficients are combined to form a discrete enhancement weight map, which can be a two-dimensional matrix. The elements in the two-dimensional matrix correspond to a pixel or pixel block in the facial image and contain an enhancement weight value. Based on the discrete enhancement weight map of the face, a continuous enhancement weight map of the face corresponding to the face image is obtained. The discrete enhancement weight map may contain some discontinuous or abrupt weight values, which may lead to unnatural edges or artifacts during enhancement. Therefore, interpolation or other smoothing techniques can be used to convert the discrete enhancement weight map into a continuous enhancement weight map, thereby making the enhancement process smoother. The face image is enhanced based on the continuous enhancement weight map of the face to obtain the face enhanced image. The original face image is then enhanced based on the continuous enhancement weight map. The enhancement process can include brightness adjustment, contrast enhancement, and sharpening. During the enhancement process, the degree of enhancement for each pixel or pixel block is determined by the weight value corresponding to each pixel or pixel block. Areas with higher weight values should be more prominent or clearer after enhancement. However, in related technology two, a general processing method is used for any facial image, which lacks specificity, results in poor processing effects, and fails to achieve an enhancement effect that generates realistic textures.
[0053] In related technology three, the edge device performs super-resolution reconstruction (SR) on the encoded and decoded video footage. It uses deep learning methods to upscale the current image, converting a low-resolution image into a high-resolution one. See also... Figure 3 The diagram illustrates a facial image processing procedure. A low-resolution facial image undergoes shallow feature extraction (SFE) via a shallow feature extraction module. Then, based on a progressive feature enhancement and upsampling module, feature enhancement is performed step-by-step through feature enhancement units (FEUs), from FEU1 to FEUn. A high-resolution face generation (HEFG) module is then used to generate a high-resolution face image. However, in related technology three, SR technology is constrained by the computing power and model performance of the edge device. If the computing power is insufficient or the model performance is not ideal, it may not achieve a significant improvement in user experience on the edge device. Furthermore, if the low-resolution image lacks details such as facial wrinkles, the super-resolution algorithm cannot recover these details and may produce a skin-smoothing and beautifying effect, resulting in a mismatch between the obtained high-resolution image and the corresponding face.
[0054] This application provides a method for displaying video images, which can improve the image quality of matched facial images based on a database storing multiple facial feature information. See also Figure 4 , Figure 4 This is a schematic diagram illustrating an implementation environment for a video display method provided in this application embodiment. The implementation environment includes an end-side device 401. Optionally, the end-side device 401 is the receiving side of the video feed from a video conference and is used to display the video feed. The end-side device 401 includes a display screen for displaying the video feed, or the end-side device 401 is connected to a display screen for displaying the video feed.
[0055] In this application embodiment, the type of the end-side device 401 is not limited. For example, the end-side device 401 can be a terminal or a server. Optionally, the terminal can be a smart device such as a mobile phone, tablet computer, or personal computer. The server can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center. The terminal and the server establish a communication connection through a wired or wireless network. Optionally, the terminal can be any electronic product that can interact with the user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device, such as a personal computer (PC), mobile phone, smartphone, personal digital assistant (PDA), wearable device, pocket PC (PPC), tablet computer, smart car system, smart TV, smart speaker, etc. The server can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center.
[0056] Those skilled in the art should understand that the above-described end-side device 401, terminal, and server are merely examples. Other existing or future end-side devices 401, terminals, or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0057] This application provides a method for displaying video images. The method is illustrated using an end-side device as an example. The end-side device can be... Figure 4 The end-side device 401 shown can also be a component on the end-side device 401, such as a single board or line card, or a functional module on the end-side device 401, or a chip used to implement the method provided in this application. This application embodiment does not specifically limit the scope. Figure 5 As shown, the method includes, but is not limited to, the following steps 501-504.
[0058] Step 501: Receive the first video frame and obtain the first facial image in the first video frame.
[0059] This application does not limit the application scenario for displaying video images, such as video conferencing, video chat, or live streaming. Taking a video conferencing scenario as an example, the end device can be a computer or mobile phone participating in the video conference. The end device receives the video image sent by the conference initiator and displays the received video image. The first video image in this application embodiment can be any video frame in the video conference. This application embodiment does not limit the number of first facial images; there can be one or multiple first facial images, and the processing method for each of the multiple first facial images is the same. This application embodiment uses a single first facial image as an example for explanation.
[0060] This application does not limit the method of acquiring the first facial image in the first video frame; the first facial image only needs to include one face. Optionally, the process of acquiring the first facial image in the first video frame may include: recognizing a face in the first video frame; determining image positioning information based on the recognized face; the image positioning information includes at least one of the position of the recognized face in the first video frame or the size of an image frame; and acquiring the first facial image in the first video frame based on the image positioning information. For example, the face in the first video frame may be recognized based on a facial recognition algorithm, which includes, but is not limited to, geometric feature-based facial recognition algorithms, template-based facial recognition algorithms, or model-based facial recognition algorithms.
[0061] The image positioning information is used to accurately locate and extract facial images from the video frame. For example, the position indicates the center point of the image frame, and the image frame size indicates the boundary of the image frame. The facial image can then be located based on the center point and the boundary. The image frame size can be flexibly set according to requirements. For example, the image frame size can be a fixed size, or it can be a size that matches the size of the face, such as being slightly larger than the face size. In one possible implementation, obtaining the first facial image from the first video frame based on the image positioning information may include either cropping the image indicated by the image positioning information from the first video frame to obtain the first facial image, or copying the image indicated by the image positioning information from the first video frame to obtain the first facial image.
[0062] For example, after receiving the first video frame, the edge device initiates a facial recognition algorithm to identify faces present in the first video frame. During the recognition process, the edge device can scan the entire first video frame based on the facial recognition algorithm, locate the facial position and size by calculating and analyzing the features of pixel data, and then determine the image positioning information based on the size of the facial position. The image indicated by the image positioning information in the first video frame is then extracted to obtain the first facial image. Optionally, the process of acquiring the first facial image in the first video frame can occur in the edge device or in the system management controller (SMC) corresponding to the edge device, and be executed by the SMC, thereby reducing the resources required by the edge device.
[0063] Step 502: Obtain first facial feature information from the database that matches the first facial image. The database includes multiple reference facial feature information, each of which corresponds to a multiple reference facial image.
[0064] In this embodiment, the database is pre-established and contains multiple reference facial feature information corresponding to multiple reference facial images. The reference facial images can be high-quality facial images. The database may or may not include these multiple reference facial images. Taking a video conferencing scenario as an example, the multiple reference facial images can be high-definition facial images of the users participating in the meeting, i.e., high-definition facial images of the attendees. In one possible implementation, the database is established by the SMC. For example, the SMC collects multiple reference facial images and extracts the multiple reference facial feature information corresponding to each reference facial image, storing it in the database. This embodiment does not limit the method by which the SMC collects the reference facial images. Taking a video conference within a company as an example, the database established by the SMC may include facial feature information of facial images of all members within the company. The SMC may collect facial images by collecting member ID badge images, member ID photos, or facial images actively uploaded by members, etc.
[0065] The content of the reference facial feature information is not limited in this embodiment. For example, the reference facial feature information may include at least one of facial geometry, facial texture, facial color, or facial feature vectors. Facial geometry can indicate the shape or size of the face. Facial texture can indicate the details and surface features of facial skin, such as wrinkles, acne scars, and skin luster. Facial color can indicate the overall tone of the face and color differences in different parts, such as skin tone, blush, or dark circles. Facial feature vectors are a specific and quantitative description method; vectors can be used to represent the position, size, and orientation of facial features. Over time, the reference facial feature information in the database may change, for example, due to changes in member facial features or the addition of new members. Therefore, it is necessary to regularly update and maintain the information in the database to ensure accuracy and effectiveness. Optionally, user identification information (ID) can be set in the database. One user ID can correspond to one or more reference facial feature information. Multiple reference facial feature information corresponding to one user ID are determined based on different photos of the same member, thereby facilitating the management of reference facial feature information.
[0066] This application does not limit the storage location of the database. For example, the database can be stored in the edge device, or it can be stored in another server, which saves resources of the edge device. In the case where the database is stored in the SMC, the process of obtaining the first facial feature information matching the first facial image in the database may include obtaining multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image; and obtaining the first facial feature information from the multiple reference facial feature information based on the multiple first similarities.
[0067] Optionally, obtaining multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image may include extracting the facial feature information of the first facial image, calculating the similarity between the facial feature information of the first facial image and each reference facial feature information, and obtaining multiple first similarities. Any first similarity indicates the degree of similarity between the first facial image and a reference facial image in the database. This application does not limit the method of determining the first similarity. For example, if the facial feature information of the first facial image indicates that the eye length is 22 mm, and a certain reference facial feature information also indicates that the eye length is 22 mm, then the similarity in the eye length is 100%, and the first similarity is obtained by combining multiple similarities from the multiple facial feature information. Alternatively, the distance between the facial feature information of the first facial image and any reference facial feature information can be calculated, such as Euclidean distance or Manhattan distance; the first similarity is determined based on the distance, for example, the larger the distance, the smaller the similarity, and the smaller the distance, the larger the similarity.
[0068] There are two scenarios for different first facial images: one is that the member corresponding to the first facial image has a facial image stored in the database, and the other is that the member corresponding to the first facial image does not have a facial image stored in the database. Based on this, this application embodiment sets a similarity threshold. A similarity threshold greater than the threshold indicates that the member corresponding to the first facial image has a facial image stored in the database, and a similarity threshold less than the threshold indicates that the member corresponding to the first facial image does not have a facial image stored in the database. Accordingly, the process of obtaining first facial feature information from multiple reference facial feature information based on multiple first similarities includes: if the maximum similarity among the multiple first similarities is greater than the similarity threshold, obtaining the first facial feature information based on the reference facial feature information corresponding to the maximum similarity; or, if the maximum similarity is less than or equal to the similarity threshold, obtaining the first facial feature information based on the reference facial feature information corresponding to the second similarity, where the second similarity is greater than any of the multiple first similarities except for the second similarity itself. That is, the second similarity is ranked in the top N positions among the multiple first similarities in descending order, where N can be any positive integer.
[0069] Specifically, if the maximum similarity is greater than the similarity threshold, it indicates that among all the reference facial feature information, there is one reference facial feature information that is sufficiently similar to the image feature information of the first facial image; that is, this sufficiently similar reference facial feature information corresponds to the same member as the image feature information of the first facial image. If the maximum similarity is less than or equal to the similarity threshold, it indicates that among all the reference facial feature information, there is no reference facial feature information that is sufficiently similar to the facial feature information of the first facial image; that is, the facial feature information of the first facial image does not correspond to the same member as any reference facial feature in the database.
[0070] In this embodiment, the second similarity can be one or more, and the number of second similarities can be freely set by the user. Although the reference facial feature information corresponding to the second similarity may not point to the same member as the facial feature information of the first facial image, the second similarity can be one or more of the highest similarities among multiple first similarities. The reference facial feature information corresponding to the second similarity can provide a reference for processing the first facial image. If the number of second similarities is one, the reference facial feature information corresponding to that second similarity can be used as the first facial feature information. If the number of second similarities is multiple, the first facial feature information can be determined based on multiple reference facial features corresponding to multiple second similarities.
[0071] Optionally, determining the first facial feature information based on multiple reference facial features corresponding to multiple second similarities may include determining the first facial feature information based on the average value of the multiple reference facial features corresponding to multiple second similarities. Alternatively, determining the first facial feature information based on multiple reference facial features corresponding to multiple second similarities may include performing a weighted calculation on the multiple reference facial feature information corresponding to multiple second similarities based on multiple second similarities, and obtaining the first facial feature information based on the weighted calculation result.
[0072] In the weighted calculation, reference facial features with higher similarity values have a larger weight in the weighted calculation, while reference facial features with lower similarity values have a smaller weight. According to the weighted calculation strategy, a weight value is assigned to each selected second similarity level. This weight value is determined based on the second similarity level; the higher the second similarity, the larger the weight value. For example, if multiple second similarities are 50%, 60%, and 70%, then the weight values for the multiple reference facial features corresponding to these second similarities are 5 / 18, 6 / 18, and 7 / 18, respectively. The reference facial feature corresponding to each second similarity level is multiplied by its corresponding weight value to obtain a weighted value. The first facial feature information is obtained based on the average of these weighted values.
[0073] Optionally, in addition to weighted calculation of the reference facial feature information, normalization can also be applied to it. Normalization transforms data of different ranges and dimensions into the same range for more accurate comparison and analysis. In facial recognition, normalization of reference facial feature information can reduce bias caused by differences in features during comparison. Normalization can be applied to different dimensions or components of a reference facial feature information. For example, if a reference facial feature information includes multiple facial feature points, such as the position and size of the eyes, nose, and mouth, the value of each feature point can be normalized separately to ensure that the value of each feature point does not interfere with the evaluation of the reference facial feature information due to large differences in a particular feature.
[0074] Normalization can be combined with weighted calculation to improve the accuracy of facial recognition. For example, reference facial feature information can be normalized first, and then weighted according to different second similarities. The first facial feature information can then be obtained based on the weighted calculation result. Optionally, the obtained first facial feature information may be normalized facial feature information; therefore, it can be de-normalized to obtain the desired first facial feature information.
[0075] In cases where the database is stored on a separate server, the end device needs to interact with the server storing the database. The process of obtaining the first facial feature information matching the first facial image in the database includes: sending the first facial image to the server storing the database; the first facial image is used by the server to obtain multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image; obtaining the first facial feature information from the multiple reference facial feature information based on the multiple first similarities; and receiving the first facial feature information sent by the server.
[0076] This application does not limit the method of sending the first facial image; for example, it can be sent using Hypertext Transfer Protocol (HTTP) or other network protocols. This application does not limit the format of the first facial image; for example, the first facial image can be in image format such as Joint Photographic Experts Group (JPEG) or Portable Network Graphics (PNG). The server obtains multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image. The method of obtaining the first facial feature information from the multiple reference facial feature information based on the multiple first similarities can refer to the method of obtaining the first facial feature information using SMC, and will not be repeated here.
[0077] In one possible implementation, the database may include multiple reference facial feature information entries for a member, with each entry corresponding to a style, such as reference facial feature information for ID photos or reference facial feature information for casual photos. In other words, the same user ID in the database may correspond to multiple reference facial feature entries and various styles, allowing different styles of reference facial feature information to be selected for different scenarios.
[0078] The process of obtaining first facial feature information from multiple reference facial feature information based on multiple first similarities includes: obtaining at least one second facial feature information from multiple reference facial feature information based on multiple first similarities; if the style information corresponding to at least one second facial feature information includes the required style of the first facial image, obtaining third facial feature information whose style information in the at least one second facial feature information is the required style; if the style information corresponding to at least one second facial feature information does not include the required style of the first facial image, obtaining third facial feature information after converting at least one facial feature information to the required style; and obtaining first facial feature information based on the third facial feature information.
[0079] Based on the first similarity score, at least one second reference facial feature is selected that is relatively similar to the first facial image. If there are multiple similar second reference facial features, these multiple second reference facial features may represent different styles of facial features of the same person. Having obtained at least one second facial feature, it is determined whether the style information corresponding to that at least one second facial feature includes the required style of the first facial image. The required style can be freely selected by the user; for example, in a more formal video conference scenario, the required style could be reference facial features of an ID photo type.
[0080] If at least one of the style information corresponding to the second facial feature information includes the required style of the first facial image, the second facial feature information corresponding to the required style is directly selected as the third facial feature information. If none of the second facial feature information meets the required style, the second facial feature information needs to be further processed to convert it into facial feature information under the required style, i.e., the third facial feature information. This application embodiment does not limit the method of style conversion, for example, aesthetic algorithm processing.
[0081] If the style information obtained is the third facial feature information corresponding to the desired style, the third facial feature information can be directly used as the first facial feature information. The method of obtaining at least one second facial feature information from multiple reference facial feature information based on multiple first similarities can be referenced in the process of obtaining first facial feature information from multiple reference facial feature information based on multiple first similarities, and will not be elaborated here.
[0082] See Figure 6 The diagram illustrates a facial image processing procedure. It describes a process where facial detection is performed on a first video frame to determine if the frame includes a face. A desired style is input, and facial recognition (i.e., acquiring the first facial image) is then performed. If the maximum similarity in the first similarity score is less than or equal to a similarity threshold, the face is not registered, and recognition fails. Then, N facial feature information points with high similarity (i.e., multiple reference facial feature information points corresponding to second similarity scores) are matched, and a weighted average is taken to obtain a feature vector (i.e., the first facial feature information). If the maximum similarity in the first similarity score is greater than the similarity threshold, recognition succeeds. The facial feature information with the highest similarity score is matched, and the feature vector (i.e., the first facial feature information) is directly obtained. Alternatively, if no facial image of the desired style is found, aesthetic algorithms are used to process the image, and a feature extractor is used to obtain the feature vector (i.e., the first facial feature information).
[0083] Step 503: Process the first facial image based on the first facial feature information to obtain a second facial image. The image quality of the second facial image is higher than that of the first facial image.
[0084] The process of processing a first facial image based on first facial feature information can be as follows: acquiring facial feature information from the first facial image, and processing the facial feature information based on the first facial feature information to achieve the effect of processing the first facial image. During the processing of the first facial image, this can be manifested in rotating, scaling, or translating the facial image to improve the similarity between the facial feature information of the first facial image and the first facial feature information. Alternatively, contrast enhancement, sharpening, and noise reduction can be performed based on the first facial feature information to improve facial details and clarity. Based on the above processing, a new facial image, namely a second facial image, is obtained. The second facial image is superior to the first facial image in image quality. Image quality includes, but is not limited to, clarity, color, or texture. For example, the clarity of the second facial image is higher than that of the first facial image.
[0085] See Figure 7 The diagram illustrates a process for processing a first facial image. Taking database storage as an example, the process involves: acquiring a first facial image from a first video frame; performing facial recognition (i.e., determining the first facial image) based on the first facial image; cropping the first facial image; retrieving first facial feature information (i.e., high-definition facial feature information) matching the first facial image from the database; and caching the high-definition facial feature information in the edge device. The facial feature information of the first facial image is then obtained through a U-net encoder. Based on the facial feature information of the first facial image and the cached high-definition facial feature information, the edge device uses a generative encoder to obtain a high-definition facial image (i.e., a second facial image).
[0086] For example, after a first facial image is captured in any video frame of a video conference, facial matching is performed using a database to accurately identify the user ID corresponding to the first facial image. That is, the user ID corresponding to the first facial feature information matching the first facial image in the database is cached in high-definition facial feature information of the user ID on the edge device. For subsequent facial images corresponding to the user ID captured in video frames, high-definition facial images can be directly obtained based on the cached high-definition facial feature information of the user ID.
[0087] In one possible implementation, processing of the first facial image can divide it into multiple component images, such as a nose component image, an eye component image, or an eyebrow component image. This allows for processing as needed and enables selective processing of only a subset of component facial images when the processing power of the edge device is insufficient, thus reducing the load on the edge device. When the first facial feature information includes multiple component facial feature information, the process of processing the first facial image based on the first facial feature information to obtain a second facial image includes: acquiring at least one component facial image from the first facial image, each component facial image corresponding to one of the multiple component facial feature information; processing the multiple component facial images based on the multiple component facial feature information to obtain processed multiple component facial images; and obtaining the second facial image based on the processed multiple component facial images.
[0088] The method of extracting at least one component facial image corresponding to these component facial feature information from the first facial image includes extracting at least one component facial image corresponding to the component facial feature information based on facial landmarks of the first facial image. Facial landmarks are used to locate and represent key facial regions, such as eyes, nose, mouth, eyebrows, and jawline. After extracting the component facial images, each component image is processed individually based on the previously extracted component facial feature information. The process of processing each component facial image is similar to the processing process of the first facial image, and will not be repeated here.
[0089] See Figure 8 The diagram illustrates another process for processing the first facial image. In this process, a high-resolution facial image (i.e., the facial image corresponding to the first facial feature information) is processed by a feature extractor to obtain the first facial feature information. Based on the facial feature information obtained from the feature extractor, the component features are cropped according to facial landmarks using region of interest alignment (RoI Align) to obtain feature vectors for the eyes, nose, lips, etc. The first facial image is then enhanced using a generative neural network to obtain the second facial image. The feature extractor for the high-resolution facial image can be VggFace, a deep face recognition system based on the VGGnet architecture. The visual geometry group network (VGGNet) is a deep convolutional neural network. The feature extractor for the first facial image can be a convolutional neural network block feature extractor (CNNBLOCK).
[0090] Step 504: Obtain the second video frame based on the second facial image and the first video frame, and display the second video frame.
[0091] In one possible implementation, the process of obtaining the second video frame based on the second facial image and the first video frame includes updating the first facial image in the first video frame to the second facial image based on image position information, thereby obtaining the second video frame. Updating the first facial image in the first video frame to the second facial image can either involve the second facial image overlaying the first facial image, or the second facial image replacing the first facial image. To better blend the superimposed image with the background video, color correction and adjustments can be performed. This ensures that the color, brightness, contrast, and other attributes of the second facial image and the first video frame match, thereby reducing any jarring effect.
[0092] For ease of understanding, this application provides a schematic diagram of a video display scene, see [link / reference]. Figure 9 The SMC registers facial data, creates a facial group, and uploads it to the facial server. Registration is successful (i.e., database establishment). When a user participates in a video conference, their face is detected (i.e., a first facial image is acquired). This first facial image is sent to the SMC, which then sends it to the facial recognition server (the server where the database resides). The facial recognition server returns the recognized high-resolution face (i.e., first facial feature information) to the SMC. The SMC returns the high-resolution face and its corresponding user ID to the terminal (i.e., the device). When the terminal subsequently acquires a first facial image corresponding to the user ID, it uses a super-resolution algorithm (i.e., super-resolution reconstruction) based on the high-resolution face corresponding to the user ID to obtain a second facial image, which is then displayed on the screen.
[0093] In summary, the video display method provided in this application obtains first facial feature information matching a first facial image by using reference facial feature information stored in a database. This first facial feature information is then used to improve the image quality of the first facial image, thereby improving the image quality of the facial image in the video frame, i.e., increasing the clarity of the facial image in the video frame. Since the reference facial feature information is stored in the database beforehand, the computational power required to obtain the first facial feature information used to improve the clarity of the first facial image is relatively small, thus reducing the computational power requirements on the video display device (i.e., the end-side device). Furthermore, the obtained first facial feature information matches the first facial image, enabling a more accurate improvement in the clarity of the first facial image.
[0094] The above describes a method for displaying video frames provided in the embodiments of this application. Corresponding to the above method, the embodiments of this application also provide a device for displaying video frames. Figure 10As shown, the video display device provided in this application embodiment includes the following modules.
[0095] Receiver module 1001 is used to receive the first video frame;
[0096] The acquisition module 1002 is used to acquire the first facial image in the first video frame;
[0097] The acquisition module 1002 is also used to acquire first facial feature information that matches the first facial image in the database. The database includes multiple reference facial feature information, and the multiple reference facial feature information corresponds to multiple reference facial images respectively.
[0098] Processing module 1003 is used to process the first facial image based on the first facial feature information to obtain a second facial image, wherein the image quality of the second facial image is higher than that of the first facial image;
[0099] The acquisition module 1002 is also used to acquire the second video frame based on the second facial image and the first video frame;
[0100] Display module 1004 is used to display the second video image.
[0101] In one possible implementation, the acquisition module 1002 is used to acquire multiple first similarities between multiple reference facial feature information and facial feature information of a first facial image; and to acquire the first facial feature information from the multiple reference facial feature information based on the multiple first similarities.
[0102] In one possible implementation, the acquisition module 1002 includes a sending submodule and a receiving submodule. The sending submodule is used to send a first facial image to a server storing a database. The first facial image is used by the server to obtain multiple first similarities between multiple reference facial feature information and the facial feature information of the first facial image, and to obtain the first facial feature information from the multiple reference facial feature information based on the multiple first similarities. The receiving submodule is used to receive the first facial feature information sent by the server.
[0103] In one possible implementation, the acquisition module 1002 is used to acquire first facial feature information based on reference facial feature information corresponding to the maximum similarity when the maximum similarity among a plurality of first similarities is greater than a similarity threshold; or, when the maximum similarity is less than or equal to the similarity threshold, to acquire first facial feature information based on reference facial feature information corresponding to a second similarity, wherein the second similarity is greater than any of the first similarities other than the second similarity among a plurality of first similarities.
[0104] In one possible implementation, the acquisition module 1002 is used to perform weighted calculation on multiple reference facial feature information corresponding to multiple second similarities based on multiple second similarities when there are multiple second similarities, and to obtain the first facial feature information based on the weighted calculation result.
[0105] In one possible implementation, the database further includes style information corresponding to multiple reference facial feature information; the acquisition module 1002 is used to acquire at least one second facial feature information from multiple reference facial feature information based on multiple first similarities; if the style information corresponding to at least one second facial feature information includes the required style of the first facial image, acquire third facial feature information whose style information in at least one second facial feature information is the required style; if the style information corresponding to at least one second facial feature information does not include the required style of the first facial image, acquire third facial feature information after converting at least one facial feature information to the required style; and acquire first facial feature information based on the third facial feature information.
[0106] In one possible implementation, the first facial feature information includes multiple component facial feature information; the processing module 1003 is used to acquire at least one component facial image in the first facial image, the at least one component facial image corresponding to the multiple component facial feature information respectively; process the multiple component facial images based on the multiple component facial feature information to obtain the processed multiple component facial images, and acquire a second facial image based on the processed multiple component facial images.
[0107] In one possible implementation, the acquisition module 1002 is used to identify faces in the first video frame, determine image positioning information based on the identified faces, the image positioning information including at least one of the position of the identified faces in the first video frame or the size of the image frame; and acquire a first facial image in the first video frame based on the image positioning information.
[0108] In one possible implementation, the acquisition module 1002 is used to update the first facial image in the first video frame to a second facial image based on the image location information, thereby obtaining the second video frame.
[0109] It should be understood that the above Figure 10 The device shown, in performing its function, possesses beneficial effects and Figure 5 The methods shown have the same beneficial effects. Figure 10The device illustrated here is only an example of the division of the above-described functional modules to demonstrate its functions. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the device and method embodiments provided in the above examples belong to the same concept, and their specific implementation processes are detailed in the method embodiments, which will not be repeated here.
[0110] In summary, the video display device provided in this application obtains first facial feature information matching a first facial image by using reference facial feature information stored in a database. This first facial feature information is then used to improve the image quality of the first facial image, thereby improving the image quality of the facial image in the video frame, i.e., increasing the clarity of the facial image in the video frame. Since the reference facial feature information is stored in the database beforehand, the computational power required to obtain the first facial feature information used to improve the clarity of the first facial image is relatively small, thus reducing the computational power requirements on the video display device (i.e., the end-side device). Furthermore, the obtained first facial feature information matches the first facial image, enabling a more accurate improvement in the clarity of the first facial image.
[0111] This application provides a video display device, which includes a memory and a processor. The memory stores at least one computer instruction, which is loaded and executed by the processor to enable the video display device to function. Figure 5 The method for displaying the video footage shown.
[0112] Figure 11 This is a schematic diagram of a server structure provided in an embodiment of this application. The server can vary significantly due to differences in configuration or performance. It may include one or more processors 1301 and one or more memories 1302. The one or more memories 1302 store at least one computer program, which is loaded and executed by the one or more processors 1301 to enable the server to implement the video display methods provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.
[0113] Figure 12This is a schematic diagram of the structure of a terminal provided in an embodiment of this application, enabling the terminal to implement the video display methods provided in the above-described method embodiments. The terminal may be, for example, a smartphone, tablet computer, media player, laptop computer, or desktop computer. The terminal may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other names.
[0114] Typically, a terminal includes a processor 1401 and a memory 1402.
[0115] Processor 1401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0116] The memory 1402 may include one or more computer-readable storage media, which may be non-transitory. The memory 1402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1402 are used to store at least one instruction, which is executed by the processor 1401 to cause the terminal to implement the video display method provided in the method embodiments of this application.
[0117] In some embodiments, the terminal may also optionally include: a peripheral device interface 1403 and at least one peripheral device. The processor 1401, memory 1402, and peripheral device interface 1403 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1403 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1404, a display screen 1405, a camera assembly 1406, an audio circuit 1407, and a power supply 1408.
[0118] Peripheral device interface 1403 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1401 and memory 1402. In some embodiments, processor 1401, memory 1402 and peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1401, memory 1402 and peripheral device interface 1403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0119] The radio frequency (RF) circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1404 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1404 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0120] Display screen 1405 is used to display a UI (User Interface). This UI may include graphics, text, icons, video, and any other combination thereof. When display screen 1405 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1401 for processing. In this case, display screen 1405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 1405 can be a single screen, located on the front panel of the terminal; in other embodiments, display screen 1405 can be at least two screens, respectively located on different surfaces of the terminal or in a folded design; in other embodiments, display screen 1405 can be a flexible display screen, located on a curved or folded surface of the terminal. Furthermore, display screen 1405 can be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 1405 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0121] The camera assembly 1406 is used to acquire images or videos. Optionally, the camera assembly 1406 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0122] The audio circuit 1407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1401 for processing, or input to the radio frequency circuit 1404 to achieve voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1407 may also include a headphone jack.
[0123] Power supply 1408 is used to power the various components in the terminal. Power supply 1408 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1408 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0124] In some embodiments, the terminal further includes one or more sensors 1409. The one or more sensors 1409 include, but are not limited to: an accelerometer 1410, a gyroscope 1411, a pressure sensor 1412, an optical sensor 1413, and a proximity sensor 1414.
[0125] Accelerometer 1410 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by the terminal. For example, accelerometer 1410 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1401 can control display screen 1405 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1410. Accelerometer 1410 can also be used for games or for acquiring user motion data.
[0126] The gyroscope sensor 1411 can detect the terminal's orientation and rotation angle. The gyroscope sensor 1411 can work in conjunction with the accelerometer sensor 1410 to collect the user's 3D movements on the terminal. Based on the data collected by the gyroscope sensor 1411, the processor 1401 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0127] The pressure sensor 1412 can be disposed on the side bezel of the terminal and / or the lower layer of the display screen 1405. When the pressure sensor 1412 is disposed on the side bezel of the terminal, it can detect the user's grip signal on the terminal, and the processor 1401 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1412. When the pressure sensor 1412 is disposed on the lower layer of the display screen 1405, the processor 1401 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0128] Optical sensor 1413 is used to collect ambient light intensity. In one embodiment, processor 1401 can control the display brightness of display screen 1405 based on the ambient light intensity collected by optical sensor 1413. Specifically, when the ambient light intensity is high, the display brightness of display screen 1405 is increased; when the ambient light intensity is low, the display brightness of display screen 1405 is decreased. In another embodiment, processor 1401 can also dynamically adjust the shooting parameters of camera assembly 1406 based on the ambient light intensity collected by optical sensor 1413.
[0129] The proximity sensor 1414, also known as a distance sensor, is typically installed on the front panel of the terminal. The proximity sensor 1414 is used to detect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1414 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1401 controls the display screen 1405 to switch from a screen-on state to a screen-off state; when the proximity sensor 1414 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1401 controls the display screen 1405 to switch from a screen-off state to a screen-on state.
[0130] Those skilled in the art will understand that Figure 12 The structure shown does not constitute a limitation on the terminal and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0131] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to enable the computer to perform... Figure 5 The method for displaying the video footage shown.
[0132] This application also provides a computer program (product) that, when executed by a computer, causes the processor or computer to perform... Figure 5 The method for displaying the video footage shown.
[0133] This application also provides a chip, including a processor, for calling and executing instructions stored in memory, causing a computer with the chip installed to perform... Figure 5 The method for displaying the video footage shown.
[0134] This application embodiment also provides another chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through internal connection paths. The processor is used to execute code in the memory. When the code is executed, it causes a computer with the chip installed to perform... Figure 5 The method for displaying the video footage shown.
[0135] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to this application are generated, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk), etc.
[0136] Those skilled in the art will recognize that the method steps and modules described in conjunction with the embodiments disclosed herein can be implemented in software, hardware, firmware, or any combination thereof. To clearly illustrate the interchangeability of hardware and software, the steps and components of each embodiment have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0137] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0138] When implemented using software, it can be implemented wholly or partially as a computer program product. This computer program product includes one or more computer program instructions. As an example, the methods of this application embodiment can be described in the context of machine-executable instructions, such as program modules that execute on a device on a real or virtual processor of the target. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., which perform specific tasks or implement specific abstract data structures. In various embodiments, the functionality of program modules can be combined or divided among the described program modules. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside on both local and remote storage media.
[0139] Computer program code used to implement the methods of the embodiments of this application may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the computer or other programmable data processing apparatus, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a standalone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0140] In the context of the embodiments of this application, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc.
[0141] Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0142] A machine-readable medium can be any tangible medium that contains or stores programs for or relating to an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. More detailed examples of machine-readable storage media include electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0143] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be found in the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0144] In the embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, devices, or modules, or they may be electrical, mechanical, or other forms of connection.
[0145] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0146] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0147] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0148] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with substantially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or order of execution. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another.
[0149] It should also be understood that, in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0150] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. For example, multiple first data packets refer to two or more first data packets. The terms "system" and "network" are often used interchangeably herein.
[0151] It should be understood that the terminology used in the description of the various examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0152] It should also be understood that the term "and / or" as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. The term "and / or" describes an association between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects are in an "or" relationship.
[0153] It should also be understood that the term “comprising” (also referred to as “includes”, “including”, “comprises” and / or “comprising”) as used in this specification specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0154] It should also be understood that the terms “if” and “if” can be interpreted as meaning “when” or “upon”, or “in response to determination” or “in response to detection”. Similarly, depending on the context, the phrases “if determination…” or “if detection [the stated condition or event]” can be interpreted as meaning “when determination…”, or “in response to determination…”, or “when detection [the stated condition or event]” or “in response to detection [the stated condition or event]”.
[0155] It should be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information.
[0156] It should also be understood that the phrases "an embodiment," "an embodiment," and "a possible implementation" used throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment or implementation is included in at least one embodiment of this application. Therefore, the phrases "in an embodiment," "an embodiment," or "a possible implementation" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0157] The above description is only an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for displaying video images, characterized in that, The method includes: Receive the first video frame and acquire the first facial image in the first video frame; Obtain first facial feature information from the database that matches the first facial image. The database includes multiple reference facial feature information, each of which corresponds to a multiple reference facial image. The first facial image is processed based on the first facial feature information to obtain a second facial image, and the image quality of the second facial image is higher than that of the first facial image. A second video frame is obtained based on the second facial image and the first video frame, and then the second video frame is displayed.
2. The method according to claim 1, characterized in that, The step of obtaining the first facial feature information in the database that matches the first facial image includes: Obtain multiple first similarities between the multiple reference facial feature information and the facial feature information of the first facial image; The first facial feature information is obtained from the plurality of reference facial feature information based on the plurality of first similarities.
3. The method according to claim 1, characterized in that, The step of obtaining the first facial feature information in the database that matches the first facial image includes: The first facial image is sent to the server storing the database. The first facial image is used by the server to obtain multiple first similarities between the multiple reference facial feature information and the facial feature information of the first facial image, and to obtain the first facial feature information from the multiple reference facial feature information according to the multiple first similarities. Receive the first facial feature information sent by the server.
4. The method according to claim 2 or 3, characterized in that, The step of obtaining the first facial feature information from the plurality of reference facial feature information based on the plurality of first similarities includes: If the maximum similarity among the plurality of first similarities is greater than a similarity threshold, the first facial feature information is obtained based on the reference facial feature information corresponding to the maximum similarity; or, If the maximum similarity is less than or equal to the similarity threshold, the first facial feature information is obtained based on the reference facial feature information corresponding to the second similarity, wherein the second similarity is greater than any of the plurality of first similarities except the second similarity.
5. The method according to claim 4, characterized in that, The step of obtaining the first facial feature information based on the reference facial feature information corresponding to the second similarity includes: When there are multiple second similarities, the multiple reference facial feature information corresponding to the multiple second similarities are weighted and calculated based on the multiple second similarities, and the first facial feature information is obtained according to the weighted calculation result.
6. The method according to claim 2 or 3, characterized in that, The database also includes style information corresponding to the plurality of reference facial feature information; obtaining the first facial feature information from the plurality of reference facial feature information based on the plurality of first similarities includes: At least one second facial feature information is obtained from the plurality of reference facial feature information based on the plurality of first similarities; If the style information corresponding to the at least one second facial feature information includes the required style of the first facial image, then obtain the third facial feature information in the at least one second facial feature information whose style information is the required style; If the style information corresponding to the at least one second facial feature information does not include the required style of the first facial image, obtain the third facial feature information after the at least one facial feature information is converted to the required style; The first facial feature information is obtained based on the third facial feature information.
7. The method according to any one of claims 1-6, characterized in that, The first facial feature information includes multiple components of facial feature information; the step of processing the first facial image based on the first facial feature information to obtain a second facial image includes: At least one component facial image is obtained from the first facial image, and the at least one component facial image corresponds to the plurality of component facial feature information respectively; The multiple component facial images are processed based on the multiple component facial feature information to obtain the processed multiple component facial images, and the second facial image is obtained based on the processed multiple component facial images.
8. The method according to any one of claims 1-7, characterized in that, The step of obtaining the first facial image from the first video frame includes: Identify faces in the first video frame, and determine image positioning information based on the identified faces. The image positioning information includes at least one of the position of the identified faces in the first video frame or the size of the image frame. The first facial image in the first video frame is obtained based on the image positioning information.
9. The method according to claim 8, characterized in that, The step of obtaining the second video frame based on the second facial image and the first video frame includes: Based on the image location information, the first facial image in the first video frame is updated to the second facial image to obtain the second video frame.
10. A display device for video images, characterized in that, The device includes: The receiving module is used to receive the first video frame; The acquisition module is used to acquire the first facial image in the first video frame; The acquisition module is further configured to acquire first facial feature information that matches the first facial image in a database. The database includes multiple reference facial feature information, and the multiple reference facial feature information corresponds to multiple reference facial images respectively. The processing module is used to process the first facial image based on the first facial feature information to obtain a second facial image, wherein the image quality of the second facial image is higher than that of the first facial image. The acquisition module is further configured to acquire a second video frame based on the second facial image and the first video frame; The display module is used to display the second video frame.
11. A video display device, characterized in that, The device includes a memory and a processor; the memory stores at least one computer instruction, which is loaded and executed by the processor to enable the device to implement the video display method according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer storage medium stores at least one instruction, which is loaded and executed by a processor to enable the computer to implement the video display method as described in any one of claims 1-9.
13. A computer program product, characterized in that, The computer program product includes: computer program code, which is loaded and executed by a computer to enable the computer to implement the video display method according to any one of claims 1-9.