Image and video processing method and device
By simultaneously acquiring and processing media streams from the front and rear cameras, and combining neural networks for iris region processing and lighting adaptation, the problem of lack of artistic expression of human eye light source and unnatural iris tracking in existing technologies has been solved, achieving high-quality integrated shooting effects of motion and still images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-12
AI Technical Summary
In selfies or portrait photography, existing technologies lack visual appeal and artistic expression due to the light reflected from the human eye. They also lack intelligent scene purification capabilities, and the iris processing and tracking mechanisms are simple, resulting in unnatural or jerky images. Furthermore, the integration of content with ambient light and shadow is not natural enough.
By simultaneously acquiring media streams from the front and rear cameras, background optimization, iris region localization and segmentation are performed. Combined with neural networks, perspective transformation and lighting adaptation are carried out to generate a fourth media stream, thereby achieving intelligent scene purification and high-precision iris tracking.
It achieves natural integration of iris content with the environment in image and video capture, enhancing visual realism and adaptability, supporting integrated shooting of static and dynamic images, resulting in higher image quality, and the integrated content moves naturally and smoothly with eye movements.
Smart Images

Figure CN122023632A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic equipment technology, and in particular to a method and apparatus for processing images and videos. Background Technology
[0002] With the widespread adoption of smartphones, tablets, and other mobile devices, and the continuous improvement of camera performance, users are increasingly pursuing more fun, creativity, and aesthetic appeal in their photography. In traditional selfies or portrait photography, the human eye typically reflects light directly from the shooting device itself or the surrounding environment, lacking visual appeal and artistic expression. To enhance the fun and aesthetics of captured images, the industry has begun exploring technologies that integrate different image content into the iris of the human eye, allowing the eye to "reflect" another scene, thus creating a unique "viewing the scene through the eye" visual effect.
[0003] However, existing technologies have drawbacks such as limited functionality, support only static image modes, lack of intelligent scene purification capabilities, simple iris processing and tracking mechanisms that may result in unnatural or jerky movements, and unnatural matching of the fused content with the curvature of the eyeball and the surrounding environment. Summary of the Invention
[0004] In view of this, embodiments of this application provide an image and video processing method and apparatus to eliminate or improve one or more defects existing in the prior art.
[0005] One aspect of this application provides a method for processing images and videos, the method comprising the following steps: Simultaneously acquire a first media stream and a second media stream; wherein the first media stream is acquired by a front-facing camera of the electronic device, and the second media stream is acquired by a rear-facing camera of the electronic device; The second media stream is subjected to background optimization processing to generate the corresponding third media stream; The first media stream is subjected to region localization processing to obtain the iris region; and the iris region is subjected to iris segmentation processing to generate a corresponding mask set. Based on the mask set, the third media stream is subjected to perspective transformation and lighting adaptation processing to generate a fourth media stream; The fourth media stream can be displayed in real time or statically.
[0006] In some embodiments of this application, background optimization processing is performed on the second media stream to generate a corresponding third media stream, including: The second media stream is subjected to noise reduction, color correction, sharpening, and fisheye distortion processing to generate the third media stream.
[0007] In some embodiments of this application, when the shooting mode is image shooting mode, the second media stream includes multiple second images; the third media stream includes each third image corresponding to each of the second images; Correspondingly, after performing background optimization processing on the second media stream to generate the corresponding third media stream, the process further includes: The third image with the least occlusion is selected from multiple third images as the reference image; The reference image is detected and target masked based on the first neural network to obtain a mask vector; Based on the second neural network, the mask vector is used to perform mask region image completion processing on the reference image to generate a third media stream with the target removed.
[0008] In some embodiments of this application, the first neural network includes an improved YOLOv8 network, whose total loss function includes CIoU loss, Focal loss, and Dice loss.
[0009] In some embodiments of this application, the second neural network includes a Transformer-based image completion network, whose total loss function includes L1 loss, perceptual loss representing the difference in semantic features between the third media stream representing the removal target and the reference image, and style loss representing the consistency in visual style features between the third media stream representing the removal target and the reference image.
[0010] In some embodiments of this application, when the shooting mode is video shooting mode, the mask set of each frame constitutes a temporal mask sequence; Correspondingly, after performing region localization processing on the first media stream to obtain the iris region; and performing iris segmentation processing on the iris region to generate a corresponding mask set, the method further includes: Calculate the iris ellipse model parameters based on the aforementioned temporal mask sequence; By combining motion prediction and weighted smoothing, the parameters of the iris ellipse model are temporally corrected to obtain the smoothed iris ellipse model parameters.
[0011] In some embodiments of this application, the iris segmentation process is performed using a third neural network based on a U-Net structure, and its total loss function is a weighted sum of Dice Loss and cross-entropy loss.
[0012] In some embodiments of this application, the fisheye distortion processing is performed using a polynomial distortion model, and the calculation formula is expressed as follows: in, This represents the distance from the original image point to the center of distortion. Represents the x-coordinate of the original pixel. Represents the ordinate of the original pixel. The x-coordinate represents the center of distortion. The ordinate representing the center of distortion. This represents the recommended value for the distortion coefficient, calibrated based on human eye physiological imaging data. This represents the distance from the corresponding pixel after distortion to the center of distortion. This represents the x-coordinate of the corresponding pixel after distortion. This represents the ordinate of the corresponding pixel after distortion.
[0013] In some embodiments of this application, the perspective transformation processing is performed based on the perspective matrix calculated from the three-dimensional face pose.
[0014] Another aspect of this application provides an electronic device including a processor and a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the image and video processing method described above.
[0015] A third aspect of this application provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the aforementioned image and video processing method.
[0016] A fourth aspect of this application provides a computer program product comprising a computer program that, when executed by a processor, implements the aforementioned image and video processing method.
[0017] This application discloses an image and video processing method, comprising the following steps: simultaneously acquiring a first media stream and a second media stream; wherein the first media stream is acquired by a front-facing camera of an electronic device, and the second media stream is acquired by a rear-facing camera of the electronic device; performing background optimization processing on the second media stream to generate a corresponding third media stream; performing region positioning processing on the first media stream to obtain an iris region; and performing iris segmentation processing on the iris region to generate a corresponding mask set; performing perspective transformation processing and illumination adaptation processing on the third media stream based on the mask set to generate a fourth media stream; and displaying the fourth media stream in real time or statically. This method supports integrated dynamic and static shooting with comprehensive functional modes; possesses intelligent scene understanding and purification capabilities, resulting in higher image quality; achieves high-precision, time-smooth iris dynamic tracking; and enhances the realism and adaptability of fusion.
[0018] Additional advantages, objectives, and features of this application will be set forth in part in the description which follows, and will in part become apparent to those skilled in the art upon review of the following description, or may be learned by practice of the application. The objectives and other advantages of this application can be realized and obtained by means of the structures specifically pointed out in the specification and drawings.
[0019] Those skilled in the art will understand that the purposes and advantages that can be achieved with this application are not limited to those specifically described above, and that the above and other purposes that this application can achieve will be more clearly understood from the following detailed description. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, do not constitute a limitation thereof. The components in the drawings are not drawn to scale but are merely for illustrating the principles of this application. For ease of illustration and description of certain parts of this application, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to this application. In the drawings: Figure 1 This is a schematic diagram of a first process of image and video processing method in one embodiment of this application.
[0021] Figure 2 This is a schematic diagram of a second process for processing images and videos in one embodiment of this application.
[0022] Figure 3 This is a schematic diagram of a third process for processing images and videos in one embodiment of this application.
[0023] Figure 4 This is a schematic diagram of the fourth process of image and video processing method in one embodiment of this application.
[0024] Figure 5 This is a schematic diagram illustrating a specific example of the image and video processing method used in this application.
[0025] Figure 6 The image and video processing device used in this application is illustrated in the diagram below as a specific example.
[0026] Figure 7 This is a flowchart illustrating one aspect of the image and video processing method in the image capture mode, as a specific example of this application.
[0027] Figure 8 This is a flowchart illustrating a video shooting mode of the image and video processing method in a specific example of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain this application, but are not intended to limit it.
[0029] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.
[0030] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0031] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0032] In the following description, embodiments of the present application will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0033] It should be noted that existing image capture methods and devices are functionally limited, supporting only static image modes. Current technologies primarily focus on capturing and fusing single images, without addressing the real-time acquisition, processing, and synthesis of continuous video streams. This fails to meet users' needs for capturing dynamic "viewing the scene from the eye" videos, limiting application scenarios. Furthermore, they lack intelligent scene cleanup capabilities. During image capture, if irrelevant foreground obstructions (such as pedestrians or objects) exist in the background scene captured by the rear camera, this solution directly distorts the entire background containing the obstructions and integrates it into the iris, resulting in cluttered content displayed in the iris, obscuring the subject and affecting visual appeal. It lacks the ability to automatically identify and remove foreground interference. The iris processing and tracking mechanism is simplistic, potentially leading to unnatural or jerky movements. For videos or situations requiring multi-frame processing, it does not elaborate on how to ensure the stability and smoothness of iris region positioning between consecutive frames. Simple single-frame detection and fusion may cause abrupt jumps or misalignments in the scene content within the iris in the video, disrupting visual continuity and realism. Furthermore, the realism and adaptability of the fusion effect of this scheme need to be improved. Although this image shooting method has implemented fisheye distortion and added enhancement effects such as shadows and highlights, it has not delved into more refined aspects, such as adaptive perspective transformation based on the specific spatial posture of the iris (three-dimensional angle, elliptical shape) and dynamic adjustment of the lighting and color of the integrated content based on the local lighting of the iris. This may result in the fusion content not matching the light and shadow of the curved surface of the eyeball and the surrounding environment naturally. Based on this, the inventors of this application first conceived of automatically identifying and removing foreground occlusions (especially people) in the background during image shooting mode through multi-frame optimization and neural network-based image completion technology, generating a clean background image with a prominent subject for fusion, improving the quality of the final image, and achieving intelligent scene purification and content completion; in video shooting mode, a lightweight neural network is introduced for frame-by-frame iris segmentation, combined with ellipse parameter estimation and weighted smoothing algorithms based on motion models, to achieve high-precision and time-stable tracking of the iris position, shape, and orientation, ensuring that the fused content changes naturally and smoothly with eye movement, thereby ensuring the accuracy and smoothness of iris tracking in the video; by introducing a perspective transformation model based on three-dimensional facial pose, the content integrated into the iris is made to better fit the three-dimensional curved surface of the eyeball; through illumination mean matching technology, the brightness and color of the fused content are dynamically adjusted to make it consistent with the lighting environment of the skin around the iris, improving the visual realism and adaptability of the fusion; by providing a unified image and video shooting framework, it can simultaneously process high-resolution still images and high-speed continuous video streams, meeting the creative shooting needs of users in different scenarios.
[0034] The following examples will provide a detailed description.
[0035] This application provides an image and video processing method, see [link to relevant documentation]. Figure 1 The method includes the following steps: Step 100: Synchronously acquire the first media stream and the second media stream; wherein the first media stream is acquired by the front-facing camera of the electronic device, and the second media stream is acquired by the rear-facing camera of the electronic device; In step 100, the first media stream includes a first image and a first video; the first image can be an image containing a face, and the first video can be a video stream captured by the front-facing camera. The second media stream includes a second image, a second image group, and a second video; the second image can be a single image captured by the rear-facing camera, and the second image group can be multiple images captured by the rear-facing camera; the second video can be a video stream captured by the rear-facing camera. The system listens for user input operations on the user's authorized device. These user input operations can be a single click on the camera button or a gesture command for image capture, or a long press on the record button or switching to video mode for video capture. Upon responding to the operation, the system immediately launches the front-facing and rear-facing cameras synchronously via the device's underlying API (e.g., Android Camera2 API's createCaptureSession, iOS's AVCaptureMultiCamSession). When the shooting mode is image capture mode, the front-facing camera captures an image containing a face as the first image, and the rear-facing camera captures one or more images as the second image or a second image group, for subsequent generation of fused content. When the shooting mode is in image capture mode, the front and rear cameras capture time-aligned video streams using the same preset frame rate and timestamp synchronization mechanism. The video stream captured by the front camera is designated as the first video, and the video stream captured by the rear camera is designated as the second video. The system ensures that each frame of the first and second videos corresponds strictly in timestamp, establishing a temporal basis for subsequent frame-by-frame fusion. The preset frame rate can be 30fps.
[0036] Step 200: Perform background optimization processing on the second media stream to generate the corresponding third media stream; In step 200, the third media stream includes a third image, a third image group, and a third video.
[0037] Step 300: Perform region localization processing on the first media stream to obtain the iris region; and perform iris segmentation processing on the iris region to generate a corresponding mask set; In step 300, face and iris region localization is performed on the first media stream (each frame of the first image or video). Lightweight models such as MTCNN (Multi-task Cascaded Convolutional Networks) can be used to quickly detect the positions of all faces in each frame of the first image or video and output a set of bounding boxes. The image and the bounding boxes of each face are input into a third neural network to segment the exposed / partially exposed iris region and generate an iris mask. For image mode, an iris mask set is obtained; for video mode, a mask set for each frame is obtained, forming a mask sequence that changes over time.
[0038] Step 400: Based on the mask set, perform perspective transformation and lighting adaptation processing on the third media stream to generate a fourth media stream; In step 400, based on the spatial pose of the iris and the three-dimensional pose information of the face, the processed background image is scaled, rotated, and transformed by perspective to adapt it to the visual effect of the iris surface. The spatial pose of the iris includes its center position, radius, and orientation angle. Simultaneously, by calculating the average illumination of the iris region and the background image, illumination consistency is adjusted to ensure that the fused content blends naturally with the surrounding environment in terms of color and brightness.
[0039] Step 500: Display the fourth media stream in real time or statically.
[0040] In step 500, the final fused fourth media stream is displayed in real-time or statically in the viewfinder using a rendering framework provided by the device's operating system, supporting user preview, confirmation, and saving. The fourth media stream includes a fourth image and a fourth video. The fourth image can be the third image processed through the above steps, and the fourth video can be the third video processed through the above steps.
[0041] As described above, the image and video processing methods provided in this application, by providing a unified image and video shooting framework, can simultaneously process high-resolution still images and high-speed continuous video streams, meeting users' creative shooting needs in different scenarios. In image shooting mode, through multi-frame optimization and neural network-based image completion technology, foreground occlusions (especially people) in the background are automatically identified and removed, generating a clean background image with a prominent subject for fusion, improving the quality of the final image, and achieving intelligent scene purification and content completion. In video shooting mode, a lightweight neural network is introduced. The system performs frame-by-frame iris segmentation and combines ellipse parameter estimation and weighted smoothing algorithms based on motion models to achieve high-precision and temporally stable tracking of the iris's position, shape, and orientation. This ensures that the fused content changes naturally and smoothly with eye movements, thereby ensuring the accuracy and smoothness of iris tracking in the video. By introducing a perspective transformation model based on 3D facial pose, the content integrated into the iris is made to better fit the three-dimensional curvature of the eyeball. Through illumination mean matching technology, the brightness and color of the fused content are dynamically adjusted to make it consistent with the lighting environment of the skin around the iris, improving the visual realism and adaptability of the fusion.
[0042] To further enhance the visual realism and adaptability of the fused images, an image and video processing method provided in this application embodiment is described below. Figure 2 Step 200 includes: Step 210: Perform noise reduction, color correction, sharpening, and fisheye distortion processing on the second media stream to generate the third media stream.
[0043] In step 210, the second media stream (each frame of the second image, second image group, or second video) is first subjected to Gaussian filtering / median filtering for noise reduction, color correction (brightness / contrast / saturation adjustment), and sharpening. Then, fisheye transformation is performed to make it conform to the visual distortion characteristics of the human eye's iris, ensuring that the image clarity, color reproduction, and distortion effect all meet the display requirements of the human eye, thus generating the third media stream (third image, third image group, or third video). In video mode, the processing is synchronized with the acquisition, ensuring real-time performance in video mode.
[0044] To further enhance the visual realism and adaptability of the fused images, an image and video processing method is provided in an embodiment of this application, see [link to relevant documentation]. Figure 3 When the shooting mode is image shooting mode, the second media stream includes multiple third images; the third media stream includes each third image corresponding to each of the second images; Correspondingly, after step 200, the following steps are also included: Step 600: Select the third image with the least occlusion from the multiple third images as the reference image; Step 700: Detect and mask the reference image based on the first neural network to obtain a mask vector; Step 800: Based on the second neural network, the mask vector is used to perform mask region image completion processing on the reference image to generate a third media stream with the target removed.
[0045] In one or more embodiments of this application, in image capture mode, a multi-frame optimization mechanism is used to select the reference image with the least occlusion from multiple images captured by the rear camera after processing in step 200. The first neural network detects and extracts the person mask in the image, and then the second neural network completes the content of the mask area, removes the foreground occlusion, and generates a clean background image as a third media stream for subsequent fusion.
[0046] To further enhance the visual realism and adaptability of the fusion, in an image and video processing method provided in this application embodiment, the first neural network includes an improved YOLOv8 network, whose total loss function includes CIoU loss, Focal loss and Dice loss.
[0047] In one or more embodiments of this application, the first neural network may be an improved YOLOv8 object detection network, which can be represented as shown in formula (1): Total loss = Bounding box loss ( ) + Classification Loss ( ) + Mask loss ( (1) Among them, the bounding box loss ( CIoU loss can be used, taking into account overlap, center point distance, and aspect ratio. Classification loss ( Focal Loss can be used to address sample imbalance. Masking loss ( DiceLoss can be used to optimize the segmentation boundary accuracy.
[0048] The bounding box loss is shown in Equation (2): (2) in, To predict the intersection-union ratio (IoU) between the bounding box and the ground truth bounding box, The Euclidean distance is the center point. The length of the bounding box diagonal. These are the weighting coefficients. For aspect ratio consistency parameters, To predict the coordinates of the center point of the bounding box, These are the coordinates of the center point of the actual bounding box.
[0049] Classification loss ( As shown in formula (3): (3) in, For category weights, To predict probabilities, The focus parameter can be set to a value of 2.
[0050] The mask loss is shown in formula (4): (4) in, The pixel position number. For the first i The predicted mask pixel value at each pixel location For the first i The actual mask pixel value at each pixel location For smoothing coefficients, .
[0051] To further enhance the visual realism and adaptability of the fusion, in an image and video processing method provided in this application embodiment, the second neural network includes a Transformer-based image completion network, whose total loss function includes L1 loss, perceptual loss representing the difference in semantic features between the third media stream of the removal target and the reference image, and style loss representing the consistency in visual style features between the third media stream of the removal target and the reference image.
[0052] In one or more embodiments of this application, the second neural network can be a Transformer-based image completion network, which can be represented as shown in formula (5): Total loss = pixel loss ( ) + Perceived loss ( ) + Style loss ( (5) Among them, pixel loss ( L1 loss can be used to ensure that the completed region closely approximates the real background. Perceptual loss ( The difference can be calculated based on the intermediate layer features of VGG16. Style loss ( It can maintain style consistency based on the VGG16 feature map Gram matrix.
[0053] The pixel loss is shown in formula (6): (6) in, To complete the image pixel values, The pixel values are those of a real image without a target person. , Let h be the image size, h be the pixel index of the image in the height direction, and w be the pixel index of the image in the width direction.
[0054] Perceived loss ( As shown in formula (7): (7) in, This is the feature output of the m-th layer of VGG16. , , The feature map size and number of channels are given by M, where M represents the total number of feature layers selected for calculating the perceptual loss in the pre-trained convolutional neural network (such as VGG16), and m represents the feature of the m-th layer currently being calculated.
[0055] Style loss ( As shown in formula (8): (8) in, For feature map Gram matrix.
[0056] To further enhance the visual realism and adaptability of the fused images, an image and video processing method is provided in an embodiment of this application, see [link to relevant documentation]. Figure 4 When the shooting mode is video shooting mode, the mask set of each frame constitutes a temporal mask sequence; Correspondingly, after step 300, the following steps are also included: Step 010: Calculate the iris ellipse model parameters based on the temporal mask sequence; In step 010, based on the iris mask of consecutive frames, the elliptical model parameters of the iris are estimated. These parameters can be the center coordinates, the major and minor axes, and the deflection angle. The elliptical model parameters are calculated as follows: in, Let be the iris mask of the k-th face in frame t. The coordinates of the iris center are given, and the major and minor semi-axis of the iris ellipse are given. and deflection angle ,in, The rotation parameters of the ellipse are obtained through an ellipse fitting algorithm, which involves fitting an ellipse using the pixels at the edge of the iris and calculating the angle between the major axis of the ellipse and the horizontal direction.
[0057] Step 020: Combine motion prediction and weighted smoothing to perform time-series correction on the iris ellipse model parameters to obtain the smoothed iris ellipse model parameters.
[0058] In step 020, temporal smoothing tracking of the iris position is achieved through motion prediction and weighted smoothing algorithms to ensure that the fused content in the video remains synchronized with the iris motion. Based on the actual iris position and motion trend of the previous frame, the theoretical position of the iris in the current frame is predicted as follows: in, and Let K be the actual center coordinates of the k-th iris in frame t-1. and Let be the predicted center coordinates of the k-th iris in frame t. For the first The iris motion velocity of the frame, Given a frame time interval, the iris position estimated in the current frame is used to correct the theoretical position through a weighted average, as shown in the following formula: in, The smoothing factor, with a value ranging from 0 to 1, is used to correct the iris deflection angle in the same way. and long and short half shafts This ensures that the iris ellipse model changes smoothly across consecutive frames.
[0059] To further enhance the visual realism and adaptability of the fusion, in an image and video processing method provided in this application embodiment, iris segmentation is performed using a third neural network based on the U-Net structure, and its total loss function is a weighted sum of Dice Loss and cross-entropy loss.
[0060] In one or more embodiments of this application, when the shooting mode is video shooting mode, in step 300, the iris region in each frame is segmented in real time using a lightweight face detection model and a third neural network to extract the iris mask sequence. The total loss function of the third neural network combines Dice Loss and Cross Entropy Loss to balance positive and negative samples and improve the accuracy of the segmentation boundary. The formula for the total loss function of the third neural network is as follows: in, This is the weighting coefficient, with a reference value of 0.5. This represents the cross-entropy loss.
[0061] To further enhance the visual realism and adaptability of the fused images, in an image and video processing method provided in this application embodiment, fisheye distortion processing is performed using a polynomial distortion model, and the calculation formula is as follows: in, This represents the distance from the original image point to the center of distortion. Represents the x-coordinate of the original pixel. Represents the ordinate of the original pixel. The x-coordinate represents the center of distortion. The ordinate representing the center of distortion. This represents the recommended value for the distortion coefficient, calibrated based on human eye physiological imaging data. This represents the distance from the corresponding pixel after distortion to the center of distortion. This represents the x-coordinate of the corresponding pixel after distortion. This represents the ordinate of the corresponding pixel after distortion.
[0062] To further enhance the visual realism and adaptability of the fusion, in an image and video processing method provided in this application embodiment, perspective transformation processing is performed based on the perspective matrix calculated from the three-dimensional face pose.
[0063] In one or more embodiments of this application, the processed background image is scaled, rotated, and perspective-transformed based on the spatial pose of the iris (center position, radius, and orientation angle) and the three-dimensional pose information of the face to adapt it to the visual effect of the iris surface. Simultaneously, by calculating the average illumination of the iris region and the background image, illumination consistency adjustment is performed to ensure that the fused content blends naturally with the surrounding environment in terms of color and brightness. In image capture mode, the third image after removing the target is scaled, rotated, and perspective-transformed to obtain an image adapted to the iris region. Specifically, the scaling factor is calculated based on the iris radius, and the formula is as follows: in, To remove the third image size of the target, For the iris radius, the scaling formula is expressed as: in, The image is scaled up. To remove the third image size of the target, This is the scaling factor.
[0064] Based on orientation angle The scaled image is rotated, and the rotation matrix is represented as follows: The rotated coordinates can be represented as: The corresponding pixel value is obtained through bilinear interpolation, as shown in the following formula: in, This represents the rotated image. For adjacent integer coordinates, These are the interpolation weights.
[0065] Based on the 3D pose of the face in the first image (calculated from facial key points), using the perspective matrix... Adjustment The perspective is adjusted to match the human eye's perspective after integration; the transformation formula is as follows: .
[0066] Calculate the mean illumination of the iris region in the first image. ,as well as Average light intensity Adjust using the following formula Light: The final fusion yields the fourth image. : ,in, This represents the union operation of multiple iris masks, ensuring that the masking is integrated only in the iris region. Content in," "" indicates taking the union of the iris masks of multiple individuals.
[0067] For specific examples of the image and video processing methods in this application, see [link to example]. Figure 5 and Figure 6The method includes the following steps: Listening for user input operations, which can be a single click of the shutter button or a gesture command for image capture, or a long press of the record button or switching to video mode for video capture. Upon responding to the operation, the front and rear cameras are immediately launched synchronously via the device's underlying API (e.g., Android Camera2 API's createCaptureSession, iOS's AVCaptureMultiCamSession). In image capture mode, the front camera captures an image containing a face as the first image (denoted as...). The rear camera captures one or more images as a second image (denoted as...). ) or the second image group (denoted as This is used to generate fused content later. In video shooting mode, the front and rear cameras capture time-aligned video streams at the same preset frame rate (e.g., 30fps) and timestamp synchronization mechanism. The video stream captured by the front camera is denoted as the first video stream. The video stream captured by the rear camera is the second video stream. The system ensures and Each frame corresponds strictly to a timestamp, establishing a temporal foundation for subsequent frame-by-frame fusion. For the second media stream ( , or Every frame (t is the frame number) First, Gaussian filtering / median filtering is used for noise reduction, color correction (brightness / contrast / saturation adjustment), and sharpening. Then, fisheye transformation is performed to make it conform to the visual distortion characteristics of the human iris, ensuring that the image clarity, color reproduction, and distortion effect all meet the display requirements of the human eye, and generating a third media stream. , or The processing is synchronized with the acquisition process to ensure real-time performance in video mode.
[0068] A polynomial distortion model is adopted, and acceleration is achieved through pre-calculated coordinate mapping tables and bilinear interpolation. The core transformation formula is as follows: in, These are the original pixel coordinates. It is the distortion center (usually the image center). The distance from the preimage point to the center. This is the distance after distortion. Recommended values for distortion coefficients based on human eye physiological imaging data: , , .
[0069] In image capture mode, see Figure 7 ,Will Input the first neural network, and output the bounding boxes and masks of all characters. ), select the 1-3 people with the largest area in the center of the image as the target people, and retain their masks ( The mask is in vector form. The first neural network can be an improved YOLOv8 object detection network: total loss = bounding box loss ( ) + Classification Loss ( ) + Mask loss ( ); Bounding box loss ( Using CIoU loss, considering overlap, center point distance, and aspect ratio, the formula is: in, To perform an intersection-union comparison between the predicted bounding box and the ground truth bounding box, The Euclidean distance is the center point. The length of the bounding box diagonal. These are the weighting coefficients. This is a parameter for aspect ratio consistency.
[0070] Classification loss ( Focal Loss is used to address sample imbalance. The formula is: in, For category weights, To predict probabilities, (Focus parameters).
[0071] Mask loss ( The Dice Loss algorithm is used to optimize the segmentation boundary accuracy. The formula is: in, To predict the mask pixel values, These are the actual mask pixel values. (Smoothing coefficient).
[0072] Will and Inputting the data into the network completes the mask coverage area, generating a third image with the target person removed. If a second set of images was captured, then the corresponding third set of images is entered.
[0073] The second neural network can be a Transformer-based image completion network, where the total loss equals the pixel loss ( ) + Perceived loss ( ) + Style loss ( ) Pixel loss ( L1 loss ensures that the filled region closely approximates the real background; the formula is: in, To complete the image pixel values, The pixel values are those of a real image without a target person. , This refers to the image size.
[0074] Perceived loss ( The difference is calculated based on the VGG16 intermediate layer features, using the following formula: in, This is the feature output of the m-th layer of VGG16. , , This defines the feature map size and number of channels for this layer.
[0075] Style loss ( ): Maintaining style consistency based on the VGG16 feature map Gram matrix, the formula is: in, For feature map Gram matrix.
[0076] For the first media stream ( or Every frame (t is the frame number) Face and iris region localization is performed; this step is common to both image and video modes. Lightweight models such as MTCNN (Multi-task Cascaded Convolutional Networks) are used to quickly detect... or The image is processed to locate all faces, and a set of bounding boxes is output. The image and the bounding boxes of each face are then input into a third neural network to segment the exposed / partially exposed iris regions, generating iris masks. For each image mode, a set of iris masks is obtained. For video mode, the mask set for each frame is obtained. This forms a mask sequence that varies over time. Its total loss function combines Dice Loss and Cross Entropy Loss to balance positive and negative samples and improve segmentation boundary accuracy, as shown in the following formula: in, This is the weighting coefficient, with a reference value of 0.5. This represents the cross-entropy loss.
[0077] In video recording mode, see Figure 8 For the first Frame number Iris mask of an individual's face Calculate its elliptical model parameters and the coordinates of the iris center. Represented as: The major and minor semi-axis of the iris ellipse and deflection angle ,in These are the ellipse rotation parameters, obtained through an ellipse fitting algorithm.
[0078] Based on the actual position and motion trend of the iris in the previous frame, the theoretical position of the iris in the current frame is predicted by the following formula: in For the first The iris motion velocity of the frame, Given a frame time interval, and combining the estimated iris position with the current frame, the theoretical position is corrected using a weighted average as follows: in The smoothing factor, with a value ranging from 0 to 1, is used to correct the iris deflection angle in the same way. and long and short half shafts This ensures that the iris ellipse model changes smoothly across consecutive frames.
[0079] The processed third-party media stream content (in image mode) Or the current frame in video mode Based on the spatial posture of the iris, it is precisely integrated into the first media stream ( or The iris region. This step is common to both image and video modes, but the video mode uses smoothed elliptical model parameters. .
[0080] Taking image mode as an example: (1) For image patterns, the iris mask set is obtained. , For the first Iris mask of an individual's face For the second image size, extract the first image using the following formula. Key parameters of an iris: Iris center coordinates : , ; Iris radius : ; Iris facing angle : By fitting an ellipse using the pixels at the edge of the iris, the angle between the major axis of the ellipse and the horizontal direction is calculated using the following formula: ,in, The major and minor semi-axes of the ellipse, These are the rotation parameters for the ellipse; right Scaling, rotation, and perspective transformations are performed to obtain an image adapted to the iris region. Specifically, based on the iris radius Calculate scaling factor ( for (Size), scaling formula is ; based on Rotate the scaled image using the rotation matrix. Coordinates after rotation The corresponding pixel value is obtained through bilinear interpolation: ,in, For adjacent integer coordinates, For interpolation weights; Based on the 3D pose of the face in the first image (calculated from facial key points), using the perspective matrix... Adjustment The perspective is adjusted to match the human eye's perspective after integration; the transformation formula is as follows: ; Calculate the mean illumination of the iris region in the first image. ,as well as Average light intensity Adjust using the following formula Light: ; The final fusion yields the fourth image. : ,in" "" represents the union operation of multiple iris masks, ensuring that the masking is integrated only in the iris region. Content in," "" indicates taking the union of the iris masks of multiple individuals.
[0081] The composite media stream is rendered in real time in the viewfinder using the UI rendering framework provided by the device's operating system (such as TextureView / SurfaceView for Android, and AVCaptureVideoPreviewLayer / CALayer for iOS).
[0082] In video mode, low-latency real-time preview is achieved, allowing users to intuitively see the dynamic effects of "viewing the scenery with their own eyes" and trigger recording to save the synthesized first video stream to the memory.
[0083] In image mode, the final composite image is displayed, and it is saved as an image file after user confirmation.
[0084] This application also provides an electronic device that may include at least one front-facing camera, at least one rear-facing camera, a display screen, a processor, a memory, a receiver, and a transmitter. The processor is used to execute the image and video processing methods mentioned in the above embodiments. The processor and memory can be connected via a bus or other means, taking a bus connection as an example. The receiver can be connected to the processor and memory via wired or wireless means.
[0085] The processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, an image processor (ISP), a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations thereof.
[0086] Memory (RAM / ROM), as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the image and video processing methods described in the embodiments of this application. The processor executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the image and video processing methods described in the above method embodiments.
[0087] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0088] The one or more modules are stored in the memory, and when executed by the processor, they perform the image and video processing method described in the embodiment.
[0089] In some embodiments of this application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected via a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.
[0090] As one implementation method, the functions of the receiver and transmitter in this application can be implemented by transceiver circuits or dedicated transceiver chips, and the processor can be implemented by dedicated processing chips, processing circuits or general-purpose chips.
[0091] As another implementation approach, the server provided in this application embodiment can be implemented using a general-purpose computer. That is, the program code implementing the processor, receiver, and transmitter functions is stored in memory, and the general-purpose processor implements the processor, receiver, and transmitter functions by executing the code in memory.
[0092] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned image and video processing method. The computer-readable storage medium can be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.
[0093] This application also provides a computer program product, including a computer program that, when executed by a processor, controls a camera to acquire data synchronously, schedules the operation of various neural network models, executes a series of algorithms such as fisheye transformation, geometric calculation, illumination adjustment, and pixel fusion, and manages the preview and storage interface to implement the steps of the aforementioned image and video processing methods.
[0094] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave.
[0095] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0096] In this application, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0097] The above description is merely a preferred embodiment of this application and is not intended to limit this application. For those skilled in the art, various modifications and variations can be made to the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for processing images and videos, characterized in that, The method includes: Simultaneously acquire a first media stream and a second media stream; wherein the first media stream is acquired by a front-facing camera of the electronic device, and the second media stream is acquired by a rear-facing camera of the electronic device; The second media stream is subjected to background optimization processing to generate the corresponding third media stream; The first media stream is subjected to region localization processing to obtain the iris region; and the iris region is subjected to iris segmentation processing to generate a corresponding mask set. Based on the mask set, the third media stream is subjected to perspective transformation and lighting adaptation processing to generate a fourth media stream; The fourth media stream can be displayed in real time or statically.
2. The method according to claim 1, characterized in that, The second media stream undergoes background optimization processing to generate a corresponding third media stream, including: The second media stream is subjected to noise reduction, color correction, sharpening, and fisheye distortion processing to generate the third media stream.
3. The method according to claim 1, characterized in that, When the shooting mode is image shooting mode, the second media stream includes multiple second images; the third media stream includes each third image corresponding to each of the second images. Correspondingly, after performing background optimization processing on the second media stream to generate the corresponding third media stream, the process further includes: The third image with the least occlusion is selected from multiple third images as the reference image; The reference image is detected and target masked based on the first neural network to obtain a mask vector. Based on the second neural network, the mask vector is used to perform mask region image completion processing on the reference image to generate a third media stream with the target removed.
4. The method according to claim 3, characterized in that, The first neural network includes an improved YOLOv8 network, whose total loss function includes CIoU loss, Focal loss, and Dice loss.
5. The method according to claim 3, characterized in that, The second neural network includes a Transformer-based image completion network, whose total loss function includes L1 loss, perceptual loss representing the difference in semantic features between the third media stream representing the removal target and the reference image, and style loss representing the consistency in visual style features between the third media stream representing the removal target and the reference image.
6. The method according to claim 1, characterized in that, When the shooting mode is video shooting mode, the mask set of each frame constitutes a temporal mask sequence; Correspondingly, after performing region localization processing on the first media stream to obtain the iris region; and performing iris segmentation processing on the iris region to generate a corresponding mask set, the method further includes: Calculate the iris ellipse model parameters based on the aforementioned temporal mask sequence; By combining motion prediction and weighted smoothing, the parameters of the iris ellipse model are temporally corrected to obtain the smoothed iris ellipse model parameters.
7. The method according to claim 1, characterized in that, The iris segmentation process is performed using a third neural network based on the U-Net structure, and its total loss function is a weighted sum of Dice Loss and cross-entropy loss.
8. The method according to claim 2, characterized in that, The fisheye distortion processing is performed using a polynomial distortion model, and the calculation formula is as follows: in, This represents the distance from the original image point to the center of distortion. Represents the x-coordinate of the original pixel. Represents the ordinate of the original pixel. The x-coordinate represents the center of distortion. The ordinate representing the center of distortion. This represents the recommended value for the distortion coefficient, calibrated based on human eye physiological imaging data. This represents the distance from the corresponding pixel after distortion to the center of distortion. This represents the x-coordinate of the corresponding pixel after distortion. This represents the ordinate of the corresponding pixel after distortion.
9. The method according to claim 1, characterized in that, The perspective transformation process is based on the calculation of the perspective matrix of the three-dimensional human face pose.
10. An electronic device comprising a processor and a memory, characterized in that, When the processor executes the running program stored in the memory, it implements the image and video processing method as described in any one of claims 1 to 9.