Image processing method, image processing chip, and electronic device

WO2026175352A1PCT designated stage Publication Date: 2026-08-27VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/079233
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-20
Filing Date
2026-02-13
Publication Date
2026-08-27

Smart Images

  • Figure CN2026079233_27082026_PF_FP_ABST
    Figure CN2026079233_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and discloses an image processing method, an image processing chip, and an electronic device. The image processing method comprises: when a video recording interface is displayed, receiving a first input; in response to the first input, performing video recording on a first scene by means of a first camera, and performing video recording on a second scene by means of a second camera; and on the basis of the first scene and the second scene, performing image fusion and then displaying a fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing methods, image processing chips and electronic devices

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510191282.6, filed in China on February 20, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application belongs to the field of image processing technology, specifically relating to an image processing method, an image processing chip, and an electronic device. Background Technology

[0004] With the continuous development of digital technology and the widespread use of electronic devices, people have increasingly higher requirements for the intelligence level of electronic devices. In daily life and work, users often need to use cameras to record videos.

[0005] Currently, when a user records video using their phone's camera app, the phone triggers the corresponding camera to record separately based on the user's selected recording mode, resulting in a separate video recording. For example, when a user takes a selfie, the phone triggers the front-facing camera to record separately, resulting in a separate video recording of the face.

[0006] However, when taking pictures with an electronic device that contains multiple cameras, the image captured is taken by a single camera, meaning the image is a picture of a specific scene. Therefore, the image effect obtained by the electronic device through the camera is relatively simple and the shooting method is not flexible enough. Summary of the Invention

[0007] The purpose of this application is to provide an image processing method, an image processing chip, and an electronic device that can improve the flexibility of electronic devices in taking pictures.

[0008] In a first aspect, embodiments of this application provide an image processing method, which includes: receiving a first input while displaying a video recording interface; in response to the first input, recording a video of a first scene using a first camera and recording a video of a second scene using a second camera; performing image fusion based on the first scene and the second scene, and displaying the fused image.

[0009] Secondly, embodiments of this application provide an image processing chip, comprising: a first frame interpolation module, a second frame interpolation module, and a fusion module, wherein the first frame interpolation module and the second frame interpolation module are both connected to the fusion module; wherein the first frame interpolation module is used to perform frame interpolation processing on image frame data of a first scene acquired by a first camera to obtain first image data; the second frame interpolation module is used to perform frame interpolation processing on image frame data of a second scene acquired by a second camera to obtain second image data; and the fusion module is used to perform fusion processing on the first image data and the second image data to obtain fused image data. Thirdly, embodiments of this application provide an electronic device, comprising: a first camera, a second camera, and the image processing chip as described in the second aspect; both the first camera and the second camera are connected to the image processing chip.

[0010] Fourthly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.

[0011] Fifthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0012] In a sixth aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0013] In a seventh aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0014] In this embodiment, when a video recording interface is displayed, a first input is received; in response to the first input, video recording of a first scene is performed using a first camera, and video recording of a second scene is performed using a second camera; image fusion is performed based on the first scene and the second scene, and the fused image is displayed. In this solution, since the electronic device can perform image fusion based on the first scene and the second scene to display the fused image, and the first scene is the scene recorded by the first camera, and the second scene is the scene recorded by the second camera, the electronic device can perform fusion processing on the image data of the first scene recorded by the first camera and the image data of the second scene recorded by the second camera, and display the image corresponding to the fused image data obtained from the fusion processing. Thus, in this way, when shooting with an electronic device containing multiple cameras, the electronic device can respond to the first input, simultaneously run multiple cameras to shoot, and perform fusion processing on the video images captured by the multiple cameras to display the fused image, making the image captured by the electronic device a fused image of multiple scenes, improving the flexibility of the electronic device in shooting. Attached Figure Description

[0015] Figure 1 is a flowchart of one of the image processing methods provided in the embodiments of this application;

[0016] Figure 2 is a structural schematic diagram of the dual-folding screen mobile phone provided in an embodiment of this application;

[0017] Figure 3 is a schematic diagram of the folding angle of the dual-folding screen mobile phone provided in the embodiment of this application;

[0018] Figure 4 is a second flowchart of the image processing method provided in the embodiment of this application;

[0019] Figure 5 is a flowchart of the third image processing method provided in the embodiment of this application;

[0020] Figure 6 is a schematic diagram of the structure of the fusion network model provided in an embodiment of this application;

[0021] Figure 7 is a schematic diagram of the fusion module provided in an embodiment of this application;

[0022] Figure 8 is a flowchart of the image processing method provided in the embodiment of this application;

[0023] Figure 9 is a schematic diagram of the structure of an image processing chip provided in an embodiment of this application;

[0024] Figure 10 is a fifth flowchart of the image processing method provided in the embodiments of this application;

[0025] Figure 11 is a second schematic diagram of the structure of the image processing chip provided in an embodiment of this application;

[0026] Figure 12 is a schematic diagram of the hardware structure of the dual-folding screen mobile phone provided in the embodiment of this application;

[0027] Figure 13 is a third schematic diagram of the image processing chip provided in an embodiment of this application;

[0028] Figure 14 is a fourth schematic diagram of the image processing chip provided in an embodiment of this application;

[0029] Figure 15 is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0030] Figure 16 is a second schematic diagram of the structure of the electronic device provided in the embodiment of this application;

[0031] Figure 17 is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0033] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0034] The terms "at least one," "at least one," etc., in this application refer to any one, any two, or a combination of two or more of the included objects. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more, and its meaning is similar to that of "at least one."

[0035] The image processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0036] This application's embodiments can be applied to scenarios where users need to simultaneously record video using multiple cameras on an electronic device and need to fuse the video images recorded by the multiple cameras. For example, a user needs to simultaneously record video using the camera on the inner screen and the camera on the outer screen of a dual-folding screen phone, and needs to fuse the face image recorded by the camera on the inner screen and the background image recorded by the camera on the outer screen.

[0037] Currently, when a dual-folding phone is folded, it has one camera on the outer screen and a set of rear cameras on the other side of the cover. When the phone is unfolded, it also has a front-facing camera on the inner screen. When a user unfolds the dual-folding phone and uses its camera app to record video, the phone will trigger the corresponding camera to record separately based on the user's selected recording mode, resulting in a separate video recording. For example, when a user fully unfolds the phone and takes a selfie, the dual-folding phone will trigger the front-facing camera on the inner screen to record separately, resulting in a separate video recording of the face.

[0038] However, in the above-mentioned methods, the multiple cameras of a dual-folding screen phone cannot work together to merge the recorded scenes. For example, it is impossible to merge and display the facial image recorded by the camera on the inner screen and the background image recorded by the camera on the outer screen. As a result, the image quality captured by the camera on the electronic device is relatively simple, and the shooting method is not flexible enough.

[0039] This application provides an image processing method. When a video recording interface is displayed, a first input is received; in response to the first input, video recording of a first scene is performed using a first camera, and video recording of a second scene is performed using a second camera; image fusion is performed based on the first scene and the second scene, and the fused image is displayed. In this solution, since the electronic device can perform image fusion based on the first scene and the second scene, and display the fused image, where the first scene is the scene recorded by the first camera and the second scene is the scene recorded by the second camera, the electronic device can perform fusion processing on the image data of the first scene recorded by the first camera and the image data of the second scene recorded by the second camera, and display the image corresponding to the fused image data obtained from the fusion processing. Thus, in this way, when shooting with an electronic device containing multiple cameras, the electronic device can respond to the first input, simultaneously operate multiple cameras to shoot, and perform fusion processing on the video images captured by the multiple cameras to display the fused image, making the image captured by the electronic device a fused image of multiple scenes, improving the flexibility of the electronic device in shooting.

[0040] The image processing method provided in this application can be executed by an image processing device, which can be an electronic device, or a functional module or functional entity within an electronic device. The following description uses an electronic device as an example to illustrate the technical solution provided in this application.

[0041] Figure 1 shows a flowchart of an image processing method provided in an embodiment of this application. As shown in Figure 1, the image processing method provided in this embodiment may include the following steps 201 to 203.

[0042] Step 201: The electronic device receives the first input while displaying the video recording interface.

[0043] In this embodiment of the application, the first input is used to trigger the electronic device to record video of the first scene through the first camera, and to record video of the second scene through the second camera, and to perform image fusion based on the first scene and the second scene, and to display the fused image.

[0044] Optionally, in this embodiment, the first input may include, but is not limited to, any of the following: touch input by the user through a touch device such as a finger or stylus, a voice command input by the user, a specific gesture input by the user, or other feasible input. The specific input can be determined according to actual usage needs, and this embodiment does not impose any limitations.

[0045] Optionally, in the embodiments of this application, the above-mentioned touch input can be single-click input, double-click input, drag input, or any number of clicks, or it can be long-press input or short-press input.

[0046] Optionally, in the embodiments of this application, the aforementioned specific gesture can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, or a double-tap gesture.

[0047] Optionally, in this embodiment of the application, the video recording interface may include an image fusion control, which is used to indicate the recording mode for selecting the image fusion mode.

[0048] Optionally, in this embodiment of the application, the video recording interface may include an image stitching control, which is used to indicate the recording mode for selecting the image stitching mode.

[0049] Optionally, in this embodiment of the application, the first input can specifically be the user's touch input to the image fusion control.

[0050] Step 202: The electronic device responds to the first input by recording video of the first scene through the first camera and recording video of the second scene through the second camera.

[0051] In this embodiment of the application, the first scene refers to the environment or view in which video is recorded by the first camera. The first scene may include a first set of objects, people, activities or environmental features that the user needs to record.

[0052] In this embodiment of the application, the second scene refers to the environment or view in which video is recorded by the second camera. The second scene may include a second set of objects, people, activities or environmental features that the user needs to record.

[0053] Optionally, in this embodiment of the application, the above-mentioned electronic device is an electronic device with a foldable screen, and the above-mentioned step 202 can be implemented by the following step 202a.

[0054] Step 202a: In response to the first input, the electronic device records video of the first scene through the first camera and the second scene through the second camera when the folding angle of the first and second screens of the foldable screen is greater than or equal to the first threshold.

[0055] In this embodiment of the application, the electronic device can detect the folding angle using an accelerometer and a gyroscope.

[0056] Optionally, in this embodiment of the application, the first threshold can be 90° or 60°, and the specific value is determined based on the requirements.

[0057] For example, as shown in Figure 2, a dual-folding screen phone includes a folding outer screen 141 and a folding inner screen 151. The folding outer screen 141 includes a first camera 111, and the folding inner screen 151 includes a second camera 121. As shown in Figure 3, the folding angle of the folding screen is θ in the figure.

[0058] For example, referring to Figure 2, when the folding angle of the outer folding screen 141 and the inner folding screen 151 is greater than or equal to 90°, the dual-folding screen phone records video of a first scene through the first camera 111, such as recording video of the surrounding environment, and records video of a second scene through the second camera 121, such as recording video of the user's selfie.

[0059] Step 203: The electronic device performs image fusion based on the first scene and the second scene, and displays the fused image.

[0060] Understandably, image fusion is a crucial step in image processing. It combines multiple images from different sensors, perspectives, or time points to generate a single, richer, and higher-quality image. The image fusion process extracts valuable information from each image and preserves it to the maximum extent possible, ultimately producing an image that provides a more comprehensive and clear description of the scene.

[0061] In this embodiment of the application, the electronic device can extract some or all of the user's required information from the image corresponding to the first scene and the image corresponding to the second scene, and perform image fusion based on the extracted information to display the fused image.

[0062] For example, if the first camera 111 records a video of the surrounding environment and the second camera 121 records a video of the user's selfie angle, then the recorded video displayed in the video recording interface can be a video obtained by image fusion of the video recorded by the first camera 111 and the video recorded by the second camera 121. For example, the fused video is obtained by fusing the background image video after removing the human figure from the surrounding environment video recorded by the first camera 111 with the human figure image video after removing the background from the user's selfie angle video recorded by the second camera 121.

[0063] This application provides an image processing method. Since the electronic device can perform image fusion based on a first scene and a second scene, and display the fused image, where the first scene is a scene recorded by a first camera and the second scene is a scene recorded by a second camera, the electronic device can perform fusion processing on the image data of the first scene recorded by the first camera and the image data of the second scene recorded by the second camera, and display the image corresponding to the fused image data obtained from the fusion processing. In this way, when shooting with an electronic device containing multiple cameras, the electronic device can respond to a first input, simultaneously operate multiple cameras to shoot, and perform fusion processing on the video images captured by the multiple cameras to display the fused image. This makes the image captured by the electronic device a fused image from multiple scenes, improving the flexibility of the electronic device in shooting.

[0064] Optionally, in this embodiment of the application, the above-mentioned electronic device includes an image processing chip. Referring to FIG1 and FIG4, the "electronic device performs image fusion based on the first scene and the second scene" in step 203 can be specifically implemented by the following steps 301 to 303. The "displaying the fused image" in step 203 is illustrated in FIG4 as step 203b.

[0065] It should be noted that for a detailed description of the above image processing chip, please refer to the description of the following embodiments, which will not be repeated here.

[0066] Step 301: The electronic device performs frame interpolation processing on the image frame data of the first scene captured by the first camera through the first frame interpolation module in the image processing chip to obtain the first image data.

[0067] Optionally, in the embodiments of this application, before step 301 above, the image processing method provided in the embodiments of this application may further include the following steps 304 and 305, and the above step 301 may be specifically implemented by the following step 301a.

[0068] Step 304: The electronic device extracts features from the image frame data of the first scene and the historical image frame data captured by the first camera through the frame interpolation network model in the first frame interpolation module to obtain the first feature map.

[0069] In this embodiment of the application, the first feature map represents the motion information of the image frame data of the first scene captured by the first camera relative to the historical image frame data.

[0070] In this embodiment, the frame interpolation network model described above can be an implementation of Artificial Intelligence Motion Estimation and Motion Compensation (AIMEMC) technology. By utilizing deep learning technology, the frame interpolation network model can learn the motion patterns and changes between image frames in a video. The electronic device can simultaneously input historical image frame data and the current image frame data of the first scene into the frame interpolation network model. This model extracts features from the historical and current image frame data through convolution operations, calculates the motion relationship between them, and inserts intermediate image frame data that conforms to the motion relationship between the two frames, thereby increasing the video's frame rate. In this way, the smoothness of low frame rate videos can be significantly improved, making the video appear smoother and more coherent.

[0071] As we can understand, AIMEMC is a video processing technology that uses artificial intelligence for motion estimation and motion compensation. It predicts the trajectory of objects by analyzing changes in the frame between consecutive frames and inserts one or more new "compensation frames" between two frames to fill the gaps in the frame rate. Through intelligent analysis and processing, AIMEMC technology can accurately predict and compensate for the trajectory of objects, making the video look more natural and smooth.

[0072] Step 305: The electronic device performs feature processing on the first feature map through the residual network model in the first frame interpolation module to obtain intermediate image frame data.

[0073] Understandably, a residual network is a deep convolutional neural network architecture that addresses the vanishing and exploding gradient problems in deep neural network training by introducing residual connections (or skip connections). These residual connections allow the network to learn the residual between the input and output, rather than directly learning the entire mapping. In frame interpolation, the residual network model is used to further process the extracted first feature map. By analyzing the information in the first feature map, it learns how to generate intermediate image frames that conform to the motion patterns.

[0074] In this embodiment, the residual network model in the first frame interpolation module performs a series of convolutional operations and residual connection processing on the input first feature map. These operations aim to extract deeper features and learn how to generate intermediate image frame data based on these features. After processing by the residual network model, the electronic device outputs one or more intermediate image frame data. These intermediate frames are used to interpolate between historical image frames and the current image frame. They conform to motion laws and maintain visual coherence and consistency with the original frames.

[0075] Step 301a: The electronic device uses the first frame interpolation module to perform frame interpolation processing on the image frame data of the first scene captured by the first camera, based on the intermediate image frame data, to obtain the first image data.

[0076] In this embodiment, after obtaining the intermediate image frame data, the intermediate image frame data is inserted between the current image frame data and the historical image frame data of the first scene captured by the first camera. By inserting these intermediate image frame data, the frame rate of the original video stream is effectively improved. This makes the video appear smoother during playback, reducing motion blur and stuttering.

[0077] It should be noted that after the first frame interpolation module performs frame interpolation processing on the current image frame data of the first scene captured by the first camera, it stores the image frame data captured by the first camera, i.e., the current image frame data, into a Double Data Rate (DDR) storage unit as historical image frame data for the next frame interpolation processing. Therefore, when the first frame interpolation module performs frame interpolation processing on the image frame data of the first scene captured by the first camera, it reads historical image frame data from the DDR storage unit, reducing the latency of reading and writing to the memory.

[0078] Step 302: The electronic device performs frame interpolation processing on the image frame data of the second scene captured by the second camera through the second frame interpolation module in the image processing chip to obtain the second image data.

[0079] Optionally, in the embodiments of this application, before step 302 above, the image processing method provided in the embodiments of this application may further include the following steps 306 and 307, and the above step 302 may be specifically implemented by the following step 302a.

[0080] Step 306: The electronic device extracts features from the image frame data of the second scene and the historical image frame data captured by the second camera through the frame interpolation network model in the second frame interpolation module to obtain the second feature map.

[0081] In this embodiment of the application, the second feature map represents the motion information of the image frame data of the second scene captured by the second camera relative to the historical image frame data.

[0082] Step 307: The electronic device performs feature processing on the second feature map through the residual network model in the second frame interpolation module to obtain intermediate image frame data.

[0083] Step 302a: The electronic device uses the second frame interpolation module to perform frame interpolation processing on the image frame data of the second scene captured by the second camera, based on the intermediate image frame data, to obtain the second image data.

[0084] It should be noted that the specific implementation methods of steps 306, 307 and 302a can be found in the descriptions of steps 304, 305 and 301a above, and will not be repeated here.

[0085] Step 303: The electronic device performs fusion processing on the first image data and the second image data through the fusion module in the image processing chip to obtain fused image data.

[0086] Optionally, in this embodiment of the application, referring to FIG4 and FIG5, the above step 303 can be specifically implemented by the following steps 401 to 403.

[0087] Step 401: The electronic device performs convolution processing on the first image data through the fusion network model in the fusion module to obtain the third feature map, and performs convolution processing on the second image data to obtain the fourth feature map.

[0088] Optionally, in this embodiment of the application, the step 401 above, "the electronic device performs convolution processing on the first image data through the fusion network model in the fusion module to obtain the third feature map", can be specifically implemented through the following steps 501 to 505.

[0089] Step 501: The electronic device downsamples the first image data through the first convolutional module in the fusion network model and outputs the first feature vector.

[0090] Step 502: The electronic device downsamples the (i-1)th feature vector through the i-th convolutional module in the fusion network model and outputs the i-th feature vector.

[0091] In this embodiment of the application, i∈[2,N], and i is an integer.

[0092] In this embodiment, N represents the number of convolutional modules in the aforementioned fusion network model.

[0093] Step 503: The electronic device upsamples the Nth feature vector and the (N-1)th feature vector through the first deconvolution module in the fusion network model to obtain the first content feature vector.

[0094] In this embodiment of the application, the above-mentioned content feature vector includes a foreground feature vector and a background feature vector.

[0095] As we can understand, the foreground typically refers to the part of an image or scene that is closest to the observer, most prominent, and most attention-grabbing, such as a flower, a stone, or a person. They stand out because of their proximity, vivid color, and strong size contrast. The foreground feature vector mentioned above refers to the feature vector corresponding to the image data of the foreground portion. The background typically refers to the part of an image or scene that is behind the foreground, farther from the observer, and usually less prominent. The background provides the environment or context for the foreground, helping the observer understand the nature, time, and location of the scene. The background feature vector mentioned above refers to the feature vector corresponding to the image data of the background portion.

[0096] Step 504: The electronic device upsamples the Nj-th feature vector and the (j-1)-th content feature vector through the j-th deconvolution module in the fusion network model, and outputs the j-th content feature vector.

[0097] In this embodiment of the application, j∈[2,N-1], and j is an integer.

[0098] Step 505: The electronic device upsamples the first image data and the (N-1)th content feature vector through the Nth deconvolution module in the fusion network model, outputs the Nth content feature vector, and determines the Nth content feature vector as the third feature map.

[0099] In this embodiment, the electronic device can use N convolutional modules to perform N downsampling processes on the feature map corresponding to the first image data to extract the semantic information of the image corresponding to the first image data, i.e., the aforementioned feature vector. Then, the electronic device can use N deconvolutional modules to perform N upsampling processes on the Nth feature vector to obtain image segmentation information, i.e., content feature vector. This content feature vector includes a foreground feature vector and a background feature vector, which are used to subsequently distinguish foreground image data and background image data in the image corresponding to the first image data. In the last upsampling process, the electronic device obtains the Nth content feature vector based on the first image data and the (N-1)th content feature vector, which can distinguish foreground image data and background image data in the image corresponding to the first image data, and determines the Nth content feature vector as the third feature map.

[0100] In this embodiment of the application, the semantic information of the above-mentioned image may include, but is not limited to, at least one of the following: basic information such as the outline, category, and positional relationship of the image elements contained in the image.

[0101] In this embodiment, during the downsampling process of the feature map corresponding to the input first image data by the convolution module, the electronic device can increase the number of convolution kernels and the number of output channels. Increasing the number of convolution kernels can extract more types of features, and increasing the number of output channels can increase the number of feature maps. This provides richer and more diverse feature inputs for subsequent layers, helping them to better process and understand the input data.

[0102] In this embodiment, during the upsampling process of the Nth feature vector by the deconvolution module, the electronic device can reduce the number of convolution kernels and the number of output channels. By reducing the number of model parameters, the complexity of the model and the risk of overfitting are reduced, and the model can focus more on learning the features most useful to the task, thereby improving the model's performance.

[0103] Optionally, in this embodiment of the application, the step 401 above, "the electronic device performs convolution processing on the second image data through the fusion network model in the fusion module to obtain the fourth feature map", can be specifically implemented through the following steps 506 to 510.

[0104] Step 506: The electronic device downsamples the first image data through the first convolutional module in the fusion network model and outputs the first feature vector.

[0105] Step 507: The electronic device performs downsampling processing on the (i-1)th feature vector through the i-th convolutional module in the fusion network model and outputs the i-th feature vector.

[0106] In this embodiment of the application, i∈[2,N], and i is an integer.

[0107] In this embodiment, N represents the number of convolutional modules in the aforementioned fusion network model.

[0108] Step 508: The electronic device upsamples the Nth feature vector and the (N-1)th feature vector through the first deconvolution module in the fusion network model to obtain the first content feature vector.

[0109] In this embodiment of the application, the above-mentioned content feature vector includes a foreground feature vector and a background feature vector.

[0110] Step 509: The electronic device upsamples the Nj-th feature vector and the (j-1)-th content feature vector through the j-th deconvolution module in the fusion network model, and outputs the j-th content feature vector.

[0111] In this embodiment of the application, j∈[2,N-1], and j is an integer.

[0112] Step 510: The electronic device upsamples the first image data and the (N-1)th content feature vector through the Nth deconvolution module in the fusion network model, outputs the Nth content feature vector, and determines the Nth content feature vector as the fourth feature map.

[0113] It should be noted that the purpose of "the electronic device performs convolution processing on the second image data through the fusion network model in the fusion module to obtain the fourth feature map" and "the electronic device performs convolution processing on the first image data through the fusion network model in the fusion module to obtain the third feature map" are both to identify the foreground image data and background image data in the image corresponding to the image data, so that one of the foreground image data and background image data can be extracted and fused in the subsequent process. Therefore, the processing of steps 506 to 510 is similar to the processing of steps 501 to 505.

[0114] For example, taking a dual-folding screen mobile phone as an example, assume that the first camera is the camera on the outer folding screen of the phone, and the second camera is the camera on the inner folding screen of the phone. The image frame data captured by the first camera includes images of the surrounding environment, and the image frame data captured by the second camera includes selfie images. After frame interpolation processing is performed on the image frame data captured by the first camera and the image frame data captured by the second camera through the first frame interpolation module and the second frame interpolation module respectively, the first image data and the second image data are obtained. The first image data and the second image data are then input into the fusion network model. As shown in Figure 6, the fusion network model includes two branches. The portrait segmentation steps in branches one and two are the same. Branch one processes the second image data, i.e., the image captured by the camera on the foldable inner screen. It continuously performs convolution and pooling operations on the second image data to obtain feature maps of different scales. For example, there are three convolution and deconvolution modules. The second image data corresponds to a feature map with a scale and number of channels of 512×512×3. The first convolution module in branch one of the fusion network model performs convolution and pooling operations on the feature map corresponding to the second image data, outputting a first feature vector with a scale and number of channels of 224×224×128. The second convolution module in branch one performs convolution and pooling operations on the first feature vector, outputting a second feature vector with a scale and number of channels of 168×168×212. Then... The third convolution module in branch one performs convolution and pooling operations on the second feature vector, outputting a third feature vector with a scale and number of channels of 64×64×256. Then, the first deconvolution module concatenates and deconvolves the third and second feature vectors to obtain a first content feature vector with a scale and number of channels of 168×168×212. The second deconvolution module concatenates and deconvolves the first feature vector and the first content feature vector, outputting a second content feature vector with a scale and number of channels of 224×224×128. Finally, the third deconvolution module concatenates and deconvolves the second image data and the second content feature vector, outputting a third content feature vector with a scale and number of channels of 512×512×4. This third content feature vector is then used as the fourth feature map. This process enables image detection and portrait segmentation, extracting pixel-level information corresponding to the selfie taker.Branch 2 is used to process the first image data, that is, the image captured by the camera on the foldable outer screen. The first image data is continuously subjected to convolution and pooling operations to obtain feature maps of different scales. The portrait segmentation step in Branch 2 is the same as the portrait segmentation step in Branch 1. After performing multiple convolution and pooling operations, stitching and deconvolution operations on the first image data, the third feature map is obtained. In this way, the functions of image detection and portrait segmentation are realized, thereby extracting the pixel-level information corresponding to the human image contained in the image captured by the camera on the foldable outer screen.

[0115] Step 402: The electronic device performs feature processing on the third feature map through the fusion network model to obtain background image data, and performs feature processing on the fourth feature map to obtain foreground image data.

[0116] Optionally, in this embodiment of the application, the step 402 above, "the electronic device performs feature processing on the third feature map through a fusion network model to obtain background image data", can be specifically implemented through the following step 601.

[0117] Step 601: The electronic device performs foreground removal and background generation processing on the third feature map through the convolution module in the fusion network model to obtain background image data.

[0118] For example, as shown in Figure 6, after obtaining the third feature map through the portrait segmentation step in branch two, multiple convolution and pooling operations are performed on the third feature map to obtain feature maps of different scales, such as feature maps with scales and number of channels of 256×256×8, 198×198×128, 64×64×256 and 512×512×3, respectively. Thus, according to the parameter configuration, the human image contained in the image captured by the camera on the folding outer screen can be erased, and the pure background image can be retained.

[0119] Optionally, in this embodiment of the application, the step 402 above, "the electronic device performs feature processing on the fourth feature map through the fusion network model to obtain foreground image data", can be specifically implemented through the following steps 602 and 603.

[0120] Step 602: The electronic device performs background removal processing on the fourth feature map through the convolution module in the fusion network model to obtain the initial foreground image data.

[0121] For example, as shown in Figure 6, after obtaining the fourth feature map through the portrait segmentation step, branch one performs multiple convolution and pooling operations on the fourth feature map to obtain feature maps of different scales, such as feature maps with scales and number of channels of 256×256×8, 168×168×64, 128×128×128 and 512×512×4, respectively. This can eliminate the background image other than the selfie image contained in the image captured by the camera on the folding inner screen, and obtain the initial portrait image data.

[0122] Step 603: The electronic device performs edge feathering processing on the initial foreground image data through the edge attention module in the fusion network model to obtain the foreground image data.

[0123] In this embodiment, before the final image fusion process, the electronic device can perform edge feathering on the initial foreground image data through the edge attention module in the fusion network model. Edge feathering softens or blurs the edges of the image corresponding to the foreground image data, thus better blending it with the background image. Edge feathering can reduce the harsh edges generated when extracting the image corresponding to the initial foreground image data, making the final fusion effect smoother and more natural. For example, when moving a person from one image to another, feathering allows the edges of the person to blend better with the new background, reducing abruptness.

[0124] For example, as shown in Figure 6, after obtaining the initial character image data through the background removal step, the image corresponding to the initial character image data is feathered through multiple convolution and pooling operations to obtain feature maps of different scales. For example, convolution and pooling operations are performed on the fourth feature map obtained after the character segmentation step and the initial character image data obtained after the background removal step to obtain a feature map with a scale and number of channels of 256×256×24. Then, convolution and pooling operations are performed on the feature map to obtain a feature map with a scale and number of channels of 512×512×3, thereby obtaining character image data with a gradual edge effect without abruptness.

[0125] Optionally, in this embodiment of the application, the electronic device can perform feature processing on the third feature map to obtain foreground image data and perform feature processing on the fourth feature map to obtain background image data through a fusion network model.

[0126] Optionally, in this embodiment of the application, the electronic device can perform feature processing on the third feature map to obtain foreground image data through a fusion network model, and perform feature processing on the fourth feature map to obtain foreground image data.

[0127] It should be noted that after the electronic device obtains the foreground image data corresponding to the third feature map and the foreground image data corresponding to the fourth feature map, assuming that both foreground image data are human image data, when the electronic device displays the corresponding human image, the user can trigger the electronic device to adjust the distance between the human figures and the human figures' poses, and can also customize the background image for the human figures.

[0128] Optionally, in this embodiment of the application, the electronic device can perform feature processing on the third feature map to obtain background image data through a fusion network model, and perform feature processing on the fourth feature map to obtain background image data.

[0129] It should be noted that the specific implementation of feature processing on the third or fourth feature map to obtain foreground image data or background image data can be found in the description of the above embodiments, and will not be repeated here.

[0130] Step 403: The electronic device performs fusion processing on the background image data and the foreground image data to obtain fused image data.

[0131] In this embodiment of the application, after obtaining the background image data and the foreground image data, the electronic device can perform fusion processing on the background image data and the foreground image data through the image fusion module in the fusion network model to obtain fused image data.

[0132] For example, as shown in Figure 6, after the two branches obtain a clean background image and a person image respectively, the background image and the person image are input into the image fusion module. Through multiple convolution and pooling operations of different scales, feature maps of different scales are obtained, such as feature maps with scales and number of channels of 512×512×6, 256×256×16, 512×512×3 and 512×512×3 respectively, thereby realizing the fusion processing of the background image and the person image, for example, the person image is integrated into a certain area of ​​the background image to form a complete image.

[0133] This application provides an image processing method. The electronic device can fuse first image data and second image data using a fusion module in its image processing chip. The first image data is obtained by interpolating image frames from a first scene captured by a first camera, and the second image data is obtained by interpolating image frames from a second scene captured by a second camera. In other words, the electronic device can fuse the image frames from the first scene captured by the first camera and the image frames from the second scene captured by the second camera to obtain fused image data. Therefore, when shooting with an electronic device containing multiple cameras, the device can simultaneously operate multiple cameras to capture images and fuse the video images obtained from the multiple cameras, resulting in a fused image from multiple scenes, thus improving the flexibility of the electronic device in shooting.

[0134] It should be noted that, as shown in Figure 7, the fusion module 23 is internally configured with two multiply-accumulate operation arrays, namely the first multiply-accumulate operation array 232 and the second multiply-accumulate operation array 233, to process the two branches of the fusion network model in the fusion module 23 in parallel. The scheduler 231 will start two tasks to control the processing flow of the two branches respectively. The front-end portrait segmentation networks of the two branches are similar, but the parameters are different, and the processing speed is also comparable. In the post-elimination part, the tasks are different, and in addition to the parameters, the network structure is also different, which will lead to different processing speeds. In particular, branch one has additional edge feathering effect processing at the end, while branch two has additional background generation function. Therefore, the processing results of the two branches need to be temporarily stored in the shared cache space 235 to wait for the two branches to finish processing simultaneously before fusion processing. During the parallel processing of the two branches, some networks will call the scalar operation unit 234. For example, the scalar operation unit 234 will be called when performing edge feathering processing. At this time, the scheduler 231 will allocate the scalar operation resources of the scalar operation unit 234 according to the order of calling. If there is a conflict of operation resources, the processing needs to be paused to wait for the resources to be released. After both branches output valid data, scheduler 231 will call on idle multiply-accumulate operation resources to process the final fusion network. The use of scalar operation resources in the fusion network is shared with the branch operations.

[0135] Optionally, in this embodiment of the application, referring to FIG4 and FIG8, the image processing method provided in this embodiment of the application may further include the following step 701, and the above step 301 may be implemented by the following step 301b.

[0136] Step 701: The electronic device performs noise reduction processing on the image frame data of the first scene captured by the first camera through the first noise reduction module in the image processing chip to obtain the third image data.

[0137] In this embodiment, the electronic device performs noise reduction processing on the image frame data of the first scene captured by the first camera through the noise reduction network model in the first noise reduction module to obtain the third image data.

[0138] Optionally, in this embodiment of the application, the above-mentioned noise reduction network model can be a deep learning model based on neural networks. The original image is the image frame data of the first scene collected by the sensor of the first camera. The noise reduction network model achieves a pyramid-like noise reduction function by designing convolution kernels of different sizes and pooling layers, which can retain sufficient image details while removing high-frequency noise.

[0139] Optionally, in this embodiment of the application, the image frame data of the first scene collected by the sensor of the first camera can be processed by the noise reduction network model in the first noise reduction module in the form of a row data stream to achieve the noise reduction function of the image.

[0140] For example, as shown in FIG9, the first camera sensor 11 sends the raw uncompressed image file RAW data to the first noise reduction module 24 of the image processing chip 20 for noise reduction processing through the first interface and the third interface.

[0141] Step 301b: The electronic device performs frame interpolation processing on the third image data through the first frame interpolation module in the image processing chip to obtain the first image data.

[0142] In this embodiment of the application, the electronic device performs noise reduction processing on the image frame data of the first scene captured by the first camera through the first noise reduction module in the image processing chip to obtain the third image data. After obtaining the third image data, the electronic device inputs the third image data into the first frame interpolation module, and then performs frame interpolation processing on the third image data through the first frame interpolation module.

[0143] It should be noted that the specific implementation of frame interpolation processing can be found in the description of the above embodiments, and will not be repeated here.

[0144] Optionally, in this embodiment of the application, referring to FIG4 and FIG10, the image processing method provided in this embodiment of the application may further include the following step 702, and the above step 302 may be implemented by the following step 302b.

[0145] Step 702: The electronic device performs noise reduction processing on the image frame data of the second scene captured by the second camera through the second noise reduction module in the image processing chip to obtain the fourth image data.

[0146] In this embodiment, the electronic device performs noise reduction processing on the image frame data of the second scene captured by the second camera through the noise reduction network model in the second noise reduction module to obtain the fourth image data.

[0147] It should be noted that the noise reduction network model in the second noise reduction module is similar to that in the first noise reduction module. However, due to the different sensor selections of the first and second cameras, the network parameters during noise reduction processing will differ. For specific implementation methods, please refer to relevant technologies, which will not be elaborated here.

[0148] Optionally, in this embodiment of the application, the image frame data of the second scene collected by the sensor of the second camera can be processed by the noise reduction network model in the second noise reduction module in the form of a row data stream to achieve the noise reduction function of the image.

[0149] For example, as shown in FIG9, the second camera sensor 12 sends the original uncompressed image file RAW data to the second noise reduction module 25 of the image processing chip 20 for noise reduction processing through the second interface and the fourth interface.

[0150] Step 302b: The electronic device performs frame interpolation processing on the fourth image data through the second frame interpolation module in the image processing chip to obtain the second image data.

[0151] In this embodiment, the electronic device performs noise reduction processing on the image frame data of the second scene captured by the second camera through the second noise reduction module in the image processing chip to obtain the fourth image data. Then, the fourth image data is input to the second frame interpolation module, and the electronic device performs frame interpolation processing on the fourth image data through the second frame interpolation module.

[0152] It should be noted that the specific implementation of frame interpolation processing can be found in the description of the above embodiments, and will not be repeated here.

[0153] In this embodiment, the electronic device performs noise reduction processing on the image frame data of the first scene captured by the first camera and the image frame data of the second scene captured by the second camera through the first noise reduction module and the second noise reduction module of the image processing chip, respectively. This avoids noise signals introduced by the imperfections of the device that collects the image frame data and removes noise signals mixed in with the conventional image information.

[0154] Optionally, in this embodiment of the application, when the electronic device displays a video recording interface and the current recording mode is image fusion mode, if the folding angle of the first screen and the second screen of the folding screen is detected to be greater than or equal to a first threshold, the first image data and the second image data are fused by the fusion module in the image processing chip to obtain fused image data.

[0155] For example, the dual-folding screen phone currently displays the video recording interface of the camera application, and the user selects the image fusion mode for recording. The user fully unfolds the phone, that is, the folding angle of the screen is 180°, which is greater than 90°. The first camera 111 records the video of the surrounding environment, and the second camera 121 records the video of the user's selfie angle. Then the recorded video displayed in the video recording interface can be a video obtained by fusing the video recorded by the first camera 111 and the video recorded by the second camera 121. For example, the background image video after removing the human figure from the surrounding environment video video recorded by the first camera 111 is fused with the human figure image video after removing the background from the user's selfie angle video recorded by the second camera 121 to obtain the fused video.

[0156] It should be noted that, as shown in Figure 9, if the current recording mode of the electronic device is image fusion mode, the electronic device performs frame interpolation processing on the image frame data of the first scene collected by the first camera sensor 11 through the first frame interpolation module 21, and then inputs the obtained first image data into the fusion module 23. The electronic device performs frame interpolation processing on the image frame data of the second scene collected by the second camera sensor 12 through the second frame interpolation module 22, and then inputs the obtained second image data into the fusion module 23. The electronic device performs fusion processing on the first image data and the second image data through the fusion module 23 to obtain fused image data. The electronic device then sends the fused image data to the application processor 10 through the fifth interface and the sixth interface for rendering and displaying the image corresponding to the fused image data.

[0157] As is understandable, an application processor is an integrated circuit chip that is a core component in smart devices (such as smartphones and tablets) and is responsible for processing and managing various software applications, graphics rendering, multimedia playback, and data transmission.

[0158] In this embodiment, when the electronic device is displaying a video recording interface, if it detects that the folding angle of the first and second screens of the foldable screen is greater than or equal to a first threshold, the electronic device can perform fusion processing on the first image data and the second image data through the fusion module in the image processing chip. The first image data is obtained by interpolating the image frame data of the first scene captured by the first camera, and the second image data is obtained by interpolating the image frame data of the second scene captured by the second camera. That is, the electronic device can perform fusion processing on the image frame data of the first scene captured by the first camera and the image frame data of the second scene captured by the second camera through the fusion module in the image processing chip to obtain fused image data, and display the image corresponding to the fused image data for the user to view. In this way, when shooting with an electronic device containing multiple cameras, the electronic device can simultaneously run multiple cameras to shoot and perform fusion processing on the video images captured by multiple cameras, so that the image captured by the electronic device is a fused image of multiple scenes, which improves the flexibility of the electronic device in shooting, realizes the multi-view camera recording function of the foldable screen phone, further improves the user experience and the playability of the camera of the foldable screen phone, and allows users to have more ways to play with camera recording and better create images.

[0159] Optionally, in this embodiment of the application, when the electronic device displays a video recording interface and the current recording mode is image stitching mode, if it detects that the folding angle of the first screen and the second screen of the folding screen is greater than or equal to a first threshold, the electronic device directly outputs the first image data and the second image data through the image processing chip, without needing to input the first image data and the second image data into the fusion module for fusion processing.

[0160] For example, the dual-folding screen phone currently displays the video recording interface of the camera application, and the user selects the image stitching mode for recording. The user fully unfolds the phone, that is, the folding angle of the screen is 180°, which is greater than 90°. The first camera 111 records the video of the surrounding environment, and the second camera 121 records the video of the user's selfie angle. Then the recorded video displayed in the video recording interface can be a video that stitches the video recorded by the first camera 111 and the video recorded by the second camera 121. For example, it can be a video that simply stitches the video of the surrounding environment recorded by the first camera 111 with the video of the user's selfie angle recorded by the second camera 121. The stitching method can be top-to-bottom stitching, left-to-right stitching, etc.

[0161] It should be noted that, as shown in Figure 9, if the current recording mode of the electronic device is image stitching mode, the electronic device performs frame interpolation processing on the image frame data of the first scene collected by the first camera sensor 11 through the first frame interpolation module 21, and then directly inputs the obtained first image data to the application processor 10. The electronic device performs frame interpolation processing on the image frame data of the second scene collected by the second camera sensor 12 through the second frame interpolation module 22, and then directly inputs the obtained second image data to the application processor 10 for rendering and displaying the images corresponding to their respective image data.

[0162] This application provides an image processing chip. Figure 11 shows a schematic diagram of the structure of an image processing chip provided in this application. As shown in Figure 11, the image processing chip 20 provided in this application may include: a first frame interpolation module 21, a second frame interpolation module 22, and a fusion module 23.

[0163] In this embodiment of the application, the first frame interpolation module 21 and the second frame interpolation module 22 are both connected to the fusion module 23.

[0164] In this embodiment of the application, the first frame interpolation module 21 is used to perform frame interpolation processing on the image frame data of the first scene captured by the first camera to obtain the first image data.

[0165] In this embodiment of the application, the second frame interpolation module 22 is used to perform frame interpolation processing on the image frame data of the second scene captured by the second camera to obtain the second image data.

[0166] In this embodiment of the application, the fusion module 23 is used to perform fusion processing on the first image data and the second image data to obtain fused image data.

[0167] It should be noted that, in this embodiment of the application, the fusion module 23 may have a built-in small-capacity cache unit to store the image data that arrives at the fusion module 23 first from the first image data and the second image data. After the other image data arrives at the fusion module 23, the fusion processing flow is started. In this way, the impact of the image data output of the first frame interpolation module 21 and the second frame interpolation module 22 not being strictly synchronized on the image fusion is reduced.

[0168] In this embodiment of the application, both the first frame interpolation module 21 and the second frame interpolation module 22 use a frame interpolation network model to perform frame interpolation processing on the image frame data.

[0169] In this embodiment of the application, the fusion module 23 uses a fusion network model to perform fusion processing on the first image data and the second image data.

[0170] It should be noted that for a detailed explanation of the above frame interpolation network model and fusion network model, please refer to the description in the above embodiments, which will not be repeated here.

[0171] It should be noted that the aforementioned fusion module can be applied to a Neural Processing Unit (NPU). As can be understood, an NPU is a specially designed processor intended to accelerate workloads in artificial intelligence and machine learning. The NPU simulates the workings of neurons in the human brain, significantly accelerating complex computational processes in neural networks, including large-scale matrix operations and convolution operations, through efficient hardware architecture and optimized algorithms.

[0172] This application provides an image processing chip. Since the fusion module can fuse first image data and second image data, where the first image data is obtained by interpolating image frames from a first scene captured by a first camera, and the second image data is obtained by interpolating image frames from a second scene captured by a second camera, the fusion module can fuse the image frames from the first scene captured by the first camera and the image frames from the second scene captured by the second camera to obtain fused image data. In this way, when shooting with an electronic device containing multiple cameras, the electronic device can simultaneously operate multiple cameras to shoot and fuse the video images obtained from the multiple cameras, resulting in a fused image from multiple scenes, thus improving the flexibility of the electronic device in shooting.

[0173] For example, Figure 12 shows a schematic diagram of the hardware structure of a dual-folding screen mobile phone provided in an embodiment of this application. As shown in Figure 12, the dual-folding screen mobile phone may include: an application processor 10, an image processing chip 20, a first camera sensor 11, a second camera sensor 12, a folding outer screen 13, a folding inner screen 14, a power module 15, a storage module 16, sensors such as accelerometer / gyroscope / compass 17, buttons 18, and an audio module 19. The first camera sensor 11 is the sensor corresponding to the camera on the folding outer screen 13, and the second camera sensor 12 is the sensor corresponding to the camera on the folding inner screen 14. Those skilled in the art will understand that the hardware structure of the dual-folding screen mobile phone shown in Figure 12 does not constitute a limitation on the dual-folding screen mobile phone, and may include more or fewer components than shown, or combine some of the modules below, or have different component arrangements.

[0174] In the aforementioned dual-folding screen phone, the image processing chip 20 provided in this application embodiment includes a first frame interpolation module, a second frame interpolation module, and a fusion module. The first frame interpolation module and the second frame interpolation module are both connected to the fusion module. The first frame interpolation module is used to perform frame interpolation processing on the image frame data of the first scene captured by the first camera to obtain first image data. The second frame interpolation module is used to perform frame interpolation processing on the image frame data of the second scene captured by the second camera to obtain second image data. The fusion module is used to perform fusion processing on the first image data and the second image data to obtain fused image data.

[0175] It should be noted that the image frame data of the first scene captured by the first camera is the image frame data of the first scene captured by the first camera sensor 11, and the image frame data of the second scene captured by the second camera is the image frame data of the second scene captured by the second camera sensor 12. After processing by the image processing chip 20, a fused image data is obtained from the image frame data of the first scene captured by the first camera sensor 11 and the image frame data of the second scene captured by the second camera sensor 12. During the fusion process, the mobile phone can extract the background image data from the image frame data of the first scene captured by the first camera sensor 11 and extract the face image data from the image frame data of the second scene captured by the second camera sensor 12. The background image data and the face image data are then fused, and the fused image data is sent to the application processor 10 for displaying the image corresponding to the fused image data on the mobile phone screen. In this way, video recording is achieved simultaneously by the camera on the folding outer screen 13 and the camera on the folding inner screen 14, and the face image recorded by the camera on the folding inner screen 14 and the background image recorded by the camera on the folding outer screen 13 are fused and displayed.

[0176] Optionally, in this embodiment of the application, as shown in FIG11 and FIG13, the image processing chip 20 provided in this embodiment of the application may further include: a first noise reduction module 24 and a second noise reduction module 25.

[0177] In this embodiment of the application, the first noise reduction module 24 is connected to the first camera and the first frame interpolation module 21.

[0178] In this embodiment, the second noise reduction module 25 is connected to the second camera and the second frame interpolation module 22.

[0179] In this embodiment of the application, the first noise reduction module 24 is used to perform noise reduction processing on the image frame data of the first scene captured by the first camera.

[0180] In this embodiment of the application, the second noise reduction module 25 is used to perform noise reduction processing on the image frame data of the second scene captured by the second camera.

[0181] It is understood that the first noise reduction module 24 is connected to the first camera, that is, the first noise reduction module 24 is connected to the sensor corresponding to the first camera. The second noise reduction module 25 is connected to the second camera, that is, the second noise reduction module 25 is connected to the sensor corresponding to the second camera. Therefore, as shown in Figure 13, the first noise reduction module 24 is connected to the first camera sensor 11, and the second noise reduction module 25 is connected to the second camera sensor 12.

[0182] In this embodiment, both the first noise reduction module 24 and the second noise reduction module 25 use a noise reduction network model to perform noise reduction processing on the image frame data.

[0183] It should be noted that for a detailed explanation of the above noise reduction network model, please refer to the description in the above embodiments, which will not be repeated here.

[0184] It should be noted that, as shown in Figure 13, the image processing path composed of the first camera sensor 11, the first noise reduction module 24, and the first frame interpolation module 21 is an independent hardware processing path from the image processing path composed of the second camera sensor 12, the second noise reduction module 25, and the second frame interpolation module 22. Therefore, when the image frame data collected by the first camera sensor 11 and the image frame data collected by the second camera sensor 12 are simultaneously input to the image processing chip 20, the image processing chip 20 can process the two data streams simultaneously. In this way, the data buffer in the image processing chip is reduced by parallel processing, thereby reducing the latency of image processing. Secondly, for a single processing path, such as an image processing path consisting of a first camera sensor 11, a first noise reduction module 24, and a first frame interpolation module 21, the first noise reduction module 24 and the first frame interpolation module 21 are independent hardware units. The image frame data acquired by the first camera sensor 11 is input to the first noise reduction module 24 in the form of a line data stream. Once each line of image data is captured, it is immediately sent to the subsequent processing unit instead of being stored in a buffer first. This image processing path can use a pipeline processing mechanism to process the image frame data. That is, when the first noise reduction module 24 performs noise reduction processing on the current image frame data acquired by the first camera sensor 11, the first frame interpolation module 21 can simultaneously perform frame interpolation processing on the historical image frame data that has already undergone noise reduction processing by the first noise reduction module 24. In this way, the efficiency of image processing is improved through concurrent processing. In addition, the fusion module 23 is also an independent hardware unit. Therefore, in the image fusion processing process, the fusion module 23, which ultimately achieves fusion, and the front-end module can also use concurrent processing to perform image processing. In this way, no data caching is required in the entire image processing path, achieving low latency and low caching in the image processing chip, thus improving the image processing effect.

[0185] Understandably, in a pipeline processing mechanism, preceding and subsequent modules can process image data concurrently. This means that while a current module is processing a row of data, a subsequent module can process the data already processed by the previous module in parallel. This concurrent processing method can fully utilize the computing resources of the processing system and improve processing efficiency.

[0186] In this embodiment, the first noise reduction module and the second noise reduction module are connected to the first camera and the second camera, respectively. They perform noise reduction processing on the original image frame data of the first scene captured by the first camera and the original image frame data of the second scene captured by the second camera, respectively. These two noise reduction modules serve as preprocessing steps in the image processing chip. By removing noise from the original images, they effectively improve the quality of the images transmitted to the first and second frame interpolation modules. Simultaneously, during the noise reduction process, the first and second noise reduction modules can also retain sufficient image detail information, further improving the accuracy and efficiency of image processing.

[0187] Optionally, in this embodiment of the application, as shown in FIG13 and FIG14, the image processing chip 20 provided in this embodiment of the application may further include: a double data rate storage unit 26.

[0188] In this embodiment of the application, the double data rate storage unit 26 is connected to the first frame interpolation module 21, the second frame interpolation module 22 and the fusion module 23.

[0189] In this embodiment of the application, the double data rate storage unit 26 is used to store image data obtained by frame interpolation processing and image data obtained by fusion processing.

[0190] It is understandable that Double Data Rate (DDR) memory cells are a type of Dynamic Random Access Memory (DRAM), widely used in computers and other electronic devices. Image processing chips need to frequently access and store large amounts of image data when processing it. DDR memory cells, with their high speed, large capacity, and low power consumption, provide reliable data storage and caching services for image processing chips. This helps ensure that image processing chips can quickly access the required data when processing image data, reducing the latency of reading and writing to memory, thereby improving processing speed and efficiency. Furthermore, in image processing chips, the image processing workflow typically includes steps such as image acquisition, preprocessing, feature extraction, image enhancement, and encoding. DDR memory cells can support the parallel processing of these steps, optimizing the image processing workflow by providing fast data transfer and access capabilities, enabling image processing chips to remain efficient and stable when handling complex image tasks.

[0191] It should be noted that the aforementioned DDR memory unit is a dedicated memory unit specifically for storing data for the aforementioned image processing chip and is not provided for use by other external devices. This improves the reliability of data transmission by the image processing chip.

[0192] This application provides an electronic device. Figure 15 shows a schematic diagram of the structure of an electronic device provided in this application embodiment. As shown in Figure 15, the electronic device 30 provided in this application embodiment may include: a first camera, a second camera, and an image processing chip 20.

[0193] In this embodiment, both the first camera and the second camera are connected to the image processing chip 20.

[0194] It is understood that both the first camera and the second camera are connected to the image processing chip 20, meaning that the sensors corresponding to the first camera and the second camera are both connected to the image processing chip 20. Therefore, as shown in Figure 15, where the first camera sensor 11 represents the first camera and the second camera sensor 12 represents the second camera, both the first camera sensor 11 and the second camera sensor 12 are connected to the image processing chip 20.

[0195] This application provides an electronic device including a first camera, a second camera, and an image processing chip. The image processing chip has a fusion module that can fuse first image data and second image data. The first image data is obtained by interpolating image frames from a first scene captured by the first camera, and the second image data is obtained by interpolating image frames from a second scene captured by the second camera. In other words, the fusion module can fuse the image frames from the first scene captured by the first camera and the image frames from the second scene captured by the second camera to obtain fused image data. Therefore, when shooting with an electronic device containing multiple cameras, the device can simultaneously operate multiple cameras to capture images and fuse the video images obtained from the multiple cameras, resulting in a fused image from multiple scenes, thus improving the flexibility of the electronic device in shooting.

[0196] Optionally, as shown in FIG16, this application embodiment also provides an electronic device 1600, including a processor 1601 and a memory 1602. The memory 1602 stores a program or instructions that can run on the processor 1601. When the program or instructions are executed by the processor 1601, they implement the various steps of the above-described image processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0197] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0198] Figure 17 is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application.

[0199] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, processor 110, and image processing chip.

[0200] It should be noted that the specific description of the image processing chip can be found in the above embodiments, and will not be repeated here.

[0201] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for powering various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The electronic device structure shown in Figure 17 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0202] The user input unit 107 is configured to receive a first input when the video recording interface is displayed. The processor 110 is configured to, in response to the first input, record video of a first scene using a first camera and record video of a second scene using a second camera. The processor 110 is also configured to perform image fusion based on the first scene and the second scene. The display unit 106 is configured to display the fused image.

[0203] Optionally, the above-mentioned electronic device is an electronic device with a foldable screen. The processor 110 is specifically used to respond to the first input and, when the folding angle of the first screen and the second screen of the foldable screen is greater than or equal to a first threshold, record video of the first scene through the first camera and record video of the second scene through the second camera.

[0204] Optionally, the aforementioned electronic device includes an image processing chip. The processor 110 is specifically configured to: perform frame interpolation processing on image frame data of a first scene captured by the first camera using a first frame interpolation module in the image processing chip to obtain first image data; perform frame interpolation processing on image frame data of a second scene captured by the second camera using a second frame interpolation module in the image processing chip to obtain second image data; and perform fusion processing on the first image data and the second image data using a fusion module in the image processing chip to obtain fused image data.

[0205] Optionally, the processor 110 is further configured to, before performing frame interpolation processing on the image frame data of the first scene acquired by the first camera through the first frame interpolation module in the image processing chip to obtain the first image data, extract features from the image frame data of the first scene acquired by the first camera and historical image frame data through the frame interpolation network model in the first frame interpolation module to obtain a first feature map, the first feature map representing the motion information of the image frame data of the first scene acquired by the first camera relative to the historical image frame data; and perform feature processing on the first feature map through the residual network model in the first frame interpolation module to obtain intermediate image frame data. Specifically, the processor 110 is configured to perform frame interpolation processing on the image frame data of the first scene acquired by the first camera based on the intermediate image frame data through the first frame interpolation module to obtain the first image data.

[0206] Optionally, the processor 110 is further configured to perform noise reduction processing on the image frame data of the first scene captured by the first camera through the first noise reduction module in the image processing chip to obtain third image data. Specifically, the processor 110 is configured to perform frame interpolation processing on the third image data through the first frame interpolation module in the image processing chip to obtain the first image data.

[0207] Optionally, the processor 110 is further configured to perform noise reduction processing on the image frame data of the second scene captured by the second camera through the second noise reduction module in the image processing chip to obtain fourth image data. Specifically, the processor 110 is configured to perform frame interpolation processing on the fourth image data through the second frame interpolation module in the image processing chip to obtain the second image data.

[0208] Optionally, the processor 110 is specifically configured to perform convolution processing on the first image data through the fusion network model in the fusion module to obtain a third feature map, and perform convolution processing on the second image data to obtain a fourth feature map; and perform feature processing on the third feature map through the fusion network model to obtain background image data, and perform feature processing on the fourth feature map to obtain foreground image data; and perform fusion processing on the background image data and the foreground image data to obtain the fused image data.

[0209] Optionally, the processor 110 is specifically configured to: downsample the first image data using the first convolutional module in the fusion network model to output a first feature vector; and downsample the (i-1)th feature vector using the i-th convolutional module in the fusion network model to output an i-th feature vector, where i ∈ [2, N], i is an integer, and N is the number of convolutional modules in the fusion network model; and upsample the N-th and (N-1)th feature vectors using the first deconvolutional module in the fusion network model to obtain... The first content feature vector includes a foreground feature vector and a background feature vector; and, through the j-th deconvolution module in the fusion network model, the Nj-th feature vector and the (j-1)-th content feature vector are upsampled to output the j-th content feature vector, where j∈[2,N-1] and j is an integer; and, through the N-th deconvolution module in the fusion network model, the first image data and the (N-1)-th content feature vector are upsampled to output the N-th content feature vector, and the N-th content feature vector is determined as the aforementioned third feature map.

[0210] Optionally, the processor 110 is specifically used to perform foreground removal and background generation processing on the third feature map through the convolution module in the fusion network model to obtain the background image data.

[0211] Optionally, the processor 110 is specifically used to perform background removal processing on the fourth feature map through the convolution module in the fusion network model to obtain initial foreground image data; and to perform edge feathering processing on the initial foreground image data through the edge attention module in the fusion network model to obtain the foreground image data.

[0212] This application provides an electronic device that can perform image fusion based on a first scene and a second scene to display the fused image. The first scene is a scene recorded by a first camera, and the second scene is a scene recorded by a second camera. Specifically, the electronic device can fuse image data from the first scene recorded by the first camera and image data from the second scene recorded by the second camera, and display the image corresponding to the fused image data. In this way, when shooting with an electronic device containing multiple cameras, the device can respond to a first input, simultaneously operate multiple cameras to shoot, and fuse the video images captured by the multiple cameras to display the fused image. This makes the image captured by the electronic device a fused image from multiple scenes, improving the flexibility of the electronic device in shooting.

[0213] The electronic device provided in this application embodiment can implement all the processes implemented in the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here. The beneficial effects of the various implementation methods in this embodiment can be found in the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, it will not be described again here.

[0214] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0215] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0216] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0217] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0218] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0219] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0220] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0221] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described image processing method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0222] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0223] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0224] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image processing method applied to an electronic device, the method comprising: Receive the first input while the video recording interface is displayed; In response to the first input, video is recorded of the first scene using the first camera, and video is recorded of the second scene using the second camera; The images are fused based on the first scene and the second scene, and the fused image is displayed.

2. The method according to claim 1, wherein, The electronic device is an electronic device with a foldable screen. The step of recording video of a first scene using a first camera and recording video of a second scene using a second camera, in response to the first input, includes: In response to the first input, when the folding angle of the first screen and the second screen of the foldable screen is greater than or equal to a first threshold, the first scene is recorded by the first camera and the second scene is recorded by the second camera.

3. The method according to claim 1, wherein, The electronic device includes an image processing chip, and the image fusion based on the first scene and the second scene includes: The first frame interpolation module in the image processing chip performs frame interpolation processing on the image frame data of the first scene captured by the first camera to obtain the first image data. The second frame interpolation module in the image processing chip performs frame interpolation processing on the image frame data of the second scene captured by the second camera to obtain the second image data. The first image data and the second image data are fused together by the fusion module in the image processing chip to obtain fused image data.

4. The method according to claim 3, wherein, Before the first image data is obtained by interpolating the image frame data of the first scene captured by the first camera through the first frame interpolation module in the image processing chip, the method further includes: The first feature map is obtained by extracting features from the image frame data and historical image frame data of the first scene captured by the first camera through the frame interpolation network model in the first frame interpolation module. The first feature map represents the motion information of the image frame data of the first scene captured by the first camera relative to the historical image frame data. The residual network model in the first frame interpolation module is used to perform feature processing on the first feature map to obtain intermediate image frame data. The step of interpolating the image frame data of the first scene captured by the first camera through the first frame interpolation module in the image processing chip to obtain the first image data includes: The first frame interpolation module performs frame interpolation processing on the image frame data of the first scene captured by the first camera based on the intermediate image frame data to obtain the first image data.

5. The method according to claim 3, wherein, The method further includes: The first noise reduction module in the image processing chip performs noise reduction processing on the image frame data of the first scene captured by the first camera to obtain the third image data. The step of interpolating the image frame data of the first scene captured by the first camera through the first frame interpolation module in the image processing chip to obtain the first image data includes: The first image data is obtained by interpolating the third image data using the first frame interpolation module in the image processing chip.

6. The method according to claim 3, wherein, The method further includes: The second noise reduction module in the image processing chip performs noise reduction processing on the image frame data of the second scene captured by the second camera to obtain the fourth image data. The step of interpolating the image frame data of the second scene captured by the second camera through the second frame interpolation module in the image processing chip to obtain the second image data includes: The second image data is obtained by interpolating the fourth image data through the second frame interpolation module in the image processing chip.

7. The method according to any one of claims 3 to 6, wherein, The process of fusing the first image data and the second image data through the fusion module in the image processing chip to obtain fused image data includes: The first image data is convolved using the fusion network model in the fusion module to obtain a third feature map, and the second image data is convolved to obtain a fourth feature map. The third feature map is processed using the fusion network model to obtain background image data, and the fourth feature map is processed to obtain foreground image data. The background image data and the foreground image data are fused together to obtain the fused image data.

8. The method according to claim 7, wherein, The step of performing convolution processing on the first image data through the fusion network model in the fusion module to obtain the third feature map includes: The first image data is downsampled using the first convolutional module in the fusion network model to output the first feature vector. The i-th convolutional module in the fusion network model performs downsampling on the (i-1)-th feature vector and outputs the i-th feature vector, where i ∈ [2, N], i is an integer, and N is the number of convolutional modules in the fusion network model. The first deconvolution module in the fusion network model upsamples the Nth feature vector and the (N-1)th feature vector to obtain the first content feature vector, which includes a foreground feature vector and a background feature vector. The j-th deconvolution module in the fusion network model performs upsampling on the Nj-th feature vector and the (j-1)-th content feature vector to output the j-th content feature vector, where j∈[2,N-1] and j is an integer; The first image data and the (N-1)th content feature vector are upsampled using the Nth deconvolution module in the fusion network model to output the Nth content feature vector, and the Nth content feature vector is determined as the third feature map.

9. The method according to claim 7, wherein, The step of performing feature processing on the third feature map through the fusion network model to obtain background image data includes: The background image data is obtained by performing foreground removal and background generation on the third feature map through the convolution module in the fusion network model.

10. The method according to claim 7, wherein, The step of performing feature processing on the fourth feature map through the fusion network model to obtain foreground image data includes: The fourth feature map is processed by the convolution module in the fusion network model to remove background, thereby obtaining the initial foreground image data. The initial foreground image data is feathered by the edge attention module in the fusion network model to obtain the foreground image data.

11. An image processing chip, comprising: The system comprises a first frame interpolation module, a second frame interpolation module, and a fusion module, wherein the first frame interpolation module and the second frame interpolation module are both connected to the fusion module. The first frame interpolation module is used to perform frame interpolation processing on the image frame data of the first scene captured by the first camera to obtain the first image data; the second frame interpolation module is used to perform frame interpolation processing on the image frame data of the second scene captured by the second camera to obtain the second image data. The fusion module is used to perform fusion processing on the first image data and the second image data to obtain fused image data.

12. The image processing chip according to claim 11, wherein, The image processing chip further includes: a first noise reduction module and a second noise reduction module; The first noise reduction module is connected to the first camera and the first frame interpolation module; the second noise reduction module is connected to the second camera and the second frame interpolation module. The first noise reduction module is used to perform noise reduction processing on the image frame data of the first scene captured by the first camera; The second noise reduction module is used to perform noise reduction processing on the image frame data of the second scene captured by the second camera.

13. The image processing chip according to claim 11 or 12, wherein, The image processing chip also includes: a double data rate storage unit; The double data rate storage unit is connected to the first frame interpolation module, the second frame interpolation module, and the fusion module. The double data rate storage unit is used to store image data obtained from frame interpolation and image data obtained from fusion processing.

14. An electronic device comprising: The first camera, the second camera, and the image processing chip as described in any one of claims 11 to 13; Both the first camera and the second camera are connected to the image processing chip.

15. An electronic device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the image processing method as claimed in any one of claims 1 to 10.

16. A computer program product, said program product being executed by at least one processor to implement the image processing method as claimed in any one of claims 1 to 10.

17. A user equipment (UE) comprising a UE configured to perform an image processing method as claimed in any one of claims 1 to 10.

18. A chip comprising a processor and a communication interface coupled to the processor, the processor being configured to run a program or instructions which, when executed by the processor, implement the steps of the image processing method as claimed in any one of claims 1 to 10.