A method and apparatus for synthesizing images

By adaptively adjusting the number of image frames and exposure parameters, and synthesizing images based on the semantic complexity of the target frame, the problems of resource consumption and imbalance between imaging effect and efficiency in multi-frame fusion technology are solved, and efficient image synthesis is achieved.

CN122138038APending Publication Date: 2026-06-02LENOVO (BEIJING) LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LENOVO (BEIJING) LTD
Filing Date
2026-03-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing multi-frame fusion technology consumes a lot of hardware resources when synthesizing images, affecting the performance of electronic devices, and it is difficult to balance imaging effect and efficiency.

Method used

By acquiring the semantic region and semantic label of the target frame, the number of image frames and exposure parameters are adaptively adjusted, and the image acquisition strategy is dynamically adjusted according to the semantic complexity to synthesize the captured image.

Benefits of technology

While ensuring image quality, it reduces hardware resource consumption, improves processing efficiency and imaging effect, and adapts to the complexity of different shooting scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138038A_ABST
    Figure CN122138038A_ABST
Patent Text Reader

Abstract

This application provides a method and apparatus for synthesizing images. The method includes: in response to a photo-taking command, acquiring semantic regions included in a target frame and semantic tags corresponding to the semantic regions; each pixel in the same semantic region corresponds to a semantic tag, and adjacent semantic regions have different semantic tags; determining photo-taking parameters based on semantic parameters of the target frame; the photo-taking parameters include at least the number of image frames representing the synthesized photo image; wherein the semantic parameters represent the complexity of the semantic regions and / or semantic tags; and synthesizing the photo image based on the image frames obtained from the photo-taking parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for synthesizing images. Background Technology

[0002] Multi-frame fusion technology combines multiple frames captured from the same shooting scene into a single frame using algorithms. This improves image quality by achieving effects such as noise reduction, increased sharpness, expanded dynamic range, and increased depth of field. The conventional approach involves capturing a fixed number of image frames and then combining them into a single image using algorithms. However, running these algorithms consumes significant hardware resources and can negatively impact the performance of electronic devices, potentially causing them to overheat. Summary of the Invention

[0003] This application provides a method for synthesizing images, applied to an electronic device. The method includes: in response to a photo-taking command, acquiring semantic regions and corresponding semantic tags of a target frame; each pixel in the same semantic region corresponds to a semantic tag, and adjacent semantic regions have different semantic tags; determining photo-taking parameters based on semantic parameters of the target frame; the photo-taking parameters include at least the number of image frames representing the synthesized photo image; wherein the semantic parameters represent the complexity of the semantic regions and / or semantic tags; and synthesizing the photo image based on the image frames obtained from the photo-taking parameters.

[0004] In some embodiments, semantic parameters include the number of semantic regions and / or the number of categories of semantic tags.

[0005] In some embodiments, the image capture parameters also include the number of image frames under different exposure parameters.

[0006] In some embodiments, in response to a photo capture command, obtaining the semantic regions included in the target frame and the semantic labels corresponding to the semantic regions includes: in response to a photo capture command, inputting the target frame into a preset model to obtain a semantic segmentation result; the semantic segmentation result characterizes the category corresponding to each pixel included in the target frame; based on the semantic segmentation result, aggregating pixels that correspond to the same category and are spatially adjacent into a semantic region, and using the category as the semantic label corresponding to the semantic region.

[0007] In some embodiments, the method further includes: determining the number of semantic regions included in the target frame; and determining the number of categories of semantic tags based on the semantic tags corresponding to the target frame.

[0008] In some embodiments, in response to a shooting command, obtaining the semantic region included in the target frame and the semantic tag corresponding to the semantic region includes: in response to a shooting command, obtaining the semantic region included in the focus region of the target frame and the semantic tag corresponding to the semantic region.

[0009] In some embodiments, the method further includes: if the target semantic region of the target frame satisfies the target condition, the number of frames represented by the shooting parameters is a first frame number; if the target semantic region of the target frame does not satisfy the target condition, the number of frames represented by the shooting parameters is a second frame number; the target semantic region is the semantic region with the largest area located within the focus area, and the first frame number and the second frame number are different.

[0010] In some embodiments, the number of first frames is less than the number of second frames, and the target condition is selected from one of the following combinations: the proportion of the target semantic region to the target frame exceeds a preset proportion threshold; the area of ​​the target semantic region is not less than an area threshold.

[0011] In some embodiments, determining the shooting parameters based on the semantic parameters of the target frame includes: determining the current shooting scene based on the semantic parameters of the target frame; if the current shooting scene represents a complex shooting scene, determining the shooting parameters as the first shooting parameters; if the current shooting scene represents a simple shooting scene, determining the shooting parameters as the second shooting parameters.

[0012] In some embodiments, the current shooting scene is determined based on the semantic parameters of the target frame, and is selected from one of the following combinations: if the number of semantic regions is less than a first threshold, the current shooting scene is determined to represent a simple shooting scene; if the number of semantic regions is not less than the first threshold, the current shooting scene is determined to represent a complex shooting scene; and / or, if the number of semantic tag categories is less than a second threshold, the shooting scene is determined to represent a simple shooting scene; if the number of semantic tag categories is not less than the second threshold, the shooting scene is determined to represent a complex shooting scene.

[0013] In some embodiments, different exposure parameters include high exposure parameters, base exposure parameters, and low exposure parameters; based on the semantic parameters of the target frame, the image capture parameters are determined, selected from one of the following combinations: based on the number of semantic regions and a first mapping relationship, the image capture parameters are determined; wherein, the first mapping relationship represents a one-to-one correspondence between the number of multiple semantic regions and multiple image capture parameters; the number of semantic regions is positively correlated with the number of image frames in the synthesized image capture, and / or the number of semantic regions is positively correlated with the number of image frames under the base exposure parameters; or, based on the number of semantic label categories and a second mapping relationship, the image capture parameters are determined; wherein, the second mapping relationship represents a one-to-one correspondence between the number of categories of multiple semantic labels and multiple image capture parameters; the number of semantic label categories is positively correlated with the number of image frames in the synthesized image capture, and / or the number of semantic label categories is positively correlated with the number of image frames under the base exposure parameters.

[0014] In some embodiments, the target frame is an image frame captured at a first moment, which is located before the second moment when the photo-taking instruction is received, and the duration between the first moment and the second moment is less than a preset duration.

[0015] This application provides an image synthesis device, including a processor and an image acquisition device. The processor is configured to, in response to a photo capture command, acquire the semantic regions included in the target frame and the semantic tags corresponding to the semantic regions; each pixel in the same semantic region corresponds to a semantic tag, and adjacent semantic regions have different semantic tags; determine photo capture parameters based on the semantic parameters of the target frame; the photo capture parameters include at least the number of image frames representing the synthesized photo image; wherein, the semantic parameters represent the complexity of the semantic regions and / or semantic tags; the processor is also configured to control the image acquisition device to acquire image frames according to the photo capture parameters, and synthesize the photo image based on the image frames acquired by the image acquisition device. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a method for synthesizing images provided in an embodiment of this application; Figure 2 This is a flowchart illustrating another method for synthesizing images provided in an embodiment of this application; Figure 3 This is a flowchart illustrating another method for synthesizing images provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the determination of a semantic region according to an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a device for synthesizing images provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

[0019] Multi-frame fusion technology combines multiple frames captured from the same shooting scene into a single frame using algorithms. This improves image quality by achieving effects such as noise reduction, increased sharpness, expanded dynamic range, and increased depth of field. The conventional approach involves capturing a fixed number of image frames and then combining them into a single image using algorithms. However, running these algorithms consumes significant hardware resources, such as central processing unit (CPU) resources, graphics processing unit (GPU) resources, and memory resources. It can also negatively impact the performance of electronic devices, potentially causing them to overheat.

[0020] To address this, embodiments of this application provide a method and apparatus for synthesizing images, capable of adaptively acquiring image frames for synthesizing a photographed image based on the semantic complexity of the target frame. For example, when the semantic complexity of the target frame is high, more image frames are acquired for synthesizing the photographed image. When the semantic complexity of the target frame is low, fewer image frames are acquired for synthesizing the photographed image. The semantic complexity of the target frame can be used to indicate the scene complexity of the current shooting scene, which can, for example, characterize the richness of the objects being photographed in the current shooting scene. Therefore, embodiments of this application can adaptively adjust the number of image frames used for synthesizing the photographed image according to the complexity of the current shooting scene, improving the imaging quality of high-complexity shooting scenes while reducing the processing overhead of low-complexity shooting scenes, thus balancing image synthesis effect and processing efficiency.

[0021] The image synthesis method provided in this application can be executed by an electronic device, such as a mobile phone, tablet computer, or smartwatch. The following describes a method and apparatus for image synthesis provided in this application, with reference to specific embodiments and accompanying drawings. Figure 1 This is a flowchart illustrating a method for synthesizing images provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes: S101, in response to the photo capture command, acquires the semantic region included in the target frame and the semantic label corresponding to the semantic region.

[0022] In this embodiment, the target frame can be a preview frame. The preview frame is an image frame captured in real-time and displayed on the preview interface during the camera preview stage. Its purpose is to help the user observe the shooting scene and confirm composition and focus. Specifically, the acquisition time of the target frame can be, for example, a first moment, meaning the target frame is a preview frame captured at the first moment. The moment the shooting command is received is a second moment, where the first moment precedes the second moment and the duration between the first and second moments is less than a preset duration. The preset duration can be, for example, 0.25 seconds or 0.5 seconds, etc., without specific limitation.

[0023] For example, the electronic device includes a camera application. In response to a launch command, the electronic device launches the camera application, captures a preview frame, and displays it on the preview interface of the camera application. Then, in response to a photo capture command, the preview frame whose capture time is before the time the photo capture command is received and whose duration between the capture time and the time the photo capture command is received is less than a preset duration is used as the target frame.

[0024] In this embodiment, the photo-taking command can be issued by the user through various means such as pressing a button, voice, or gesture. Alternatively, the photo-taking command can be issued by other applications installed on the electronic device, or by other electronic devices connected to the electronic device; this application does not specifically limit this.

[0025] In this embodiment, a semantic region is a set of pixels of the same category in a target frame. Here, the category can be understood as a semantic label. For example, all pixels within the same semantic region correspond to the same semantic label; that is, the semantic region is composed of pixels with the same semantic label. In other words, the semantic region is composed of pixels of the same category. Adjacent semantic regions have different semantic labels. In other words, adjacent semantic regions include pixels of different categories.

[0026] In this embodiment, semantic tags are used to identify the semantic category to which a semantic region belongs, in order to distinguish different semantic regions. Semantic tags can be, for example, sky, grass, buildings, roads, faces, arms, bodies, legs, etc.

[0027] After determining the semantic region of the target frame and the corresponding semantic tag, the electronic device can determine the image capture parameters based on the semantic complexity of the target frame. For example, the electronic device can execute S102.

[0028] S102, determine the image capture parameters based on the semantic parameters of the target frame.

[0029] In some embodiments, the semantic parameters of a target frame characterize the semantic complexity of the target frame. For example, the semantic parameters of a target frame characterize the complexity of the semantic regions included in the target frame. For instance, the semantic parameters of a target frame include the number of semantic regions. The more semantic regions there are, the higher the complexity of the semantic regions; the fewer semantic regions there are, the lower the complexity of the semantic regions. As another example, the semantic parameters of a target frame characterize the complexity of the semantic tags corresponding to the target frame. For example, the semantic parameters of a target frame include the number of semantic tag categories. The more categories of semantic tags there are, the higher the complexity of the semantic tags; the more singular the categories of semantic tags, the lower the complexity of the semantic tags. As yet another example, the semantic parameters of a target frame characterize both the complexity of the semantic regions included in the target frame and the complexity of the semantic tags. For example, the semantic parameters of a target frame include the number of semantic regions and the number of semantic tag categories.

[0030] In this embodiment of the application, the photographing parameters characterize the acquisition information of the image frames used to synthesize the photographed image. This acquisition information may include, for example, the number of frames, exposure parameters, acquisition order, and other information.

[0031] In some embodiments, the photographic parameters include at least the number of image frames representing the synthesized image. For example, the number of image frames representing the synthesized image may be 4, 6, or 7. In this embodiment, the more complex the semantic region and / or semantic label indicated by the semantic parameters of the target frame, the more image frames the photographic parameters represent in the synthesized image; conversely, the simpler the semantic region and / or semantic label indicated by the semantic parameters of the target frame, the fewer image frames the photographic parameters represent in the synthesized image. The image frames of the synthesized image may, for example, be image frames under different exposure parameters.

[0032] In some embodiments, the shooting parameters can be determined based on the semantic parameters and the preset mapping relationship. This will be described in detail below with reference to specific embodiments, and will not be repeated here.

[0033] Optionally, after determining the shooting parameters, the electronic device can acquire image frames based on the determined shooting parameters. For example, if the shooting parameters indicate that the number of image frames in the composite image is 4, the electronic device can acquire 4 image frames. As another example, if the shooting parameters indicate that the composite image consists of 1 image frame with an exposure level of EV+, 5 image frames with an exposure level of EV0, and 1 image frame with an exposure level of EV-, the electronic device can acquire 1 image frame with an exposure level of EV+, 5 image frames with an exposure level of EV0, and 1 image frame with an exposure level of EV-. Then, the electronic device can execute S103.

[0034] S103 synthesizes the captured image based on the image frames obtained from the shooting parameters.

[0035] For example, an electronic device can use a fusion algorithm to synthesize a photographed image based on the image frames obtained from the photographing parameters.

[0036] As can be seen, this application can adaptively adjust the number of image frames used to synthesize the captured image based on the semantic complexity of the target frame, thus balancing image synthesis effect and processing efficiency.

[0037] In some embodiments, the shooting parameters may further include the number of image frames under different exposure parameters. Different exposure parameters may include, for example, high exposure parameters, basic exposure parameters, and low exposure parameters. Exposure parameters may be, for example, parameters characterizing the degree of exposure, such as exposure level or exposure duration. Taking exposure level as an example, different exposure parameters may be, for example, high exposure level (EV+), basic exposure level (EV0), and low exposure level (EV-). For example, the shooting parameters may characterize the composite image as having 1 image frame with an exposure level of EV+, 2 image frames with an exposure level of EV0, and 1 image frame with an exposure level of EV-. As another example, the shooting parameters may characterize the composite image as having 1 image frame with an exposure level of EV+, 5 image frames with an exposure level of EV0, and 1 image frame with an exposure level of EV-.

[0038] In this embodiment, the more complex the semantic region indicated by the semantic parameters of the target frame, the more image frames in the synthesized image in the shooting parameters, and / or the more image frames under the basic exposure parameters. Conversely, the simpler the semantic region indicated by the semantic parameters of the target frame, the fewer image frames in the synthesized image in the shooting parameters, and / or the fewer image frames under the basic exposure parameters. Similarly, the more complex the semantic label indicated by the semantic parameters of the target frame, the more image frames in the synthesized image in the shooting parameters, and / or the more image frames under the basic exposure parameters. Finally, the simpler the semantic label indicated by the semantic parameters of the target frame, the fewer image frames in the synthesized image in the shooting parameters, and / or the fewer image frames under the basic exposure parameters.

[0039] In some embodiments, the electronic device determines the semantic regions and corresponding semantic labels of the target frame using a neural network model. For example, in response to a photo-taking command, the electronic device inputs the target frame into a preset model to obtain a semantic segmentation result. The preset model is a preset neural network model, which may be, for example, a semantic segmentation model. The semantic segmentation model may be, for example, at least one of a fully convolutional network model, a U-shaped network model, or a segmentation network model. The semantic segmentation result characterizes the category corresponding to each pixel in the target frame. That is, the semantic segmentation result includes information about the category corresponding to each pixel in the target frame.

[0040] Electronic devices can obtain the semantic regions included in the target frame and the corresponding semantic labels based on the semantic segmentation results. For example, based on the semantic segmentation results, pixels corresponding to the same category and spatially adjacent are aggregated into semantic regions, and the category is used as the semantic label corresponding to that semantic region. In other words, a semantic region is obtained by aggregating pixels of the same category and spatially adjacent, and the semantic label of the semantic region is the category of the pixels included in that semantic region. For example, pixels corresponding to the category "sky" and spatially adjacent are aggregated into a semantic region, and the semantic label of this semantic region is "sky". As another example, pixels corresponding to the category "grass" and spatially adjacent are aggregated into a semantic region, and the semantic label of this semantic region is "grass".

[0041] In some embodiments, after determining the semantic regions included in the target frame and the corresponding semantic tags, the electronic device can further determine the number of semantic regions and the number of semantic tag categories. The number of semantic regions represents the total number of independent semantic regions included in the target frame, and the number of semantic tag categories represents the total number of semantic tag categories corresponding to the target frame. In other words, the number of semantic regions is the number of independent semantic regions included in the target frame. The number of semantic tag categories is the number of types of semantic tags corresponding to the semantic regions included in the target frame. For example, the target frame includes 5 semantic regions, corresponding to the semantic tags for sky, grass, road, building, and building, respectively. The number of corresponding semantic regions is 5, and the number of semantic tag categories is 4.

[0042] In this embodiment, the number of semantic regions in a target frame can characterize the semantic complexity of the target frame. The number of semantic tag categories can also characterize the semantic complexity of the target frame. Semantic complexity characterizes the richness of the captured objects in the target frame. Semantic complexity is positively correlated with the number of semantic regions; for example, the more semantic regions, the higher the semantic complexity of the target frame. Alternatively, semantic complexity is positively correlated with the number of semantic tag categories; for example, the more semantic tag categories, the higher the semantic complexity of the target frame. Or, semantic complexity is positively correlated with both the number of semantic regions and the number of semantic tag categories; for example, the more semantic regions and the more semantic tag categories, the higher the semantic complexity of the target frame.

[0043] In some embodiments, when determining the shooting parameters based on the semantic parameters of the target frame, the electronic device can first determine whether the current shooting scene is a complex shooting scene or a simple shooting scene, and different shooting scenes correspond to different determination logic.

[0044] For example, the electronic device can first determine the current shooting scene based on the semantic parameters of the target frame. For instance, the electronic device can determine the current shooting scene based on the number of semantic regions included in the target frame. If the number of semantic regions included in the target frame is less than a first threshold, the electronic device determines that the current shooting scene is a simple shooting scene, i.e., the current shooting scene represents a simple shooting scene; if the number of semantic regions included in the target frame is not less than the first threshold, the electronic device determines that the current shooting scene is a complex shooting scene, i.e., the current shooting scene represents a complex shooting scene. Here, the first threshold is a preset value, for example, the first threshold is 5, 6, or 7.

[0045] For example, an electronic device determines the current shooting scene based on the number of semantic tag categories corresponding to the target frame. If the number of semantic tag categories is less than a second threshold, the electronic device determines that the current shooting scene is a simple shooting scene; if the number of semantic tag categories is not less than the second threshold, the electronic device determines that the current shooting scene is a complex shooting scene. The second threshold is a preset value, such as 4, 5, or 6.

[0046] For example, an electronic device determines the current shooting scene based on the number of semantic regions and the number of semantic tag categories included in the target frame. Specifically, if the number of semantic regions is less than a first threshold and the number of semantic tag categories is less than a second threshold, the current shooting scene is determined to be a simple shooting scene. If the number of semantic regions is not less than the first threshold or the number of semantic tag categories is not less than the second threshold, the current shooting scene is determined to be a complex shooting scene.

[0047] If the current shooting scene is complex, then the shooting parameters are determined to be the first shooting parameters. If the current shooting scene is simple, then the shooting parameters are determined to be the second shooting parameters. The first and second shooting parameters are different. For example, the number of image frames in the composite image represented by the first shooting parameters is different from the number of image frames in the composite image represented by the second shooting parameters. Another example is that the number of image frames in the composite image represented by the first and second shooting parameters is the same, but the number of image frames under at least one exposure parameter is different. Yet another example is that the number of image frames in the composite image represented by the first and second shooting parameters is different, and the number of image frames under at least one exposure parameter is different. For example, the number of image frames in the composite image represented by the first shooting parameters is greater than the number of image frames in the composite image represented by the second shooting parameters, and / or, the number of image frames at the base exposure level represented by the first shooting parameters is greater than the number of image frames at the base exposure level represented by the second shooting parameters.

[0048] For example, the second image capture parameter is determined based on the semantic parameters of the target frame. The first image capture parameter is a preset parameter, such as the original parameter. In this embodiment, for simple shooting scenarios, determining the image capture parameter based on semantic parameters enables adaptive adjustment of the parameter, optimizing imaging efficiency and reducing power consumption and processing time while ensuring imaging quality. For complex shooting scenarios, using the preset first image capture parameter (the original parameter) ensures the stability and compatibility of the imaging process, avoids imaging risks caused by parameter adjustments in complex scenarios, and improves imaging reliability. By distinguishing between simple and complex shooting scenarios and using adaptively determined and preset parameters respectively, the adaptive optimization capability for simple scenarios and the imaging stability for complex scenarios are balanced, improving imaging quality and user experience while ensuring reliable and efficient operation of the electronic device.

[0049] For example, both the first and second image capture parameters are determined based on the semantic parameters of the target frame. In other words, regardless of whether the shooting scene is complex or simple, the image capture parameters can be determined based on the semantic parameters of the target frame.

[0050] In other words, there are two methods for determining the first and second shooting parameters. The first method is to determine the shooting parameters based on the semantic parameters of the target frame if the current shooting scene is simple. If the current shooting scene is complex, the preset shooting parameters are used as the first shooting parameters. The second method is to determine the shooting parameters based on the semantic parameters of the target frame regardless of whether the shooting scene is simple or complex. If the current shooting scene is simple, the shooting parameters determined based on the semantic parameters of the target frame are the second shooting parameters; if the current shooting scene is complex, the shooting parameters determined based on the semantic parameters of the target frame are the first shooting parameters.

[0051] The following describes, with specific examples, how to determine the image capture parameters based on the semantic parameters of the target frame.

[0052] In some embodiments, the semantic parameters of the target frame include the number of semantic regions and / or the number of semantic tag categories.

[0053] In some embodiments, the image capture parameters are determined based on the semantic parameters of the target frame, specifically based on the number of semantic regions. For example, the electronic device determines the image capture parameters based on the number of semantic regions and a first mapping relationship. The first mapping relationship represents a one-to-one correspondence between the number of multiple semantic regions and multiple image capture parameters. The image capture parameters can represent the number of image frames in the synthesized image. For example, the first mapping relationship can be 1 semantic region corresponding to 4 image frames, 2 semantic regions corresponding to 4 image frames, 3 semantic regions corresponding to 5 image frames, 4 semantic regions corresponding to 5 image frames, 5 semantic regions corresponding to 7 image frames, etc.

[0054] Alternatively, the shooting parameters can represent the number of image frames under different exposure parameters. For example, the first mapping relationship can be: 1 semantic region corresponds to 1 image frame with an exposure level of EV+, 1 image frame with an exposure level of EV- and 2 image frames with an exposure level of EV0; 2 semantic regions correspond to 1 image frame with an exposure level of EV+, 1 image frame with an exposure level of EV- and 2 image frames with an exposure level of EV0; 3 semantic regions correspond to 1 image frame with an exposure level of EV+, 1 image frame with an exposure level of EV- and 4 image frames with an exposure level of EV0; 4 semantic regions correspond to 1 image frame with an exposure level of EV+, 1 image frame with an exposure level of EV- and 4 image frames with an exposure level of EV0; 5 semantic regions correspond to 1 image frame with an exposure level of EV+, 1 image frame with an exposure level of EV- and 5 image frames with an exposure level of EV0, etc. This example uses one image frame each for exposure levels EV- and EV+. It should be understood that the number of image frames for exposure levels EV- and EV+ can be designed as needed, as long as the number of semantic regions is positively correlated with the number of image frames under the basic exposure parameters.

[0055] In some embodiments, a larger number of semantic regions indicates a greater number of objects in the current shooting scene. Different elements are likely to require different exposure parameters to achieve a better presentation. To fully capture the image information of the objects, a larger number of image frames need to be acquired to enhance the clarity and fidelity of the elements in the captured image. Conversely, a smaller number of semantic regions indicates fewer objects in the current shooting scene, and a smaller number of image frames are sufficient to achieve the required clarity and fidelity of the elements, thus improving processing efficiency. In other words, the number of semantic regions is positively correlated with the number of image frames required to synthesize the captured image.

[0056] For example, target frame A includes grassland, lake, trees, sky, sun, and swans in the lake. To fully capture image information from different subjects, more image frames need to be acquired. This image information can include brightness, color, texture, and contour information, among others. Synthesizing a photograph based on more image frames can effectively improve the dynamic range, signal-to-noise ratio, and sharpness of the photographed image. Furthermore, it can enrich the image's detail, suppress image noise, and enhance color reproduction and depth. For example, if target frame A includes 7 semantic regions, 10 image frames can be acquired. On the other hand, target frame B includes a white wall and calligraphy / paintings, requiring fewer image frames to be acquired. This reduces data processing while maintaining image quality and improving shooting response speed. For example, if target frame B includes 2 semantic regions, 4 image frames can be acquired.

[0057] In some embodiments, image frames at the base exposure parameters serve as reference frames for image synthesis. For example, image frames at the base exposure parameters possess advantages such as reasonable exposure, complete image information, moderate noise levels, and stable overall image quality. Therefore, they are used as reference frames in multi-frame synthesis processing for registration, fusion, noise reduction, and detail enhancement with image frames at different exposure parameters, thereby ensuring the overall quality and stability of the synthesized image. The more semantic regions there are, the more objects are represented in the current shooting scene. To fully capture the image information of different objects in the current shooting scene, more image frames at the base exposure parameters need to be acquired. This allows for the capture of richer image information and improves image quality. Conversely, the fewer objects are represented in the current shooting scene, the required image frames at the base exposure level only need to meet basic imaging requirements. Therefore, the number of semantic regions is positively correlated with the number of image frames at the base exposure parameters.

[0058] Continuing with the example of target frame A and target frame B from the previous text, more image frames can be captured for target frame A under more basic exposure parameters. For example, target frame A includes 7 semantic regions, so 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 8 image frames at exposure level EV0 can be captured. For target frame B, fewer image frames can be captured under fewer basic exposure parameters. For example, target frame B includes 2 semantic regions, so 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 2 image frames at exposure level EV0 can be captured.

[0059] In some embodiments, the number of semantic regions is positively correlated with the number of image frames in the synthesized image and the number of image frames under the base exposure parameters. That is, the more semantic regions there are, the more image frames are collected for synthesizing the image, and the more image frames are under the base exposure parameters; the fewer semantic regions there are, the fewer image frames are collected for synthesizing the image, and the fewer image frames are under the base exposure parameters.

[0060] In other embodiments, the imaging parameters are determined based on the semantic parameters of the target frame, specifically based on the number of semantic tag categories. For example, the electronic device determines the imaging parameters based on the number of semantic tag categories and a second mapping relationship. The second mapping relationship represents a one-to-one correspondence between the number of semantic tag categories and multiple imaging parameters. The imaging parameters can represent the number of image frames in the synthesized image. For example, the second mapping relationship can be 1 semantic tag corresponding to 4 image frames, 2 semantic tags corresponding to 4 image frames, 3 semantic tags corresponding to 5 image frames, 4 semantic tags corresponding to 5 image frames, and 5 semantic tags corresponding to 7 image frames. Alternatively, the imaging parameters can represent the number of image frames under different exposure parameters. For example, the second mapping relationship could be: 1 semantic tag corresponding to 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 2 image frames at exposure level EV0; 2 semantic tags corresponding to 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 2 image frames at exposure level EV0; 3 semantic tags corresponding to 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 4 image frames at exposure level EV0; 4 semantic tags corresponding to 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 4 image frames at exposure level EV0; 5 semantic tags corresponding to 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 5 image frames at exposure level EV0. This example uses 1 image frame each at exposure level EV- and EV+. It should be understood that the number of image frames at exposure levels EV- and EV+ can be designed as needed, as long as the number of semantic tag categories is positively correlated with the number of image frames under the basic exposure parameters.

[0061] In some embodiments, the more categories of semantic tags, the richer the representation of the subjects in the current shooting scene. Similarly, to fully capture image information of different subjects in the current shooting scene, a larger number of image frames can be acquired to enhance the clarity and fidelity of the captured image. Conversely, the fewer categories of semantic tags, the more singular the representation of the subjects in the current shooting scene, and the number of image frames required only needs to meet basic imaging requirements. In other words, the number of semantic tag categories is positively correlated with the number of image frames required to synthesize the captured image.

[0062] Continuing with the example of target frame A and target frame B from the previous text, more image frames can be acquired for target frame A. For example, target frame A corresponds to 7 semantic tags, so 10 image frames can be acquired. For target frame B, in order to reduce data processing and improve shooting response speed while ensuring image quality, fewer image frames can be acquired. For example, target frame B corresponds to 2 semantic tags, so 4 image frames can be acquired.

[0063] In some embodiments, image frames with base exposure parameters serve as reference frames for image synthesis. For example, image frames with base exposure parameters possess advantages such as reasonable exposure, complete image information, moderate noise levels, and stable overall image quality. Therefore, they are used as reference frames in the image synthesis process for registration, fusion, noise reduction, and detail enhancement with image frames under different exposure parameters, thereby ensuring the overall quality and stability of the synthesized image. Similarly, the more categories of semantic tags there are, the richer the subjects in the current shooting scene, allowing for the acquisition of more image frames with base exposure parameters. Conversely, fewer categories of semantic tags indicate a more singular subject in the current shooting scene, requiring only a sufficient number of image frames with base exposure parameters to meet basic imaging requirements. Therefore, the number of semantic tag categories is positively correlated with the number of image frames with base exposure parameters.

[0064] Continuing with the example of target frame A and target frame B from the previous text, for target frame A, more image frames can be captured under the basic exposure parameters. For example, target frame A corresponds to 7 semantic tags, so 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 8 image frames at exposure level EV0 can be captured. For target frame B, in order to reduce data processing and improve shooting response speed while ensuring image quality, fewer image frames can be captured. For example, target frame B corresponds to 2 semantic tags, so 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 2 image frames at exposure level EV0 can be captured.

[0065] In some embodiments, the number of semantic tag categories is positively correlated with the number of image frames in the synthesized image and the number of image frames under the base exposure parameters. That is, the more semantic tag categories there are, the more image frames are collected for synthesizing the image, and the more image frames are under the base exposure parameters; the fewer semantic tag categories there are, the fewer image frames are collected for synthesizing the image, and the fewer image frames are under the base exposure parameters.

[0066] In some embodiments, after determining the shooting parameters based on the semantic parameters of the target frame, the shooting parameters can also be adjusted by combining the semantic regions within the focus area of ​​the target frame. For example, as shown... Figure 2As shown, S101 can be replaced by S201.

[0067] S201, in response to the shooting command, obtains the semantic region included in the focus area of ​​the target frame and the semantic label corresponding to the semantic region.

[0068] The focus area is the region within the target frame used for focus calculations and determining focus sharpness; it is the sharpest area in the target frame. The focus area can be the entire target frame or a specific region. The focus area can be selected by the user; for example, a user can specify a specific region in the image frame as the focus area through touch controls in the camera application's preview interface, causing the camera to focus on the user-specified area. The focus area can also be automatically selected by the electronic device; for example, the electronic device can automatically analyze the shooting scene, identify the subject, and then use the area containing the subject as the focus area for focusing.

[0069] For example, an electronic device can obtain the semantic segmentation result of a target frame through a preset model, as described above. The electronic device can obtain the position and range of the focus area of ​​the target frame. For instance, it can obtain the focus position and / or focus range specified by the user in the preview interface of a camera application. If only the user-specified focus position is obtained, the electronic device uses that focus position as the center and selects an area of ​​preset size or shape as the focus area. For another example, in autofocus mode, the electronic device uses a focus algorithm to perform sharpness statistics on the target frame, determines the area with the highest sharpness as the focus position, and uses that focus position as the center and selects an area of ​​preset size or shape as the focus area. Then, based on the semantic segmentation result of the target frame and the position and range of the focus area, the electronic device obtains the semantic regions included in the focus area of ​​the target frame and the corresponding semantic tags. For example, the electronic device first filters out pixels in the target frame located within the focus area based on the position and range of the focus area. For pixels located within the focus area, the electronic device, based on the semantic segmentation results of these pixels, aggregates spatially adjacent pixels of the same category into semantic regions, and uses the category as the semantic label corresponding to the semantic region. In this way, the electronic device can obtain the semantic regions within the focus area of ​​the target frame and the semantic labels corresponding to the semantic regions.

[0070] S202, determine the image capture parameters based on the semantic parameters of the target frame.

[0071] The electronic device determines the image capture parameters as either the first image capture parameter or the second image capture parameter based on the semantic parameters of the target frame. S201 is the same as S102 and can be referred to the method described above, so it will not be repeated here. Next, the electronic device can also adjust the determined image capture parameters based on the target semantic region of the target frame to obtain the final image capture parameters. For example, the electronic device can execute S203 or S204.

[0072] S203, if the target semantic region of the target frame satisfies the target condition, determine the frame number represented by the shooting parameters as the first frame number.

[0073] S204, If the target semantic region of the target frame does not meet the target conditions, determine the frame number represented by the shooting parameters as the second frame number.

[0074] In some embodiments, the target semantic region is the semantic region with the largest area located within the focus area of ​​the target frame; in other words, the target semantic region is the semantic region with the largest area within the focus area of ​​the target frame. For example, after acquiring the semantic region of the focus area of ​​the target frame and the semantic tag corresponding to the semantic region, the electronic device can determine the semantic region with the largest area from at least one semantic region included in the focus area as the target semantic region.

[0075] In some embodiments, the target condition includes at least one of the following: the proportion of the target semantic region to the target frame exceeds a preset proportion threshold, and the area of ​​the target semantic region is not less than an area threshold. The "the proportion of the target semantic region to the target frame exceeds the preset proportion threshold" can be understood as: the area of ​​the target semantic region to the total area of ​​the target frame exceeds a preset proportion. The preset proportion could be, for example, 70%, 80%, or 90%. The "the area of ​​the target semantic region is not less than an area threshold" can be understood as: the area of ​​the target region is greater than or equal to a preset area threshold.

[0076] In some embodiments, the first frame count is less than the second frame count. Specifically, the first frame count and the second frame count can be the total number of frames in the synthesized image, or the first frame count and the second frame count can be the number of image frames at the base exposure level.

[0077] For example, if the target semantic region satisfies the target condition, it means that there is a semantic region with a large proportion and prominent subject within the focus area, and the semantic content within the focus area is simple, clearly structured, and concentrated in elements. In this case, the electronic device can reduce the number of frames in the determined shooting parameters to obtain the final shooting parameters, such as the third shooting parameters. The number of frames represented by the third shooting parameters is the first frame number.

[0078] For example, an electronic device reduces the total number of frames in the first or second shooting parameters to obtain the third shooting parameters. Another example is reducing the number of frames at the base exposure parameter in the first or second shooting parameters to obtain the third shooting parameters. Taking a shooting parameter (first or second) of capturing 4 image frames as an example, the third shooting parameter could be capturing 2 or 3 image frames. Taking a shooting parameter of capturing 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 4 image frames at exposure level EV0 as an example, the third shooting parameter could be capturing 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 2 or 3 image frames at exposure level EV0.

[0079] If the target semantic region does not meet the target conditions, it indicates that there is no prominent main semantic region within the focus area. The semantic regions are numerous, small in area, scattered, and diverse in content, with complex semantic structures and chaotic elements within the focus area. In this case, the electronic device can use the determined imaging parameters as the final imaging parameters, i.e., the fourth imaging parameters. That is, the electronic device uses the first or second imaging parameters determined based on the semantic parameters of the target frame as the fourth imaging parameters. The fourth imaging parameter represents the number of frames in the synthesized image, which is the second frame number.

[0080] Alternatively, the electronic device can increase the number of frames in the determined shooting parameters to obtain a fourth shooting parameter. The electronic device can increase the total number of frames in the first or second shooting parameters to obtain the fourth shooting parameter. For example, the electronic device can increase the number of frames under the base exposure parameters in the first or second shooting parameters to obtain the fourth shooting parameter. Taking the shooting parameters (first or second shooting parameters) as capturing 4 image frames as an example, the fourth shooting parameter could be capturing 5 or 6 image frames. Taking the shooting parameters as capturing 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 4 image frames at exposure level EV0 as an example, the fourth shooting parameter could be capturing 1 image frame at exposure level EV+, 1 image frame at exposure level EV-, and 5 or 6 image frames at exposure level EV0.

[0081] In this embodiment, the shooting parameters are first initially determined based on the semantic parameters of the target frame, and then the shooting parameters are further adjusted in combination with the largest semantic region within the focus area. This enables two-level fine adjustment of the shooting parameters, so that the final shooting parameters can not only reflect the semantic complexity of the overall scene, but also fully combine the salience of the main target within the focus area, further improving the matching degree between the shooting parameters and the current shooting scene. This ensures image clarity while improving the accuracy and adaptability of the imaging effect.

[0082] In some embodiments, the first or second shooting parameters can be adjusted based on ambient light brightness information. For example, the electronic device can acquire ambient light brightness information, such as the ambient light brightness value, in the current shooting scene. If the ambient light brightness value is greater than a preset brightness value, the number of image frames under the low exposure parameter in the first or second shooting parameters is increased. If the ambient light brightness value is not greater than the preset brightness value, the number of image frames under the high exposure parameter in the first or second shooting parameters is increased. Wherein, an ambient light brightness value greater than the preset brightness value indicates sufficient light in the current shooting scene, such as during the day. In this case, acquiring more image frames under the low exposure parameter helps improve the dynamic range, sharpness, and color reproduction accuracy of the final image, thereby obtaining a more detailed and nuanced photographic image in strong light or high brightness scenes. An ambient light brightness value not greater than the preset brightness value indicates insufficient light in the current shooting scene, such as at night or on a cloudy day. In this case, acquiring more image frames under the high exposure parameter helps improve the brightness, sharpness, and shadow detail of the image.

[0083] Similarly, the third or fourth shooting parameters can be adjusted based on ambient light brightness information. When the ambient light brightness is greater than a preset brightness value, the number of image frames under the low exposure parameter in the third or fourth shooting parameters is increased. When the ambient light brightness is not greater than the preset brightness value, the number of image frames under the high exposure parameter in the third or fourth shooting parameters is increased.

[0084] The following describes a method and apparatus for synthesizing images provided in this application, using a mobile phone as an example and specific embodiments. Existing photo-taking processes are adapted to a fusion algorithm, but the operation of this algorithm consumes significant CPU, GPU, and memory resources, posing a considerable challenge to the phone's performance and heat dissipation. Current solutions involve capturing a fixed number of image frames during photo taking, and then synthesizing them into a single image using a fusion algorithm.

[0085] This application utilizes a neural network model to employ different acquisition strategies for different images, reducing imaging time and overall power consumption, thereby improving the user experience of taking photos. The core of this application is: using a neural network model to divide the image, such as the target frame mentioned above, into blocks, i.e., to divide semantic regions, and adjusting the acquisition strategy based on the block information.

[0086] Mobile phones typically include a camera application, a camera hardware abstraction layer, and a camera. The camera, for example, may include an image acquisition sensor. This image acquisition sensor is used to capture image frames, which may be, for example, raw (RAW) images. The following section combines... Figure 3 This application introduces a method for synthesizing images according to an embodiment.

[0087] S301, the camera application displays a preview interface.

[0088] S302, the camera application receives the photo-taking command.

[0089] S303, the camera application sends acquisition commands to the camera hardware abstraction layer.

[0090] This acquisition command instructs the camera hardware abstraction layer to drive the camera to acquire image frames.

[0091] S304, the camera hardware abstraction layer determines the semantic regions included in the target frame and the semantic tags corresponding to the semantic regions.

[0092] The camera hardware abstraction layer can input a target frame into a preset model to obtain the semantic segmentation result of the target frame. Based on the semantic segmentation result, it determines the semantic regions included in the target frame and the semantic labels corresponding to those regions. In some embodiments, the camera hardware abstraction layer can further determine the semantic regions included in the focus area of ​​the target frame and the semantic labels corresponding to those regions. Please refer to the preceding description; further details are omitted here. For example, as shown... Figure 4 As shown, the target frame can include 12 semantic regions.

[0093] S305, the camera hardware abstraction layer determines the shooting parameters based on the number of semantic regions and / or the number of semantic tag types.

[0094] For example, the camera hardware abstraction layer determines the imaging parameters as either the first imaging parameter or the second imaging parameter based on the number of semantic regions and / or the number of semantic tag types. Furthermore, the camera hardware abstraction layer can also adjust the frame number of the image frame in the first or second imaging parameter based on the target semantic region of the target frame to obtain the final imaging parameters, such as the third or fourth imaging parameter.

[0095] S306, the camera hardware abstraction layer sends the shooting parameters to the camera through the driver.

[0096] The camera hardware abstraction layer sends the shooting parameters to the camera through the driver. These shooting parameters can be, for example, one of the first shooting parameters, the second shooting parameters, the third shooting parameters, or the fourth shooting parameters.

[0097] S307, the camera captures image frames according to the shooting parameters.

[0098] S308: The camera sends image frames acquired based on the image capture parameters to the camera hardware abstraction layer via the driver.

[0099] The camera captures image frames according to the shooting parameters and sends the captured image frames to the camera hardware abstraction layer.

[0100] S309, the camera hardware abstraction layer calls a fusion algorithm to synthesize the captured image.

[0101] The camera hardware abstraction layer can also compress captured images into storage formats such as JPEG and store them in a preset area.

[0102] In conventional technology, when a camera application sends a capture request to the camera hardware abstraction layer (HAL), the current approach is for the camera application to send a capture request representing a fixed frame to the HAL. Upon receiving this request, the HAL then requests the corresponding number and required number of frames from the driver. After receiving these frames, the HAL sends them back to the algorithm for processing, and finally compresses them into a JPEG. In this embodiment, the capture request sent by the camera application to the HAL may not carry image frame capture information. The HAL adaptively generates image frame capture information based on the semantic complexity of the target frame and instructs the camera to capture image frames based on this information. This balances image quality and processing efficiency.

[0103] The method embodiments of this application have been described above with reference to specific examples and accompanying drawings. The apparatus embodiments of this application will now be described below with reference to the accompanying drawings. For example, as shown... Figure 5 As shown, the image synthesis device 500 may include a processor 510 and an image acquisition device 520. The image acquisition device 520 may be, for example, the camera mentioned above.

[0104] The processor 510 is used to respond to a photo capture command to acquire the semantic regions included in the target frame and the semantic tags corresponding to the semantic regions; the pixels contained in the same semantic region correspond to a semantic tag, and the semantic tags between adjacent semantic regions are different; based on the semantic parameters of the target frame, the photo capture parameters are determined; the photo capture parameters include at least the number of image frames representing the synthesized photo capture image; wherein, the semantic parameters represent the complexity of the semantic regions and / or semantic tags.

[0105] The processor 510 is also used to control the image acquisition device 520 to acquire image frames according to the photographing parameters, and to synthesize a photographed image based on the image frames acquired by the image acquisition device.

[0106] In some embodiments, the processor 510 is further configured to respond to a photo capture command by inputting the target frame into a preset model to obtain a semantic segmentation result; the semantic segmentation result characterizes the category corresponding to each pixel in the target frame; based on the semantic segmentation result, pixels corresponding to the same category and spatially adjacent are aggregated into a semantic region, and the category is used as the semantic label corresponding to the semantic region.

[0107] In some embodiments, the processor 510 is further configured to, in response to a shooting command, acquire the semantic region included in the focus area of ​​the target frame and the semantic tag corresponding to the semantic region.

[0108] In some embodiments, the processor 510 is further configured to, if the target semantic region of the target frame satisfies the target condition, use a first frame number to represent the number of frames in the shooting parameters; and if the target semantic region of the target frame does not satisfy the target condition, use a second frame number to represent the number of frames in the shooting parameters; the target semantic region is the semantic region with the largest area located within the focus area, and the first frame number is different from the second frame number.

[0109] In some embodiments, the processor 510 is further configured to determine the current shooting scene based on the semantic parameters of the target frame; if the current shooting scene represents a complex shooting scene, determine the shooting parameters as the first shooting parameters; if the current shooting scene represents a simple shooting scene, determine the shooting parameters as the second shooting parameters.

[0110] In some embodiments, the processor 510 is further configured to determine the number of semantic regions included in the target frame; and to determine the number of semantic tag categories based on the semantic tags corresponding to the target frame.

[0111] In some embodiments, the processor 510 is further configured to determine that the current shooting scene represents a simple shooting scene if the number of semantic regions is less than a first threshold; determine that the current shooting scene represents a complex shooting scene if the number of semantic regions is not less than the first threshold; and / or determine that the shooting scene represents a simple shooting scene if the number of semantic tag categories is less than a second threshold; and determine that the shooting scene represents a complex shooting scene if the number of semantic tag categories is not less than the second threshold.

[0112] In some embodiments, the processor 510 is further configured to determine photographing parameters based on the number of semantic regions and a first mapping relationship; wherein the first mapping relationship represents a one-to-one correspondence between the number of multiple semantic regions and multiple photographing parameters; the number of semantic regions is positively correlated with the number of image frames in the synthesized photographed image, and / or the number of semantic regions is positively correlated with the number of image frames under the basic exposure parameters; or, the processor 510 determines photographing parameters based on the number of semantic label categories and a second mapping relationship; wherein the second mapping relationship represents a one-to-one correspondence between the number of multiple semantic label categories and multiple photographing parameters; the number of semantic label categories is positively correlated with the number of image frames in the synthesized photographed image, and / or the number of semantic label categories is positively correlated with the number of image frames under the basic exposure parameters.

[0113] Optional, such as Figure 5 As shown, device 500 may further include memory 530, on which a program is stored, which can be executed by processor 510 to cause processor 510 to perform the methods described in the preceding method embodiments. Memory 530 may be independent of processor 510 or integrated into processor 510.

[0114] Optionally, device 500 may also include transceiver 540. Processor 510 can communicate with other devices or chips via transceiver 540. For example, processor 510 can send and receive data with other devices or chips via transceiver 540.

[0115] This application provides a computer-readable storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the methods described in any of the above embodiments.

[0116] It should be noted that the descriptions of the computer-readable storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the computer-readable storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0117] The aforementioned processor can be at least one of the following: application-specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field-programmable gate array (FPGA), central processing unit (CPU), controller, microcontroller, and microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.

[0118] The aforementioned computer-readable storage medium / memory can be a read-only memory, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD ROM), etc.

[0119] This application provides a computer program including computer-readable code. When the computer-readable code runs in an electronic device, the processor in the electronic device executes some or all of the steps in the above-described method.

[0120] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer-readable storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0121] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0122] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0123] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0124] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0125] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0126] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0127] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an in-vehicle terminal (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0128] The above are merely embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

[0129] The above embodiments are merely preferred embodiments provided to fully illustrate this application, and the scope of protection of this application is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on this application are all within the scope of protection of this application.

Claims

1. A method for synthesizing images, applied to an electronic device, the method comprising: In response to a photo capture command, the semantic regions included in the target frame and the semantic tags corresponding to the semantic regions are acquired; pixels contained in the same semantic region correspond to one semantic tag, and the semantic tags between adjacent semantic regions are different; Based on the semantic parameters of the target frame, the image capture parameters are determined; the image capture parameters include at least the number of image frames representing the synthesized image capture; wherein, the semantic parameters represent the complexity of the semantic region and / or the semantic tag. The captured image is synthesized based on the image frames obtained from the captured parameters.

2. The method according to claim 1, wherein the semantic parameters include the number of semantic regions and / or the number of categories of the semantic tags.

3. The method according to claim 1, wherein the photographing parameters further include the number of image frames under different exposure parameters.

4. The method according to claim 1, wherein the step of acquiring the semantic region included in the target frame and the semantic tag corresponding to the semantic region in response to the photo-taking command includes: In response to the photo capture command, the target frame is input into a preset model to obtain a semantic segmentation result; The semantic segmentation result represents the category corresponding to each pixel in the target frame; Based on the semantic segmentation results, pixels that correspond to the same category and are spatially adjacent are aggregated into semantic regions, and the category is used as the semantic label corresponding to the semantic region.

5. The method according to claim 1, wherein the step of acquiring the semantic region included in the target frame and the semantic tag corresponding to the semantic region in response to the photo-taking command includes: In response to a photo capture command, the semantic region included in the focus area of ​​the target frame and the semantic tag corresponding to the semantic region are obtained.

6. The method according to claim 5, further comprising: If the target semantic region of the target frame satisfies the target condition, the frame number represented by the shooting parameters is the first frame number; If the target semantic region of the target frame does not meet the target conditions, the frame number represented by the shooting parameters is the second frame number; The target semantic region is the semantic region with the largest area located within the focus area, and the first frame number is different from the second frame number.

7. The method of claim 6, wherein the first number of frames is less than the second number of frames, and the target condition is selected from one of the following combinations: The proportion of the target semantic region in the target frame exceeds a preset proportion threshold; The area of ​​the target semantic region is not less than the area threshold.

8. The method according to any one of claims 1-7, wherein determining the image capture parameters based on the semantic parameters of the target frame includes: Based on the semantic parameters of the target frame, the current shooting scene is determined; If the current shooting scene represents a complex shooting scene, the shooting parameters are determined to be the first shooting parameters; If the current shooting scene represents a simple shooting scene, the shooting parameters are determined to be the second shooting parameters.

9. The method according to any one of claims 1-7, wherein the target frame is an image frame acquired at a first moment, the first moment being prior to the second moment of receiving the photo-taking instruction, and the duration between the first moment and the second moment being less than a preset duration.

10. An apparatus for synthesizing images, comprising a processor and an image acquisition device; The processor is used to respond to a photo capture command to acquire the semantic region included in the target frame and the semantic tag corresponding to the semantic region; the pixels contained in the same semantic region correspond to a semantic tag, and the semantic tags between adjacent semantic regions are different; Based on the semantic parameters of the target frame, the image capture parameters are determined; the image capture parameters include at least the number of image frames representing the synthesized image; wherein, The semantic parameters characterize the complexity of the semantic region and / or the semantic label; The processor is also used to control the image acquisition device to acquire image frames according to the photographing parameters, and to synthesize a photographed image based on the image frames acquired by the image acquisition device.