Shooting method, equipment and medium

Through the combination of infrared sensors and image sensors, the target motion level information of the stage scene is determined and the exposure time is adjusted, which solves the problem of poor image quality in the stage scene and achieves high-quality shooting in complex light environments.

CN120302160APending Publication Date: 2025-07-11VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510473127.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The high dynamics and lighting complexity of stage scenes have an impact on mobile device shooting, resulting in poor image quality. The existing exposure strategy for calculating motion amplitude based on preview frames is inaccurate under the influence of strobe lights and fast moving laser beams.

Method used

Using the combination of infrared sensors and image sensors, the characteristics are fused through the cross attention mechanism to determine the target motion level information, and then adjust the exposure time to improve image quality.

Benefits of technology

Reduce light interference and improve the quality of the captured images, especially in stage scenes, especially in highly dynamic scenes such as concerts and dramas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302160A_ABST
    Figure CN120302160A_ABST
Patent Text Reader

Abstract

The invention discloses a shooting method, equipment and a medium, and belongs to the technical field of camera shooting. The shooting method applied to first electronic equipment comprises the following steps: receiving target motion level information sent by second electronic equipment; wherein the target motion grade information is determined by the second electronic equipment according to a first image frame sequence obtained by carrying out image acquisition on a shooting scene by an image sensor and a second image frame sequence obtained by carrying out image acquisition on the shooting scene by an infrared sensor; determining a target exposure duration according to the target motion level information; and shooting the shooting scene based on the target exposure duration to obtain a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of imaging technology, and specifically relates to a shooting method, device, and medium. Background Art

[0002] With the rapid development of mobile Internet and smart terminal technologies, mobile devices such as smartphones and tablets have been deeply integrated into people's lives, and the rapid iteration of their shooting performance is reshaping the way of imaging content creation.

[0003] In the field of stage performing arts, using mobile devices to shoot performance content by the audience has become an important part of the interactive experience: on the one hand, the popularity of social media has driven a surge in users' demand for sharing high-definition direct stage shots and close-up shots; on the other hand, professional performers also rely on the immediacy and portability of mobile device shooting to capture rehearsal movements and review performance effects. Some mobile devices are equipped with telephoto lenses to meet the needs of users' telephoto shooting. However, the high dynamic range of stage scenes (e.g., fast movement, multiple targets intersecting), the complexity of lighting (e.g., stroboscopic lights, follow spot switching), and the interference of shooting device jitter will affect the shooting of mobile devices, resulting in poor quality of the captured images.

[0004] To improve the image quality, in related technologies, an exposure strategy based on calculating the motion amplitude from preview frames is adopted to improve the image quality. However, when shooting a stage scene, the rapid change in brightness between preview frames caused by stage stroboscopic lights and fast-moving laser beams will lead to inaccurate calculation of the motion amplitude, thus affecting the image quality. Summary of the Invention

[0005] The objective of the embodiments of this application is to provide a shooting method, device, and medium, which can reduce the interference of light during shooting by an electronic device and improve the quality of the image obtained by the electronic device.

[0006] In a first aspect, the embodiments of this application provide a shooting method, which is applied to a first electronic device. The method includes:

[0007] Receiving target motion level information sent by a second electronic device; where the target motion level information is determined by the second electronic device based on a first image frame sequence obtained by the image sensor for image acquisition of the shooting scene and a second image frame sequence obtained by the infrared sensor for image acquisition of the shooting scene;

[0008] Determining a target exposure duration according to the target motion level information;

[0009] Shooting the shooting scene based on the target exposure duration to obtain a target image.

[0010] Second aspect, an embodiment of the present application provides a shooting method, which is applied to a second electronic device. The second electronic device includes an image sensor and an infrared sensor. The method includes:

[0011] Obtain a first image frame sequence obtained by the image sensor collecting images of the shooting scene and a second image frame sequence obtained by the infrared sensor collecting images of the shooting scene;

[0012] Determine target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence;

[0013] Send the target motion level information to a first electronic device; wherein, the target motion level information is used for the first electronic device to determine the exposure duration and shoot the shooting scene based on the exposure duration.

[0014] Third aspect, an embodiment of the present application provides a first electronic device, including:

[0015] A receiving module, configured to receive the target motion level information sent by the second electronic device; wherein, the target motion level information is determined by the second electronic device according to a first image frame sequence obtained by the image sensor collecting images of the shooting scene and a second image frame sequence obtained by the infrared sensor collecting images of the shooting scene;

[0016] A first determination module, configured to determine a target exposure duration according to the target motion level information;

[0017] A shooting module, configured to shoot the shooting scene based on the target exposure duration to obtain a target image.

[0018] Fourth aspect, an embodiment of the present application provides a second electronic device, including:

[0019] An image sensor, configured to collect images of the shooting scene to obtain a first image frame sequence;

[0020] An infrared sensor, configured to collect images of the shooting scene to obtain a second image frame sequence;

[0021] A second acquisition module, configured to acquire the first image frame sequence and the second image frame sequence;

[0022] A second determination module, configured to determine target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence;

[0023] A sending module, configured to send the target motion level information to the first electronic device; wherein, the target motion level information is used for the first electronic device to determine the exposure duration and shoot the shooting scene based on the exposure duration to obtain a target image.

[0024] Fifth aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the shooting method provided by the embodiment of the present application are implemented.

[0025] Sixth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the shooting method provided by the embodiment of the present application are implemented.

[0026] Seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps of the shooting method provided by the embodiment of the present application.

[0027] Eighth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the steps of the shooting method provided by the embodiment of the present application.

[0028] In the embodiment of the present application, the second electronic device determines the motion level through the infrared thermal imaging map collected by the infrared sensor and the image collected by the image sensor, and then sends the determined motion level to the first electronic device. The first electronic device determines the exposure duration based on the motion level, and then performs shooting based on the exposure duration. Since the infrared thermal imaging map is not affected by the change of light, the light interference during the shooting of the first electronic device can be reduced, and thus the image quality of the image obtained by the first electronic device can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic flowchart of the shooting method applied to the first electronic device provided by the embodiment of the present application;

[0030] Figure 2 is a schematic structural diagram of the motion level information determination model provided by the embodiment of the present application;

[0031] Figure 3 is a schematic flowchart of the shooting method applied to the second electronic device provided by the embodiment of the present application;

[0032] Figure 4 is a schematic structural diagram of the first electronic device provided by the embodiment of the present application;

[0033] Figure 5 is a schematic structural diagram of the second electronic device provided by the embodiment of the present application;

[0034] Figure 6It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0035] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0036] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.

[0037] The terms "first", "second", etc. in the specification of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0038] Next, in conjunction with the accompanying drawings, the shooting methods, devices, and media provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.

[0039] Figure 1 It is a schematic flowchart of a shooting method applied to a first electronic device provided by an embodiment of the present application. The shooting method applied to the first electronic device may include:

[0040] Step 101: Receive target motion level information sent by a second electronic device; wherein, the target motion level information is determined by the second electronic device according to a first image frame sequence obtained by an image sensor for image acquisition of a shooting scene and a second image frame sequence obtained by an infrared sensor for image acquisition of the shooting scene;

[0041] In some possible implementations of the embodiments of the present application, the first electronic device in the embodiments of the present application may be an electronic device such as a mobile phone or a tablet computer. The second electronic device in the embodiments of the present application may be an embedded camera, which is equipped with an image sensor, an infrared sensor, a processor, and a communication module. Among them, the communication module may be a Bluetooth communication module or a Wi-Fi communication module. The shooting scenarios in the embodiments of the present application include but are not limited to stage scenarios such as concerts, dramas, and dances. It can be understood that the first image frame sequence is an RGB image sequence, and the second image frame sequence is an infrared thermal imaging map sequence. The image frame sequence in the embodiments of the present application may include the current frame image and N frames of images before the current frame image, where N is a positive integer and N can be set according to actual needs.

[0042] When shooting a stage scene, the second electronic device is installed at the midline position directly in front of the stage, and the height is the same as the stage height to ensure that the entire stage can be covered.

[0043] The first electronic device can join the device cluster by scanning the stage QR code or accessing the dedicated Service Set Identifier (SSID) to complete the binding with the second electronic device.

[0044] The second electronic device can record a video stream at a certain resolution (for example, 1080P) and a certain frame rate (for example, 30 frames per second), and through the matching of Inertial Measurement Unit (IMU) data and Oriented FAST and Rotated BRIEF (ORB) features, compensate for mechanical vibrations in real time to generate a stable video sequence.

[0045] In some possible implementations of the embodiments of the present application, the second electronic device may obtain the first image frame sequence obtained by the image sensor for image acquisition of the shooting scene and the second image frame sequence obtained by the infrared sensor for image acquisition of the shooting scene, determine the target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence, and then send the determined target motion level information to the first electronic device.

[0046] In some possible implementations of the embodiments of the present application, the second electronic device may input the first image frame sequence and the second image frame sequence into a motion level information determination model to obtain the target motion level information.

[0047] In some possible implementations of the embodiments of the present application, the motion level information determination model provided by the embodiments of the present application may include a backbone network and an infrared branch network. The second electronic device may input the first image frame sequence into the backbone network to obtain first image features; input the second image frame sequence into the infrared branch network to obtain second image features; use the Cross Attention mechanism to perform feature fusion on the first image features and the second image features to obtain fused features; process the fused features using global average pooling to obtain a probability matrix; and determine the motion level information corresponding to the position with the maximum probability in the probability matrix as the target motion level information.

[0048] In some possible implementations of the embodiments of the present application, each position in the probability matrix corresponds to a motion level information. The larger the value of a certain position in the probability matrix, the higher the probability of the motion level information corresponding to that position.

[0049] In some possible implementations of the embodiments of the present application, when the second electronic device inputs the first image frame sequence into the backbone network to obtain first image features, it may use an object detection model based on deep learning (You Only Look Once, YOLO) to determine the region of interest (ROI) of each frame of the first image frame sequence; determine the mask image corresponding to the region of interest; splice the three-channel matrix corresponding to the first image frame sequence with the single-channel matrix corresponding to the mask image to obtain a spliced matrix; and input the spliced matrix into the backbone network to obtain first image features.

[0050] In some possible implementations of the embodiments of the present application, when the second electronic device inputs the second image frame sequence into the infrared branch network to obtain second image features, it may convert each frame of the second image frame sequence into a grayscale image; splice the single-channel matrices corresponding to the grayscale images and input them into the infrared branch network to obtain second image features.

[0051] Exemplarily, assume that the above N takes the value of 4, that is, both the above first image frame sequence and the second image frame sequence include 5 images. Use the YOLO model to determine the ROI corresponding to each of the 5 RGB images included in the first image frame sequence, calculate the mask corresponding to each ROI, splice the single-channel matrices corresponding to the 5 masks after the three-channel matrix of the 5 RGB images to obtain a 20-channel spliced matrix, and use this 20-channel spliced matrix as the input to the backbone network; normalize the 5 thermal imaging images included in the second image frame sequence to obtain 5 grayscale images, splice the single-channel matrices corresponding to the 5 grayscale images to obtain a 5-channel spliced matrix, and use this 5-channel spliced matrix as the input to the infrared branch network.

[0052] The input of the backbone network passes through the first residual block (Resblock) of the backbone network to obtain feature a1. The first ResBlock of the backbone network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*20, and 32 channels, a batch normalization (Batch Normalization, BN) layer, a rectified linear unit (ReLU) layer, a convolutional layer with a stride of 1, a size of 3*3*32, and 32 channels, and a convolutional layer with a stride of 2, a size of 3*3*32, and 64 channels. Feature a1 passes through the second Resblock of the backbone network to obtain feature a2. The second ResBlock of the backbone network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*64, and 64 channels, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*64, and 64 channels, and a convolutional layer with a stride of 2, a size of 3*3*64, and 128 channels. Feature a2 passes through the third Resblock of the backbone network to obtain feature a3. The third ResBlock of the backbone network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*128, and 128 channels, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*128, and 128 channels, and a convolutional layer with a stride of 2, a size of 3*3*128, and 256 channels. Feature a3 passes through the fourth Resblock of the backbone network to obtain feature a4. The fourth ResBlock of the backbone network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*256, and 256 channels, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*256, and 256 channels, and a convolutional layer with a stride of 2, a size of 3*3*256, and 256 channels.

[0053] The input of the infrared branch network passes through the first Resblock of the infrared branch network to obtain the feature b1. The first ResBlock of the infrared branch network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*5 and 32 in number, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*32 and 32 in number, and a convolutional layer with a stride of 2, a size of 3*3*32 and 64 in number. The feature b1 passes through the second Resblock of the infrared branch network to obtain the feature b2. The second ResBlock of the infrared branch network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*64 and 64 in number, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*64 and 64 in number, and a convolutional layer with a stride of 2, a size of 3*3*64 and 128 in number. The feature b2 passes through the third Resblock of the infrared branch network to obtain the feature b3. The third ResBlock of the infrared branch network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*128 and 128 in number, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*128 and 128 in number, and a convolutional layer with a stride of 2, a size of 3*3*128 and 256 in number. The feature b3 passes through the fourth Resblock of the infrared branch network to obtain the feature b4. The fourth ResBlock of the infrared branch network consists of five parts: a convolutional layer with a stride of 1, a size of 3*3*256 and 256 in number, a BN layer, a ReLU layer, a convolutional layer with a stride of 1, a size of 3*3*256 and 256 in number, and a convolutional layer with a stride of 2, a size of 3*3*256 and 256 in number.

[0054] The features a4 and b4 are fused using the cross-attention mechanism to obtain the feature a5.

[0055] The feature a5 passes through the fifth Resblock of the backbone network to obtain the feature a6. The fifth ResBlock of the backbone network consists of four parts: a convolutional layer with a stride of 1, a size of 3*3*256 and 256 in number, a BN layer, a ReLU layer, and a convolutional layer with a stride of 1, a size of 3*3*256 and 256 in number.

[0056] The features a6 and b4 are fused using the cross-attention mechanism to obtain the feature a7.

[0057] Feature a7 passes through the sixth ResBlock of the backbone network to obtain feature a8. The sixth ResBlock of the backbone network consists of four parts: a convolutional layer with a stride of 1, a size of 3*3*256, and 256 channels, a BN layer, a ReLU layer, and a convolutional layer with a stride of 1, a size of 3*3*256, and 256 channels.

[0058] Feature a8 undergoes global average pooling (GlobalAverage Pooling, GAP) and a 1*1 convolution to obtain a probability matrix with 1 row and M columns, where M is the number of motion level information. Among them, the probability matrix with 1 row and M columns includes M positions, each position corresponding to a motion level information. The larger the value of a certain position in the probability matrix, the higher the probability of the motion level information corresponding to that position. Select the motion level information with the highest probability as the target motion level information.

[0059] In the embodiment of the present application, global average pooling and 1*1 convolution can not only retain key features but also significantly reduce parameters and computational complexity, thereby significantly improving the model processing speed and efficiency.

[0060] In some possible implementations of the embodiment of the present application, the probability matrix can also be normalized using the normalization indicator function (softmax).

[0061] Figure 2 It is a schematic structural diagram of the motion level information determination model provided by the embodiment of the present application.

[0062] Among them, Figure 2 the infrared thermal image in is the second image collected by the infrared sensor, and the video frame is the first image collected by the image sensor. Figure 2 the detection box in is the ROI region.

[0063] In some possible implementations of the embodiment of the present application, before inputting the first image frame sequence and the second image frame sequence into the motion level information determination model to obtain the target motion level information, the motion level information determination model can also be trained. The process of training the motion level information determination model is described below.

[0064] First, obtain multiple frames of RGB images and the corresponding thermal imaging images for the multiple frames of RGB images. Use the YOLO model to determine the ROI regions of each frame of RGB image in the multiple frames of RGB images, where the values within the ROI regions are 1 and the values outside the ROI regions are 0. Calculate the mask corresponding to each ROI region, splice the single-channel matrices corresponding to multiple masks after the three-channel matrix of the multiple frames of RGB images to obtain a first spliced matrix, and use this first spliced matrix as the input of the backbone network; normalize the multiple frames of thermal imaging images into 8-bit grayscale images, splice the single-channel matrices corresponding to multiple grayscale images to obtain a second spliced matrix, and use this second spliced matrix as the input of the infrared branch network; use the motion level information corresponding to the last frame of RGB image in the pre-labeled multiple frames of RGB images as the expected motion level information to train the motion level information determination model.

[0065] In some possible implementations of the embodiments of the present application, the cross entropy can be used as the loss function in the motion level information determination model in the embodiments of the present application.

[0066] Step 102: Determine the target exposure duration according to the target motion level information.

[0067] In some possible implementations of the embodiments of the present application, step 102 may include: determining the exposure duration corresponding to the target motion level information according to the correspondence between the motion level information and the exposure duration as the target exposure duration.

[0068] Exemplarily, the correspondence between the motion level information and the exposure duration is shown in Table 1 below.

[0069] Table 1

[0070] Motion level Exposure duration (unit: millisecond) 1 50 2 25 3 10 4 5 5 2

[0071] Assume that the motion level information sent by the second electronic device is 3, and the target exposure duration is determined to be 10 milliseconds according to Table 1 above.

[0072] In some possible implementations of the embodiments of the present application, when determining the target exposure duration, it can also be determined in combination with the brightness. Based on this, before step 102, the shooting method applied to the first electronic device provided in the embodiments of the present application may further include: obtaining the current brightness information of the shooting scene; correspondingly, step 102 may include: determining the exposure duration corresponding to the target motion level information and the current brightness information according to the correspondence between the brightness information, the motion level information and the exposure duration as the target exposure duration.

[0073] Exemplarily, the correspondence between the brightness information, the motion level information and the exposure duration is shown in Table 2 below.

[0074] Table 2

[0075] Motion level Luminance information (unit: lux) Exposure duration (unit: millisecond) 1 300 50 1 400 40 1 500 30 2 300 25 2 400 20 2 500 15 3 300 10 3 400 9 3 500 8 4 300 5 4 400 4 4 500 3 5 300 2 5 400 1.5 5 500 1

[0076] Assume that the motion level information sent by the second electronic device is 3, and the brightness of the current shooting scene is 400 lux. According to Table 2 above, the target exposure duration is determined to be 9 milliseconds.

[0077] In some possible implementations of the embodiments of the present application, when there is no exposure duration corresponding to the current brightness information in the correspondence table of brightness information, motion level information, and exposure duration, at this time, the correspondence stored in the correspondence table of brightness information, motion level information, and exposure duration can be used, and the exposure duration corresponding to the current brightness information can be interpolated by the interpolation method.

[0078] Exemplarily, assume that the motion level information sent by the second electronic device is 3, and the current brightness of the shooting scene is 450 lux. There is no exposure duration corresponding to the motion level information 3 and the brightness 450 lux in the correspondence table shown in Table 2 above. At this time, the correspondence stored in Table 2 above can be used to interpolate and calculate the exposure duration of 8.5 milliseconds corresponding to the motion level information 3 and the brightness 450 lux.

[0079] Step 103: Shoot the shooting scene based on the target exposure duration to obtain a target image.

[0080] In the embodiments of the present application, the second electronic device determines the motion level through the infrared thermal imaging map collected by the infrared sensor and the RGB image collected by the image sensor, and then sends the determined motion level to the first electronic device. The first electronic device determines the exposure duration based on this motion level, and then shoots based on this exposure duration. Since the infrared thermal imaging map is not affected by the change of light, it can reduce the light interference during the shooting of the first electronic device, and thus improve the image quality of the image obtained by the first electronic device.

[0081] The embodiments of the present application also provide a shooting method applied to the second electronic device. The second electronic device includes an image sensor and an infrared sensor. As Figure 3 shown, Figure 3 is a schematic flowchart of the shooting method applied to the second electronic device provided by the embodiments of the present application. The shooting method applied to the second electronic device may include:

[0082] Step 301: Obtain a first image frame sequence obtained by the image sensor collecting images of the shooting scene and a second image frame sequence obtained by the infrared sensor collecting images of the shooting scene;

[0083] Step 302: Determine the target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence;

[0084] Step 303: Send the target motion level information to the first electronic device; the target motion level information is used for the first electronic device to determine the target exposure duration and capture a target image of the shooting scene based on the target exposure duration.

[0085] In the embodiment of the present application, the second electronic device determines the motion level through the infrared thermal imaging map collected by the infrared sensor and the RGB image collected by the image sensor, and then sends the determined motion level to the first electronic device. The first electronic device determines the exposure duration according to the motion level, and then performs shooting based on the exposure duration. Since the infrared thermal imaging map is not affected by the change of light, it can reduce the light interference during the shooting of the first electronic device, and thus improve the image quality of the image captured by the first electronic device.

[0086] In some possible implementations of the embodiment of the present application, step 302 may include:

[0087] Input the first image frame sequence and the second image frame sequence into the motion level information determination model to obtain the target motion level information.

[0088] In some possible implementations of the embodiment of the present application, the motion level information determination model includes: a backbone network and an infrared branch network;

[0089] Correspondingly, step 302 may include:

[0090] Input the first image frame sequence into the backbone network to obtain the first image feature;

[0091] Input the second image frame sequence into the infrared branch network to obtain the second image feature;

[0092] Use the cross-attention mechanism to fuse the first image feature and the second image feature to obtain the fused feature;

[0093] Process the fused feature by global average pooling to obtain the probability matrix;

[0094] Determine the motion level information corresponding to the position with the maximum probability in the probability matrix as the target motion level information.

[0095] In some possible implementations of the embodiment of the present application, inputting the first image frame sequence into the backbone network to obtain the first image feature may include:

[0096] Use the object detection model based on deep learning to determine the region of interest of each frame of the first image frame sequence;

[0097] Determine the mask image corresponding to the region of interest;

[0098] Concatenate the three-channel matrix corresponding to the first image frame sequence with the single-channel matrix corresponding to the mask image to obtain a concatenated matrix;

[0099] Input the concatenated matrix into the backbone network to obtain the first image feature.

[0100] In some possible implementations of the embodiments of the present application, inputting the second image frame sequence into the infrared branch network to obtain the second image feature may include:

[0101] Convert each frame of the second image frame sequence into a grayscale image;

[0102] Concatenate the single-channel matrices corresponding to the grayscale images and then input them into the infrared branch network to obtain the second image feature.

[0103] It should be noted that the implementation processes of the steps in the shooting method applied to the second electronic device provided in the embodiments of the present application can refer to the descriptions in the embodiments of the shooting method applied to the first electronic device provided in the embodiments of the present application. To avoid repetition, they will not be elaborated here.

[0104] Figure 4 It is a schematic structural diagram of the first electronic device provided in the embodiments of the present application. The first electronic device 400 may include:

[0105] A receiving module 401, configured to receive the target motion level information sent by the second electronic device; wherein, the target motion level information is determined by the second electronic device according to the first image frame sequence collected by the image sensor for the shooting scene and the second image frame sequence collected by the infrared sensor for the shooting scene;

[0106] A first determination module 402, configured to determine the target exposure duration according to the target motion level information;

[0107] A shooting module 403, configured to shoot the shooting scene based on the target exposure duration to obtain a target image.

[0108] In the embodiments of the present application, the second electronic device determines the motion level through the infrared thermal imaging map collected by the infrared sensor and the RGB image collected by the image sensor, and then sends the determined motion level to the first electronic device. The first electronic device determines the exposure duration according to this motion level, and then shoots based on this exposure duration. Since the infrared thermal imaging map is not affected by light changes, it can reduce the light interference during the shooting of the first electronic device, and thus improve the image quality of the image obtained by the first electronic device.

[0109] In some possible implementations of the embodiments of the present application, the first determination module 402 is specifically configured to:

[0110] According to the correspondence between the motion level information and the exposure duration, the exposure duration corresponding to the target motion level information is determined as the target exposure duration.

[0111] In some possible implementations of the embodiments of the present application, the first electronic device provided by the embodiments of the present application further includes:

[0112] A first acquisition module, configured to acquire the current brightness information of the shooting scene;

[0113] Correspondingly, the first determination module 402 is specifically configured to:

[0114] According to the correspondence between the brightness information, the motion level information, and the exposure duration, the exposure duration corresponding to the target motion level information and the current brightness information is determined as the target exposure duration.

[0115] The first electronic device in the embodiments of the present application may be a terminal or other devices other than the terminal. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. It may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0116] The first electronic device in the embodiments of the present application may be an electronic device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0117] The first electronic device provided by the embodiments of the present application can implement Figures 1 to 2 each process implemented by the shooting method embodiment. To avoid repetition, it will not be elaborated here.

[0118] Figure 5 is a schematic structural diagram of a second electronic device provided by the embodiments of the present application. The second electronic device 500 may include:

[0119] An image sensor 501 for collecting images of a shooting scene to obtain a first image frame sequence;

[0120] An infrared sensor 502 for collecting images of a shooting scene to obtain a second image frame sequence;

[0121] A second acquisition module 503 for acquiring the first image frame sequence and the second image frame sequence;

[0122] A second determination module 504 for determining target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence;

[0123] A sending module 505 for sending the target motion level information to a first electronic device; wherein, the target motion level information is used for the first electronic device to determine a target exposure duration and capture a target image of the shooting scene based on the target exposure duration.

[0124] In an embodiment of the present application, the second electronic device determines the motion level through an infrared thermal imaging map collected by an infrared sensor and an RGB image collected by an image sensor, and then sends the determined motion level to the first electronic device. The first electronic device determines the exposure duration according to the motion level, and then captures based on the exposure duration. Since the infrared thermal imaging map is not affected by light changes, it is possible to reduce the light interference during shooting by the first electronic device, and thus improve the image quality of the image captured by the first electronic device.

[0125] In some possible implementations of the embodiment of the present application, the second determination module 504 is specifically configured to:

[0126] Input the first image frame sequence and the second image frame sequence into a motion level information determination model to obtain the target motion level information.

[0127] In some possible implementations of the embodiment of the present application, the motion level information determination model includes: a backbone network and an infrared branch network;

[0128] The second determination module 504 may include:

[0129] A first extraction sub-module for inputting the first image frame sequence into the backbone network to obtain first image features;

[0130] A second extraction sub-module for inputting the second image frame sequence into the infrared branch network to obtain second image features;

[0131] A fusion sub-module for fusing the first image features and the second image features by using a cross-attention mechanism to obtain fusion features;

[0132] A processing sub-module, configured to process the fused features by global average pooling to obtain a probability matrix;

[0133] A determination sub-module, configured to determine the motion level information corresponding to the position with the maximum probability in the probability matrix as the target motion level information.

[0134] In some possible implementations of the embodiments of the present application, the first extraction sub-module is specifically configured to:

[0135] Use a deep learning-based object detection model to determine the region of interest of each frame of image in the first image frame sequence;

[0136] Determine a mask image corresponding to the region of interest;

[0137] Concatenate the three-channel matrix corresponding to the first image frame sequence with the single-channel matrix corresponding to the mask image to obtain a concatenated matrix;

[0138] Input the concatenated matrix into the backbone network to obtain the first image feature.

[0139] In some possible implementations of the embodiments of the present application, the second extraction sub-module is specifically configured to:

[0140] Convert each frame of image in the second image frame sequence into a grayscale image;

[0141] Concatenate the single-channel matrices corresponding to the grayscale images and then input them into the infrared branch network to obtain the second image feature.

[0142] The second electronic device in the embodiments of the present application may be an electronic device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0143] The second electronic device provided in the embodiments of the present application can implement Figure 3 each process implemented by the shooting method embodiment. To avoid repetition, it will not be elaborated here.

[0144] Optionally, as Figure 6 shown, the embodiments of the present application further provide an electronic device 600, including a processor 601 and a memory 602. A program or instruction is stored on the memory 602 and can run on the processor 601. When the program or instruction is executed by the processor 601, it implements each step of the shooting method embodiment provided in the embodiments of the present application and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0145] Figure 7 is a schematic hardware structure diagram of the electronic device in the embodiments of the present application.

[0146] The electronic device 700 includes, but is not limited to, components such as a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, and a processor 710.

[0147] Those skilled in the art can understand that the electronic device 700 may further include a power source (such as a battery) for supplying power to each component. The power source can be logically connected to the processor 710 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. Figure 7 The structure of the electronic device shown does not limit the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0148] Among them, the radio frequency unit 701 is used to: receive the target motion level information sent by the second electronic device; where the target motion level information is determined by the second electronic device based on the first image frame sequence obtained by the image sensor for image acquisition of the shooting scene and the second image frame sequence obtained by the infrared sensor for image acquisition of the shooting scene;

[0149] The processor 710 can be used to: determine the target exposure duration according to the target motion level information; and perform shooting on the shooting scene based on the target exposure duration to obtain a target image.

[0150] In the embodiments of the present application, the second electronic device determines the motion level through the infrared thermal imaging map collected by the infrared sensor and the RGB image collected by the image sensor, and then sends the determined motion level to the electronic device. The electronic device determines the exposure duration according to the motion level, and then performs shooting based on the exposure duration. Since the infrared thermal imaging map is not affected by light changes, it can reduce the light interference during shooting of the electronic device, and thus improve the image quality of the image obtained by the electronic device.

[0151] In some possible implementations of the embodiments of the present application, the processor 710 is specifically used to:

[0152] According to the correspondence between the motion level information and the exposure duration, determine the exposure duration corresponding to the target motion level information as the target exposure duration.

[0153] In some possible implementations of the embodiments of the present application, the processor 710 may also be used to obtain the current brightness information of the shooting scene; according to the correspondence between the brightness information, the motion level information, and the exposure duration, determine the exposure duration corresponding to the target motion level information and the current brightness information as the target exposure duration.

[0154] In some possible implementations of the embodiments of the present application, the sensor 705 may include an image sensor and an infrared sensor;

[0155] The image sensor is configured to collect images of the shooting scene to obtain a first image frame sequence;

[0156] The infrared sensor is configured to collect images of the shooting scene to obtain a second image frame sequence;

[0157] The processor 710 may be configured to: obtain the first image frame sequence and the second image frame sequence; determine target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence; send the target motion level information to the first electronic device; wherein, the target motion level information is used for the first electronic device to determine a target exposure duration and perform shooting on the shooting scene based on the target exposure duration to obtain a target image.

[0158] In the embodiments of the present application, the electronic device determines the motion level through the infrared thermal imaging map collected by the infrared sensor and the RGB image collected by the image sensor, and then sends the determined motion level to the first electronic device. The first electronic device determines the exposure duration according to the motion level, and then performs shooting based on the exposure duration. Since the infrared thermal imaging map is not affected by changes in light, it is possible to reduce the light interference during shooting by the first electronic device, and thus improve the image quality of the image obtained by the first electronic device.

[0159] In some possible implementations of the embodiments of the present application, the processor 710 is specifically configured to:

[0160] Input the first image frame sequence and the second image frame sequence into a motion level information determination model to obtain target motion level information.

[0161] In some possible implementations of the embodiments of the present application, the motion level information determination model includes: a backbone network and an infrared branch network; the processor 710 is specifically configured to:

[0162] Input the first image frame sequence into the backbone network to obtain first image features;

[0163] Input the second image frame sequence into the infrared branch network to obtain second image features;

[0164] Use a cross-attention mechanism to perform feature fusion on the first image features and the second image features to obtain fused features;

[0165] Process the fused features using global average pooling to obtain a probability matrix;

[0166] Determine the motion level information corresponding to the position with the maximum probability in the probability matrix as the target motion level information.

[0167] In some possible implementations of the embodiments of the present application, the processor 710 is specifically configured to:

[0168] Using an object detection model based on deep learning, determine the region of interest of each frame of the first image frame sequence;

[0169] Determine the mask image corresponding to the region of interest;

[0170] Stitch the three-channel matrix corresponding to the first image frame sequence and the single-channel matrix corresponding to the mask image to obtain a stitched matrix;

[0171] Input the stitched matrix into the backbone network to obtain the first image feature.

[0172] In some possible implementations of the embodiments of the present application, the processor 710 is specifically configured to:

[0173] Convert each frame of the second image frame sequence into a grayscale image;

[0174] Stitch the single-channel matrices corresponding to the grayscale images and input them into the infrared branch network to obtain the second image feature.

[0175] It should be understood that in the embodiments of the present application, the input unit 704 may include a Graphics Processing Unit (GPU) 7041 and a microphone 7042. The graphics processor 7041 processes the image data of static pictures or videos obtained by an image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 706 may include a display panel 7061, and the display panel 7061 may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include two parts: a touch detection device and a touch controller. The other input devices 7072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0176] The memory 709 can be used to store software programs and various data. The memory 709 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 709 may include a volatile memory or a non-volatile memory, or the memory 709 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically Erasable PROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 709 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0177] The processor 710 may include one or more processing units; optionally, the processor 710 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 710 either.

[0178] The embodiments of the present application also provide a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, the various processes of the shooting method embodiments provided by the embodiments of the present application are implemented, and the same technical effects can be achieved. To avoid repetition, details are not described here again.

[0179] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes a computer-readable storage medium. Examples of computer-readable storage media include non-transitory computer-readable media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0180] An embodiment of the present application further provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the shooting method embodiment provided by the embodiment of the present application, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0181] It should be understood that the chip mentioned in the embodiment of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0182] An embodiment of the present application further provides a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the shooting method embodiment provided by the embodiment of the present application, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0183] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.

[0184] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0185] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A shooting method, characterized in that, Applied to a first electronic device, the method includes: Receiving target motion level information sent by a second electronic device; wherein, the target motion level information is determined by the second electronic device based on a first image frame sequence obtained by an image sensor collecting images of a shooting scene and a second image frame sequence obtained by an infrared sensor collecting images of the shooting scene; Determining a target exposure duration according to the target motion level information; Shooting the shooting scene based on the target exposure duration to obtain a target image.

2. The method according to claim 1, wherein The determining the target exposure duration according to the target motion level information includes: According to the correspondence between the motion level information and the exposure duration, determining the exposure duration corresponding to the target motion level information as the target exposure duration.

3. The method according to claim 1, wherein Before the determining the target exposure duration according to the target motion level information, the method further includes: Obtaining the current brightness information of the shooting scene; The determining the target exposure duration according to the target motion level information includes: According to the correspondence between the brightness information, the motion level information and the exposure duration, determining the exposure duration corresponding to the target motion level information and the current brightness information as the target exposure duration.

4. A shooting method, characterized in that, Applied to a second electronic device, the second electronic device includes an image sensor and an infrared sensor, and the method includes: Obtaining a first image frame sequence obtained by the image sensor collecting images of a shooting scene and a second image frame sequence obtained by the infrared sensor collecting images of the shooting scene; Determining target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence; Sending the target motion level information to a first electronic device; wherein, the target motion level information is used for the first electronic device to determine a target exposure duration corresponding to the shooting scene and shoot the shooting scene based on the target exposure duration to obtain a target image.

5. The method according to claim 4, characterized in that The determining the target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence includes: Inputting the first image frame sequence and the second image frame sequence into a motion level information determination model to obtain the target motion level information.

6. The method according to claim 5, wherein The motion level information determination model includes: a backbone network and an infrared branch network; The inputting the first image frame sequence and the second image frame sequence into the motion level information determination model to obtain the target motion level information includes: Inputting the first image frame sequence into the backbone network to obtain a first image feature; Inputting the second image frame sequence into the infrared branch network to obtain a second image feature; Using a cross-attention mechanism to perform feature fusion on the first image feature and the second image feature to obtain a fused feature; Processing the fused feature by global average pooling to obtain a probability matrix; Determining the motion level information corresponding to the position with the maximum probability in the probability matrix as the target motion level information.

7. The method according to claim 6, wherein Inputting the first image frame sequence into the backbone network to obtain first image features includes: Using an object detection model based on deep learning to determine the region of interest of each frame of the first image frame sequence; Determining a mask image corresponding to the region of interest; Concatenating the three-channel matrix corresponding to the first image frame sequence and the single-channel matrix corresponding to the mask image to obtain a concatenated matrix; Inputting the concatenated matrix into the backbone network to obtain the first image features.

8. The method according to claim 6, wherein Inputting the second image frame sequence into the infrared branch network to obtain second image features includes: Converting each frame of the second image frame sequence into a grayscale image; Concatenating the single-channel matrices corresponding to the grayscale images and inputting the result into the infrared branch network to obtain the second image features.

9. A first electronic device, characterized in that, The first electronic device includes: A receiving module, configured to receive target motion level information sent by a second electronic device; wherein, the target motion level information is determined by the second electronic device according to a first image frame sequence collected by an image sensor for a shooting scene and a second image frame sequence collected by an infrared sensor for the shooting scene; A first determination module, configured to determine a target exposure duration according to the target motion level information; A shooting module, configured to shoot the shooting scene based on the target exposure duration to obtain a target image.

10. A second electronic device, characterized in that, The second electronic device includes: An image sensor, configured to collect an image of a shooting scene to obtain a first image frame sequence; An infrared sensor, configured to collect an image of the shooting scene to obtain a second image frame sequence; A second acquisition module, configured to acquire the first image frame sequence and the second image frame sequence; A second determination module, configured to determine target motion level information corresponding to the shooting scene according to the first image frame sequence and the second image frame sequence; A sending module, configured to send the target motion level information to the first electronic device; wherein, the target motion level information is used by the first electronic device to determine a target exposure duration and shoot the shooting scene based on the target exposure duration to obtain a target image.

11. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the shooting method according to any one of claims 1-8 are implemented.

12. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the shooting method according to any one of claims 1-8 are implemented.