Image processing method, electronic equipment and computer readable medium

Through inter-frame differential processing and frame rate switching technology, the problem of waste of resources and high energy consumption when shooting animals by fixed frame rate cameras is solved, and more accurate animal behavior capture and resource utilization is achieved.

CN119996604AActive Publication Date: 2025-05-13ADDX (BEIJING) TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510130861.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-13
Estimated Expiration
2045-02-05

AI Technical Summary

Technical Problem

In the prior art, when shooting animals, cameras with fixed frame rates cannot accurately distinguish different behaviors of animals, resulting in waste of resources and high energy consumption, and cannot capture picture details when the animal is moving at a high intensity, and there is duplicate information when it is still.

Method used

By acquiring the image set and initial shooting frame rate information of the target animal, inter-frame difference processing is performed to obtain the motion information, and switching the shooting frame rate based on the motion information and the frame rate information.

Benefits of technology

It improves the accuracy of the camera's frame rate switching in different scenarios, reduces resource waste and energy consumption, and reduces the damage rate of the camera.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996604A_ABST
    Figure CN119996604A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the following steps: acquiring a first animal image set and first shooting frame rate information of a target animal set; performing inter-frame difference processing on the first animal image set to obtain a motion information set of a first animal in the target animal set; and according to the motion information set of the first animal and the first shooting frame rate information, carrying out shooting frame rate switching processing to obtain shooting frame rate switching information. According to the embodiment, the accuracy of camera switching can be improved, the stability of the camera is improved, and the damage rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to an image processing method, an electronic device, and a computer-readable medium. Background Art

[0002] At present, with the development of Internet of Things technology, monitoring and shooting of pets or wild animals has become more and more common. Since monitoring and shooting usually use fixed frame rates for shooting, a lot of resources are wasted. For animal information storage based on frame rate switching cameras, the usual method is to use a fixed frame rate camera to collect animal images.

[0003] However, the inventors have found that when the above method is used to switch the camera's animal information storage based on the frame rate, the following technical problems often occur:

[0004] Due to the use of a fixed frame rate for image acquisition, the different behaviors of animals cannot be accurately distinguished and a large amount of storage resources are wasted. When the animal's movement intensity is high, the image cannot be accurately captured and a large number of details are missed. When the animal is still, the animal images taken contain a large amount of repeated information, resulting in high energy consumption and damage rate of the camera.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the invention

[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.

[0007] Some embodiments of the present disclosure propose an image processing method, an electronic device, and a computer-readable medium to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide an image processing method for camera shooting, including: acquiring a first animal image set and first shooting frame rate information of a target animal set; performing inter-frame difference processing on the first animal image set to obtain a motion information set for the first animal in the target animal set; performing shooting frame rate switching processing based on the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information.

[0009] Optionally, the method further includes: controlling the camera to shoot an animal image set for the target animal set under the shooting frame rate switching information; performing animal identification on the animal image set to obtain animal identification information; and performing time rolling update on the animal identification information and the animal image set and storing them in the target storage end.

[0010] Optionally, the above-mentioned shooting frame rate switching processing is performed according to the motion information set of the above-mentioned first animal and the above-mentioned first shooting frame rate information to obtain shooting frame rate switching information, including: performing status information filtering processing on the motion information set of the above-mentioned first animal to obtain filtered animal status information; performing shooting frame rate switching processing according to the above-mentioned filtered animal status information and the above-mentioned first shooting frame rate information to obtain shooting frame rate switching information.

[0011] Optionally, the above-mentioned shooting frame rate switching processing is performed based on the motion information set of the above-mentioned first animal and the above-mentioned first shooting frame rate information to obtain shooting frame rate switching information, including: performing intensity information filtering processing on the above-mentioned animal motion information set to obtain the filtered animal motion intensity information; in response to determining that the above-mentioned filtered animal motion intensity information is not greater than a preset motion intensity threshold, performing shooting frame rate switching processing on the above-mentioned first shooting frame rate information to obtain second shooting frame rate information as shooting frame rate switching information.

[0012] Optionally, the above-mentioned shooting frame rate switching processing is performed based on the motion information set of the above-mentioned first animal and the above-mentioned first shooting frame rate information to obtain shooting frame rate switching information, including: performing intensity information filtering processing on the above-mentioned animal motion information set to obtain the filtered animal motion intensity information; in response to determining that the above-mentioned filtered animal motion intensity information is greater than the above-mentioned preset motion intensity threshold, performing shooting frame rate switching processing on the above-mentioned first shooting frame rate information to obtain third shooting frame rate information as shooting frame rate switching information.

[0013] Optionally, the above-mentioned animal recognition is performed on the above-mentioned animal captured image set to obtain animal recognition information, including: in response to determining that the above-mentioned animal captured image set is an image set based on a first preset shooting frame rate, performing image preprocessing on the above-mentioned animal captured image set to obtain a preprocessed animal image set; inputting the above-mentioned preprocessed animal image set into the animal feature extraction network included in the animal recognition model to obtain a first animal feature information set, wherein the above-mentioned animal recognition model also includes: a multi-scale fusion network and a head recognition output network, and the above-mentioned animal feature extraction network includes: multiple hybrid attention networks, multiple first feature extraction networks, multiple maximum pooling layers, multiple cascade aggregation networks and multiple lightweight aggregation networks; inputting the above-mentioned first animal feature information set into the above-mentioned multi-scale fusion network to obtain a second animal feature information set, wherein the above-mentioned multi-scale fusion network includes: multiple first feature extraction networks, multiple second feature extraction networks, multiple lightweight aggregation networks and multiple content-aware prediction networks; inputting the above-mentioned second animal feature information set into the above-mentioned head recognition output network to obtain animal recognition information, wherein the above-mentioned head recognition output network includes: multiple third feature extraction networks.

[0014] Optionally, the above-mentioned animal recognition is performed on the above-mentioned animal captured image set to obtain animal recognition information, including: in response to determining that the above-mentioned animal captured image set is an image set based on a second preset shooting frame rate, downsampling convolution processing is performed on the above-mentioned animal captured image set to obtain a first animal posture feature information set; the first animal posture feature information set is input into the feature extraction network included in the animal posture recognition model to obtain a second posture feature information set, wherein the above-mentioned animal posture recognition model also includes: a first branch parallel fusion network, a second branch parallel fusion network, a third branch parallel fusion network and a convolution layer, and the above-mentioned feature extraction network includes: a residual convolution network, A plurality of paths are fused into a convolutional network; the second posture feature information set is input into the first branch parallel fusion network to obtain a first multi-scale posture feature information set; the first multi-scale posture feature information set is input into the second branch parallel fusion network to obtain a second multi-scale posture feature information set; the second multi-scale posture feature information set is input into the third branch parallel fusion network to obtain a third multi-scale posture feature information set; the third multi-scale posture feature information set is input into the convolutional layer to obtain animal posture recognition information as animal recognition information to determine whether the animal posture recognition information of the target animal set is flight posture information.

[0015] Optionally, the first branch parallel fusion network includes: a branch parallel expansion network, multiple multi-scale feature extraction convolution networks and multiple channel space attention networks; and the second posture feature information set is input into the branch parallel fusion network to obtain the first multi-scale posture feature information set, including: inputting the second posture feature information set into the first branch parallel expansion network to obtain a fourth multi-scale posture feature information set; inputting the fourth multi-scale posture feature information into the multiple multi-scale feature extraction convolution networks to obtain a fifth multi-scale posture feature information set; inputting the fifth multi-scale posture feature information set into the multiple channel space attention networks to obtain a sixth multi-scale posture feature information set; performing sampling and fusion processing on the sixth multi-scale posture feature information set to obtain a seventh multi-scale posture feature information set; performing nonlinear processing on the seventh multi-scale posture feature information set to obtain the first multi-scale posture feature information set.

[0016] Optionally, before obtaining the first animal image set and the first shooting frame rate information of the target animal set, the method further includes: in response to detecting infrared signal information sent by a passive infrared sensor, performing signal detection processing on the infrared signal information to obtain infrared detection information; and according to the infrared detection information, controlling the camera to shoot animal images to obtain the first animal image set for the target animal set in the shooting area.

[0017] Optionally, the method further includes: in response to determining that the infrared detection information indicates that no motion information is detected, controlling the camera to enter a standby state.

[0018] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the above-mentioned first aspect.

[0019] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any one of the implementation modes in the first aspect above is implemented.

[0020] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the image processing method of some embodiments of the present disclosure, when applied to camera shooting, can improve the accuracy of camera switching, improve the stability of the camera and reduce the damage rate. Specifically, the reasons for the high energy consumption and damage rate of the camera are: due to the use of a fixed frame rate for image acquisition, the different behaviors of the animal cannot be accurately distinguished and a large amount of storage resources are wasted, and when the animal's movement intensity is high, the picture cannot be accurately captured, a large number of details are missed, and when the animal is stationary, there is a large amount of repeated information in the animal image captured, resulting in high energy consumption and damage rate of the camera. Based on this, the image processing method of some embodiments of the present disclosure, when applied to camera shooting, can first obtain the first animal image set and the first shooting frame rate information of the target animal set. Here, the first animal image set obtained facilitates the subsequent determination of the motion information of the current target animal set, and the first shooting frame rate information obtained facilitates the subsequent switching of the shooting frame rate of the camera. Then, the above-mentioned first animal image set is subjected to inter-frame difference processing to obtain the motion information set of the first animal in the above-mentioned target animal set. Here, the motion information of the target animal can be accurately determined so as to subsequently determine the switching of the shooting frame rate of the camera. Finally, according to the motion information set of the first animal and the first shooting frame rate information, the shooting frame rate switching process is performed to obtain the shooting frame rate switching information. Here, the accuracy of the camera shooting frame rate switching in different scenes can be improved to reduce the loss of the camera and the waste of power resources. It can be obtained that the image processing method, applied to camera shooting, can accurately control the switching of the camera shooting frame rate through the motion information of the target animal, and can realize switching of different shooting frame rates in different situations to capture animal images that are more in line with the current scene, that is, it can realize the capture of more details when the target animal moves violently, reduce the collection of a large amount of repeated information when the target animal is stationary, reduce the loss of camera performance and electric energy, and reduce the damage rate of the camera. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0022] Figure 1 is a schematic diagram of an application scenario of the image processing method according to some embodiments of the present disclosure;

[0023] Figure 2 is a flow chart of some embodiments of the image processing method according to the present disclosure;

[0024] Figure 3is a schematic diagram of an image captured of an animal in some embodiments of the image processing method disclosed herein;

[0025] Figure 4 is a schematic diagram of a cascade aggregation network in some embodiments of the image processing method according to the present disclosure;

[0026] Figure 5 are flowcharts of other embodiments of the image processing method according to the present disclosure;

[0027] Figure 6 It is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0029] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0030] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0031] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0033] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0034] Figure 1 It is a schematic diagram of an application scenario applied to camera shooting according to the image processing method of some embodiments of the present disclosure.

[0035] exist Figure 1In the application scenario, the electronic device 101 can first obtain the first animal image set 102 and the first shooting frame rate information 103 of the target animal set. Then, the inter-frame difference processing is performed on the first animal image set 102 to obtain the motion information set 104 of the first animal in the target animal set. Finally, according to the motion information set 104 of the first animal and the first shooting frame rate information 103, the shooting frame rate switching processing is performed to obtain the shooting frame rate switching information 105.

[0036] It should be noted that the electronic device 101 can be hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or it can be implemented as a single server or a single terminal device. When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, for example, or it can be implemented as a single software or software module. No specific limitation is made here.

[0037] It should be understood that Figure 1 The number of electronic devices in the embodiment is only for illustration. Any number of electronic devices may be provided according to implementation requirements.

[0038] Continue to refer Figure 2 , shows a process 200 of some embodiments of the image processing method according to the present disclosure, which is applied to camera shooting. The image processing method, applied to camera shooting, comprises the following steps:

[0039] Step 201: Acquire a first animal image set and first shooting frame rate information of a target animal set.

[0040] In some embodiments, the above image processing method is applied to an execution subject (for example, Figure 1 The electronic device 101 shown) can obtain the first animal image set and the first shooting frame rate information of the target animal set through a wired connection method or a wireless connection method. Among them, the above-mentioned camera can be a camera with shooting frame rate switching for shooting the target animal set. The target animals in the above-mentioned target animal set can be animals that are close to the feeder and captured by the camera. For example, the above-mentioned target animals can be birds. The first animal image in the above-mentioned first animal image set can be an image of the target animal set shot by the camera at the current moment with the shooting frame rate corresponding to the above-mentioned first shooting frame rate information. The above-mentioned first shooting frame rate information can be information on the number of images that can be shot per second when the camera shoots the target animal at the current moment.

[0041] In some optional implementations of some embodiments, before obtaining the first animal image set and the first shooting frame rate information of the target animal set, the method may further include the following steps:

[0042] In the first step, in response to detecting the infrared signal information sent by the passive infrared sensor, the infrared signal information is subjected to signal detection processing to obtain infrared detection information. The passive infrared sensor may be a sensor that determines the degree of motion of the target animal through infrared signals. The infrared signal information may be information of infrared signals emitted and collected by the passive infrared sensor. The infrared detection information may represent information of the degree of motion of the target animal.

[0043] As an example, the execution subject may first filter the infrared signal information to obtain filtered infrared signal information, and then perform spectrum analysis on the filtered infrared signal information using a Fourier transform algorithm to obtain infrared frequency information as infrared detection information.

[0044] The second step is to control the camera to shoot animal images according to the infrared detection information to obtain a first animal image set for the target animal set in a shooting area, wherein the shooting area may be an area that can be shot by the camera.

[0045] As an example, the above-mentioned execution entity can determine the current shooting frame rate information as the shooting frame rate of the above-mentioned camera in response to determining that the infrared frequency corresponding to the above-mentioned infrared detection information is greater than or equal to a preset frequency threshold, and control the above-mentioned camera to capture animal images at the shooting frame rate corresponding to the above-mentioned current shooting frame rate information to obtain a first animal image set for the above-mentioned target animal in the shooting area.

[0046] Step 202: perform inter-frame difference processing on the first animal image set to obtain a motion information set for the first animal in the target animal set.

[0047] In some embodiments, the execution subject may perform inter-frame difference processing on the first animal image set to obtain a motion information set for the first animal in the target animal set. The animal motion information in the animal motion information set may represent information about the motion of the target animal. For example, the animal motion information set may include but is not limited to at least one of the following: stillness, flight, standing, and motion intensity. The first animal may be any animal in the target animal set.

[0048] As an example, the execution subject may use a visual algorithm to determine a pixel difference set of any adjacent first animal images in the first animal image set for the target animal set as the animal motion information set. The visual algorithm may be an algorithm that uses a target detection model to identify and track the position of the target animal.

[0049] In the process of adopting technical solutions to solve the above-mentioned technical problem 1, the following technical problem 2 is often accompanied: when determining the motion intensity of the target animal through the inter-frame pixel difference of the first animal image set, the background image and light will have an impact on the inter-frame difference calculation, resulting in low accuracy of the inter-frame difference calculation, which in turn causes loss and damage to the performance of the camera. For the above-mentioned technical problem 2, the conventional solution is generally: using the GrabCut algorithm to determine the animal foreground image set and the animal background image set of the first animal image set, and then determining the sum of the pixel differences of the animal foreground image set and the animal background image set as the animal motion intensity information set. However, the above conventional solution still has the following technical problems: Since the GrabCut algorithm only uses the foreground and background images with significant texture and color differences between the first animal images of the front and back frames for segmentation, when the difference between the foreground pixels and the background pixels is small or the background is more complex, the segmentation accuracy of the image segmentation is low, and the operation time of the GrabCut algorithm is long, resulting in an increase in the segmentation operation time of the image segmentation, and there is a large amount of redundant information of the foreground pixels and background pixels, resulting in a long time for the inter-frame difference operation, and the accuracy of the calculation of the animal motion intensity information is low, resulting in low stability of the camera performance, a large amount of camera resources wasted and an increase in the damage rate of the camera. The inventor, taking into account the shortcomings of the conventional solution and combining the advantages / technical status of the inter-frame difference algorithm owned by the inventor, can decide to adopt the following solution:

[0050] In some optional implementations of some embodiments, performing inter-frame difference processing on the first animal image set to obtain an animal motion intensity information set for the target animal set may include the following steps:

[0051] In the first step, for each first animal image in the first animal image set, the following foreground and background segmentation steps are performed:

[0052] Sub-step 1, determining an animal depth image of the first animal image. Wherein, each pixel value in the animal depth image can represent a distance value from the camera. The execution subject can determine the animal depth image for the first animal image through a depth sensor.

[0053] Sub-step 2, performing morphological filtering on the above-mentioned animal depth image to obtain a filtered animal depth image. In practice, the above-mentioned execution subject can use the morphological opening and closing operation algorithm to first grayscale the above-mentioned animal depth image and then perform morphological filtering to obtain a filtered animal depth image. It should be noted that the opening operation in the morphological opening and closing operation algorithm can remove part of the background noise in the above-mentioned animal depth image, so that the contour of the obtained animal depth image is smoother. The closing operation in the above-mentioned morphological opening and closing operation algorithm can fill the broken parts in the contour of the above-mentioned animal depth image. Therefore, the above-mentioned morphological filtering process can effectively avoid the problem of over-segmentation of the image caused by the presence of irregular details and noise in the above-mentioned animal depth image.

[0054] Sub-step 3, performing a first image segmentation on the filtered animal depth image to obtain a segmented animal depth image. In practice, the execution entity may filter out at least one pixel value greater than or equal to a preset pixel threshold from each pixel value in the filtered animal depth image as a pixel value set to be removed. The preset pixel threshold may be a pre-set pixel threshold for image segmentation. For example, the preset pixel threshold may be 160. The pixel value set to be removed is removed from each pixel value to obtain a segmented animal depth image.

[0055] Sub-step 4: fusing the segmented animal depth image and the first animal image to obtain a fused animal image.

[0056] Sub-step 5, performing super-pixel segmentation processing on the above-mentioned fused animal image to obtain the animal image after super-pixel segmentation. Among them, the animal image after super-pixel segmentation can be a block image. The above-mentioned super-pixel segmentation processing can be a super-pixel segmentation processing of the SLIC (simple linear iterative cluster) algorithm that integrates depth information. The super-pixel distance value in the SLIC algorithm that integrates depth information can be a distance value including the depth information, color information and position space information of the pixel. The above-mentioned super-pixel distance value can be obtained by the following steps: First, determine the square of the difference between the pixel values ​​of any two pixels corresponding to the above-mentioned fused animal image as the pixel distance value. Secondly, determine the sum of the square of the difference between the horizontal axis values ​​of the above-mentioned any two pixels, the square of the difference between the vertical axis values ​​and the square of the difference between the depth values, as the depth distance value. Again, determine the sum of the square of the difference between the brightness of the above-mentioned any two pixels, the square of the difference between the first color, and the square of the difference between the second color, as the color distance value. Among them, the above-mentioned first color can be the A color in the CIELAB (CIELab color model) color, and the colors included in the A color are from dark green to gray to bright pink. The second color can be the B color in CIELAB color, and the B color includes colors ranging from bright blue to gray and then to yellow. Then, determine the position distance value of any two pixels. Finally, the arithmetic square root of the sum of the ratio of the pixel distance value, the depth distance value, the color distance value to the square of the preset threshold, and the ratio of the position distance value to the square of the cluster distance value is determined as the superpixel distance value. Among them, the preset threshold can be any value in [1,40]. The cluster distance can be the arithmetic square root of the number of pixels included in the fused animal image and the number of cluster pixel centers included in the cluster pixel center set in the SLIC algorithm.

[0057] Sub-step 6, constructing a graph model for the animal image after superpixel segmentation, and obtaining an undirected weighted graph for the animal image after superpixel segmentation. The nodes in the undirected weighted graph may be superpixels in the animal image after superpixel segmentation. The edges in the undirected weighted graph may be edges between superpixels. The weight values ​​corresponding to the edges in the undirected weighted graph may be a value with e as the base, and the inverse of the sum of the ratio of the difference between the color values ​​of any two pixels to a preset color threshold, the ratio of the difference between the depth values ​​to a preset depth threshold, and the ratio of the difference between the normalized average values ​​of the surface normals of the superpixels to a preset superpixel threshold as the exponent.

[0058] Sub-step 7, according to the above-mentioned undirected weighted graph, the animal image after superpixel segmentation is subjected to foreground prior processing to obtain an animal background salient image and an animal foreground salient image. Among them, the above-mentioned animal foreground salient image can be a background salient map obtained by arbitrarily selecting the upper boundary, lower boundary, left boundary and right boundary as the background query point set on the edge background corresponding to the fused animal image, and using a popular sorting algorithm. The above-mentioned edge background can be an image received by the user selected outside the area framed by the above-mentioned fused animal image. The above-mentioned animal foreground salient image can be a foreground salient map obtained by using a popular sorting algorithm with the central moment of the above-mentioned fused animal image as the foreground query point.

[0059] As an example, the execution subject may utilize a popular sorting algorithm to perform foreground-background prior processing on the animal image after superpixel segmentation according to the undirected weighted graph to obtain an animal background salient image and an animal foreground salient image.

[0060] Sub-step 8: performing image fusion on the animal foreground salient image and the animal background salient image to obtain a fused animal salient image.

[0061] Sub-step 9, by using the GrabCut algorithm that integrates the depth information, the above-mentioned fused animal salient image is segmented to obtain the animal foreground image and the animal background image. Among them, the energy function in the GrabCut algorithm that integrates the depth information can be an energy function that integrates the depth information as a constraint item in the energy function of the original GrabCut algorithm. The energy function in the GrabCut algorithm that integrates the depth information can be a product of adding a first weight value and a regional data item to the regional data item in the original energy function, and adding a second weight value and a smoothing item that integrates the depth information to the smoothing item. The above-mentioned first weight value can be obtained by the following steps: First, determine the KL (Kullback-Leibler divergence) divergence function value of the foreground Gaussian mixture model corresponding to the color information and the background Gaussian mixture model corresponding to the color information as the first KL divergence function value. Secondly, determine the KL divergence function value of the foreground Gaussian mixture model corresponding to the depth information and the background Gaussian mixture model corresponding to the depth information as the second KL divergence function value. Then, determine the sum of the first KL divergence function value and the second KL divergence function value as the third KL divergence function value. Finally, the ratio of the first KL divergence function value to the third KL divergence function value is determined as the first weight value. The above-mentioned second weight value can be obtained by the following steps: First, determine the smoothing term function value of the fused depth information and the transparency information of the superpixel as the depth smoothing term function value. Secondly, determine the sum of the function value corresponding to the original smoothing term function and the depth smoothing term function value as the smoothing term function value. Finally, determine the ratio of the function value corresponding to the original smoothing term function to the smoothing term function value as the second weight value.

[0062] The second step is to determine the foreground pixel difference set between each obtained animal foreground image and the previous frame animal foreground image, and to determine the background pixel difference set between each obtained animal background image and the previous frame animal background image.

[0063] The third step is to perform weighted summation on each foreground pixel difference in the foreground pixel difference set and the corresponding background pixel difference in the background pixel difference set to obtain an inter-frame pixel difference set as an animal motion intensity information set.

[0064] The above-mentioned technical scheme and its related contents, combined with step "step 203" as an inventive point of an embodiment of the present disclosure, solve the second technical problem mentioned in the background technology: "Because the GrabCut algorithm only uses the foreground and background images with significant texture and color differences in the first animal images of the front and back frames for segmentation, when the difference between the foreground pixels and the background pixels is small or the background is more complex, the segmentation accuracy of the image segmentation is low, and the operation time of the GrabCut algorithm is long, resulting in an increase in the segmentation operation time of the image segmentation, and there is a large amount of redundant information of the foreground pixels and the background pixels, resulting in a long time for the inter-frame difference operation, and the accuracy of the calculation of the animal motion intensity information is low, resulting in low stability of the camera performance, a large amount of waste of camera resources and an increase in the damage rate of the camera." The factors that lead to low stability of camera performance, waste of a large number of camera resources and increase of camera damage rate are often as follows: Since the GrabCut algorithm only uses the foreground and background images with significant texture and color differences in the first animal image of the front and back frames for segmentation, when the difference between the foreground pixels and the background pixels is small or the background is more complex, the segmentation accuracy of the image segmentation is low, and the operation time of the GrabCut algorithm is long, resulting in an increase in the segmentation operation time of the image segmentation, and there is a large amount of redundant information of the foreground pixels and the background pixels, resulting in a long time for the inter-frame difference operation, and the accuracy of the calculation of the animal motion intensity information is low. If the above factors are solved, the effect of improving the stability of camera performance, reducing the waste of camera resources and reducing the waste of camera damage rate can be achieved. In order to achieve this effect, the present disclosure first grays the acquired animal depth image, filters it and roughly segments it, which can reduce the area of ​​the background pixels on the basis of retaining the foreground information, so that the computing resources can be reduced and the operation time can be shortened when performing image segmentation later. Secondly, the fused animal image after fusion of the depth information and the first animal image is subjected to superpixel segmentation, and a pixel set with high similarity can be formed into superpixels, which can reduce the operation time, and the fusion of depth information can improve the accuracy of superpixel segmentation. Then, the fused animal image is subjected to foreground prior processing and background prior processing based on depth information, and an animal salient image fused with depth information can be obtained, which can further highlight the difference between foreground pixels and background pixels. Afterwards, the image segmentation of the fused animal salient image is performed through the GrabCut algorithm fused with depth information, which can improve the accuracy of image segmentation and shorten the operation time, thereby reducing the redundant information between foreground pixels and background pixels when compressing animal images, and reducing the waste of limited storage space. Finally, through the accurately segmented animal foreground image set and animal background image set, a more accurate animal motion intensity information set is determined, and the frame rate switching of the camera can be accurately controlled, the waste of camera resources can be reduced, the performance stability of the camera can be improved, and the damage rate of the camera can be reduced.

[0065] Step 203: Perform shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information.

[0066] In some embodiments, the execution subject may perform shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information. The shooting frame rate switching information may be information obtained after changing the current shooting frame rate information.

[0067] As an example, in response to determining that the intensity information representing the flight posture exists in the animal motion intensity information set, the execution subject may perform frame rate enhancement processing on the shooting frame rate corresponding to the current shooting frame rate information to obtain enhanced frame rate information as shooting frame rate switching information. The enhanced frame rate information may be frame rate information corresponding to the animal motion intensity information obtained by collecting a large amount of shooting frame rate and motion intensity information data for statistics.

[0068] In some optional implementations of some embodiments, performing shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information may include the following steps:

[0069] The first step is to filter the state information of the first animal's motion information set to obtain filtered animal state information. The filtered animal state information may be information in the motion state in the motion information set. The motion state may include but is not limited to at least one of the following: stationary, flying, standing, and walking.

[0070] The second step is to perform shooting frame rate switching processing according to the filtered animal state information and the first shooting frame rate information to obtain shooting frame rate switching information.

[0071] As an example, the execution subject may, in response to determining that the filtered animal state information represents state information in motion, perform shooting frame rate switching processing on the first shooting frame rate information to obtain fourth shooting frame rate information as shooting frame rate switching information. The fourth shooting frame rate information may be shooting frame rate information that can accurately capture the motion details of the first animal. The fourth shooting frame rate information may be a critical value for the motion state obtained by performing regression analysis on a large amount of animal motion state information collected and the shooting frame rate of the camera.

[0072] In some optional implementations of some embodiments, performing shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information may include the following steps:

[0073] The first step is to perform intensity information screening on the exercise information set to obtain filtered exercise intensity information, wherein the filtered exercise intensity information may be the information with the highest exercise intensity.

[0074] In the second step, in response to determining that the motion intensity information after the screening is not greater than the preset motion intensity threshold, the first shooting frame rate information is subjected to shooting frame rate switching processing to obtain second shooting frame rate information as shooting frame rate switching information. The preset motion intensity threshold may be a preset value. The preset motion intensity threshold may be a critical value between strong and weak motion intensities obtained by collecting a large amount of motion intensity information and shooting frame rates. The second shooting frame rate information may be a preset frame rate information for shooting a target animal with a lower motion intensity. For example, the first shooting frame rate information may be 10 frames per second.

[0075] In some optional implementations of some embodiments, performing shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information may include the following steps:

[0076] The first step is to perform intensity information screening on the above motion information set to obtain screened motion intensity information.

[0077] In the second step, in response to determining that the above-mentioned filtered motion intensity information is greater than the above-mentioned preset motion intensity threshold, the above-mentioned first shooting frame rate information is subjected to shooting frame rate switching processing to obtain third shooting frame rate information as shooting frame rate switching information. Among them, the above-mentioned third shooting frame rate information may be a pre-set frame rate that can clearly capture the posture of the target animal. For example, the above-mentioned first shooting frame rate information may be 60 frames per second. It should be noted that the camera switches different shooting frame rates for different motion intensities, which can take into account the needs of different scenes, thereby reducing the operating cost and resource energy consumption of the camera, improving the operating stability of the camera, and meeting the needs of target animal identification in different scenes.

[0078] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the image processing method of some embodiments of the present disclosure, when applied to camera shooting, can improve the accuracy of camera switching, improve the stability of the camera and reduce the damage rate. Specifically, the reasons for the high energy consumption and damage rate of the camera are: due to the use of a fixed frame rate for image acquisition, the different behaviors of the animal cannot be accurately distinguished and a large amount of storage resources are wasted, and when the animal's movement intensity is high, the picture cannot be accurately captured, a large number of details are missed, and when the animal is stationary, there is a large amount of repeated information in the animal image captured, resulting in high energy consumption and damage rate of the camera. Based on this, the image processing method of some embodiments of the present disclosure, when applied to camera shooting, can first obtain the first animal image set and the first shooting frame rate information of the target animal set. Here, the first animal image set obtained facilitates the subsequent determination of the motion information of the current target animal set, and the first shooting frame rate information obtained facilitates the subsequent switching of the shooting frame rate of the camera. Then, the above-mentioned first animal image set is subjected to inter-frame difference processing to obtain the motion information set of the first animal in the above-mentioned target animal set. Here, the motion information of the target animal can be accurately determined so as to subsequently determine the switching of the shooting frame rate of the camera. Finally, according to the motion information set of the first animal and the first shooting frame rate information, the shooting frame rate switching process is performed to obtain the shooting frame rate switching information. Here, the accuracy of the camera shooting frame rate switching in different scenes can be improved to reduce the loss of the camera and the waste of power resources. It can be obtained that the image processing method, applied to camera shooting, can accurately control the switching of the camera shooting frame rate through the motion information of the target animal, and can realize switching of different shooting frame rates in different situations to capture animal images that are more in line with the current scene, that is, it can realize the capture of more details when the target animal moves violently, reduce the collection of a large amount of repeated information when the target animal is stationary, reduce the loss of camera performance and electric energy, and reduce the damage rate of the camera.

[0079] Further references Figure 3 , shows a process 300 of another embodiment of the image processing method according to the present disclosure, which is applied to camera shooting. The image processing method, applied to camera shooting, comprises the following steps:

[0080] Step 301: Acquire a first animal image set and first shooting frame rate information of a target animal set.

[0081] Step 302: perform inter-frame difference processing on the first animal image set to obtain a motion information set for the first animal in the target animal set.

[0082] Step 303: Perform shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information.

[0083] In some embodiments, the specific implementation of steps 301-303 and the technical effects thereof can be referred to in Figure 2 The steps 201-203 in the corresponding embodiment are not described in detail here.

[0084] Step 304 , controlling the camera to shoot an animal image set for the target animal set under the shooting frame rate switching information.

[0085] In some embodiments, the execution subject may control the camera to capture an animal image set for the target animal set under the capture frame rate switching information. The animal images in the animal image set may be images of the target animal set captured by the camera at the capture frame rate corresponding to the capture frame rate switching information. The animal images may be as follows: Figure 4 Schematic diagram shown.

[0086] Step 305: perform animal identification on the animal image set to obtain animal identification information.

[0087] In some embodiments, the execution subject may perform animal recognition on the set of animal images to obtain animal recognition information. The animal recognition information may be information characterizing the posture and category identification of the target animal set. For example, the animal recognition information may include but is not limited to at least one of the following: category information, posture information, and key point information of the target animal.

[0088] As an example, the execution subject may input the above-mentioned animal image set into a target detection model to obtain animal identification information. The target detection model may be a pre-trained SSD (Single Shot MultiBox Detector) with a residual network or VisionTransformer as the main network.

[0089] In some optional implementations of some embodiments, performing animal identification on the above-mentioned animal photographed image set to obtain animal identification information may include the following steps:

[0090] The first step is to perform image preprocessing on the above-mentioned animal shot image set in response to determining that the above-mentioned animal shot image set is an image set based on a first preset shooting frame rate, and obtain a preprocessed animal image set. Among them, the above-mentioned first preset shooting frame rate can be a shooting frame rate that is sufficient to capture the movement of the target animal when the movement degree of the target animal is relatively gentle and does not require a high frame rate. For example, the above-mentioned first preset shooting frame rate can be the above-mentioned second shooting frame rate information. The above-mentioned image preprocessing may include: mosaic data enhancement processing, adaptive anchor frame calculation and adaptive image scaling processing. The above-mentioned mosaic data enhancement processing may be a data enhancement processing of randomly scaling, cropping and arranging any four animal shot images in the above-mentioned animal shot image set to generate a new image. The above-mentioned adaptive anchor frame calculation may be a calculation that determines the location of the target action through a clustering algorithm. The above-mentioned adaptive image scaling processing may be an image size processing that scales the image size of the above-mentioned animal shot image set to an image size that meets the input size of the animal recognition model.

[0091] In the second step, the pre-processed animal image set is input into the animal feature extraction network included in the animal recognition model to obtain the first animal feature information set, wherein the animal recognition model further includes: a multi-scale fusion network and a head recognition output network, and the animal feature extraction network includes: multiple first feature extraction networks, multiple maximum pooling layers, multiple cascade aggregation networks, multiple lightweight aggregation networks and a pyramid pooling network. The animal recognition model can be a neural network model that identifies the category of the animal in the input pre-processed animal image set and outputs it. The animal feature extraction network can be a network that performs multi-scale feature extraction on the input pre-processed animal image set to output multi-scale animal feature information. The multi-scale fusion network can be a network that fuses multi-scale animal feature information to obtain feature information containing more detailed information. The head recognition output network can be a network that performs target detection classification and regression prediction on the fused feature information output by the multi-scale fusion network to output the category to which the target animal belongs. The first feature extraction network may be a network including a convolution layer with a convolution kernel of 3*3 and a stride of 2, a BN (Batch Normalization) layer, and a LeakyReLU (Leaky Rectified Linear Unit) layer.

[0092] The above-mentioned multiple lightweight aggregation networks may be networks with a residual structure that perform two-channel parallel feature extraction on the input feature information. The above-mentioned lightweight aggregation network may be a network including multiple second feature extraction networks and multiple third feature extraction networks. The above-mentioned second feature extraction network may be a network including a convolution layer with a convolution kernel of 3*3 and a step size of 1, a BN layer and a LeakyReLU activation function. The above-mentioned third feature extraction network may be a network including a convolution layer with a convolution kernel of 1*1 and a step size of 1, a BN layer and a LeakyReLU activation function. The above-mentioned lightweight aggregation network may be a network that extracts features from the input animal feature information through the following steps: first, inputting the input animal feature information into the third feature extraction network included in the first branch to obtain the first animal input feature information. Among them, the above-mentioned first animal input feature information may be information that characterizes the features of the animal image in the form of a graph. Secondly, the above-mentioned first animal input feature information is sequentially input into the two second feature extraction networks included in the first branch to obtain the second animal input feature information and the third animal input feature information. Next, the input animal feature information is input into the third feature extraction network included in the second branch to obtain fourth animal input feature information. Then, channel feature splicing is performed on the first animal input feature information, the second animal input feature information, the third animal input feature information and the fourth animal input feature information to obtain input channel splicing feature information. Finally, the input channel splicing feature information is input into the third feature extraction network to obtain fifth animal input feature information.

[0093] The cascaded aggregation network in the above-mentioned multiple cascaded aggregation networks can be a network that expands the channel technology through multiple group convolutions and performs channel shuffling and merging. The above-mentioned cascaded aggregation network can be a network that adds a hybrid attention network to the third feature extraction network and the second feature extraction network in the first branch of the above-mentioned lightweight aggregation network, changes the second second feature extraction network to the third feature extraction network, adds a channel shuffling network in the middle, and adds a residual connection network in the first branch. The network structure diagram of the above-mentioned cascaded aggregation network is shown in FIG. Figure 5 shown. Figure 5 The two branches in the model split the input feature channels into equal parts. The upper branch keeps the original features to avoid gradient vanishing due to excessive attention, and the lower branch uses the idea of ​​dense residual network to extract spatial and channel features. Figure 5The channel shuffle network (ChannelShuffle) in the method may be a network that rearranges the input feature information between groups of channels to enhance the information interaction between channels. The hybrid attention network may be a network that introduces feature focusing at the spatial level and the channel level. The hybrid attention network may be a network that extracts features by the following steps: First, the input animal feature information is input into SENet (Squeeze-and-Excitation Networks) to obtain animal channel feature information. Secondly, the animal channel feature information is input into the channel global average pooling network to obtain animal channel attention feature information. Among them, the channel global average pooling network may be a network that inputs the input animal channel feature information into the global average pooling layer (Global Average Pooling) to perform global average pooling along the channel direction, and then passes through a fully connected layer and uses the Sigmoid activation function to input the animal channel attention feature information. Thirdly, the animal channel attention feature information and the input animal feature information are globally weighted along the channel direction to obtain animal channel weighted feature information. Next, the above-mentioned animal channel weighted feature information is respectively input into the maximum pooling layer and the average pooling layer to obtain the animal maximum pooling feature information and the animal average pooling feature information. Subsequently, the above-mentioned animal maximum pooling feature information and the above-mentioned animal average pooling feature information are channel-spliced ​​to obtain the animal pooling splicing feature information. Afterwards, the above-mentioned animal pooling splicing feature information is input into a 1*1 convolutional layer and a Sigmod activation function to obtain the animal spatial attention feature information. Among them, the above-mentioned 1*1 convolutional layer can be a convolutional layer that reduces the animal pooling splicing feature information to a single channel. Finally, the above-mentioned animal spatial attention feature information and the above-mentioned input animal feature information are dot-multiplied to obtain the animal spatial channel feature information.

[0094] The pyramid pooling network may be a network that uses different pooling kernels to perform pooling operations to fuse feature information of different scales. The pyramid pooling network may perform feature pooling processing by the following steps: First, input the input animal feature information into the second feature extraction network to obtain input feature information. Secondly, input the input feature information into the branch pooling network and the branch convolution network respectively to obtain pooling feature information and convolution feature information. Among them, the branch pooling network may be a branch network that performs feature splicing on the input feature information in parallel with the maximum pooling kernels of 5*5, 9*9 and 13*13, and then inputs it into the second feature extraction network for feature extraction. The branch convolution network may be a branch network that inputs the input feature information into the second feature extraction network. Finally, the pooling feature information and the convolution feature information are spliced ​​with channel features and then input into the second feature extraction network to obtain animal pooling feature information.

[0095] In the third step, the first animal feature information set is input into the multi-scale fusion network to obtain the second animal feature information set, wherein the multi-scale fusion network includes: a plurality of first feature extraction networks, a plurality of second feature extraction networks, a plurality of lightweight aggregation networks and a plurality of content-aware prediction networks. The content-aware prediction network (CARAFE, Content-Aware ReAssembly of FEatures) in the plurality of content-aware prediction networks may be a network for upsampling features by aggregating contextual information and adaptive kernels in the perception domain. The content-aware prediction network may include a network composed of a kernel prediction module and a content-aware feature reassembly module. The content-aware feature reassembly module may be composed of three parts: a channel compression module, a content encoder and a kernel normalization module. The channel compression module may be a compression module that uses 1*1 convolution compression to reduce the channels of the input feature map. The content encoder may be a module that takes the compressed feature map as input and encodes the content to generate a reassembly kernel. The kernel normalization module is a module that applies a Softmax function to each reassembly kernel for activation. The content-aware reorganization module adds the weights obtained by the weighted kernel prediction module to each target reorganized region, and then splices the modules of all reorganized target regions. For example, the animal feature information input to the content-aware prediction network may be C*W*H, and the output of the content-aware prediction network may be C*aW*aH. C may represent the number of channels. C may represent the number of channels. W may represent the feature length. H may represent the feature width. a may represent the upsampling ratio.

[0096] The fourth step is to input the second animal feature information set into the head recognition output network to obtain animal identification information, wherein the head recognition output network includes: a plurality of third feature extraction networks. The animal identification information may be information of an animal category label to which the target animal belongs.

[0097] In some optional implementations of some embodiments, performing animal identification on the above-mentioned animal photographed image set to obtain animal identification information may include the following steps:

[0098] The first step is to perform downsampling convolution processing on the above-mentioned animal shot image set in response to determining that the above-mentioned animal shot image set is an image set based on a second preset shooting frame rate, and obtain an animal posture feature information set. Among them, the above-mentioned second preset shooting frame rate can be a shooting frame rate when the target animal's movement is more intense and a higher shooting frame rate is required to capture more detailed posture information of the target animal. For example, the above-mentioned second preset shooting frame rate can be the above-mentioned third shooting frame rate information. The animal posture feature information in the above-mentioned animal posture feature information set can be information about the posture of the target animal in the animal shot image. The above-mentioned downsampling convolution processing can be a downsampling processing of sequentially inputting the above-mentioned animal shot image set into two posture recognition convolution networks. The above-mentioned posture recognition convolution network can be a network including a convolution layer with a convolution kernel of 3*3, a step size of 2, an edge padding of 1, and an output channel of 64, a BN layer, and a ReLU activation function.

[0099] In the second step, the above-mentioned animal posture feature information set is input into the feature extraction network included in the animal posture recognition model to obtain the second posture feature information set, wherein the above-mentioned animal posture recognition model also includes: a first branch parallel fusion network, a second branch parallel fusion network, a third branch parallel fusion network and a convolution layer, and the above-mentioned feature extraction network includes: a residual convolution network, a plurality of path fusion convolution networks. Among them, the above-mentioned animal posture recognition model can be a deep neural network model that recognizes the animal posture of the input animal photographed image and outputs the animal posture information. The above-mentioned feature extraction network can be a network that extracts features of different levels and scales from the input first animal posture feature information. The above-mentioned residual convolution network can be a network of residual structure including the first posture feature extraction network, the second posture feature extraction network, and the third posture feature extraction network as the first branch, and the third posture feature extraction network as the second branch. The above-mentioned first posture feature extraction network can be a network including a convolution layer with a convolution kernel of 1*1, a step size of 1, and an output channel of 64, a BN layer, and a ReLU activation function. The second posture feature extraction network may be a network including a standard convolution layer with a convolution kernel of 3*3, a convolution layer with a convolution kernel of 1*3 for capturing features in the horizontal direction, and a convolution layer with a convolution kernel of 3*1 for capturing features in the vertical direction, and performing feature splicing, BN layer and ReLU activation function on the posture feature information obtained by the three convolution layers. The third posture feature extraction network may be a network including a convolution layer with a convolution kernel of 1*1, a step size of 1, and an output channel of 256, a BN layer and a ReLU activation function. The multiple path fusion convolution network may be a network after removing the second branch in the residual convolution network. The first branch parallel fusion network may be a network that adds one more animal posture feature information feature extraction branch and performs multi-scale posture feature fusion. The initial feature information of the branch added in the first branch parallel fusion network may be feature information of a size half of the size of the second posture feature information obtained by outputting a convolution layer and a BN layer including a convolution kernel of 3*3, a step size of 1, and an output channel number twice that of the previous branch. The second branch parallel fusion network may be a network having one more feature extraction branch than the first branch parallel fusion network, and the size of the initial feature information of the extra branch may be one quarter of the size of the second posture feature information. The third branch parallel fusion network may be a network having one more feature extraction branch than the second branch parallel fusion network, and the size of the initial feature information of the extra branch may be one eighth of the size of the second posture feature information. The convolution layer may be a convolution layer including a convolution kernel of 1*1, a step size of 2, an edge padding of 1, and an output channel of 17. The animal posture recognition information may be information of a recognition result for determining whether the target animal is flying.

[0100] The third step is to input the second posture feature information set into the first branch parallel fusion network to obtain a first multi-scale posture feature information set. The first multi-scale posture feature information set may include posture feature information of different scales.

[0101] The fourth step is to input the first multi-scale posture feature information set into the second branch parallel fusion network to obtain a second multi-scale posture feature information set.

[0102] In the fifth step, the second multi-scale posture feature information set is input into the third branch parallel fusion network to obtain a third multi-scale posture feature information set.

[0103] In the sixth step, the third multi-scale posture feature information set is input into the convolution layer to obtain animal posture recognition information as animal recognition information to determine whether the animal posture recognition information of the target animal set is flight posture information. The animal posture recognition information may be information about the posture of the identified target animal.

[0104] In some optional implementations of some embodiments, the first branch parallel fusion network includes: a branch parallel expansion network, multiple multi-scale feature extraction convolution networks and multiple channel space attention networks. Among them, the first branch parallel expansion network can be a network including an added posture feature extraction branch network that uses the second posture feature information set for 2-fold downsampling as the initial posture information, and a branch network that originally extracts the second posture feature information from the first posture feature information. The 2-fold downsampling can be downsampling through a convolution layer with a convolution kernel of 3*3, a step size of 1, and an output channel number that is twice that of the previous branch. The multi-scale feature extraction convolution network in the multiple multi-scale feature extraction convolution networks can be a network in which multiple convolution layers with different convolution kernels combine posture feature information to extract richer feature information and reduce the number of parameters. The multi-scale feature extraction convolution network can be a network that extracts features by following the following steps: First, the input posture feature information is respectively input into the convolution layer with a convolution kernel of 3*3 and the atrous convolution layer with a convolution kernel of 3*3, and the posture feature information after the convolution kernel is 3*3 is added to obtain the added posture feature information. Then, the added posture feature information is globally averaged and pooled to obtain global posture feature information of length L. Subsequently, the above-mentioned global posture feature information is input into the fully connected layer for aggregation and compression to obtain compressed posture feature information. Next, the above-mentioned compressed posture feature information is respectively input into two fully connected layers to obtain two tiled posture feature information of length L. After that, the two tiled posture feature information are input into the Softmax activation function to obtain two posture weight values. Finally, the above-mentioned two posture weight values ​​are weighted and summed with the corresponding posture feature information input through the convolution layer and the hole convolution layer. The channel spatial attention network in the above-mentioned multiple channel spatial attention networks can be a network that extracts channel attention and spatial attention from the posture feature information to suppress noise in the animal shot image and improve the target detection accuracy. The above-mentioned channel spatial attention network can include a channel attention module and a spatial attention module. The above-mentioned channel attention module can be a network that obtains two 1*1*C feature maps by subjecting the input posture feature information to global maximum pooling and global average pooling, and then inputs the two 1*1*C posture feature information into a multi-layer perceptron (MLP) with shared weights to sum the two 1*1*C posture feature information element by element and output the channel weight coefficient, and finally multiplies the channel weight coefficient and the input posture feature information to obtain the output posture feature information.The above-mentioned spatial attention module can be a channel attention module that performs channel-based maximum pooling and average pooling on the output posture feature information output by the channel attention module to obtain pooled posture feature information, and then compresses the pooled posture feature information to 1 through a convolution layer with a convolution kernel of 7*7, and obtains the spatial weight coefficient through a Sigmoid activation function layer. Finally, the spatial weight coefficient and the output posture feature information are feature multiplied to output the convolution network of the posture feature information.

[0105] Optionally, the step of inputting the second posture feature information set into the branch parallel fusion network to obtain the first multi-scale posture feature information set may include the following steps:

[0106] The first step is to input the second posture feature information set into the branch parallel expansion network to obtain a fourth multi-scale posture feature information set.

[0107] The second step is to input the fourth multi-scale posture feature information set into the multiple multi-scale feature extraction convolutional networks to obtain a fifth multi-scale posture feature information set. The input may be an input of a fourth multi-scale posture feature information into a multi-scale feature extraction convolutional network.

[0108] The third step is to input the fifth multi-scale posture feature information set into the multiple channel spatial attention networks to obtain a sixth multi-scale posture feature information set. The input may be an input of the fifth multi-scale posture feature information into a channel spatial attention network.

[0109] The fourth step is to perform sampling and fusion processing on the sixth multi-scale posture feature information set to obtain the seventh multi-scale posture feature information set. The sampling processing may include upsampling and downsampling. The sampling processing may be a sampling process in which the posture feature information finally output by each branch is upsampled and downsampled and each branch is subjected to feature fusion. In practice, the execution subject may perform feature fusion on the posture feature information obtained by the second branch by performing a 2-fold upsampling process on the posture feature information obtained by the second branch and the posture feature information output by the first branch to obtain the seventh multi-scale posture feature information; and perform feature fusion on the posture feature information obtained by the first branch by performing a 2-fold downsampling process on the posture feature information obtained by the first branch and the posture feature information output by the second branch to obtain the seventh multi-scale posture feature information.

[0110] In the fifth step, nonlinear processing is performed on the seventh multi-scale posture feature information set to obtain a first multi-scale posture feature information set. The nonlinear processing may be performed by inputting the seventh multi-scale posture feature information set into a ReLU activation function.

[0111] Step 306: The animal identification information and the animal photographed image set are updated in a time rolling manner and stored in a target storage end.

[0112] In some embodiments, the execution entity may update the animal identification information and the animal image set in a rolling manner and store them in a target storage end, wherein the target storage end may be a local database or a cloud database for storing the animal identification information and the animal image set.

[0113] As an example, the execution subject may first determine the image capture time of the animal image set. Then, according to the image capture time, the animal image set and the historical animal image set are stored in the target storage end in reverse order based on time, so as to access the latest animal image in the target storage end. Finally, the animal identification information is stored in the storage space corresponding to the animal image set in the target storage end.

[0114] In some optional implementations of some embodiments, after step 306, the method may further include the following steps:

[0115] In response to determining that the infrared detection information indicates that no motion information is detected, the camera is controlled to enter a standby state, wherein the standby state may be a state in which the camera stops image capture and image recognition.

[0116] from Figure 3 It can be seen that Figure 2 Compared with the description of some corresponding embodiments, Figure 3 In the corresponding image processing method of some embodiments, the process 300 applied to camera shooting embodies the steps of determining the animal shooting image set of the target animal set to be shot according to the shooting frame rate switching information, and performing animal recognition on the animal shooting image set. Because the scheme described in these embodiments can first control the above-mentioned camera to shoot the animal shooting image set for the above-mentioned target animal set under the above-mentioned shooting frame rate switching information, it can improve the capture of the target animal movement details by the animal shooting image set, and remove a large amount of repeated information when the target animal is stationary. Then, the animal shooting image set is subjected to animal recognition to obtain animal recognition information. Since the animal shooting image set accurately captures different movement information of the target animal, the accuracy of animal recognition can be improved, so as to accurately grasp the behavior information of the target animal. Finally, the above-mentioned animal recognition information and the above-mentioned animal shooting image set are time-rolled and updated and stored in the target storage end. Here, the waste of storage resources can be reduced, and it is convenient for users to access and grasp the situation of the target animal. It can be obtained that the target recognition accuracy for different shooting frame rates in different scenes can be improved, and the waste of storage resources can be reduced.

[0117] Reference below Figure 6 , which shows an electronic device (eg, Figure 1 Schematic diagram of the structure of the electronic device 101)600. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0118] like Figure 6 As shown, the electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0119] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 6 Each block shown in the figure may represent one device, or may represent multiple devices as required.

[0120] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.

[0121] It should be noted that the computer-readable medium in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0122] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0123] The computer-readable medium may be included in the electronic device; or it may exist independently without being installed in the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains the first animal image set and the first shooting frame rate information of the target animal set; performs inter-frame difference processing on the first animal image set to obtain the motion information set for the first animal in the target animal set; performs shooting frame rate switching processing according to the motion information set of the first animal and the first shooting frame rate information to obtain shooting frame rate switching information.

[0124] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0126] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0127] Some embodiments of the present disclosure further provide a computer program product, including a computer program, which implements any of the above-mentioned image processing methods when executed by a processor and is applied to camera shooting.

[0128] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with the technical features with similar functions disclosed in the embodiments of the present disclosure (but not limited to) and the technical solutions formed.

Claims

1. An image processing method, applied to camera shooting, comprising: Acquire a first animal image set and first shooting frame rate information of a target animal set; Performing inter-frame difference processing on the first animal image set to obtain a motion information set for the first animal in the target animal set; According to the first animal's motion information set and the first shooting frame rate information, shooting frame rate switching processing is performed to obtain shooting frame rate switching information.

2. The method according to claim 1, wherein: The method further comprises: Controlling the camera to shoot a set of animal images for the target animal set under the shooting frame rate switching information; Performing animal identification on the set of animal photographed images to obtain animal identification information; The animal identification information and the animal photographed image set are updated in a rolling manner over time and stored in a target storage end.

3. The method according to claim 1, wherein: The step of performing a shooting frame rate switching process according to the first animal's motion information set and the first shooting frame rate information to obtain the shooting frame rate switching information includes: Performing state information screening processing on the movement information set of the first animal to obtain screened animal state information; According to the filtered animal state information and the first shooting frame rate information, shooting frame rate switching processing is performed to obtain shooting frame rate switching information.

4. The method according to claim 1, wherein: The step of performing a shooting frame rate switching process according to the first animal's motion information set and the first shooting frame rate information to obtain the shooting frame rate switching information includes: Performing intensity information screening processing on the motion information set to obtain screened motion intensity information; In response to determining that the filtered motion intensity information is not greater than a preset motion intensity threshold, performing a shooting frame rate switching process on the first shooting frame rate information to obtain second shooting frame rate information as shooting frame rate switching information.

5. The method according to claim 1, wherein: The step of performing a shooting frame rate switching process according to the first animal's motion information set and the first shooting frame rate information to obtain the shooting frame rate switching information includes: Performing intensity information screening processing on the motion information set to obtain screened motion intensity information; In response to determining that the filtered motion intensity information is greater than the preset motion intensity threshold, performing a shooting frame rate switching process on the first shooting frame rate information to obtain third shooting frame rate information as shooting frame rate switching information.

6. The method according to claim 2, wherein: The step of performing animal identification on the set of animal photographed images to obtain animal identification information includes: In response to determining that the animal shot image set is an image set based on a first preset shooting frame rate, performing image preprocessing on the animal shot image set to obtain a preprocessed animal image set; Inputting the preprocessed animal image set into an animal feature extraction network included in an animal recognition model to obtain a first animal feature information set, wherein the animal recognition model further includes: a multi-scale fusion network and a head recognition output network, and the animal feature extraction network includes: a plurality of hybrid attention networks, a plurality of first feature extraction networks, a plurality of maximum pooling layers, a plurality of cascade aggregation networks, and a plurality of lightweight aggregation networks; Inputting the first animal feature information set into the multi-scale fusion network to obtain a second animal feature information set, wherein the multi-scale fusion network includes: a plurality of first feature extraction networks, a plurality of second feature extraction networks, a plurality of lightweight aggregation networks, and a plurality of content-aware prediction networks; The second animal feature information set is input into the head recognition output network to obtain animal recognition information, wherein the head recognition output network includes: a plurality of third feature extraction networks.

7. The method according to claim 2, wherein: The step of performing animal identification on the set of animal photographed images to obtain animal identification information includes: In response to determining that the animal shot image set is an image set based on a second preset shooting frame rate, performing downsampling convolution processing on the animal shot image set to obtain a first animal posture feature information set; Inputting the first animal posture feature information set into a feature extraction network included in an animal posture recognition model to obtain a second posture feature information set, wherein the animal posture recognition model further includes: a first branch parallel fusion network, a second branch parallel fusion network, a third branch parallel fusion network and a convolution layer, and the feature extraction network includes: a residual convolution network and a plurality of path fusion convolution networks; Inputting the second posture feature information set into the first branch parallel fusion network to obtain a first multi-scale posture feature information set; Inputting the first multi-scale posture feature information set into the second branch parallel fusion network to obtain a second multi-scale posture feature information set; Inputting the second multi-scale posture feature information set into the third branch parallel fusion network to obtain a third multi-scale posture feature information set; The third multi-scale posture feature information set is input into the convolution layer to obtain animal posture recognition information as animal recognition information to determine whether the animal posture recognition information of the target animal set is flight posture information.

8. The method according to claim 7, wherein: The first branch parallel fusion network includes: a branch parallel expansion network, a plurality of multi-scale feature extraction convolution networks and a plurality of channel space attention networks; and The step of inputting the second posture feature information set into the first branch parallel fusion network to obtain a first multi-scale posture feature information set includes: Inputting the second posture feature information set into the branch parallel expansion network to obtain a fourth multi-scale posture feature information set; Inputting the fourth multi-scale posture feature information set into the multiple multi-scale feature extraction convolutional networks to obtain a fifth multi-scale posture feature information set; Inputting the fifth multi-scale posture feature information set into the multiple channel spatial attention networks to obtain a sixth multi-scale posture feature information set; Performing sampling and fusion processing on the sixth multi-scale posture feature information set to obtain a seventh multi-scale posture feature information set; Nonlinear processing is performed on the seventh multi-scale posture feature information set to obtain a first multi-scale posture feature information set.

9. The method according to claim 1, wherein: Before acquiring the first animal image set and the first shooting frame rate information of the target animal set, the method further includes: In response to detecting infrared signal information sent by the passive infrared sensor, performing signal detection processing on the infrared signal information to obtain infrared detection information; According to the infrared detection information, the camera is controlled to shoot animal images to obtain a first animal image set for the target animal set in a shooting area.

10. The method according to claim 9, wherein: The method further comprises: In response to determining that the infrared detection information indicates that no motion information is detected, the camera is controlled to enter a standby state.

11. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.

12. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Photographing Apparatus, And Method For Photographing Moving Object With The Same

    CN107205115A

  • Recording frame rate control method and related apparatus

    CN113475057A

  • Small target detection method for images acquired by unmanned aerial vehicle based on improved YOLOv8 algorithm

    CN118628939A

  • Lightweight network target detection method based on structure optimization and feature fusion

    CN118710883A

  • Video stream acquisition method based on deep learning

    CN119299703A