Infrared image coding method, device and equipment for pulse neural network
By constructing the grayscale and mapping range of infrared images, generating enhancement parameters and converting them into a pulse sequence set, the noise interference and waste of computing resources caused by pseudo-color mapping are solved, efficient infrared image encoding and training are achieved, and the spatiotemporal training requirements of pulse neural networks are adapted.
Patent Information
- Application Number
- CN202510825451.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-06-19
AI Technical Summary
In existing infrared image encoding methods, pseudo-color mapping may mask the original grayscale features, resulting in noise interference and waste of computing resources in model training. Especially in complex multi-target scenes, it is easy to misjudge noise as feature signals, and the feature area information of high-resolution images is lost.
By determining the grayscale range and mapping range of the infrared image, generating enhancement parameters, and constructing a hybrid mapping function, the infrared image is converted into a pulse sequence set to form a three-dimensional data set, which is then trained using a pulse neural network to avoid the pseudo-color conversion link and reduce computing resource redundancy.
It improves computing speed and related performance, reduces the pseudo-color conversion link in computer recognition tasks, enhances sensitivity to specific targets, adapts to the spatiotemporal training requirements of pulse neural networks, and reduces computing power redundancy and energy waste.
Smart Images

Figure CN120747255A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to an infrared image encoding method, apparatus, and device for a pulse neural network. Background Art
[0002] Images captured by infrared cameras are typically presented as thermal images. However, the human eye cannot see infrared light, necessitating the conversion of these images into visual images through pseudo-color mapping. This is commonly done to encode infrared images, mapping infrared signals of varying intensities to visible colors.
[0003] However, when using the above method to encode infrared images, the following technical problems often occur:
[0004] First, when manually annotating infrared image datasets, the invisible infrared signal intensity needs to be mapped through visible light channels such as RGB / HSV for pseudo-color coding. Pseudo-color coding mapping may mask the original grayscale features, aggravate noise interference in model training, and lead to a decrease in accuracy.
[0005] Second, real-time pseudo-color conversion and signal processing during the computer deployment stage will occupy a large amount of computing resources, resulting in redundant computing power and waste of energy.
[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0007] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] Some embodiments of the present disclosure propose infrared image encoding methods, devices and equipment for pulse neural networks to solve one or more of the technical problems mentioned in the above background technology section.
[0009] In a first aspect, some embodiments of the present disclosure provide an infrared image encoding method for a pulse neural network, comprising: determining the grayscale range and mapping range of the original infrared image based on an image histogram corresponding to the original infrared image, wherein the mapping range is a range used to adapt to the pulse neural network input; determining at least one enhancement range for feature enhancement in the grayscale range; generating enhancement parameters corresponding to the enhancement range for each enhancement range in the at least one enhancement range; constructing a mixed mapping function based on the grayscale range, the mapping range and the obtained enhancement parameter set; inputting each pixel value in the original infrared image into the mixed mapping function to generate a mapping pixel value to obtain a mapping image; converting each mapping pixel value in the mapping image into a pulse sequence through a Bernoulli experiment to obtain a pulse sequence set; splicing each pulse sequence in the pulse sequence set to generate three-dimensional data to obtain a three-dimensional data set as an infrared image encoding result; and inputting the infrared image encoding result into the pulse neural network for training.
[0010] In a second aspect, some embodiments of the present disclosure provide an infrared image encoding device for a pulse neural network, comprising: a determination unit, configured to determine the grayscale range and mapping range of the original infrared image according to the image histogram corresponding to the original infrared image, wherein the mapping range is a range used to adapt the pulse neural network input; an enhancement range screening unit, configured to determine at least one enhancement range to be enhanced in the grayscale range; a parameter generation unit, configured to generate, for each enhancement range in the at least one enhancement range, an enhancement parameter corresponding to the enhancement range; a function construction unit, configured to generate an enhancement parameter corresponding to the enhancement range according to the grayscale range, the mapping range, and the enhancement range. The mapping range and the obtained enhancement parameter set are used to construct a hybrid mapping function; the image mapping unit is configured to input each pixel value in the above-mentioned original infrared image into the above-mentioned hybrid mapping function to generate a mapping pixel value and obtain a mapping image; the pulse conversion unit is configured to convert each mapping pixel value in the above-mentioned mapping image into a pulse sequence through a Bernoulli experiment to obtain a pulse sequence set; the three-dimensional encoding unit is configured to splice each pulse sequence in the above-mentioned pulse sequence set to generate three-dimensional data and obtain a three-dimensional data set as the infrared image encoding result; the training input unit is configured to input the above-mentioned infrared image encoding result into the pulse neural network for training.
[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0013] The aforementioned embodiments of the present disclosure have the following beneficial effects: infrared images obtained through the infrared image encoding methods of some embodiments of the present disclosure reduce the time required for pseudo-color conversion in computer recognition tasks, thereby improving computing speed and related performance. Specifically, the reduction in computer computing speed and related performance is caused by the fact that real-time pseudo-color conversion and signal processing during the computer deployment phase consume a large amount of computing resources, resulting in redundant computing power and wasted energy. Based on this, the infrared image encoding methods for spiking neural networks of some embodiments of the present disclosure first determine the grayscale range and mapping range of the original infrared image based on the image histogram corresponding to the original infrared image. The mapping range is used to adapt the input of the spiking neural network. This avoids the problem of reduced training efficiency caused by a mismatch between the input data range and the spiking neural network, ensuring direct data format compatibility. Then, at least one enhancement range within the grayscale range is determined for feature enhancement. This focuses on key areas, provides a target range for feature enhancement, and improves sensitivity to specific targets. Next, for each of the at least one enhancement range, an enhancement parameter corresponding to the enhancement range is generated. Thus, an enhancement factor is generated for each enhancement region, providing data for constructing an enhancement function. Next, a hybrid mapping function is constructed based on the grayscale range, the mapping range, and the obtained enhancement parameter set. This achieves a balance between global grayscale adaptation and local feature enhancement through linear and nonlinear changes, dynamically optimizing the image range. Each pixel value in the original infrared image is then input into the hybrid mapping function to generate a mapped pixel value, resulting in a mapped image. This produces a contrast-enhanced image that separates key target areas from the original background, reducing noise interference during model training. Then, using a Bernoulli experiment, each mapped pixel value in the mapped image is converted into a pulse sequence, resulting in a pulse sequence set. This converts continuous pixel values into discrete pulse sequences, which simulate the activation characteristics of biological neurons and reduce computational redundancy. Next, each pulse sequence in the pulse sequence set is concatenated to generate three-dimensional data, resulting in a three-dimensional dataset as the infrared image encoding result. This creates a spatiotemporal fusion encoding structure that preserves the characteristics of the original image and meets the temporal processing requirements of the spiking neural network. Finally, the infrared image encoding result is input into the spiking neural network for training. Therefore, the spatiotemporal training characteristics of spiking neural networks can be directly utilized for training, eliminating the traditional pseudo-color conversion step. In summary, by converting the original image pixels into a coding structure with spatiotemporal time sequence to adapt the training of spiking neural networks, the pseudo-color conversion required for manual annotation can be reduced, thereby reducing redundant computing power and wasted energy. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0015] Figure 1 is a flow chart of some embodiments of an infrared image encoding method for a spiking neural network according to the present disclosure;
[0016] Figure 2 is a schematic structural diagram of some embodiments of an infrared image encoding device for a spiking neural network according to the present disclosure;
[0017] Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] refer to Figure 1, shows a process 100 of some embodiments of the infrared image encoding method for a pulse neural network according to the present disclosure. The infrared image encoding method for a pulse neural network includes the following steps:
[0025] Step 101: Determine the grayscale range and mapping range of the original infrared image according to the image histogram corresponding to the original infrared image.
[0026] In some embodiments, the execution subject (for example, a computing device) of the infrared image encoding method for the pulse neural network can be a computer. The image histogram can be a chart showing the distribution of grayscale values of all pixels, with the horizontal axis being the grayscale value and the vertical axis being the number of pixels with corresponding grayscale values. The mapping range is a range used to adapt to the input of the pulse neural network. The grayscale range can be an interval consisting of the minimum and maximum values of all grayscale values. The mapping range is to adjust the original grayscale range to the target interval through mathematical transformation to adapt to the input requirements of the pulse neural network. The pulse neural network can be a third-generation artificial neural network, a neural network that simulates the mechanism of biological neurons transmitting information through pulses.
[0027] Step 102: Determine at least one enhancement range in the grayscale range for feature enhancement.
[0028] In some embodiments, the execution entity may determine at least one enhancement range within the grayscale range for feature enhancement. The enhancement range may be a subrange within the original grayscale range. In practice, key regions or features within the grayscale range that need to be highlighted may be selected. For forest fire monitoring, the enhancement range may be a high-temperature flame region or a high grayscale value range. For example, grayscale values range from 1200 to 1600.
[0029] While employing technical solutions to address the aforementioned technical problem, the following issues often arise: Noise interference: For complex scenes with multiple targets, pseudo-color coding can easily misjudge background noise as valid signals, leading to noise interference and affecting model accuracy. Increased computational complexity: When optimizing enhanced range boundaries, simplified range fusion can lead to over-fusion for high-resolution infrared images, resulting in loss of grayscale features for targets with small temperature differences.
[0030] Conventional solutions to these problems typically employ fixed-step uniform segmentation or predefined thresholds to divide and enhance ranges, sacrificing accuracy for computational efficiency. Range fusion is performed using simplified weighted distances to reduce complexity. However, the inventors considered that fixed-step segmentation and simplified range fusion could result in the loss of key feature areas. Therefore, they decided to adopt the following solution.
[0031] In some optional implementations of some embodiments, the execution entity may determine at least one enhancement range in the grayscale range for feature enhancement, which may include the following steps:
[0032] The first step is to obtain a grayscale range based on the grayscale histogram of the original infrared image. The grayscale range can be the range between the minimum and maximum grayscale levels of the original infrared image. For example, the minimum grayscale can be 0, and the maximum grayscale can be 16383, so the grayscale range can be [0, 16383].
[0033] The second step is to traverse the pixels of the original infrared image to obtain the number of pixels corresponding to each grayscale value as a pixel frequency set. The pixel frequency set can be a set of the number of pixels corresponding to each grayscale value. In practice, each pixel value in the original image is first traversed and counted. Then, the number of occurrences of each grayscale level is counted as the pixel frequency.
[0034] The third step is to divide each pixel frequency in the pixel frequency set by the total pixel value of the original infrared image to generate a pixel probability density, thereby obtaining a pixel probability density set. The pixel probability density can be the ratio of the frequency of each grayscale value to the total number of pixels in the original image. For example, if the total number of pixels in the original image is 1 million and the frequency of a grayscale value of 1000 is 5000, then the pixel probability density corresponding to the grayscale value of 1000 is 0.005.
[0035] The fourth step is to determine the grayscale value corresponding to the maximum pixel density in the pixel probability density set as the first initial boundary value. The first initial boundary value can be the grayscale that appears most frequently in the image.
[0036] In the fifth step, the grayscale range is evenly segmented according to the first initial boundary value to obtain the remaining boundary values as a boundary value set. The boundary value set may be a result of each boundary value being a multiple of the first initial boundary value. First, the grayscale range is segmented according to a segmentation step size, which may be a preset step size. For example, the preset step size may be 10. Then, starting from the first initial boundary value, the step sizes are sequentially added to obtain the boundary value set.
[0037] Step 6: For each boundary value in the above boundary value set, perform the following steps:
[0038] The first sub-step is to determine the weighted distance between the above boundary value and the weight of each grayscale value, where the weight is the pixel probability density corresponding to the grayscale value. In practice, the weighted distance can be an indicator to measure the difference between the boundary value and the surrounding grayscale values. First, determine the domain radius and the domain range of the above boundary value. Then, traverse the grayscale values in the neighborhood, calculate the distance to the boundary value and multiply it by the corresponding weight to obtain the weighted distance.
[0039] In the second sub-step, the boundary value is optimized by performing a minimum distance optimization based on the weighted distance to obtain an optimized boundary value. The minimum distance optimization may be performed by finding the minimum value of the objective function to determine the optimal solution, or by traversing the neighborhood of the boundary value and selecting the location with the minimum weighted distance as the optimization result.
[0040] In the seventh step, the optimized boundary value set is sorted and the boundaries whose adjacent distances are less than the preset threshold are merged to obtain the enhanced sub-range boundary set. The above sorting can be ascending sorting. First, the boundary value set is sorted in ascending order to obtain a sorted queue. Then, the sorted queue is traversed through the adjacent boundary values to obtain the distances. If the distances are less than the preset threshold, the boundary values are merged and replaced with the average value. For example, the sorted boundary values can be {0, 500, 750, 1200, 16383}, and the preset threshold can be 300. Then, the distance between 500 and 750 is 250, which is less than the preset threshold 300. Then, the distance is merged and replaced with 625, the average value of the boundary values.
[0041] Step 8: For each enhancer range boundary in the enhancer range boundary set, perform the following steps:
[0042] In the third sub-step, in response to the enhancement sub-range boundary being less than the lower limit of the grayscale range, the enhancement sub-range boundary is corrected to the lower limit of the grayscale range. For example, the enhancement sub-range boundary may be {-50, 800, 16383}. If -50 is less than 0, then -50 is replaced with 0.
[0043] In a fourth sub-step, in response to the enhancement sub-range boundary being greater than the upper limit of the grayscale range, the enhancement sub-range boundary is corrected to the upper limit of the grayscale range. For example, the enhancement sub-range boundary may be {50, 800, 17000}. If 17000 is greater than 16383, 17000 is replaced with 16383.
[0044] In the ninth step, based on the grayscale range, the boundaries of the enhancement sub-range are used as the boundaries of the enhancement sub-range to obtain the enhancement range. The enhancement range can be a range division of the image processing interval based on task requirements. In practice, the task requirement can be feature enhancement. First, adjacent boundary values are traversed to generate closed intervals as a single interval range. The interval range is then verified to be within the original grayscale range and determined as the enhancement range.
[0045] Steps 1 through 9, as an inventive feature of the present disclosure, address the first technical issue mentioned in the background art: "When manually annotating infrared image datasets, the invisible infrared signal intensity needs to be mapped to a pseudo-color encoding using visible light channels such as RGB / HSV. This pseudo-color encoding may mask the original grayscale features, exacerbate noise interference during model training, and lead to reduced accuracy." The factors that contribute to the aforementioned technical issue are often as follows: In complex multi-target scenes, due to inaccurate grayscale sub-range demarcation (for example, the boundary between high-temperature areas and the background is blurred), pseudo-color encoding can easily misjudge noise as feature signals. For high-resolution infrared images, domain merging based on a fixed threshold can result in the loss of key feature region information, such as targets with small temperature differences. Therefore, the present disclosure designs a dynamic boundary optimization strategy, introducing a domain difference calculation weighted by pixel probability density, combined with unsupervised clustering to pre-divide the initial boundary values, to improve the accuracy of grayscale sub-range demarcation. An adaptive merging threshold is introduced to merge and correct the optimized boundary values for out-of-bounds, ensuring feature integrity and preserving targets with small temperature differences.
[0046] Step 103: For each enhancement range in the at least one enhancement range, generate enhancement parameters corresponding to the enhancement range.
[0047] In some embodiments, the execution entity may generate, for each enhancement range in the at least one enhancement range, an enhancement parameter corresponding to the enhancement range, wherein the enhancement parameter may be a mathematical coefficient for adjusting pixel values within the enhancement range.
[0048] As an example, the enhancement coefficient may be a contrast stretch coefficient that maps the selected enhancement range into the enhancement range, such as a coefficient that maps the enhancement range [800, 1200] to the enhancement range [0, 255].
[0049] In some optional implementations of some embodiments, the execution entity may generate enhancement parameters corresponding to the enhancement range, which may include the following steps:
[0050] The first step is to obtain the lower limit and upper limit corresponding to the enhancement range. For example, if the enhancement range is [0, 750], the corresponding lower limit may be 0 and the upper limit may be 750.
[0051] The second step is to calculate the mean and variance based on the lower limit and upper limit. The mean can be the center of the data distribution, and the variance can describe the degree of dispersion in the data distribution. In infrared image coding, the mean can be the middle value of the enhancement range. The variance can be a dynamic indicator used to control the enhancement range and enhance the intensity. For example, the variance can be calculated as one-sixth of the difference between the upper limit and the lower limit of the enhancement range.
[0052] The third step is to determine the above mean and variance as enhancement parameters. For example, the enhancement range can be [750, 5000], the mean can be 2875, and the variance can be 3.0×10 6 , then the enhancement parameters can be 2875 and 3.0×10 6 .
[0053] Step 104: construct a hybrid mapping function according to the grayscale range, the mapping range and the obtained enhancement parameter set.
[0054] In some embodiments, the execution entity may construct a hybrid mapping function based on the grayscale range, the mapping range, and the obtained enhancement parameter set, wherein the hybrid mapping function may be a composite function that combines a global grayscale mapping function with a local image enhancement function.
[0055] As an example, the hybrid mapping function may be a superposition of a global mapping function and a local enhancement function.
[0056] In forest fire monitoring, the hybrid mapping function can be a piecewise function that selects the enhanced range and original range of the infrared image for different temperature regions. For example, the original grayscale range can be [0, 16383], the target compressed range can be [0, 255], the high temperature region can be [1200, 1600], and the medium and low temperature region can be [0, 800]. The global mapping function can be a linear function that maps the grayscale range to the target compressed range, and the local enhancement function can be a function that enhances the high temperature region and the medium and low temperature region.
[0057] In some optional implementations of some embodiments, the hybrid mapping function may be a maximum function of the output value of the linear mapping function and the output value of the Gaussian enhancement function, wherein the linear mapping function is a linear function that linearly maps the grayscale range to the mapping range. For example, the linear function may be in the following form:
[0058]
[0059] O L It can be the lower limit of the grayscale range of the original image, O U It can be the upper limit of the grayscale range of the original image, P LIt can be the lower limit of the grayscale range of the image after mapping, P U It can be the lower limit of the grayscale range of the mapped image, x can be the pixel value in the image, and B(x) can be the linear mapping value of the grayscale value in the original image after linear mapping.
[0060] The value range of the linear function is limited to the mapping range. The slope of the linear function is the ratio of the first difference to the second difference, where the first difference is the difference between the upper limit of the mapping range and the lower limit of the mapping range, and the second difference is the difference between the upper limit of the grayscale range and the lower limit of the grayscale range. For example, the slope can be in the following form:
[0061]
[0062] The intercept of the linear function is the lower limit of the mapping range. The Gaussian enhancement function is a nonlinear Gaussian function that performs grayscale enhancement in at least one enhancement range. For example, the nonlinear Gaussian function can be in the following form:
[0063]
[0064] μ n It can be the mean parameter of the nth group enhancement range, σ n It can be the variance parameter of the nth group enhancement range. n (x) may be a Gaussian enhancement function corresponding to the enhancement range of the nth group, n may be the enhancement range of the nth group, and x may be a grayscale value of the original image in the grayscale interval corresponding to the enhancement range of the nth group.
[0065] The value range of the Gaussian function is the mapping range. The offset of the Gaussian function is the lower limit of the mapping range. The scaling factor of the Gaussian function is the difference between the upper limit of the mapping range and the lower limit of the mapping range. For example, the scaling factor can be P L -P U .
[0066] As an example, the hybrid mapping function can be expressed as follows:
[0067] f(x)=max{B(x), f1(x), f2(x),..., f M (x), x∈D}
[0068] max{} can be the maximum value of all functions in the brackets. M is the number of divided enhancement groups, and f(x) can be the linear mapping value B(x) and all enhancement functions f for each gray value x. n (x), and take the maximum value as the final mapping result.
[0069] Step 105: Input each pixel value in the original infrared image into a hybrid mapping function to generate a mapped pixel value, thereby obtaining a mapped image.
[0070] In some embodiments, the execution entity may input each pixel value in the original infrared image into the hybrid mapping function to generate a mapped pixel value, thereby obtaining a mapped image. The mapped image value may be a pixel value of the original infrared image, obtained by applying the hybrid mapping function. The mapped image may be an image outputted after all pixel values of the original image are processed by the hybrid function.
[0071] In forest fire monitoring, the original infrared image can be a 14-bit grayscale image captured by an infrared camera. The pixel value range of the high-temperature flame area can be [1200, 1600], the temperature pixel value range of the background trees or ground can be [0, 800], and the ultra-high temperature flame temperature can be a pixel value greater than 1600. These ranges can be mapped to a target range using a hybrid mapping function. The target range can be [0, 255]. After mapping, the pixel value range of the high-temperature flame area can be [150, 255], the temperature pixel value range of the background trees or ground can be [0, 150], and the ultra-high temperature flame temperature pixel value can be 255. The mapped pixels are arranged according to the spatial form of the original image to obtain a mapped image.
[0072] Step 106 : Convert each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment to obtain a pulse sequence set.
[0073] In some embodiments, the execution entity may convert each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment to obtain a pulse sequence set. The pulse sequence may be a binary sequence consisting of multiple time steps for a single pixel. In practice, the time step may be the number of trials performed in the Bernoulli experiment. For example, if the number of Bernoulli trials is 4, the pulse sequence for a single pixel may be [0, 1, 1, 0].
[0074] In some optional implementations of some embodiments, the execution entity may convert each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment, which may include the following steps:
[0075] The first step is to determine the lower and upper limits of the enhancement range corresponding to the mapped pixel value. The lower limit can be the grayscale boundary value at the lower limit of the enhancement range, and the upper limit can be the grayscale value at the upper limit of the enhancement range. In practice, first, all enhancement range sets are traversed to find a range that satisfies the current mapped pixel value. Then, the lower and upper limits of this range are determined.
[0076] The second step is to determine the single-step probability corresponding to the above-mentioned mapped pixel value based on the above-mentioned lower limit value and upper limit value, wherein the above-mentioned single-step probability is the ratio of the first difference to the second difference, the above-mentioned first difference is the difference between the above-mentioned pixel value and the lower limit value of the above-mentioned enhancement range, and the above-mentioned second difference is the difference between the upper limit value of the above-mentioned enhancement range and the lower limit value of the above-mentioned enhancement range.
[0077] In practice, the single-step probability can be the probability of generating a pulse at each time step, calculated based on the relative position of the pixel value in the enhancement range.
[0078] As an example, the enhancement range may be [5000, 16383], and the mapped pixel value may be 10000. Then the first difference may be 10000-5000=5000, the second difference may be 16383-5000=11383, and the single-step probability may be 5000 / 11383≈0.44.
[0079] The third step is to perform the following determination steps for each trial in the above Bernoulli experiment:
[0080] The first sub-step is to generate a random number corresponding to the above experiment. The random number can be a random variable that follows a uniform distribution in the interval [0, 1) to simulate the randomness of the Bernoulli experiment. For example, the random number can be 0.40.
[0081] In the second sub-step, in response to the random number being less than the single-step probability, determining that the experimental result of the experiment is a positive result. A positive result may be an event determined as a success in a Bernoulli experiment, corresponding to a pulse of 1. For example, the random number may be 0.40 and the single-step probability may be 0.44. Therefore, 0.40 < 0.44, and the pulse is 1.
[0082] In a third sub-step, in response to the random number being not less than the single-step probability, determining that the experimental result of the experiment is a negative result. A negative result may be an event determined as a failure in a Bernoulli experiment, corresponding to a pulse of 0. For example, the random number may be 0.60 and the single-step probability may be 0.44. Therefore, 0.60 > 0.44, and the pulse is 0.
[0083] The fourth step is to summarize the experimental results corresponding to the above-mentioned Bernoulli experiment to obtain the above-mentioned pulse sequence. The pulse sequence can be a sequence composed of the Bernoulli experiment results of the pixels of the image at multiple time steps.
[0084] Step 107 , splicing each pulse sequence in the pulse sequence set to generate three-dimensional data, thereby obtaining a three-dimensional data set as the infrared image encoding result.
[0085] In some embodiments, the execution entity may concatenate each pulse sequence in the pulse sequence set to generate three-dimensional data, thereby obtaining a three-dimensional data set as the infrared image encoding result. The pulse sequence set may be a set of pulse sequences for all pixels in the mapped image, and the three-dimensional data may be in the form of a data structure consisting of the width, height, and time step of the original infrared image.
[0086] In practice, each pulse sequence in the pulse sequence set corresponds to the pulse record of a pixel at multiple time steps, which can be in the form of a binary array. The data structure can be in the form of a three-dimensional cube composed of the width, height and time step of the original infrared image.
[0087] In some optional implementations of some embodiments, the execution entity may splice each pulse sequence in the pulse sequence set to generate three-dimensional data. Obtaining a three-dimensional data set may include the following steps:
[0088] The first step is to obtain the pixel width and height of the original infrared image and the length of each pulse sequence. The width and height can be the horizontal and vertical resolution of the image, and the number of columns and rows of the image. The length of the pulse sequence can be the number of trials in the Bernoulli experiment.
[0089] As an example, first, use the image processing library to read the width and height of the original infrared image. Finally, determine the number of experiments based on the Bernoulli experiment design or requirements. For example, the original image resolution can be 3×2 (width w, height h). The pulse sequence length (n) can be 3.
[0090] The second step is to construct an initial three-dimensional matrix based on the width, height, and length. As an example, create a three-dimensional array of all zeros with a shape of (width, height, length).
[0091] Step 3: For each of the pulse sequences in the pulse sequence set, perform the following steps:
[0092] In a first sub-step, pixel positions of the original infrared image corresponding to the pulse sequence are determined to obtain fill positions, where the fill positions include a first fill position in the width direction of the original infrared image and a second fill position in the height direction of the original infrared image. For example, the fill positions may be the coordinates of the pixels corresponding to the pulse sequence in the image (first fill position, second fill position).
[0093] As an example, the filling position is obtained by traversing each pixel in the image in row-major or column-major order.
[0094] The second sub-step is to fill the pulse sequence into the matrix position corresponding to the initial three-dimensional matrix in chronological order according to the filling position, wherein the matrix position is the time step corresponding to the depth direction, corresponding to a pulse value in the pulse sequence.
[0095] As an example, the pulse sequence of pixel (i, j) can be [1, 0, 1]. The time step can be 3. Then the matrix position (i, j) corresponding to the three-dimensional matrix M can be, for time step k = 0, M[i, j, 0] = 1. For time step k = 1, M[i, j, 1] = 0. For time step k = 2, M[i, j, 2] = 1.
[0096] The fourth step is to determine that the initial three-dimensional matrix is filled with three-dimensional data of pulse sequences of multiple time steps by filling all pulse sequences.
[0097] As an example, first, all pixels are traversed to ensure that the pulse sequence corresponding to each pixel has been filled. Finally, the three-dimensional matrix is checked to see if there are any unfilled gaps. For example, the filled three-dimensional matrix can be M = [[[1, 0, 1], [0, 1, 0]], [[1, 1, 0], [0, 0, 1]], [[0, 1, 1], [1, 0, 0]]].
[0098] The fifth step is to determine that the three-dimensional data is the three-dimensional data set.
[0099] Step 108: Input the infrared image encoding result into the spiking neural network for training.
[0100] In some embodiments, the aforementioned execution entity may input the infrared image encoding result into a spiking neural network for training. As an example, first, the infrared image encoding result is determined to be a w×h×n cube structure, where the w and h directions represent the width and height of the original infrared image, and the n direction represents the time direction of pulse generation. Next, the time direction is used as the input direction of the neural network, and n corresponding w×h sized pulse maps are generated. Finally, the pulse maps are input into the spiking network for training.
[0101] While employing technical solutions to address the second technical issue, the following challenges often arise: Real-time pseudo-color conversion and signal processing during the computer deployment phase consume significant computing power. Traditional spiking neural network training requires strictly synchronized time steps for the input of three-dimensional spatiotemporal data, resulting in low hardware resource utilization, long training times, and failure to leverage the advantages of parallel computing.
[0102] Conventional solutions to these problems typically involve downsampling high-resolution infrared images to reduce the resolution and data size, thereby alleviating the load. Separating pseudo-color encoding conversion and model training from one stage to avoid real-time conversion overhead is also a common approach. However, the inventors considered that downsampling could result in loss of grayscale information for targets with small temperature differences, thus reducing model generalization. Separating pseudo-color encoding conversion and model training from one stage to another would not meet the real-time inference requirements of dynamic scenarios, such as drone inspections. Therefore, we decided to adopt the following solution.
[0103] In some optional implementations of some embodiments, the execution entity may input the infrared image encoding result into a spiking neural network for training, which may include the following steps:
[0104] The first step is to configure the infrared image encoding result into a spatiotemporal structure. The spatiotemporal structure includes the width, height, and time step of the original infrared image, where the time step is the number of Bernoulli trials. The spatiotemporal structure can be a three-dimensional data cube with a shape of width (w) × height (h) × time step (n).
[0105] In the second step, based on the aforementioned spatiotemporal structure, the two-dimensional pulse map at each time step is used as the sequential input to the spiking neural network in chronological order. The input data for each time step is a two-dimensional matrix consisting of the pulse values of all pixels at the current time step. This two-dimensional pulse map can be a width (w) x height (h) matrix corresponding to each time step (k), where the elements in the matrix are the pulse values of all pixels at the current time step. This sequential input is the input processed by the spiking neural network according to the time step, simulating the dynamic response of a biological neural network.
[0106] The third step is to dynamically adjust the membrane potential and synaptic weights of neurons during training using an event-driven spiking neuron model. The event-driven spiking neuron model can be one in which the neuron updates its membrane potential only when it receives an input pulse, remaining silent otherwise. The membrane potential can be calculated based on the input pulse and the current potential. The synaptic weights can be adjusted based on the connection weights between neurons based on spike timing correlation. Spike timing correlation can be based on the STDP (Spike Timing Dependent Plasticity) rule.
[0107] The fourth step is to combine the asynchronous parallel computing framework to split the spatiotemporal cube data into multiple sub-blocks and distribute them to neuromorphic computing chips for distributed processing to accelerate the training process of the spiking neural network. The asynchronous parallel computing framework can be used to split the data into independent sub-blocks, each of which is processed in parallel on different computing units without the need for time synchronization. The neuromorphic computing chip can be designed for spiking computing, supporting event-driven and low power consumption. For example, the Loihi chip or the TrueNorth chip.
[0108] As an example, first, the data is partitioned into multiple sub-blocks. Next, each sub-block is assigned to a core on the chip to independently process spike events. Finally, the spikes or weight updates are aggregated and output.
[0109] The first to fourth steps mentioned above, as an inventive point of the present disclosure, solve the second technical problem mentioned in the background technology, "real-time pseudo-color conversion and signal processing in the computer deployment stage will occupy a large amount of computing resources, resulting in redundant computing power and waste of energy." The factors that lead to the above technical problems are often as follows: In complex multi-target scenes, due to the inaccurate division of grayscale sub-ranges (for example, the boundary between high-temperature areas and backgrounds is blurred), pseudo-color coding mapping can easily misjudge noise as a feature signal. For high-resolution infrared images, merging based on fixed threshold areas may lead to the loss of key feature area information. For example, targets with small temperature differences.
[0110] Therefore, the present invention designs a spatiotemporal stereo structure and a one-step parallel computing architecture. By constructing a three-dimensional spatiotemporal data structure (width × height × time step), the high-dimensional pulse sequence is divided into independent sub-blocks, and the event-driven characteristics of neuromorphic computing chips (for example, Loihi chips) are utilized to achieve asynchronous distributed parallel processing, breaking through the traditional time synchronization limitations and significantly improving computing efficiency and resource utilization.
[0111] Optionally, the above execution entity may further perform the following steps:
[0112] The first step is to obtain the training results of the pulse neural network to obtain an infrared image pulse model. The infrared image pulse model can be a model obtained by training the pulse neural network based on the pulse sequence corresponding to the infrared image.
[0113] As an example, first, a trained model file (e.g., a model weight file) can be loaded from a storage device (e.g., a hard disk). Second, the model file is loaded and the trained weight parameters are loaded into the model network structure to construct a usable infrared image pulse model.
[0114] The second step is to input the original infrared image into the infrared image pulse model to obtain a pulse result sequence. The pulse result sequence can be a three-dimensional tensor (width × height × time step) of the pulse emission at each time step output by the spiking neural network.
[0115] The third step is to perform dimensional decoupling on the pulse sequence and determine the corresponding pulse frequency in each dimension to obtain the dimensional pulse quantity. The dimensional decoupling can be performed by spatially decomposing the pulse sequence into two dimensions: height and width, while maintaining the independent dimension of the time step. The different dimensions are then processed separately. The pulse frequency can be the number of pulses emitted per unit time or per unit space in a specific dimension. The dimensional pulse quantity can be the pulse frequency information obtained by statistics in different dimensions.
[0116] As an example, first, decouple the resulting pulse sequence spatially and temporally, obtaining spatial and temporal dimensions. Next, count the total number of pulses over the entire time span at each spatial location in each spatial dimension. Divide this total number of pulses by the time step to obtain the spatial pulse frequency. At each time step, count the total number of pulses at the corresponding spatial location. Divide this total number of pulses by the total number of spatial pixels to obtain the average pulse frequency for each time step. Finally, the spatial pulse frequency and the average temporal pulse frequency are used as the dimensional pulse quantity.
[0117] In the fourth step, based on the three-dimensional data set, relative position mapping is performed on the pulse quantities in the above dimensions to obtain a relative position offset. The relative position mapping may be mapping the pulse frequency information to a position change in actual three-dimensional space. The relative position offset may be the offset of the target relative to the current camera center position.
[0118] As an example, first, load the mapping relationship between image pixel coordinates and 3D space coordinates from a 3D dataset. Filter the dimensional pulse quantity for locations with high pulse frequencies and use them as target pixel coordinates. Next, convert the angles in real space (e.g., azimuth) based on the target pixel coordinates and the mapping relationship. Convert the coordinates corresponding to the current camera to real-space angles. Calculate the difference between the two angles to obtain the relative position offset.
[0119] Step 5: Generate an offset instruction for the relative position offset to obtain a camera control instruction. The offset instruction may be an instruction for controlling the camera's rotation direction (e.g., horizontal and vertical) and rotation amount (e.g., angle and step length). The camera control instruction may be a string instruction specifically for controlling camera rotation.
[0120] As an example, the direction of camera rotation can be determined based on the relative position offset, and converted into the instruction format required for camera control based on the rotation direction.
[0121] The sixth step is to send the camera control command to the camera control terminal to activate the infrared camera. The camera control terminal can be a hardware device (eg, a cloud controller) that receives the control command and drives the infrared camera to rotate.
[0122] As an example, the camera control instruction can be sent through the communication connection (e.g., serial port and network port) of the camera control terminal. The camera control terminal can transmit the control instruction by calling the interface of the infrared camera to drive the infrared camera.
[0123] The aforementioned embodiments of the present disclosure have the following beneficial effects: infrared images obtained through the infrared image encoding methods of some embodiments of the present disclosure reduce the time required for pseudo-color conversion in computer recognition tasks, thereby improving computing speed and related performance. Specifically, the reduction in computer computing speed and related performance is caused by the fact that real-time pseudo-color conversion and signal processing during the computer deployment phase consume a large amount of computing resources, resulting in redundant computing power and wasted energy. Based on this, the infrared image encoding methods for spiking neural networks of some embodiments of the present disclosure first determine the grayscale range and mapping range of the original infrared image based on the image histogram corresponding to the original infrared image. The mapping range is used to adapt the input of the spiking neural network. This avoids the problem of reduced training efficiency caused by a mismatch between the input data range and the spiking neural network, ensuring direct data format compatibility. Then, at least one enhancement range within the grayscale range is determined for feature enhancement. This focuses on key areas, provides a target range for feature enhancement, and improves sensitivity to specific targets. Next, for each of the at least one enhancement range, an enhancement parameter corresponding to the enhancement range is generated. Thus, an enhancement factor is generated for each enhancement region, providing data for constructing an enhancement function. Next, a hybrid mapping function is constructed based on the grayscale range, the mapping range, and the obtained enhancement parameter set. This achieves a balance between global grayscale adaptation and local feature enhancement through linear and nonlinear changes, dynamically optimizing the image range. Each pixel value in the original infrared image is then input into the hybrid mapping function to generate a mapped pixel value, resulting in a mapped image. This produces a contrast-enhanced image that separates key target areas from the original background, reducing noise interference during model training. Then, using a Bernoulli experiment, each mapped pixel value in the mapped image is converted into a pulse sequence, resulting in a pulse sequence set. This converts continuous pixel values into discrete pulse sequences, which simulate the activation characteristics of biological neurons and reduce computational redundancy. Next, each pulse sequence in the pulse sequence set is concatenated to generate three-dimensional data, resulting in a three-dimensional dataset as the infrared image encoding result. This creates a spatiotemporal fusion encoding structure that preserves the characteristics of the original image and meets the temporal processing requirements of the spiking neural network. Finally, the infrared image encoding result is input into the spiking neural network for training. Therefore, the spatiotemporal training characteristics of spiking neural networks can be directly utilized for training, eliminating the traditional pseudo-color conversion step. In summary, by converting the original image pixels into a coding structure with spatiotemporal time sequence to adapt the training of spiking neural networks, the pseudo-color conversion required for manual annotation can be reduced, thereby reducing redundant computing power and wasted energy.
[0124] Further references Figure 2As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an infrared image encoding device for a pulse neural network. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the infrared image encoding device for pulse neural network can be specifically applied to various electronic devices.
[0125] like Figure 2 As shown, an infrared image encoding device 200 for a spiking neural network includes a determination unit 201, an enhancement range screening unit 202, a parameter generation unit 203, a function construction unit 204, an image mapping unit 205, a spiking conversion unit 206, a three-dimensional encoding unit 207, and a training input unit 208. The determination unit 201 is configured to determine the grayscale range and mapping range of the original infrared image based on an image histogram corresponding to the original infrared image, wherein the mapping range is used to adapt the input of the spiking neural network. The enhancement range screening unit 202 is configured to determine at least one enhancement range within the grayscale range for feature enhancement. The parameter generation unit 203 is configured to generate enhancement parameters corresponding to each enhancement range within the at least one enhancement range. The function construction unit 204 is configured to construct a hybrid mapping function based on the grayscale range, the mapping range, and the obtained enhancement parameter set. The image mapping unit 205 is configured to input each pixel value in the original infrared image into the hybrid mapping function to generate a mapped pixel value, thereby obtaining a mapped image. The pulse conversion unit 206 is configured to convert each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment, thereby obtaining a pulse sequence set. The 3D encoding unit 207 is configured to concatenate each pulse sequence in the pulse sequence set to generate 3D data, thereby obtaining a 3D data set as the infrared image encoding result. The training input unit 208 is configured to input the infrared image encoding result into the spiking neural network for training.
[0126] It is understood that the units described in the infrared image encoding device 200 for pulse neural network are similar to those in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the infrared image encoding device 200 for pulse neural network and the units included therein, and will not be repeated here.
[0127] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device (eg, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0128] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0129] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0130] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0131] It should be noted that in some embodiments of the present disclosure, the computer-readable medium mentioned above may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0132] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0133] The computer-readable medium may be included in the electronic device, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: determine a grayscale range and a mapping range of the original infrared image based on an image histogram corresponding to the original infrared image, wherein the mapping range is a range used to adapt the input of a spiking neural network; determine at least one enhancement range within the grayscale range for feature enhancement; generate enhancement parameters corresponding to the enhancement range for each enhancement range within the at least one enhancement range; construct a hybrid mapping function based on the grayscale range, the mapping range, and the obtained enhancement parameter set; input each pixel value in the original infrared image into the hybrid mapping function to generate a mapped pixel value, thereby obtaining a mapped image; convert each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment, thereby obtaining a pulse sequence set; concatenate each pulse sequence in the pulse sequence set to generate three-dimensional data, thereby obtaining a three-dimensional data set as an infrared image encoding result; and input the infrared image encoding result into a spiking neural network for training.
[0134] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0136] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The units described may also be provided in a processor. For example, they may be described as follows: a processor including a determination unit, an enhancement range screening unit, a parameter generation unit, a function construction unit, an image mapping unit, a pulse conversion unit, a three-dimensional encoding unit, and a training input unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the determination unit may also be described as a "unit for determining the grayscale range and mapping range of the original infrared image according to the image histogram corresponding to the original infrared image."
[0137] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0138] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An infrared image encoding method for a spiking neural network, comprising: Determining a grayscale range and a mapping range of the original infrared image according to an image histogram corresponding to the original infrared image, wherein the mapping range is a range used to adapt to the input of the spiking neural network; Determining at least one enhancement range to be feature enhanced in the grayscale range; For each enhancement range in the at least one enhancement range, generating an enhancement parameter corresponding to the enhancement range; constructing a hybrid mapping function according to the grayscale range, the mapping range and the obtained enhancement parameter set; Inputting each pixel value in the original infrared image into the hybrid mapping function to generate a mapped pixel value to obtain a mapped image; Converting each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment to obtain a pulse sequence set; splicing each pulse sequence in the pulse sequence set to generate three-dimensional data, thereby obtaining a three-dimensional data set as an infrared image encoding result; The infrared image encoding result is input into a pulse neural network for training.
2. The method according to claim 1, wherein The method further comprises: Obtaining the training results of the pulse neural network to obtain an infrared image pulse model; Inputting the original infrared image into the infrared image pulse model to obtain a pulse result sequence; Performing dimensional decoupling on the pulse result sequence and determining the pulse frequency corresponding to each dimension to obtain the dimensional pulse quantity; According to the three-dimensional data set, relative position mapping is performed on the dimensional pulse quantity to obtain a relative position offset; Generating an offset instruction for the relative position offset to obtain a camera control instruction; The camera control instruction is sent to the camera control terminal to mobilize the infrared camera.
3. The method according to claim 1, wherein Generating the enhancement parameters corresponding to the enhancement range includes: Obtaining a lower limit value and an upper limit value corresponding to the enhancement range; Based on the lower limit value and the upper limit value, a mean and a variance are obtained; The mean and the variance are determined as enhancement parameters.
4. The method according to claim 1, wherein The mixed mapping function is a maximum function of the output value of the linear mapping function and the output value of the Gaussian enhancement function, wherein the linear mapping function is a linear function that linearly maps the grayscale range to the mapping range; the value range of the linear function is the mapping range; the slope of the linear function is the ratio of the first difference to the second difference, the first difference is the difference between the upper limit value of the mapping range and the lower limit value of the mapping range, and the second difference is the difference between the upper limit value of the grayscale range and the lower limit value of the grayscale range; the intercept of the linear function is the lower limit value of the mapping range; the Gaussian enhancement function is a nonlinear Gaussian function that enhances the grayscale value in the at least one enhancement range; the value range of the nonlinear Gaussian function is the mapping range; the offset of the nonlinear Gaussian function is the lower limit value of the mapping range; the scaling factor of the nonlinear Gaussian function is the difference between the upper limit value of the mapping range and the lower limit value of the mapping range.
5. The method according to claim 1, wherein The step of converting each mapped pixel value in the mapped image into a pulse sequence through the Bernoulli experiment comprises: Determining a lower limit value and an upper limit value of an enhancement range corresponding to the mapped pixel value; Determining a single-step probability corresponding to the mapped pixel value based on the lower limit and the upper limit, wherein the single-step probability is a ratio of a first difference value to a second difference value, the first difference value being a difference between the pixel value and the lower limit of the enhancement range, and the second difference value being a difference between the upper limit of the enhancement range and the lower limit of the enhancement range; For each trial in the Bernoulli experiment, the following determination steps are performed: Generate a random number corresponding to the experiment; In response to the random number being less than the single-step probability, determining that the experimental result of the experiment is a positive result; In response to the random number being not less than the single-step probability, determining that the experimental result of the experiment is a negative result; The experimental results corresponding to the Bernoulli experiment are summarized to obtain the pulse sequence.
6. The method according to claim 1, wherein The step of splicing each pulse sequence in the pulse sequence set to generate three-dimensional data and obtain a three-dimensional data set includes: Obtaining the width and height of the pixels of the original infrared image, and the length of each pulse sequence; constructing an initial three-dimensional matrix according to the width, the height, and the length; For each pulse sequence in the pulse sequence set, performing the following steps: Determining pixel positions of the original infrared image corresponding to the pulse sequence to obtain filling positions, wherein the filling positions include a first filling position in a width direction of the original infrared image and a second filling position in a height direction of the original infrared image; According to the filling position, the pulse sequence is sequentially filled into the matrix position corresponding to the initial three-dimensional matrix in time order, wherein the matrix position is a time step corresponding to the depth direction and corresponds to a pulse value in the pulse sequence; By filling all pulse sequences, determining that the initial three-dimensional matrix is three-dimensional data filled with pulse sequences of multiple time steps; The three-dimensional data is determined to be the three-dimensional data set.
7. An infrared image encoding device for a spiking neural network, comprising: a determining unit configured to determine a grayscale range and a mapping range of the original infrared image based on an image histogram corresponding to the original infrared image, wherein the mapping range is a range used to adapt the spiking neural network input; an enhancement range screening unit, configured to determine at least one enhancement range to be feature enhanced in the grayscale range; a parameter generating unit configured to generate, for each enhancement range in the at least one enhancement range, an enhancement parameter corresponding to the enhancement range; a function construction unit configured to construct a hybrid mapping function according to the grayscale range, the mapping range and the obtained enhancement parameter set; an image mapping unit configured to input each pixel value in the original infrared image into the hybrid mapping function to generate a mapped pixel value to obtain a mapped image; a pulse conversion unit configured to convert each mapped pixel value in the mapped image into a pulse sequence through a Bernoulli experiment to obtain a pulse sequence set; a three-dimensional encoding unit configured to splice each pulse sequence in the pulse sequence set to generate three-dimensional data and obtain a three-dimensional data set as an infrared image encoding result; The training input unit is configured to input the infrared image encoding result into the pulse neural network for training.
8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Federal learning method of pulse neural network for flaw detection
CN117875408A
Width pulse neural network-based image classification model training method and device
CN118968151A
Method and device for enhancing image contrast
WO2020082593A1
Image processing method and apparatus, electronic device, and storage medium
WO2024198594A1