Visual feature extraction method and system based on multi-modal image enhancement

By locating key identity regions in visible light images using infrared images and performing high-frequency data compensation, the problem of identity feature loss caused by visible light image compression under low bandwidth is solved, realizing the combination of target detection and identity recognition under limited bandwidth.

CN121437908BActive Publication Date: 2026-04-21JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Under low bandwidth conditions, the high compression of visible light images in multimodal visual surveillance systems leads to the loss of texture details in key identity areas, resulting in the inability of identity recognition confidence to reach a reliable threshold.

Method used

By utilizing the thermal features of infrared images to locate key identity regions in visible light images, and combining bit-plane analysis with high-frequency data compensation, details of the key identity regions lost during compression are recovered by iteratively adjusting the compensation intensity.

Benefits of technology

This technology recovers key identity-critical region details lost due to compression in visible light images under low bandwidth conditions, improving recognition accuracy and solving the challenge of failing to confirm target identity in traditional technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121437908B_ABST
    Figure CN121437908B_ABST
Patent Text Reader

Abstract

This invention discloses a visual feature extraction method and system based on multimodal image enhancement, relating to the field of image processing technology. The method includes: receiving an infrared image, a visible light image, and compression parameters of the visible light image; obtaining thermal feature information of a target object based on the infrared image, and locating its key identity region in the visible light image accordingly, while extracting edge feature information of the region based on the infrared image; decomposing the key identity region bit plane to obtain multiple bit planes, identifying missing bit planes based on the compression parameters, and then compensating for high-frequency data of the missing bit planes with a preset compensation intensity based on the edge feature information; reconstructing the key identity region and performing identity recognition to obtain a set of recognition confidence scores; if the maximum value in the set is lower than a preset confidence threshold, adjusting the compensation intensity for iterative compensation; its beneficial effect is that it can recover details of key identity regions lost due to compression in visible light images under low bandwidth conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method and system for visual feature extraction based on multimodal image enhancement. Background Technology

[0002] With the development of security monitoring technology, multimodal visual monitoring systems have been widely used. These systems typically deploy both visible light cameras and infrared thermal imaging cameras to achieve all-weather, all-time target monitoring capabilities. Visible light cameras can capture rich texture and color information, facilitating target identification. Infrared thermal imaging cameras are not sensitive to lighting conditions and can stably detect the location and thermal characteristics of targets at night or in adverse weather conditions.

[0003] However, communication infrastructure in remote areas is often weak, with available bandwidth typically only 50-200kbps. Under such bandwidth constraints, monitoring systems need to highly compress video data to achieve real-time transmission. Infrared thermal imaging data, due to its single-channel and low-resolution characteristics, has a relatively small data volume and can be transmitted with high quality. In contrast, visible light video data has a large data volume and must be compressed at an extremely low bit rate under extremely low bandwidth. This inevitably introduces obvious block effects, blurring, and loss of texture details, resulting in quantization distortion that severely damages key visual features that can be used for identification.

[0004] However, existing technologies mostly focus on using multimodal data to improve the robustness of detection and tracking, but have not solved the fundamental problem that the identification features are damaged at the transmission end. The limitation is that there is an essential information mismatch between the thermal features provided by the infrared mode and the identity details that the visible light mode should provide. This limitation makes existing technologies face the challenge of "being able to detect the target, but unable to confirm the target's identity".

[0005] Therefore, a visual feature extraction method and system based on multimodal image enhancement is proposed. Summary of the Invention

[0006] In view of the above-mentioned prior art, this application is hereby proposed. Embodiments of this application provide a visual feature extraction method and system based on multimodal image enhancement, which can recover details of key identity regions lost due to compression in visible light images under low bandwidth conditions, thereby improving recognition accuracy.

[0007] According to one aspect of this application, a visual feature extraction method based on multimodal image enhancement is provided, comprising: S1, receiving a fully transmitted infrared image of a target region, a compressed transmitted visible light image, and compression parameters of the visible light image; S2, obtaining thermal feature information of a target object based on the infrared image; S3, locating a key identity region of the target object in the visible light image based on the thermal feature information; S4, extracting edge feature information of the key identity region based on the infrared image; S5, performing bit-plane decomposition on the key identity region to obtain multiple bit planes; S6, based on the compression parameters of the infrared image... S7. Based on the edge feature information, perform high-frequency data compensation on the missing bit planes in the multiple bit planes with a preset compensation intensity. S8. Reconstruct the identity key region based on the multiple bit planes. S9. Perform identity recognition on the reconstructed identity key region to obtain an identification confidence set. S10. Determine whether the maximum value in the identification confidence set is lower than a preset confidence threshold. If yes, adjust the compensation intensity and return to S7 for iteration; otherwise, extract the identity features of the target object based on the reconstructed identity key region.

[0008] According to another aspect of this application, a visual feature extraction system based on multimodal image enhancement is provided, comprising: a data acquisition module for receiving a fully transmitted infrared image of a target region, a compressed transmitted visible light image, and compression parameters of the visible light image; a thermal feature acquisition module for acquiring thermal feature information of a target object based on the infrared image; a key region localization module for locating a key identity region of the target object in the visible light image based on the thermal feature information; an edge feature extraction module for extracting edge feature information of the key identity region based on the infrared image; a bit-plane decomposition module for performing bit-plane decomposition on the key identity region to obtain multiple bit planes; and missing feature identification. The system comprises the following modules: a module for identifying missing bit planes due to compression in the plurality of bit planes based on the compression parameters; a compensation module for performing high-frequency data compensation on the missing bit planes in the plurality of bit planes based on the edge feature information and the compensation intensity; a reconstruction module for reconstructing the identity key region based on the plurality of bit planes; a recognition module for performing identity recognition on the reconstructed identity key region to obtain a recognition confidence set; and a decision module for determining whether the maximum value in the recognition confidence set is lower than a preset confidence threshold. If so, the compensation intensity is adjusted with the goal of increasing the recognition confidence and the result is returned to the compensation module; otherwise, the identity features of the target object are extracted based on the reconstructed identity key region.

[0009] According to another aspect of this application, an electronic device is provided, including a memory and a processor, the memory being used to store computer-executable instructions, and the processor being used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method described above.

[0010] According to another aspect of this application, a computer storage medium is provided that stores computer-executable instructions thereon, which, when executed by a processor, implement the steps of the method described above.

[0011] Compared with the prior art, the visual feature extraction method and system based on multimodal image enhancement according to the embodiments of this application can use the thermal feature information of infrared images to locate the key identity regions in visible light images, combine compression parameters to identify missing bit planes and perform high-frequency data compensation, and iteratively adjust the compensation intensity until the recognition confidence requirements are met. This solves the problem of identity feature loss caused by visible light image compression under low bandwidth conditions and has the advantage of restoring the details of key identity regions lost due to compression in visible light images under low bandwidth conditions. Attached Figure Description

[0012] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0013] Figure 1 This is a flowchart of the visual feature extraction method based on multimodal image enhancement according to the present invention.

[0014] Figure 2 This is a block diagram of the visual feature extraction system based on multimodal image enhancement according to the present invention.

[0015] Figure 3 This is a block diagram of an electronic device according to the present invention. Detailed Implementation

[0016] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0017] Exemplary methods

[0018] In existing multimodal visual surveillance methods, the high compression rate transmission of visible light images under low bandwidth conditions leads to the loss of texture details in key identity areas, making it impossible for the confidence level of identity recognition to reach a reliable threshold.

[0019] To address the aforementioned challenges, this application first considers how to leverage the complementarity between fully transmitted infrared images and compressed visible light images. The thermal features of the infrared image can stably characterize the presence of the target object, while the visible light image, although damaged by compression, still retains some identity-discriminating features. To this end, this application finds that guiding the localization of key identity regions in the visible light image using thermal feature information can overcome spatial mapping bias caused by compression noise. Furthermore, the bit-plane decomposition method can accurately locate the missing high-frequency components, and high-frequency compensation based on edge features can specifically restore texture details. Moreover, to solve the dynamic matching problem between compensation intensity and recognition confidence, a confidence feedback mechanism is introduced to form a closed-loop adjustment, avoiding feature distortion caused by insufficient or over-compensation.

[0020] To address this, this application proposes a visual feature extraction method based on multimodal image enhancement, such as... Figure 1 The process includes: S1, receiving a complete transmitted infrared image of the target area, a compressed transmitted visible light image, and compression parameters of the visible light image; S2, obtaining thermal feature information of the target object based on the infrared image; S3, locating the key identity region of the target object in the visible light image based on the thermal feature information; S4, extracting edge feature information of the key identity region based on the infrared image; S5, performing bit-plane decomposition on the key identity region to obtain multiple bit planes; S6, identifying missing bit planes in the multiple bit planes due to compression based on the compression parameters; S7, performing high-frequency data compensation on the missing bit planes in the multiple bit planes with a preset compensation intensity based on the edge feature information; S8, reconstructing the key identity region based on the multiple bit planes; S9, performing identity recognition on the reconstructed key identity region to obtain the recognition confidence score with the target object as the identity recognition result; S10, determining whether the recognition confidence score is lower than a preset confidence threshold. If so, adjusting the compensation intensity with the goal of increasing the recognition confidence score and returning to S7 for iteration; otherwise, extracting the identity features of the target object based on the reconstructed key identity region.

[0021] Among them, the complete transmitted infrared image refers to infrared thermal imaging data transmitted without compression or using lossless compression algorithms. Specifically, it can be implemented using PNG or TIFF-LZW lossless compression formats to ensure that thermal feature information is not distorted during transmission, providing a reliable data foundation for subsequent thermal profile and thermal centroid extraction.

[0022] Among them, the visible light image transmitted by compression refers to visible light image data processed by lossy compression algorithms. Specifically, it can be implemented using H.264 or HEVC low bit rate encoding standards. Under limited bandwidth, basic image information is transmitted first, but high frequency details will be lost due to the quantization process.

[0023] Thermal feature information refers to the thermal radiation distribution characteristics of the target object in the infrared image. Specifically, it can be realized through the gray value matrix collected by the thermal imager, which includes temperature gradient change characteristics and is used to characterize the outline shape and thermal distribution center of the target object.

[0024] Among them, the key identity region refers to the local area in the visible light image that carries the identity discrimination features of the target object. The texture details of this region play a decisive role in identity recognition.

[0025] Bit-plane decomposition refers to the technique of processing the binary representation of image pixels in bits. Specifically, it can be implemented using the bit-level algorithm of 8-bit grayscale images, which decomposes the image into 8 binary bit planes, where the high bit plane carries the main structural information and the low bit plane contains high-frequency details.

[0026] Among them, missing bit planes refer to the loss of bit plane data caused by the priority discarding of high-frequency information during the compression process. Specifically, this can be achieved by analyzing the quantization table and the bit rate control parameters. Low bit rate compression will cause bit planes other than high bit planes to be truncated or set to zero.

[0027] Among them, high-frequency data compensation is a technique that injects edge gradient information into the missing part plane to restore details. Specifically, it can be achieved by using a directional filter to generate high-frequency noise that matches the edge features. The compensation intensity controls the noise energy level to balance detail restoration and the risk of overfitting.

[0028] Among them, the recognition confidence geometry refers to the probability value of all recognition results output by the identity recognition model. The maximum value is lower than the threshold, which indicates that the compensated image still cannot meet the accuracy requirements.

[0029] Among them, compensation intensity adjustment refers to the mechanism of dynamically optimizing high-frequency compensation parameters based on the recognition results. Specifically, it can be achieved by using the gradient descent method or a heuristic step size adjustment strategy. This can be done by increasing the compensation intensity to enhance the detail recovery or by decreasing the intensity to suppress noise interference.

[0030] In multimodal vision, infrared and visible light images capture target information using different imaging mechanisms, forming a natural complementary relationship. Infrared images, based on thermal radiation, can completely and stably reflect the thermal distribution contour, thermal center of gravity, and overall morphological structure of the target object. This information is generally unaffected by changes in ambient light or compression loss during transmission, and its geometric structural information can maintain a high degree of accuracy and consistency with the actual structure of the target object, providing a solid foundation for the spatial positioning of the target.

[0031] In contrast, while visible light images contain rich texture and color information, they require lossy compression in low-bandwidth transmission environments, leading to the destruction or loss of crucial high-frequency details and edge information. The compression process not only results in information loss but can also cause blurring and misalignment of spatial structures, making the localization of key regions and the extraction of identity features based on visible light images highly uncertain. Particularly in identity recognition tasks, the loss of texture details severely impacts recognition accuracy, becoming a bottleneck restricting system performance.

[0032] Therefore, infrared and visible light images exhibit significant complementary characteristics at both the structural and textural levels. Infrared images provide clear and robust geometric positioning data, accurately indicating the spatial location of key identification areas; while visible light images retain some identity-related texture remnants, carrying the high-frequency detail information required for final identification. By utilizing the complete and stable thermal profile and thermal centroid of infrared images, spatial mapping deviations and texture misalignments caused by visible light image compression can be effectively avoided, providing an accurate and reliable spatial reference for subsequent high-frequency compensation and detail restoration based on bit-plane decomposition.

[0033] Furthermore, the thermal structural information carried by infrared images is irreplaceable by visible light images. This unique structural integrity ensures that the compensation mechanism, while restoring texture details, maintains the morphological consistency and edge sharpness of key target identity regions, preventing feature distortion caused by over-compensation or deviation. Through the synergy of these two modalities in spatial localization and texture compensation, effective enhancement of target identity features is achieved in low-bandwidth environments.

[0034] The core innovation of this application lies in locating key identity regions in visible light images through the thermal features of infrared images, and combining bit-plane analysis with an iterative mechanism of high-frequency data compensation to directly repair identity recognition features damaged by compression, thereby achieving synergistic enhancement of multimodal data at the feature level.

[0035] Through the above-described scheme, this application can solve the problem of identity feature loss caused by visible light image compression in low-bandwidth environments. By utilizing the thermal feature information of infrared images to locate key identity regions, and combining bit-plane analysis and high-frequency data compensation techniques, it achieves enhancement of key identity regions in compressed images. The iteratively optimized compensation strategy ensures the accuracy of identity recognition, enabling an effective combination of target detection and identity recognition under limited bandwidth. This method overcomes the challenge faced by traditional technologies that "can detect targets, but cannot confirm their identities."

[0036] In some of the solutions described above in this application, locating the key identification regions of a target object in a visible light image includes:

[0037] First, the target object is categorized based on the infrared image. For example, the thermal feature distribution pattern of the infrared image can be used to identify different categories such as people, vehicles, or animals.

[0038] Next, the thermal profile and thermal centroid of the target object are extracted based on the thermal feature information. The thermal profile can be obtained by edge detection of the infrared image, while the thermal centroid can be obtained by calculating the centroid of the thermal feature distribution.

[0039] Then, the corresponding geometric constraints are obtained according to the object category, and the position coordinates of the center of the key identity region in the visible light image are calculated based on the geometric constraints, thermal profile, and thermal centroid. Specifically, geometric transformation or statistical learning can be used to establish the coordinate mapping relationship between the infrared and visible light images. It should be noted that the geometric constraints represent the mapping rules between the thermal profile and thermal centroid of the object category and the spatial position of the key identity region. For example, for a human target, the head region is usually located above the thermal profile and offset from the thermal centroid by a certain distance.

[0040] Finally, using the location coordinates as the center, a range of preset dimensions is defined as the key identification area. The preset size can be set according to the object category and image resolution; for example, for face recognition tasks, the preset size can be set to 96×96 pixels.

[0041] Through the above technical solution, this application can accurately locate the key identity regions of a target object in a visible light image using the thermal feature information provided by infrared images. Since infrared images are unaffected by lighting conditions, this method can stably locate key identity regions in various environments, providing a reliable region localization basis for subsequent feature extraction and identity recognition. Furthermore, by introducing geometric constraints related to object categories, the accuracy of key identity region localization is improved, and the possibility of mislocalization is reduced.

[0042] In some of the solutions described above in this application, a fixed preset size range is used when locating the key identity area. However, the thermal profile size of different target objects may change due to distance or environmental factors, causing the preset size to be unable to be adaptively adjusted, resulting in inaccurate or incomplete positioning of the key identity area.

[0043] This application further proposes that before using the range of preset dimensions as the key area of ​​identity, it also includes: obtaining the standard thermal profile corresponding to the object category; calculating the ratio between the size of the standard thermal profile and the size of the thermal profile as a size coefficient; and calculating the product between the size coefficient and the preset size as a new preset size.

[0044] The standard thermal profile stores thermal imaging profile size data for different object categories at a standard observation distance. The size factor reflects the relative distance change between the target object and the camera by the ratio of the actual thermal profile to the standard thermal profile. The preset size is adjusted linearly through multiplication. For example, when the standard thermal profile size is 200 pixels and the actual detected thermal profile size is 160 pixels, the size factor is 0.8, and the area with the original preset size of 100 pixels will be adjusted to an 80-pixel range.

[0045] Through the above technical solution, this application can adaptively adjust the size of the key identification region according to the actual size of the target. This adaptive adjustment avoids the problems of incomplete key region cropping or excessive redundant background caused by fixed size, thus improving the accuracy and efficiency of subsequent identity recognition. At the same time, this method makes full use of the thermal contour information provided by infrared images, realizing effective fusion between visible light images and infrared images, enhancing the advantages of multimodal image processing.

[0046] In some of the solutions described above in this application, identifying missing bit planes due to compression among multiple bit planes based on compression parameters includes:

[0047] First, extract the bitrate and quantization parameters based on the compression parameters. For example, the bitrate is 50kbps and the quantization parameter is 40 from the compression parameters.

[0048] Then, based on the bitrate and quantization parameters, the missing probability of each bit plane in multiple bit planes is determined from a pre-defined bitrate-quantization-bit plane missing relationship table. Specifically, a bitrate-quantization-bit plane missing relationship table can be pre-established, which records the missing probability of each bit plane under different combinations of bitrate and quantization parameters. Continuing the previous example, let's assume that, based on the table, under the conditions of a bitrate of 50kbps and a quantization parameter of 40, the missing probability of bit plane 1 is 0.9, the missing probability of bit plane 2 is 0.7, the missing probability of bit plane 3 is 0.5, the missing probability of bit plane 4 is 0.3, and the missing probability of bit planes 5-8 is 0.1.

[0049] Finally, bit planes with a missing probability higher than a preset probability threshold are identified as missing bit planes. Further, the preset probability threshold can be set to 0.6. Therefore, bit planes 1 and 2, with missing probabilities higher than 0.6, are identified as missing bit planes.

[0050] Through the above technical solution, this application can identify bit planes missing due to compression, providing compensation targets for subsequent high-frequency data compensation. This missing bit plane identification method based on compression parameters avoids the waste of computational resources caused by blindly compensating all bit planes, and also prevents insufficient compensation caused by omitting key missing bit planes. Therefore, image quality restoration is achieved with limited computational resources, providing a reliable data foundation for subsequent identity recognition.

[0051] In some of the solutions described above in this application, high-frequency data compensation for missing bit planes in multiple bit planes with a preset compensation intensity includes:

[0052] First, the gradient direction for high-frequency data compensation is determined based on edge feature information. Specifically, the edge feature information is analyzed to extract the direction and intensity information of the edges. For example, the Sobel operator can be used to calculate the horizontal and vertical gradients of the image, and then the main direction of the edges is determined based on the magnitude and direction of the gradients.

[0053] Then, based on the compensation intensity, high-frequency compensated data with gradient direction is generated. Further, Gaussian noise or Laplace noise can be used as a basis, and the noise can be oriented according to the previously determined gradient direction. The compensation intensity is used to control the amplitude of the generated high-frequency data.

[0054] Finally, the high-frequency compensation data is injected into the missing bit plane to perform high-frequency data compensation. Thus, by fusing the generated high-frequency compensation data with the missing bit plane, compensation for missing high-frequency information can be achieved. The fusion process can employ operations such as addition or multiplication, depending on the characteristics of the compensation data and the properties of the missing bit plane.

[0055] Through the above technical solution, this application can recover high-frequency detail information lost due to compression, improving image clarity and detail. Because the compensation process considers the edge feature information of the original image, the compensated image can better maintain its original structure and texture features, avoiding over-smoothing or the introduction of artifacts. This targeted high-frequency data compensation method can improve the quality of compressed images, providing more reliable visual features for subsequent identity recognition tasks.

[0056] In some of the above-mentioned schemes in this application, the adjustment of the compensation intensity includes:

[0057] Determine if this is the first adjustment of the compensation intensity. If so, increase the compensation intensity according to the preset first step length. For example, if the first step length can be set to 0.1 and the initial compensation intensity is 0.5, then the compensation intensity will increase to 0.6 after the first adjustment.

[0058] If not, determine whether the maximum value in the confidence set of the current iteration is higher than the maximum value in the confidence set of the previous iteration. If so, increase the compensation strength according to the first step length. For example, increase the compensation strength from 0.6 to 0.7.

[0059] If not, the compensation intensity is reduced according to a preset second step size, which is smaller than the first step size. For example, the second step size can be set to 0.05, and the compensation intensity is reduced from 0.7 to 0.65.

[0060] Through the above technical solution, this application achieves adaptive adjustment of the compensation intensity. When the recognition effect improves, a larger step size is used to rapidly increase the compensation intensity, accelerating the convergence speed. When the recognition effect deteriorates, a smaller step size is used to fine-tune the compensation intensity, avoiding over-compensation. This adaptive adjustment mechanism can minimize the number of iterations while ensuring the recognition effect.

[0061] In some of the above-mentioned schemes in this application, there is a problem of uncontrollable iteration number in the process of adjusting the compensation intensity. After the compensation intensity is adjusted many times, the identification confidence may fluctuate, causing the iteration process to fail to converge, resulting in wasted computing resources and extended identity feature extraction time.

[0062] This application further proposes that step S10 includes: determining whether the compensation intensity in the previous iteration decreased according to the second step size; if so, determining whether the maximum value in the confidence set in the current iteration is higher than the maximum value in the confidence set in the previous iteration; if not, stopping the iteration and extracting the identity features of the target object based on the identity key region in the previous iteration.

[0063] Specifically, during the compensation intensity adjustment process, when it is detected that the previous adjustment used the second step size for negative weakening, the maximum value of the current iteration's identification confidence is compared with the previous maximum value. If the current confidence does not exceed the previous confidence, it indicates that the compensation intensity adjustment has entered a local optimum region, and the iteration process is terminated immediately.

[0064] Through the above technical solution, this application can promptly terminate invalid iteration processes and avoid feature distortion caused by overcompensation. This reduces unnecessary computational overhead while ensuring recognition performance.

[0065] Exemplary System

[0066] Figure 2The figure illustrates a visual feature extraction system based on multimodal image enhancement according to an embodiment of this application, comprising: a data acquisition module for receiving a fully transmitted infrared image of a target region, a compressed transmitted visible light image, and compression parameters of the visible light image; a thermal feature acquisition module for acquiring thermal feature information of a target object based on the infrared image; a key region localization module for locating a key identity region of the target object in the visible light image based on the thermal feature information; an edge feature extraction module for extracting edge feature information of the key identity region based on the infrared image; and a bit-plane decomposition module for performing bit-plane decomposition on the key identity region to obtain multiple... The system comprises the following modules: a bit plane identification module, a missing bit plane identification module, a compensation module, a reconstruction module, a recognition module, and a decision module. The first module identifies the missing bit planes in multiple bit planes due to compression based on compression parameters. The second module determines whether the maximum value in the recognition confidence set is lower than a preset confidence threshold. If so, the compensation intensity is adjusted and the system returns to the compensation module. Otherwise, the system extracts the identity features of the target object based on the reconstructed identity key region.

[0067] In one example, the key region localization module locates the identity key region of a target object in a visible light image by: identifying the object category of the target object based on the infrared image; extracting the thermal profile and thermal centroid of the target object based on thermal feature information; obtaining the corresponding geometric constraint relationship based on the object category, and calculating the position coordinates of the center of the identity key region in the visible light image based on the geometric constraint relationship, thermal profile, and thermal centroid, wherein the geometric constraint relationship represents the mapping rule between the thermal profile and thermal centroid of the object category and the spatial position of the identity key region; and using the position coordinates as the center, defining a range of a preset size as the identity key region.

[0068] In one example, before the key area localization module uses the preset size range as the key area of ​​identity, it also includes: obtaining the standard thermal profile corresponding to the object category; calculating the ratio between the size of the standard thermal profile and the size of the thermal profile as the size coefficient; and calculating the product between the size coefficient and the preset size as the new preset size.

[0069] In one example, the missing bit plane identification module identifies missing bit planes due to compression in multiple bit planes based on compression parameters by: extracting bit rate parameters and quantization parameters based on compression parameters; determining the missing probability of each bit plane in multiple bit planes from a preset bit rate-quantization-bit plane missing relationship table based on bit rate parameters and quantization parameters; and identifying bit planes with missing probabilities higher than a preset probability threshold as missing bit planes.

[0070] In one example, the compensation module performs high-frequency data compensation on missing bit planes in multiple bit planes with a preset compensation intensity, including: determining the gradient direction of high-frequency data compensation based on edge feature information; generating high-frequency compensation data with gradient direction based on the compensation intensity; and injecting the high-frequency compensation data into the missing bit planes for high-frequency data compensation.

[0071] In one example, the decision module adjusts the compensation intensity by: determining whether the compensation intensity is being adjusted for the first time; if yes, increasing the compensation intensity according to a preset first step length; if no: determining whether the maximum value in the confidence set in the current iteration is higher than the maximum value in the confidence set in the previous iteration; if yes, increasing the compensation intensity according to the first step length; if no, decreasing the compensation intensity according to a preset second step length, where the second step length is less than the first step length.

[0072] In one example, the decision module is also used to: determine whether the compensation intensity in the previous iteration decreased according to the second step size; if so, determine whether the maximum value in the confidence set in the current iteration is higher than the maximum value in the confidence set in the previous iteration; if not, stop the iteration and extract the identity features of the target object based on the identity key region in the previous iteration.

[0073] Exemplary electronic devices

[0074] Figure 3 A block diagram of an electronic device according to an embodiment of this application is illustrated.

[0075] like Figure 3 As shown, the electronic device includes one or more processors and memory.

[0076] A processor can be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and can control other components in an electronic device to perform desired functions.

[0077] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0078] In one example, the electronic device may also include input devices and output devices, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0079] Of course, for the sake of simplicity, Figure 3Only some of the components of the electronic device relevant to this application are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.

[0080] Exemplary computer-readable media

[0081] Embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps described in the "Exemplary Methods" section above according to the various embodiments of this application.

[0082] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0083] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0084] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0085] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0086] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0087] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A visual feature extraction method based on multimodal image enhancement, characterized in that, Includes the following steps: S1, receive the complete transmitted infrared image of the target area, the compressed transmitted visible light image, and the compression parameters of the visible light image; S2, Based on the infrared image, obtain the thermal characteristic information of the target object; S3, based on the thermal feature information, locate the key identity region of the target object in the visible light image; S4, extract the edge feature information of the key identity region based on the infrared image; S5, perform bit-plane decomposition on the identity key region to obtain multiple bit planes; S6, Identify the missing bit planes among the plurality of bit planes that are missing due to compression based on the compression parameters; S7, Based on the edge feature information, perform high-frequency data compensation on the missing bit planes in the plurality of bit planes with a preset compensation intensity; S8, Reconstruct the identity key region based on the plurality of bit planes; S9, perform identity recognition on the reconstructed key identity region to obtain a set of recognition confidence scores; S10, determine whether the maximum value in the identification confidence set is lower than the preset confidence threshold. If yes, adjust the compensation intensity and return to S7 for iteration; otherwise, extract the identity features of the target object based on the reconstructed identity key region.

2. The visual feature extraction method based on multimodal image enhancement according to claim 1, characterized in that, The process of locating the key identification region of the target object in the visible light image includes: Identify the object category of the target object based on the infrared image; Extract the thermal profile and thermal center of gravity of the target object based on the thermal feature information; Obtain the corresponding geometric constraint relationship according to the object category, and calculate the position coordinates of the center of the identity key region in the visible light image based on the geometric constraint relationship, the thermal profile and the thermal centroid. The geometric constraint relationship represents the mapping rule between the thermal profile and thermal centroid of the object category and the spatial position of the identity key region. Centered on the location coordinates, a range of preset size is defined as the key identity area.

3. The visual feature extraction method based on multimodal image enhancement according to claim 2, characterized in that, Before specifying the range of a preset size as the key identity area, the following is also included: Obtain the standard thermal profile corresponding to the object category; The ratio of the size of the standard thermal profile to the size of the thermal profile is calculated as a size factor; The product of the size factor and the preset size is calculated as the new preset size.

4. The visual feature extraction method based on multimodal image enhancement according to claim 1, characterized in that, The step of identifying the missing bit planes among the plurality of bit planes due to compression based on the compression parameters includes: Extract the bitrate parameter and quantization parameter based on the compression parameters; Based on the bitrate parameter and quantization parameter, the probability of missing each bit in the plurality of bit planes is determined from the preset bitrate-quantization-bit plane missing relationship table; The bit plane with a missing probability higher than a preset probability threshold is identified as the missing bit plane.

5. The visual feature extraction method based on multimodal image enhancement according to claim 1, characterized in that, The high-frequency data compensation for missing bit planes in the plurality of bit planes with a preset compensation intensity includes: The gradient direction of the high-frequency data compensation is determined based on the edge feature information; Based on the compensation intensity, high-frequency compensation data with the gradient direction is generated; The high-frequency compensation data is injected into the missing bit plane to perform the high-frequency data compensation.

6. The visual feature extraction method based on multimodal image enhancement according to claim 1, characterized in that, The adjustment of the compensation intensity includes: Determine whether the compensation intensity is being adjusted for the first time; If so, the compensation intensity is increased according to the preset first step length; If not: Determine whether the maximum value in the confidence set of the current iteration is higher than the maximum value in the confidence set of the previous iteration; If so, then increase the compensation strength according to the first step length; If not, the compensation intensity is reduced according to a preset second step length, where the second step length is less than the first step length.

7. The visual feature extraction method based on multimodal image enhancement according to claim 6, characterized in that, S10 also includes: Determine whether the compensation intensity in the previous iteration was reduced based on the second step size; If so, determine whether the maximum value in the confidence set in the current iteration is higher than the maximum value in the confidence set in the previous iteration. If not, stop the iteration and extract the identity features of the target object based on the identity key region in the previous iteration.

8. A visual feature extraction system based on multimodal image enhancement, characterized in that, include: The data acquisition module is used to receive the complete transmitted infrared image of the target area, the compressed transmitted visible light image, and the compression parameters of the visible light image; A thermal feature acquisition module is used to acquire thermal feature information of the target object based on the infrared image; A key region localization module is used to locate the key identity region of the target object in the visible light image based on the thermal feature information. An edge feature extraction module is used to extract edge feature information of the key identity region based on the infrared image; The bit-plane decomposition module is used to perform bit-plane decomposition on the identity key region to obtain multiple bit planes; A missing bit plane identification module is used to identify the missing bit planes among the plurality of bit planes that are missing due to compression, based on the compression parameters; The compensation module is used to perform high-frequency data compensation on the missing bit planes in the plurality of bit planes according to the edge feature information and with a preset compensation intensity. The reconstruction module is used to reconstruct the identity key region based on the multiple bit planes; The identification module is used to identify the identity in the reconstructed key identity region and obtain an identification confidence set. The decision module is used to determine whether the maximum value in the recognition confidence set is lower than a preset confidence threshold. If so, the compensation intensity is adjusted and returned to the compensation module; otherwise, the identity features of the target object are extracted based on the reconstructed identity key region.

9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method as described in any one of claims 1 to 7.

10. A computer storage medium storing computer-executable instructions thereon, characterized in that: When the computer-executable instructions are executed by a processor, they implement the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Edge feature enhancement method and device, storage medium and electronic equipment

    CN118154442A

  • Multi-spectral fusion night low-illumination image enhancement and occlusion compensation method

    CN120689257A