Heavy-duty car recorder image set compression method based on reference structure and image enhancement
By constructing a reference structure and grouping method, similar image sets from heavy-duty vehicle recorders are compressed, solving the problems of unstable redundancy utilization and weak consistency of reconstruction quality in existing technologies, and achieving stable compression and high-quality reconstruction under low bit rate conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies have unstable cross-image redundancy utilization in the compression scenario of similar image sets in heavy-duty vehicle recorders, and the low bitrate reconstruction quality consistency is weak, which makes the reconstruction results prone to visual distortions such as noise, blurring and block effects. They cannot meet the requirements of compression efficiency and reconstruction quality under the conditions of limited resources and bitrate budget of the vehicle end.
By measuring the similarity of similar image sets, constructing reference structures and grouping them, determining the compression order, using intra-frame coding and prediction and residual information coding, generating auxiliary information images and performing multi-scale fusion, and finally performing image enhancement processing, the stability and consistency of reference dependencies are ensured.
It achieves stable utilization of cross-image redundancy under low bitrate conditions, improves the consistency of reconstruction quality and compression efficiency, and meets the compression reconstruction quality and playback evidence collection requirements of heavy-duty vehicle recorders in long-term recording scenarios.
Smart Images

Figure CN121842403A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of vehicle-mounted image processing and image data compression, and particularly relates to a heavy truck recorder image set compression method based on a reference structure and image enhancement. BACKGROUND
[0002] In the prior art, the storage and transmission of heavy truck recorder image data usually rely on single-frame lossy compression or a video coding framework, which establishes a prediction relationship within / inter frames and encodes transform coefficients or residual information to meet the needs of long-term retention and playback retrieval at the vehicle-mounted end. However, for similar image sets formed by continuous collection of recorders, the existing compression organization method has some obvious deficiencies in the utilization and processing flow organization of cross-image correlation.
[0003] In actual application, the redundancy distribution within the similar image set often presents uneven strength, and the existing scheme usually adopts fixed adjacent reference or time sequence processing. The adaptability of the reference relationship and grouping strategy to factors such as scene displacement and light fluctuation is weak, which leads to unstable cross-image redundancy reduction effect and low overall compression efficiency. At the same time, under the condition that the vehicle-mounted end resources and code rate budget are limited, the reconstruction result is more likely to present visual distortion accumulation such as noise, blur and blocking effect, and the consistency of reconstruction quality is weak.
[0004] Therefore, the prior art often has problems such as unstable cross-image redundancy utilization effect in the compression scene of similar image sets of heavy truck recorders and weak consistency of low code rate reconstruction quality. This is the deficiency of the prior art.
[0005] Therefore, the present application provides a heavy truck recorder image set compression method based on a reference structure and image enhancement to solve the above-mentioned defects in the prior art, which is very necessary. SUMMARY
[0006] The present application aims to solve the above-mentioned technical problems by providing a heavy truck recorder image set compression method based on a reference structure and image enhancement.
[0007] To achieve the above-mentioned purpose, the present application provides the following technical solution: The present application provides a heavy truck recorder image set compression method based on a reference structure and image enhancement, which comprises: The similarity of the similar image set is measured, the reference structure for representing the reference relationship between images is constructed based on the similarity measurement, and the similar image set is grouped according to the reference structure; The compression order of the similar image set is determined according to the reference structure and the grouping result, and the compression order satisfies the parent-child reference order constraint defined by the reference structure; The similar image set is coded according to the compression order, the root image is coded by intra-frame coding, and the non-root image is coded by prediction and residual information coding with the parent image as the reference, to obtain the reconstructed image of the root image and the reconstructed image of each non-root image; The auxiliary information image corresponding to the non-root image is generated based on the reconstructed image of the parent image corresponding to the non-root image, and the reconstructed image of the non-root image and the auxiliary information image are multi-scale fused to obtain the fused image corresponding to the non-root image; The fused image corresponding to the non-root image is subjected to image enhancement processing to output an enhanced reconstructed image, and the enhanced reconstructed image is taken as the compression reconstructed result of the similar image set.
[0008] By using the above technical scheme, the similarity of the similar image set is measured, and the reference structure representing the reference relationship between images is constructed according to the measurement, so as to realize the structuralized description and organization of the internal redundancy distribution of the similar image set, maintain the controllability and consistency of the cross-image correlation utilization path when factors such as scene displacement and light fluctuation exist, and meet the needs of stable cross-image redundancy utilization effect and relatively optimal overall compression efficiency in the compression scene of the similar image set.
[0009] The compression order is determined according to the reference structure and the grouping result, and the parent-child reference order constraint is satisfied, so that the reference dependence is clearly constrained and consistently executed in the processing flow, the influence of the reference link uncertainty on the prediction effectiveness is reduced, and the stability and repeatability of the encoding process are maintained; on this basis, the similar image set is coded according to the compression order, the root image is coded by intra-frame coding to provide a reliable reference anchor point, and the non-root image is coded by prediction and residual information coding with the parent image as the reference to strengthen the utilization intensity of the cross-image redundancy, so that more competitive compression performance can be obtained under the same storage and transmission budget; at the same time, the auxiliary information image corresponding to the non-root image is generated based on the reconstructed image of the parent image corresponding to the non-root image, and the reconstructed image of the non-root image and the auxiliary information image are multi-scale fused, so that the reference side information cooperatively supplements the structure and detail expression at different scales, the consistency of the reconstructed image in terms of texture, edge and structure continuity is enhanced, and the needs of strong consistency of the reconstructed quality under the condition of low code rate are met; then, the fused image is subjected to image enhancement processing to output an enhanced reconstructed image, so that the fused structure information and detail information are further optimized and expressed, and the cumulative effect of compression distortion in vision is suppressed, the visual consistency and stability of the compression reconstructed result of the similar image set between different images are improved, and the needs of consistent compression reconstructed quality and stable playback experience in the long-time recording scene of the heavy-duty vehicle recorder are met.
[0010] Preferably, the step of measuring the similarity of a set of similar images includes: Stable region features are extracted from each image in the similar image set, and these stable region features are used as the dominant features for similarity measurement. Stable region features include vehicle outline features, license plate region features, and road marking features. Identify volatile regions in each image and use the features corresponding to these volatile regions as suppression features. Volatile regions include regions with changes in illumination and regions with motion interference. A first similarity metric component is formed based on the dominant feature, a second similarity metric component is formed based on the suppression feature, and the similarity metric between the two images is determined based on the first and second similarity metric components.
[0011] By adopting the above technical solution, stable region features are used as the dominant basis for similarity measurement and variable region features are used to form suppression constraints. This enables similarity evaluation to focus on key vehicle structures and road semantic information, which can reduce the traction effect of illumination fluctuations and motion interference on similarity judgment and improve the discrimination stability and consistency of similarity measurement of similar image sets.
[0012] Preferably, the step of constructing a reference structure based on a similarity metric to characterize the reference relationship between images includes: A graph structure is constructed based on a similarity metric, and the similarity metric is used as the edge weight. A tree-like reference structure is generated based on edge weights. The tree-like reference structure includes a root image and parent-child reference relationships, and the parent-child reference relationships satisfy the constraints of similarity measurement. The parent image of each non-root image is determined based on the tree reference structure.
[0013] By adopting the above technical solution, a tree-like reference structure is generated using the graph structure represented by edge weights and the parent-child reference relationship is determined, realizing the hierarchical organization of reference dependencies between images. Under similarity constraints, clearer reference paths and reference anchor point selection can be formed, improving the determinism of reference structure construction and the interpretability of reference relationships.
[0014] Preferably, the step of grouping similar image sets according to a reference structure includes: Grouping is formed by splitting the tree reference structure into subtrees or branches; Set group boundary constraints for each group. The group boundary constraints limit the parent-child reference relationships within the group to participate in the compressed dependencies of the current group, and limit the parent-child reference relationships outside the group to not participate in the compressed dependencies of the current group.
[0015] By adopting the above technical solution, using a grouping method based on subtrees or branches and setting group boundary constraints, the compressed dependencies are enclosed within the group and isolated from cross-group dependencies. This reduces the link coupling caused by reference diffusion between groups, improves the controllability of the effective range of reference relationships within the group, and enhances the stability of the compression process.
[0016] Preferably, the step of determining the compression order of similar image sets based on the reference structure and grouping results includes: Perform a traversal on the tree reference structure corresponding to each group to obtain the compression order within the group; During the traversal, the similarity measurement determines the order of parent and child nodes based on the parent-child reference sequence constraint, and determines the similarity measurement of the order of appearance of parent and child chains based on the continuity constraint of parent and child chains within the same branch. When a branch switch occurs during traversal, the reference relationship after the branch switch is limited to the group based on the group boundary constraints.
[0017] By adopting the above technical solution, the compression order within a group is determined by traversing a tree-type reference structure and the order is guided by parent-child sequence and parent-child chain continuity constraints. This achieves consistent alignment between the compression order and reference dependencies, reduces the impact of reference jumps caused by traversal switching, and improves the coherence of compression order generation and the stability of predicted dependencies.
[0018] Preferably, the steps of encoding and decoding similar image sets according to the compression order include: The current image in the set of similar images is selected sequentially according to the compression order, and the encoding result of the current image is written into the compressed bit stream. When the compression order points to the root image, intra-frame coding is performed on the root image, and the coding result of the root image is decoded to obtain the reconstructed image of the root image. When the compression order points to a non-root image, prediction information is generated based on the reconstructed image of the corresponding parent image of the non-root image, and residual information is generated based on the prediction information. The residual information is encoded and the encoding result of the residual information is written into the compressed bit stream. Decoding is performed on the reconstructed image of the parent image and the encoding result of the residual information to obtain the reconstructed image of the non-root image.
[0019] By adopting the above technical solution, the compressed bit stream is written in the compression order and the reconstruction results of the root image and non-root images are obtained simultaneously. This enables the sequential organization of the encoding results and the timely availability of the reconstruction information on the reference side, thereby improving the closed-loop consistency of the encoding and decoding process and the traceability of the bit stream organization.
[0020] Preferably, the step of generating prediction information based on the reconstructed image of the parent image corresponding to the non-root image includes: Perform local feature matching between the reconstructed image of the parent image and the non-root image, and determine the corresponding region based on the local feature matching results; Geometric transformation parameters and illumination compensation parameters are estimated based on the corresponding region; The reconstructed image of the parent image is transformed based on the geometric transformation parameters and the illumination compensation parameters to obtain the prediction information.
[0021] By adopting the above technical solution, the corresponding region is determined by local feature matching and the geometric transformation parameters and illumination compensation parameters are estimated. This enables adaptive mapping of the parent image reconstruction information to the non-root image prediction information, which can improve the alignment between the prediction information and the target content, and enhance the concentration of residual information and the stability of coding efficiency.
[0022] Preferably, the step of generating an auxiliary information image corresponding to a non-root image based on the reconstructed image of the parent image corresponding to the non-root image includes: The reconstructed image of the parent image is parametrically processed based on the geometric transformation parameters and illumination compensation parameters to obtain the initial auxiliary information image; Dynamic region identifiers are generated based on the differences between the reconstructed image of the parent image and the reconstructed image of the non-root image. Based on the dynamic region identifier, dynamic interference suppression processing is performed on the initial auxiliary information image to obtain the auxiliary information image. The dynamic interference suppression processing includes at least one of smoothing processing and replacement processing.
[0023] By adopting the above technical solution, parameterized processing is used to generate initial auxiliary information and dynamic interference suppression is achieved by combining dynamic region identification. This enables the auxiliary information to faithfully represent static structural information and suppress dynamic interference, thereby improving the effectiveness and reliability of the auxiliary information image and enhancing the structural consistency of subsequent fusion.
[0024] Preferably, the step of multi-scale fusion of the reconstructed image of the non-root image and the corresponding auxiliary information image to obtain the fused image corresponding to the non-root image includes: A scale pyramid is constructed, and feature extraction is performed on the reconstructed image and auxiliary information image of the non-root image at each scale to obtain the first feature and the second feature at each scale. Feature fusion is performed at each scale based on the spatial correspondence between the first and second features to obtain the fused features at each scale; Scale backpropagation aggregation is performed on the fusion features at each scale to obtain the fusion image corresponding to the non-root image.
[0025] By adopting the above technical solution, a scale pyramid is constructed and feature extraction, fusion and scale backpropagation aggregation are performed at each scale. This enables the collaborative integration of structural and texture information at multiple scale levels, which can improve the ability of the fused image to take into account both the level of detail and the global structure, and enhance the visual consistency of the reconstruction results at different scales.
[0026] Preferably, the step of performing image enhancement processing on the fused image corresponding to the non-root image to output an enhanced reconstructed image includes: Dense residual structure processing is performed on the fused image to obtain intermediate features; Upsampling is performed on the intermediate features, and an attention mechanism is applied between the upsampling and output convolution processes to obtain the channel weights. Edge feature maps are calculated based on the fused image, and the global statistical information of the edge feature maps and the fused image is input into the attention mechanism for processing to obtain channel weights constrained by the global statistical information and the local statistical information of the edge feature maps; Channel recalibration is performed on intermediate features based on channel weights, and output convolution processing is performed to obtain an enhanced reconstructed image.
[0027] By adopting the above technical solution, a combined enhancement method of dense residual structure and attention mechanism is used, and edge features and global statistical information are introduced to constrain channel weights. This enables adaptive enhancement of key details and structural boundaries, which can improve the edge sharpness and texture expression stability of the enhanced and reconstructed image, and reduce the risk of distortion amplification in the enhancement process.
[0028] As can be seen from the above technical solutions, the present invention has the following advantages: This application provides a method for compressing heavy-duty vehicle recorder image sets based on reference structure and image enhancement. By establishing parent-child reference relationships between images based on similarity metrics and organizing the grouping and compression order of similar image sets accordingly, cross-image redundancy can be stably utilized under constrained reference links. At the same time, auxiliary information generated by the reconstruction results of the parent image is introduced on the decoding side and fused with the reconstruction results of non-root images at multiple scales before image enhancement processing is performed. This achieves synergistic reinforcement of structural and detail information under low bitrate compression conditions, and can obtain more stable redundancy utilization and more consistent reconstruction visual quality under the same storage and transmission budget. This meets the requirements of superior compression efficiency and stable playback evidence viewing in continuous acquisition scenarios of heavy-duty vehicle recorders.
[0029] Furthermore, the design principle of this invention is reliable, the structure is simple, and it has a very wide range of application prospects.
[0030] Therefore, it is evident that the present invention has outstanding substantive features and significant progress compared with the prior art, and the beneficial effects of its implementation are also obvious. Attached Figure Description
[0031] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a flowchart of a method for compressing heavy-duty vehicle recorder image sets based on reference structure and image enhancement, provided by the present invention. Detailed Implementation
[0033] Various embodiments of this disclosure are described more fully below with reference to the accompanying drawings. This disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of this disclosure to the specific embodiments disclosed herein, but rather this disclosure should be understood to cover all adjustments, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of this disclosure.
[0034] In the following, the terms “comprising” or “may include”, which may be used in various embodiments of this disclosure, indicate the presence of the disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. Furthermore, as used in various embodiments of this disclosure, the terms “comprising,” “having,” and their cognates are intended only to indicate a particular feature, number, step, operation, element, component, or combination of the foregoing, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing, or the possibility of adding one or more combinations of the foregoing.
[0035] The terms used in the various embodiments of this disclosure (such as "first," "second," etc.) may modify various components in the various embodiments, but do not limit the corresponding components. For example, the above terms do not limit the order and / or importance of the components. The above terms are only used for the purpose of distinguishing one component from other components. For example, a first user device and a second user device refer to different user devices, although both are user devices. For example, a first component may be referred to as a second component without departing from the scope of the various embodiments of this disclosure, and similarly, a second component may also be referred to as a first component.
[0036] It should be noted in advance that, in order to facilitate a clear and accurate description of the technical solutions in the embodiments of this application, the following is a brief explanation of some terms and related technologies involved in the embodiments of this application: 1. Local feature matching: This refers to a method that extracts local key points and their descriptors that have repeatability and detectability from two images, and establishes the correspondence between key points through descriptor similarity retrieval. It is often used for corresponding region localization and subsequent geometric estimation under conditions of cross-viewpoint, scale or illumination changes.
[0037] 2. Attention mechanism: This is a type of computational structure used to adaptively allocate feature weights. It typically generates weight coefficients based on the statistical information or contextual association of the input features, thereby recalibrating the importance of channels or spatial locations, improving the network's ability to represent key information and suppressing redundant or noisy information.
[0038] 3. Scale Pyramid: This is a multi-scale representation method that constructs different scale levels by downsampling or multi-resolution of images or features. This allows the algorithm to simultaneously depict the global structure and local details at coarse to fine levels. It is often used in multi-scale fusion, matching, and reconstruction processes.
[0039] To address the issues of unstable cross-image redundancy utilization and weak reconstruction quality consistency under low bitrate conditions in the compression processing of similar image sets continuously acquired by heavy-duty vehicle recorders, which make it difficult to meet the actual needs of vehicle-mounted devices for compression efficiency and visual stability in playback and evidence collection under limited storage and transmission conditions, this application discloses a heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement. By introducing a reference structure organization mechanism for similar image sets and combining it with quality enhancement processing of compression and reconstruction results, the reference relationships between similar images are utilized in an orderly manner, and the structural and detail information is optimized collaboratively during the reconstruction stage. This improves the overall compression efficiency and enhances the visual consistency of reconstruction results under low bitrate conditions, further improving the stability and reliability of heavy-duty vehicle recorder image sets in playback and evidence collection scenarios.
[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0041] like Figure 1 As shown in the figure, this embodiment provides a method for compressing heavy-duty vehicle recorder image sets based on reference structure and image enhancement, including: Step S1: Perform similarity measurement on the similar image set, construct a reference structure based on the similarity measurement to characterize the reference relationship between images, and group the similar image set according to the reference structure; Step S2: Determine the compression order of similar image sets based on the reference structure and grouping results. The compression order satisfies the parent-child reference order constraint defined by the reference structure. Step S3: Perform encoding and decoding processing on the similar image set according to the compression order. The root image is encoded intra-frame, and the non-root images are encoded using prediction and residual information with the parent image as a reference, to obtain the reconstructed image of the root image and the reconstructed images of each non-root image. Step S4: Generate an auxiliary information image corresponding to the non-root image based on the reconstructed image of the parent image corresponding to the non-root image, and perform multi-scale fusion of the reconstructed image of the non-root image and the corresponding auxiliary information image to obtain the fused image corresponding to the non-root image. Step S5: Perform image enhancement processing on the fused image corresponding to the non-root image to output the enhanced reconstructed image, and use the enhanced reconstructed image as the compressed reconstruction result of the similar image set.
[0042] This embodiment constructs a reference structure between images based on similarity metrics and organizes similar image sets into groups and compression sequences accordingly. This allows the reference dependencies between images to participate in the encoding and decoding process in a structured form, thereby achieving stable utilization of cross-image redundancy in continuous acquisition scenarios. By using the root image as a stable reference anchor point and the parent image to guide the prediction and residual information expression of non-root images during the encoding stage, the strongly correlated content in the similar image set is effectively compressed, balancing compression efficiency and encoding stability under limited storage and bandwidth conditions. In the reconstruction stage, auxiliary information generated based on the reconstruction results of the parent image is introduced and fused with the reconstruction results of non-root images at multiple scales. This allows the reference-side structural information to collaboratively supplement details and contours at different scales, reducing the impact of compression distortion on structural consistency. Furthermore, by performing image enhancement processing on the fusion results, visual elements such as texture and edges are uniformly optimized, improving the overall consistency and visual stability of the compressed reconstruction results of similar image sets under low bitrate conditions. This meets the practical needs of heavy-duty vehicle recorders for long-term continuous recording and playback evidence collection scenarios, which require both compression efficiency and reconstruction quality.
[0043] Hereinafter, steps S1 to S5 will be specifically described according to embodiments of this application.
[0044] In step S1, the core task is to establish a computable and constrained expression of the correlation between images for the similar image set continuously acquired by the heavy-duty vehicle recorder, so that the reference relationship between images can be explicitly represented in a structured form and can be directly called by the subsequent compression process. Its input is the similar image set and its image content information, and the output is a reference structure used to represent the reference relationship between images and a grouping result consistent with the reference structure. The reference structure needs to be able to give the root image, parent-child reference relationship and the parent image of each non-root image. The grouping result needs to give the grouping boundary constraints to limit the effective range of compression dependency.
[0045] Specifically, similarity measures can be performed on sets of similar images, allowing the correlation strength between any two images to be compared under a unified dimension and to participate in structure generation.
[0046] In some embodiments of this application, to make the similarity measurement more closely match the data characteristics of heavy-duty vehicle recorder scenarios where the background is stable but local disturbances are frequent, dominant features and suppression features are introduced in the feature design of the similarity measurement: by extracting stable region features of each image in the similar image set to capture reusable structural information across images, and using stable region features as the dominant features of the similarity measurement to ensure that the measurement results are mainly driven by stable information, wherein stable region features include vehicle contour features, license plate region features and road marking features; at the same time, identifying volatile regions in each image to explicitly characterize the sources of interference that may cause "pseudo-similarity" or "pseudo-dissimilarity", and using the features corresponding to volatile regions as suppression features for the suppression mechanism, so that the similarity measurement remains robust to illumination and motion disturbances, wherein volatile regions include illumination change regions and motion disturbance regions.
[0047] Specifically, a set of similar images can be represented as For any two images and Construct dominant feature vectors respectively and and suppressing feature vectors and Among them, vehicle contour features can be realized by the shape descriptor formed by edge gradient and contour point set; license plate area features can be realized by the key point distribution and local texture descriptor within the license plate positioning box; road marking features can be realized by the line segment parameters and vanishing point consistency of lane lines / stop lines; illumination variation areas can be obtained by the brightness component change rate or illumination inconsistency mask; motion interference areas can be obtained by the foreground motion mask or optical flow anomaly areas.
[0048] Furthermore, to map the two types of features to composable similarity metric components, a first similarity metric component can be formed based on the dominant feature, for example, using cosine similarity or normalized cross-correlation.
[0049] in, This represents the first similarity measure component obtained from the dominant feature. Represents the vector dot product. Represents the L2 norm, This represents a very small positive number used to avoid a denominator of zero.
[0050] Simultaneously, a second similarity metric component is formed based on the suppression features, for example, a mapping from "difference degree to suppression amount" can be used:
[0051] in, This represents the second similarity measure component obtained from the suppressed features. Represents the L2 norm, The scale parameter controls the sensitivity of the suppressed component to the degree of difference. Based on this, the two components are combined into a final similarity measure between the two images, so that the dominant component "brings closer to true similarity," and the suppressed component "weakens the influence of perturbations." For example, a weighted fusion method can be used:
[0052] in, This represents a measure of similarity between two images. This indicates the weight of the dominant component. For example, in the daytime operation of a heavy-duty vehicle dashcam, it can be taken as... To enhance the decisiveness of the stable structure, the operating conditions at night or in backlight can be appropriately reduced. To enhance the restraining effect of the suppression component on changes in illumination.
[0053] In this embodiment of the application, after obtaining the similarity measure, it is necessary to further complete the construction of a reference structure based on the similarity measure to characterize the reference relationship between images, so that the reference relationship can satisfy the consistency constraint in the global scope and facilitate subsequent dependency calls.
[0054] Specifically, a graph structure can be used to represent candidate reference relationships between images. The graph structure is constructed based on a similarity metric, which is then used as edge weights to achieve a global connectivity representation. For example, an undirected weighted graph can be constructed. , where the vertex set Image indices and edge sets corresponding to similar image sets It consists of image pairs that satisfy a threshold constraint, i.e., when Time introduction of edge and order the border to ,in This represents a similarity threshold used to filter out weakly correlated connections to reduce structural noise.
[0055] Furthermore, the general graph structure can be converged into a tree-like reference structure to ensure that the reference links are acyclic and can be traversed sequentially. A tree-like reference structure is generated under edge weight constraints, completing the transformation from a "candidate reference set" to an "executable reference link". For example, it can be... Calculate the maximum spanning tree This maximizes the sum of edge weights on the tree and covers all vertices, thus prioritizing connections with high similarity. Alternatively, the minimum spanning tree can be calculated when the similarity metric already satisfies the distance mapping to achieve the same purpose.
[0056] In some embodiments of this application, the generated tree-like reference structure needs to satisfy explicit role and relationship definitions, including the root image and parent-child reference relationships, and use a similarity metric to constrain the validity of parent-child edges, thereby ensuring that the parent-child reference relationships satisfy the constraints of the similarity metric. For example, for any parent-child edge... Must meet ,in This represents the parent-child constraint threshold, used to ensure that the reference information contributes sufficiently to prediction and dependency compression. For root image selection, the image with the highest average similarity metric to the remaining images can be chosen. This allows it to provide strong reference coverage for most images when used as a reference starting point. After constructing the tree and determining the root, it is also necessary to determine the parent image of each non-root image based on the tree reference structure, thus grounding the reference directions on the tree onto the parent pointer of each node. For example, for any non-root image... In the tree Find the unique simple path from the root, and take the image corresponding to the node that is adjacent to it and closer to the root as the parent image. ,in This represents the parent index mapping function. When there are multiple candidate parent edges that satisfy the similarity constraint and have small weight differences, a stable region coverage constraint can be introduced. The candidate parent with higher coverage matching between vehicle contour features and road marking features is preferred to reduce reference instability caused by motion interference areas.
[0057] In this embodiment, after the reference structure is determined, it is necessary to further group similar image sets according to the reference structure, so that subsequent compression dependencies are closed within groups and isolated between groups, thereby reducing the complexity of cross-branch cache switching and dependency management from an engineering perspective. The basic action of grouping is to segment along the tree-like reference structure, so that each group corresponds to a local reference link or local branch. In this process, by forming groups based on subtrees or branches according to the tree-like reference structure, the mapping from structure to processing units can be achieved. For example, the tree... Branching can be done by dividing the tree into groups based on several first-level child nodes of the root, or by using a depth threshold on the tree. Split such that the depth of nodes within any group does not exceed This limits the reference link length and controls error accumulation; it can also introduce a packet size upper limit. Ensure that the number of images in each group does not exceed To match the computing power and caching conditions of the vehicle side.
[0058] Furthermore, to ensure that the grouping results provide executable constraints on "whether compressed dependencies are allowed across groups," boundary rules need to be assigned to each group. Therefore, group boundary constraints should be set for each group after grouping is generated. These boundary constraints need to simultaneously ensure that dependencies within the group are valid and dependencies outside the group are invalid, thus forming a dependency closure. Specifically, group boundary constraints can limit parent-child reference relationships within a group to participate in the compressed dependencies of the current group, and limit parent-child reference relationships outside the group to not participate in the compressed dependencies of the current group. For example, for grouping... Define the set of nodes within the group Set of parent and child edges within the group and define dependency indicator functions ,when Time to take Otherwise take Subsequent processes, when establishing compression dependencies, only allow the use of dependencies that satisfy the specified conditions. The father-son reference relationship, and at the same time for or Cross-group parent-child reference relationships are forcibly disabled to ensure clear group boundaries and consistent dependency management. For example, when a group corresponds to a branch and the branch length is long, the edges in the branch with similarity metrics below the threshold can be used as split points, making the overall similarity metrics within each group higher after splitting, thereby improving the effectiveness of intra-group references and reducing the risk of compression error propagation.
[0059] Thus far, step S1 has completed a unified characterization of the association strength within a similar image set through a two-component similarity measurement system dominated by stable region features and constrained by volatile region suppression. This characterization is mapped to a tree-like reference structure containing the root image and parent-child reference relationships. At the same time, group boundary constraints are introduced on the basis of subtree or branch segmentation to form an executable group dependency closure, so that the reference relationships between images can be stably expressed in a structured form. This provides a clear, constrained and reusable reference input basis for subsequent organization of compression dependencies and compression order around parent-child reference relationships.
[0060] In step S2, the core task is to construct an executable processing sequence for the similar image set under the constraints of the obtained reference structure and grouping results, so that any parent-child reference relationship is parent first and then child in time order, and the reference links within the same branch are as continuous as possible; its input is the tree reference structure and group boundary constraints corresponding to each group, and the output is the compression order for each group and the overall compression order of the similar image set spliced between groups.
[0061] In this embodiment, the compression order of similar image sets is determined based on the reference structure and grouping results. This allows for the generation of a compression order within each group, and ensures that the compression order has a strictly verifiable sequential constraint on the parent-child reference relationship in the tree-like reference structure, thus satisfying the parent-child reference sequential constraint defined by the reference structure. Specifically, the grouping can be... The tree reference structure within is represented as For any parent-child edge Introducing sequence mapping functions The parent-child reference order constraint can be written as: (The index of the image within the sequence of groups represents the position index of the image.) ,in, This indicates the index of the parent image within the compression order of the group. This represents the position index of the sub-image within the compressed order of the group. Based on this, to obtain an executable sequence that satisfies the above constraints, a traversal strategy can be performed on the tree reference structure corresponding to the group to obtain the compressed order within the group. The traversal can adopt a root-first, then-sub-first strategy, so that the parent node is output to the sequence upon access, thus naturally satisfying the constraints. .
[0062] It should be noted that, to avoid frequent backtracking during multi-branch traversal that weakens the continuity of the reference link, a joint determination of similarity measurement and link continuity constraint can be introduced when selecting the access order of child nodes at the same level. This allows the similarity measurement to determine the order of parent and child nodes based on the parent-child reference sequence constraint during traversal, while simultaneously ensuring that parent-child chains on the same branch form consecutive sequence segments. Thus, the similarity measurement of the parent-child chain appearance order is determined based on the continuity constraint of parent-child chains within the same branch. For example, in the parent node... Having multiple child node sets When, prioritize choosing with Similarity measurement Larger child nodes As the next access target, the "continuous constraint" is expressed as a reward for continuous access to the same branch. This allows for continuous output on the branch. A higher overall score is obtained; this overall score is only used to determine the order of visits within the same level and does not change the order of visits. The hard constraints ensure that the semantics of the reference structure are not diluted.
[0063] Furthermore, when traversal involves switching from one subtree to another, it's necessary to ensure that the reference relationships after the switch do not cross group boundaries and introduce invalid dependencies. Therefore, boundary constraint checks are introduced in the switch decision. When a branch switch occurs during traversal, the reference relationships after the branch switch are limited to be effective within the group based on the group boundary constraints. For example, the group boundary constraints can be implemented as an allowable set. Only if the parent-child edge to be used after the switch belongs to Only when the corresponding child node is added to the subsequent sequence of the current group is the node allowed to be added. Otherwise, the node is delayed until the traversal process of its own group to ensure that the dependencies within the group are closed and the dependencies between groups are isolated.
[0064] Thus far, step S2 constructs a tree-like reference structure on the corresponding group to satisfy... The traversal sequence is used, and the same-level access order driven by similarity metric is used to maintain the continuity of branch links. At the same time, the group boundary constraint is used to suppress the effect of cross-group references, forming a compressed order that satisfies the parent-child reference order constraint and can be directly scheduled in engineering. This provides a stable and verifiable processing clock input basis for subsequent sequential execution of prediction and residual coding.
[0065] In step S3, the core task is to encode the set of similar images into a compressed bitstream along a predetermined compression order, and simultaneously obtain a reproducible reconstructed image sequence at the encoding end, so that the prediction of subsequent non-root images strictly takes the reconstructed image of the parent image as a reference; its input is the compression order, tree reference structure and image data, and the output is the compressed bitstream, the reconstructed image of the root image and the reconstructed images of each non-root image.
[0066] Specifically, similar image sets can be encoded and decoded in the order of compression to ensure that the input of the reference link is available when used.
[0067] In some embodiments of this application, the overall compression order may be set as follows: ,in Indicates the first Each processed image index, when sequentially selecting the current image and advancing the write stream process at each time step, can sequentially select the current image from the set of similar images in the compression order to perform encoding processing, and append the encoding result of this encoding to the compressed bit stream. Simultaneously, to ensure reference consistency, the encoder needs to perform decoding loopback on key nodes, ensuring that subsequent predictions are strictly based on the reconstructed domain rather than the original domain.
[0068] Specifically, when the compression order points to the root image, since the root image is independent of the reference, it should be independently coded to generate the reconstruction result. Therefore, intra-frame coding is performed at this node, and loopback decoding is performed on the encoded result of the root image to obtain the reconstructed image. For example, the root image can be represented as... The intra-frame coding result is denoted as a bit sequence. The decoding operator is denoted as The root image reconstruction result is ,in This represents the reconstructed image of the root image and serves as the starting point for subsequent references.
[0069] Based on this, when the compression order points to a non-root image, prediction information can be generated based on the reconstructed image of its corresponding parent image, and residual information can be generated based on the prediction information. Specifically, for the current non-root image... and its parent image index The parent image reconstruction result is First, generate prediction information Then, residual information is generated. The residual information is then encoded, and the encoded result is written into a compressed bitstream. Simultaneously, loopback decoding is performed at the encoding end on the reconstructed image of the parent image and the encoded result of the residual information to obtain the reconstructed image of the current non-root image. For example, the residual encoding output can be a bit sequence. And make the reconstruction of the residuals from the decoding loop be... The reconstructed image can then be written as ,in The reconstructed image of the non-root image is used as a prediction reference for its child nodes, thereby satisfying the closure of the reference link in the reconstruction domain.
[0070] Furthermore, the generation of predictive information needs to simultaneously address local geometric offsets and illumination differences to adapt to cross-frame mismatches in heavy-duty vehicle recorders caused by changes in vehicle speed, subtle changes in viewing angle, and variations in day and night illumination. To this end, local feature matching can be performed between the reconstructed image of the parent image and the current non-root image, and the corresponding region can be determined based on the local feature matching results. Local features can be obtained from keypoints and descriptors in stable regions, and mismatches can be filtered out from the set of matched intrapoints through a consistency check. Based on this, geometric transformation parameters and illumination compensation parameters are jointly estimated within the corresponding region; the geometric transformation parameters can be represented as a homography matrix. or affine matrix The illumination compensation parameter can be expressed as gain. With bias After obtaining the parameters, a geometric transformation is performed on the reconstructed image of the parent image, and illumination compensation is applied to obtain the prediction information. For example, it can be written as:
[0071] in, Indicates the pixel position Predicted information pixel values at the location The function representing the pixel values of the reconstructed image from the parent image. This represents a coordinate mapping function determined by geometric transformation parameters. Indicates illumination compensation gain. This indicates a lighting compensation bias. For example, when vehicle movement causes a slight translation between adjacent frames and the overall lighting dims, It can be approximated as a small perturbation matrix containing translation components. A value slightly greater than 1 is acceptable. A small positive value can be taken to compensate for the decrease in brightness, thereby concentrating the residual information energy and improving coding efficiency.
[0072] Thus far, step S3 establishes an independent reconstruction starting point for the root image intra-frame coding by advancing the write stream and reconstruction loop in the compression order, and completes the closed-loop link of "prediction-residual-write stream-loop reconstruction" at non-root images with the parent image reconstruction as a reference. At the same time, joint parameter estimation of local feature matching, geometric transformation and illumination compensation is introduced in the prediction generation, so that the prediction information can fit the geometric and illumination changes of the vehicle scene, providing a stable reconstruction domain reference basis for subsequent higher compression efficiency and controllable reconstruction quality.
[0073] In step S4, the core task is to supplement the "transferable information of the parent image" and suppress dynamic interference on the reconstruction domain reference link of the non-root image, so that the subsequent enhancement stage can obtain more structurally consistent input. The input is the reconstructed image of the parent image, the reconstructed image of the non-root image, and the geometric transformation parameters and illumination compensation parameters. The output is the fused image corresponding to the non-root image. The fusion process needs to preserve structural continuity and detail differences at multiple scales.
[0074] Specifically, auxiliary information images corresponding to non-root images can be generated based on the reconstructed images of the parent images corresponding to non-root images, thereby generating transferable auxiliary information carriers in the reconstruction domain of the parent images.
[0075] In some embodiments of this application, when the non-root image index is Its parent image index is At that time, the reconstructed image of the parent image is denoted as The reconstructed image of a non-root image is denoted as The geometric transformation parameters are denoted as The illumination compensation parameter is denoted as Parameterization can be used to... Migration to Within the same geometric and illumination domain, and based on geometric transformation parameters and illumination compensation parameters, the reconstructed image of the parent image is parameterized to obtain the initial auxiliary information image, which can be written as:
[0076] in, This indicates the pixel position of the initial auxiliary information image. The pixel values at that location. Based on this, dynamic region identifiers can be generated based on the differences between the reconstructed images of the parent image and the reconstructed images of the non-root image, thereby explicitly identifying dynamic regions that are "unreliable after parent domain migration," preventing dynamic objects, motion blur, and transient occlusion from being introduced as transferable information. For example, a binary identifier map can be constructed from the differences after migration:
[0077] in, Indicates a dynamic region identifier. Indicates an indicator function, Indicates the difference threshold. Based on the dynamic region identifier, dynamic interference suppression and repair processing is performed on the initial auxiliary information image to prevent dynamic interference from being amplified in subsequent fusion, ultimately obtaining the auxiliary information image; wherein, the dynamic interference suppression processing includes at least one of smoothing processing and replacement processing. For example, when... Time can be used to Perform spatial smoothing in this region to reduce high-frequency artifacts; or apply spatial smoothing to this region. Pixel replacement at the same position is used to ensure content consistency, thereby obtaining auxiliary information images. .
[0078] After obtaining the auxiliary information image, it is necessary to complementarily synthesize the "real content of the reconstructed image" and the "transferable structure of the auxiliary information" at multiple scales. Therefore, a fusion process can be further performed, fusing the reconstructed image of the non-root image and the corresponding auxiliary information image at multiple scales to obtain the fused image corresponding to the non-root image. In this process, to simultaneously cover large-scale geometric structures and small-scale texture details, a scale pyramid can be constructed and cross-scale features extracted. Feature extraction is performed on the reconstructed image of the non-root image and the auxiliary information image at each scale to obtain the first and second features for each scale. For example, let the scale level be... ,use Indicates the first With the layer downsampling operator, the first feature and the second feature can be expressed as: , ,in , This represents the feature extraction mapping at the corresponding scale. Based on this, feature fusion is performed at each scale according to the spatial correspondence between the first and second features to obtain fused features at each scale. This ensures that aligned structural information is preferentially introduced from auxiliary information, while significantly different regions preferentially retain the reconstructed image content. For example, this can be achieved using a learnable gating graph. Weighting the two features point by point:
[0079] in, This represents element-wise multiplication. Finally, the multi-scale fusion features are scaled back to the original resolution and aggregated to obtain the final fused image corresponding to the non-root image. For example, the fusion features at each scale can be upsampled to the same resolution, summed, and then reconstructed using a convolution to obtain the fused image. This allows it to simultaneously possess structural consistency and detail recoverability.
[0080] Thus far, step S4 forms initial auxiliary information by parametric migration on the reconstruction domain of the parent image and suppresses dynamic interference by dynamic region identification constraints. Then, it completes spatial alignment fusion and backpropagation aggregation of the two features on the scale pyramid, forming a fused image for subsequent enhancement processing. This allows the structural reference information of the non-root image to be stably introduced under controlled dynamic perturbation, providing a consistent and reliable input basis for detail restoration in the subsequent enhancement stage.
[0081] In step S5, the core task is to perform learnable enhancement mapping with the fused image as input, enhance key details such as vehicle edges and road markings and suppress compression artifacts, so that the output has higher usable visual quality; its input is the fused image corresponding to the non-root image, and the output is the enhanced reconstructed image and the compressed reconstruction result of the similar image set.
[0082] Specifically, image enhancement processing can be performed on the fused image corresponding to the non-root image to output an enhanced reconstructed image, and a reconstructed result that can be directly delivered can be formed on the output side. Then, the enhanced reconstructed image can be used as the compressed reconstructed result of a similar image set.
[0083] In some embodiments of this application, to balance artifact removal and detail compensation, an augmentation network combining a dense residual structure and an attention mechanism can be used, and the fused image is denoted as... And it serves as input to the augmentation network. Specifically, dense residual structure processing can be performed on the fused image to obtain intermediate features, which can be denoted as... ,in Indicates intermediate features, This represents a feature map composed of cascaded dense residual blocks. Subsequently, scale recovery is performed on the intermediate features, and channel attention is introduced at key locations. Specifically, upsampling is performed on the intermediate features, and an attention mechanism is executed between the upsampling and output convolution processes to obtain channel weights. Based on this, to ensure that the attention weights are simultaneously constrained by global brightness statistics and local edge details, an edge feature map is further calculated based on the fused image. The edge feature map and the global statistical information of the fused image are then input into the attention mechanism for processing, resulting in channel weights constrained by both global statistical information and local statistical information of the edge feature map. For example, the edge feature map can be obtained by the gradient operator and denoted as... Global statistics can be obtained by global average pooling and denoted as... Attention mechanism outputs channel weight vector .
[0084] Furthermore, based on the channel weights, channel recalibration is performed on the intermediate features, and output convolution processing is then performed to obtain the enhanced reconstructed image. The channel recalibration can be expressed as... .
[0085] In this embodiment, to ensure that the enhanced output remains consistent with the input in content and compensates for details in the form of residuals, a global residual join can be used, so that the enhanced output consists of "enhancement mapping + input", which can be written as:
[0086] in, Indicates enhanced reconstructed image, This indicates an augmented mapping. This represents the fused image. To ensure that the augmentation mapping has an actionable optimization objective during the training phase, a loss function centered on reconstruction error can be used, which can be written as:
[0087] in, This indicates an increased loss. This represents a low-quality reconstructed image sample from the input. This represents the corresponding real image sample. This indicates enhanced network output. This represents the expected batch of data. For example, it can be set as follows: Take as The norm enhances the fidelity constraint on edge details and can increase the proportion of samples containing license plate and road marking areas in nighttime scenes of heavy-duty vehicle dashcams, so that the channel weights can be more stably improved on edge-related channels, thereby enhancing the suppression of artifacts and blur.
[0088] Thus far, step S5 extracts intermediate representations using dense residual structures, generates channel weights using an attention mechanism constrained by both global and edge local statistics, and performs channel recalibration. Combined with global residual connections, it forms a learnable enhancement link from the fused image to the enhanced reconstructed image, enabling the compressed reconstruction result to achieve synergistic improvement in artifact suppression and detail restoration, thereby forming an enhanced output that can be directly used as the compressed reconstruction result of similar image sets.
[0089] In summary, this method constructs a reference structure with clear parent-child reference relationships and forms intra-group dependency closures by using a similarity metric dominated by stable region features and suppressed by volatile region features. It combines prediction and residual coding based on reconstruction domain references with consistency transfer of geometric transformation and illumination compensation. Furthermore, it enhances structural continuity by introducing auxiliary information constrained by dynamic interference suppression and multi-scale fusion. Finally, it strengthens key details with attention enhancement constrained by both global and edge local statistics. This method can improve compression efficiency and reconstruction clarity under the conditions of light fluctuations and motion interference in heavy-duty vehicle recorders, suppress artifacts and detail loss, enhance the recognizability of elements such as license plate areas and road markings, and improve cross-frame consistency and stability, thereby improving the usability and engineering reliability of the compressed reconstruction results.
[0090] It should be noted that, although the embodiments in this application are based on... Figure 1 Steps S1 to S5 are described sequentially, but this does not mean that steps S1 to S5 must be performed in a strict order. The reason this embodiment follows this order is... Figure 1 The order in which steps S1 to S5 are described is provided to facilitate understanding of the technical solutions of the embodiments of this application by those skilled in the art. In other words, in the embodiments of this application, the order of steps S1 to S5 can be appropriately adjusted according to actual needs.
[0091] In some embodiments of this application, a heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement is applied to the offline compression scenario of similar image sets in pre-installed heavy-duty vehicle recorders. The recorder acquires continuous image sequences at fixed time intervals and forms similar image sets using adjacent / near neighbor time windows. For example, Collected continuously from the same road segment Images, with a resolution of The background road markings and the main body of the vehicle in front have a reusable structure across multiple images, with only minor displacement, local occlusion, and changes in lighting.
[0092] A complete implementation process may include the following steps: Step 1: Create a set of similar images It also completes feature extraction for stable and volatile regions, providing a unified data standard for subsequent reference relationship calculations. For example, for each image... Extracting feature vectors of stable regions The stable region covers vehicle outline features, license plate area features, and road marking features; at the same time, it identifies volatile regions and extracts suppressed feature vectors. The variable regions cover areas with varying illumination and areas affected by motion interference. To strengthen the dominance of stable regions in similarity measurement, higher weights can be given to keypoint matching pairs in stable regions during feature construction, allowing vehicle outlines, license plate areas, and road markings to contribute more to the measurement results. This adaptation strategy is used to make the reference structure more closely fit the redundant distribution.
[0093] Step two involves calculating pairwise similarity metrics for each image and generating a tree-like reference structure to characterize the reference relationships between images. For example, for any... and Constructing the first similarity measure component And construct a second similarity measure component. Then by weight The similarity measure is obtained by fusion. ,in, To balance the effects of stable structures and volatile disturbances, an example is taken in a daytime scenario. , Take as , Take as Based on this, construct a complete graph and... As edge weights, a minimum spanning tree is generated using the Prim algorithm to form a tree-like reference structure. Tree nodes represent images, tree edges represent the reference correlation strength between images, and the root image is selected as the reference starting point with the highest average similarity metric with the other images.
[0094] Step three involves grouping similar image sets according to the tree-like reference structure and determining the compression order. This ensures that reference dependencies occur continuously within a group and reduces the caching overhead caused by branch switching. For example, the minimum spanning tree is divided into groups by subtrees or branches. Each group is bounded by a boundary constraint, which ensures that parent-child reference relationships within a group participate in the compression dependency of the current group, while parent-child reference relationships outside the group do not. Then, a depth-first search is used to traverse the tree corresponding to each group to obtain the compression order within the group. The branches are processed sequentially according to the depth-first search order, so that reference dependencies within the same branch occur continuously, thereby reducing reference cache switching and better adapting to the resource-constrained automotive coding environment.
[0095] Step four: Perform encoding and decoding in compression order and close the reference link in the reconstruction domain to obtain the reconstructed image of the root image and the reconstructed images of each non-root image. For example, the root image is intra-coded to obtain a bit sequence. The root image is reconstructed by performing decoding at the encoding end. For non-root images and its parent image Reconstructing from the parent image Generate prediction information And form residual information The residual information is encoded and written into the compressed bitstream, and then loop-back decoded at the encoding end to obtain the result. Thus, non-root image reconstruction is obtained. To improve prediction consistency and residual sparsity, local feature matching is performed between the reconstructed image of the parent image and the non-root image to locate corresponding regions, and geometric transformation parameters are estimated from these corresponding regions. With illumination compensation parameters Based on right Perform the transformation to obtain the prediction information This adaptation strategy is used to reduce the cost of residual coding.
[0096] Step 5: Generate an auxiliary information image based on the reconstructed image of the parent image and fuse it with the reconstructed image of the non-root image at multiple scales to obtain the fused image corresponding to the non-root image. For example, based on... right Perform parameterization processing to obtain the initial auxiliary information image. and based on and Difference generation dynamic region identifier ,in, For example, take as (Measured using 8-bit grayscale intensity). Then based on... right Auxiliary information images are obtained by performing dynamic interference suppression processing. Dynamic interference suppression processing may include... Smoothing of the area or with Pixels at the same location are replaced. A scale pyramid is further constructed, and the replacement process is performed at each scale. and Feature extraction yields the first and second features, and feature fusion is performed at each scale based on spatial correspondence to obtain fused features. Finally, scale-based backpropagation aggregation is performed on the fused features at each scale to obtain the fused image. .
[0097] Step six: Perform image enhancement processing on the fused image to output an enhanced reconstructed image, which serves as the compressed reconstruction result for a set of similar images. For example, the fused image is denoted as the low-quality input. And an enhanced mapping is constructed based on the dense residual structure. Global residual connection is used to enhance the output of the reconstructed image. ,in, To enhance the output, the loss function of the augmentation network is exemplarily chosen as follows: ,in, For the input low-quality reconstructed image sample, To correspond to real image samples, To enhance network output, For batch data expectations.
[0098] Image enhancement processing can be achieved by combining dense residual structures with attention mechanisms. For example, enhancement mapping... use A series of dense residual blocks are cascaded, each containing [a certain number of] dense residual blocks. There are 1 convolutional layer with a kernel size of 1. The number of output channels in each convolutional layer is 1. Dense connections are used within blocks to reuse intermediate features; local feature fusion within blocks employs... Convolution compresses the concatenated features back to their original state. Channels. To adapt to the resolution alignment requirements of multi-scale fusion output, the downsampling path can adopt two layers with a stride of 2. The convolutional layer has 64 output channels in the first layer and 128 output channels in the second layer; the upsampling path can use a single layer with a stride of 2. The transposed convolution restores the features to the target resolution and outputs 64 channels. An attention mechanism is set between the upsampling and output convolution processes. This attention mechanism uses a compression-activation structure: global average pooling is performed on the feature map to obtain a length of... The channel description vector is passed through two fully connected layers to obtain the channel weights. The output dimensions of the two fully connected layers are, for example, 64 and 128 respectively; then, recalibration is performed using channel-wise multiplication, resulting in recalibrated features. The image is then processed through output convolution to generate an enhanced reconstructed image. .
[0099] To further enhance details such as vehicle edges and road markings, edge feature maps and global statistical information can be introduced into the input side of the attention mechanism, making the channel weights more sensitive to edge-related channels, thereby improving the detail restoration effect.
[0100] Through the complete implementation process described above, this method can effectively resolve cross-image redundancy in similar image sets within a group using a tree-like reference structure and grouping boundary constraints, and reduce the caching and scheduling overhead caused by branch switching by using a depth-first search-organized compression order; simultaneously, it jointly utilizes the reference links in the reconstruction domain. and Improve prediction consistency and enhance residual sparsity to reduce residual coding costs at the source; further, through dynamic region identification. Suppressing the contamination of auxiliary information by motion and illumination perturbations, and using multi-scale fusion and enhanced mapping. Targeted suppression and detail compensation of low bitrate artifacts enhance the reconstructed image. Higher clarity and cross-frame consistency are achieved in key elements such as license plate areas, road markings, and vehicle edges, thereby simultaneously improving both compression efficiency and reconstruction quality, and enhancing the engineering usability and stability of the compressed reconstruction results.
[0101] It should be understood that the step numbers identified by "Step 1, Step 2" and other similar forms in the above embodiments are only used to distinguish different steps and do not limit the steps to be executed in the order of these numbers. The specific execution order of each step can be adjusted according to its functional requirements and the inherent logic in the actual application scenario. The above step numbers should not be interpreted as a limitation on the implementation process of the embodiments of this application.
[0102] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.
Claims
1. A method for compressing heavy-duty vehicle recorder image sets based on reference structure and image enhancement, characterized in that, include: A similarity metric is performed on a set of similar images. Based on the similarity metric, a reference structure is constructed to characterize the reference relationship between images. The set of similar images is then grouped according to the reference structure. The compression order of similar image sets is determined based on the reference structure and grouping results. The compression order satisfies the parent-child reference order constraint defined by the reference structure. The similar image set is encoded and decoded in the compression order. The root image is encoded intra-frame, and the non-root images are encoded using prediction and residual information with the parent image as a reference, to obtain the reconstructed image of the root image and the reconstructed images of each non-root image. Based on the reconstructed image of the parent image corresponding to the non-root image, an auxiliary information image corresponding to the non-root image is generated, and the reconstructed image of the non-root image and the corresponding auxiliary information image are fused at multiple scales to obtain the fused image corresponding to the non-root image. Image enhancement processing is performed on the fused image corresponding to the non-root image to output the enhanced reconstructed image, and the enhanced reconstructed image is used as the compressed reconstruction result of a similar image set.
2. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 1, characterized in that, The steps for measuring the similarity of a set of similar images include: Stable region features are extracted from each image in the similar image set, and these stable region features are used as the dominant features for similarity measurement. Stable region features include vehicle outline features, license plate region features, and road marking features. Identify volatile regions in each image and use the features corresponding to these volatile regions as suppression features. Volatile regions include regions with changes in illumination and regions with motion interference. A first similarity metric component is formed based on the dominant feature, a second similarity metric component is formed based on the suppression feature, and the similarity metric between the two images is determined based on the first and second similarity metric components.
3. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 1, characterized in that, The steps for constructing a reference structure to characterize reference relationships between images based on similarity metrics include: A graph structure is constructed based on a similarity metric, and the similarity metric is used as the edge weight. A tree-like reference structure is generated based on edge weights. The tree-like reference structure includes a root image and parent-child reference relationships, and the parent-child reference relationships satisfy the constraints of similarity measurement. The parent image of each non-root image is determined based on the tree reference structure.
4. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 3, characterized in that, The steps for grouping similar image sets based on a reference structure include: Grouping is formed by splitting the tree reference structure into subtrees or branches; Set group boundary constraints for each group. The group boundary constraints limit the parent-child reference relationships within the group to participate in the compressed dependencies of the current group, and limit the parent-child reference relationships outside the group to not participate in the compressed dependencies of the current group.
5. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 4, characterized in that, The steps for determining the compression order of similar image sets based on the reference structure and grouping results include: Perform a traversal on the tree reference structure corresponding to each group to obtain the compression order within the group; During the traversal, the similarity measurement determines the order of parent and child nodes based on the parent-child reference sequence constraint, and determines the similarity measurement of the order of appearance of parent and child chains based on the continuity constraint of parent and child chains within the same branch. When a branch switch occurs during traversal, the reference relationship after the branch switch is limited to the group based on the group boundary constraints.
6. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 1, characterized in that, The steps for encoding and decoding similar image sets according to the compression order include: The current image in the set of similar images is selected sequentially according to the compression order, and the encoding result of the current image is written into the compressed bit stream. When the compression order points to the root image, intra-frame coding is performed on the root image, and the coding result of the root image is decoded to obtain the reconstructed image of the root image. When the compression order points to a non-root image, prediction information is generated based on the reconstructed image of the corresponding parent image of the non-root image, and residual information is generated based on the prediction information. The residual information is encoded and the encoding result of the residual information is written into the compressed bit stream. Decoding is performed on the reconstructed image of the parent image and the encoding result of the residual information to obtain the reconstructed image of the non-root image.
7. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 6, characterized in that, The steps for generating prediction information from the reconstructed image based on the parent image corresponding to the non-root image include: Perform local feature matching between the reconstructed image of the parent image and the non-root image, and determine the corresponding region based on the local feature matching results; Geometric transformation parameters and illumination compensation parameters are estimated based on the corresponding region; The reconstructed image of the parent image is transformed based on the geometric transformation parameters and the illumination compensation parameters to obtain the prediction information.
8. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 7, characterized in that, The steps for generating auxiliary information images corresponding to non-root images based on the reconstructed images of the parent images corresponding to non-root images include: The reconstructed image of the parent image is parametrically processed based on the geometric transformation parameters and illumination compensation parameters to obtain the initial auxiliary information image; Dynamic region identifiers are generated based on the differences between the reconstructed image of the parent image and the reconstructed image of the non-root image. Based on the dynamic region identifier, dynamic interference suppression processing is performed on the initial auxiliary information image to obtain the auxiliary information image. The dynamic interference suppression processing includes at least one of smoothing processing and replacement processing.
9. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 8, characterized in that, The steps of fusing the reconstructed image of the non-root image and the corresponding auxiliary information image at multiple scales to obtain the fused image corresponding to the non-root image include: A scale pyramid is constructed, and feature extraction is performed on the reconstructed image and auxiliary information image of the non-root image at each scale to obtain the first feature and the second feature at each scale. Feature fusion is performed at each scale based on the spatial correspondence between the first and second features to obtain the fused features at each scale; Scale backpropagation aggregation is performed on the fusion features at each scale to obtain the fusion image corresponding to the non-root image.
10. The heavy-duty vehicle recorder image set compression method based on reference structure and image enhancement as described in claim 1, characterized in that, The steps of performing image enhancement processing on the fused image corresponding to the non-root image and outputting the enhanced reconstructed image include: Dense residual structure processing is performed on the fused image to obtain intermediate features; Upsampling is performed on the intermediate features, and an attention mechanism is applied between the upsampling and output convolution processes to obtain the channel weights. Edge feature maps are calculated based on the fused image, and the global statistical information of the edge feature maps and the fused image is input into the attention mechanism for processing to obtain channel weights constrained by the global statistical information and the local statistical information of the edge feature maps; Channel recalibration is performed on intermediate features based on channel weights, and output convolution processing is performed to obtain an enhanced reconstructed image.