A multi-modal target detection and recognition method based on image processing
By adaptively adjusting the image extraction strategy and feature point optimization, the problem of region segmentation in vehicle re-identification not adapting to actual scenarios is solved, and the accuracy and robustness of multimodal re-identification are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING TOPMOO TECH
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-14
AI Technical Summary
In existing technologies, vehicle re-identification methods fail to effectively consider the actual driving state and environmental factors of vehicles, making it difficult for the region division method to adapt to complex traffic environments and affecting the multimodal re-identification effect.
Image state is determined based on viewpoint richness and feature stability. The extraction strategy is adjusted to uniform or compensated extraction. Feature points are optimized by combining feature anchor value and spatial coverage. Key points are determined based on the evaluation balance. One or two types of sub-regions are divided and the area of feature regions is adjusted to achieve adaptive region division.
It improves the accuracy and robustness of multimodal re-identification, enhances the precision and adaptability of key points, and improves the effect of vehicle re-identification.
Smart Images

Figure CN122391955A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of target recognition, and more particularly to a multimodal target detection and recognition method based on image processing. Background Technology
[0002] With the development of artificial intelligence and big data technologies, vehicle re-identification technology, with its powerful feature learning capabilities, has been widely applied in key scenarios such as suspect vehicle tracking, unmanned parking lot management, and autonomous driving. In cross-modal vehicle re-identification tasks, vehicle appearance exhibits significant intra-class differences and inter-class similarities due to factors such as viewing angle, lighting, and model, posing a severe challenge to recognition accuracy. To improve the efficiency of multimodal re-identification, existing methods typically introduce local features for fine-grained matching. However, existing technologies often employ predefined key points or fixed region divisions, which are difficult to flexibly handle the dynamic changes in vehicle appearance under complex traffic environments, resulting in insufficient local region segmentation accuracy and thus limiting the effectiveness of multimodal re-identification. Therefore, how to adaptively adjust the local region segmentation strategy according to the actual shooting conditions of the target vehicle to improve the accuracy and robustness of cross-modal re-identification has become a critical problem that urgently needs to be solved.
[0003] Chinese Patent Publication No. CN118262300A discloses a method for vehicle re-identification based on instance segmentation. The method includes the following steps: acquiring an image of a vehicle to be re-identified; obtaining the re-identification result of the vehicle image using a re-identification network model based on the image; wherein the re-identification network model includes a ResNet-50 network and a segmentation network, the ResNet-50 network being used for global feature extraction of the image to be re-identified, and the segmentation network being used for local feature extraction of the image; feature extraction of the vehicle is performed through the segmentation network, using a set of semantic feature centroids to capture the vehicle's feature information, and generating a semantic probability map by calculating the distance between each pixel-level feature vector and all feature centroids. A self-correction module is proposed, using the instance probability map to reduce erroneously activated feature regions; this improves the training efficiency of the model and effectively improves the accuracy of vehicle re-identification. However, while the above technical solution discloses the impact of region segmentation on the re-identification effect, it does not consider the influence of the vehicle's actual driving state and environmental parameters on the region segmentation method, leading to the problem that the region segmentation method is difficult to adapt to actual scenarios, thus resulting in poor vehicle re-identification performance. Summary of the Invention
[0004] To address this issue, the present invention provides a multimodal target detection and recognition method based on image processing, which overcomes the problem in the prior art that the region division method is difficult to adapt to the actual scene due to the failure to consider the influence of parameters such as the actual driving state of the vehicle and the environment during the region division process, resulting in poor vehicle re-identification performance.
[0005] To achieve the above objectives, the present invention provides a multimodal target detection and recognition method based on image processing, comprising: The image state is determined based on the richness of viewpoints and the stability of features, and the extraction strategy corresponding to the target video segment is adjusted from uniform extraction to compensated extraction based on the image state. When performing uniform sampling, the number of samples is adjusted based on the comparison between the viewpoint satisfaction and the preset viewpoint satisfaction. When performing compensated sampling, the compensated segment is determined based on the region parameter characterization value. Several reference images are obtained by extracting the target video segment according to the corresponding extraction parameters; For the extracted baseline image, the system determines whether to perform feature point optimization based on the feature anchor point representation value and spatial coverage. In feature point optimization, the system determines whether to adjust the point selection strategy from determining key points based on call frequency to determining key points based on feature evaluation value based on the evaluation balance. The sub-region category is determined as either a Class I or Class II sub-region based on the regional prominence. The distribution status is determined based on the Class I quantity and Class I clustering degree. Based on the distribution status, it is determined whether to adjust the area of the feature region.
[0006] Furthermore, for image states where the view richness is greater than the preset view richness and the feature stability is greater than the preset feature stability, the extraction strategy is determined to be uniform extraction.
[0007] Furthermore, during the uniform sampling process, when the viewpoint satisfaction is less than or equal to the preset viewpoint satisfaction, the sampling quantity is adjusted based on the viewpoint satisfaction coefficient. The viewpoint satisfaction coefficient is determined based on the viewpoint satisfaction degree and the preset viewpoint satisfaction degree. The number of samples extracted is positively correlated with the viewpoint satisfaction coefficient.
[0008] Furthermore, for image states where the viewpoint richness is less than or equal to the preset viewpoint richness or the feature stability is less than or equal to the preset feature stability, the extraction strategy is determined to be to perform compensation extraction on the compensation segment. The compensation segment is a video segment whose regional parameter characterization value is greater than the preset regional parameter characterization value.
[0009] Furthermore, for a reference image where the feature anchor point representation value is greater than the preset feature anchor point representation value and the spatial coverage is greater than the preset spatial coverage, feature point optimization is performed.
[0010] Furthermore, when the evaluation balance is greater than the preset evaluation balance, the selection strategy for the optimal location is to determine the key points based on the call frequency. When the evaluation balance is less than or equal to the preset evaluation balance, the selection strategy for the optimal location is to determine the key points based on the feature evaluation value.
[0011] Furthermore, feature evaluation values are determined based on the complexity representation values of the surrounding domain and the number of neighboring points; The feature evaluation value is positively correlated with the surrounding complexity representation value and the number of neighboring points.
[0012] Furthermore, when executing the first optimization strategy, reference models with structural similarity greater than the preset structural similarity in the historical records are selected, and points with spatial parameter similarity greater than the preset spatial parameter similarity in the reference models are recorded as reference points, the number of reference points is recorded as the call frequency, and feature points with a call frequency greater than the preset call frequency are recorded as key points. The first preferred strategy is to determine key points based on the call frequency.
[0013] Furthermore, the feature region is evenly divided into several sub-regions. When the region prominence is greater than the preset region prominence, the sub-region is determined to be a type of sub-region. Regions other than the first-class sub-regions are denoted as second-class sub-regions.
[0014] Furthermore, for a distribution state where the quantity of a certain category is greater than a preset quantity of another category and the degree of aggregation of that category is greater than a preset degree of aggregation of another category, the area of the adjustment feature region is determined based on the degree of regional salience difference.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: the technical solution of the present invention determines the image state based on view richness and feature stability. View richness characterizes the view coverage of the target vehicle within the target video segment, and the comparison result of feature stability with the preset feature stability characterizes whether the speed of the target vehicle is stable during driving. The present invention achieves multimodal fusion in the target detection process by fusing image data and the actual driving data of the target vehicle. The image state is jointly determined by multimodal fusion, and different extraction strategies are set accordingly. This avoids the problem that the single extraction method in the prior art is difficult to effectively meet the actual operating state of the vehicle, resulting in low use value of the extracted image, and thus effectively improves the multimodal re-identification effect.
[0016] Furthermore, in the technical solution of this invention, the determination of whether to optimize feature points is based on feature anchor point representation values and spatial coverage. The feature anchor point representation values characterize the registration degree of the reference image, and the spatial coverage characterizes the tolerance of feature points to local occlusion. Considering that the target vehicle may be occluded by obstacles or other pedestrians during actual driving, if the feature points cover the entire target, then even if some areas are occluded, the unoccluded feature points can still be used for feature point matching. By optimizing the key points of the feature points, the accuracy of the key points is further improved, thereby effectively improving the multimodal re-identification effect.
[0017] Furthermore, the optimization strategy for determining key points based on the comparison result between the evaluation balance degree and the preset evaluation balance degree in the technical solution of the present invention is to determine key points based on the call frequency or based on the feature evaluation value. The feature evaluation value characterizes the anti-interference ability of the feature points. The larger the feature evaluation value, the stronger the anti-interference ability. The evaluation balance value characterizes the degree of difference in the anti-interference ability of the feature points. If the degree of difference in ability is large, the key points are confirmed based on the feature evaluation value. If the degree of difference in ability is small, the key points cannot be optimized based on the feature evaluation value. In this case, the historical call frequency of points with high spatial parameter similarity in the reference model in the historical record is considered as the benchmark for key point confirmation, thereby improving the adaptability of key point screening and the accuracy of key point confirmation, and thus effectively improving the multimodal re-identification effect. Attached Figure Description
[0018] Figure 1 This is a flowchart of the multimodal target detection and recognition method based on image processing according to the present invention; Figure 2 This is a flowchart illustrating how the present invention determines the extraction strategy based on viewpoint richness and feature stability. Figure 3 This is a flowchart illustrating the process of determining the optimal positioning strategy based on the evaluation balance of the present invention. Figure 4 This is a flowchart illustrating the process of determining whether to perform feature point optimization based on feature anchor point characterization values and spatial coverage, according to the present invention. Detailed Implementation
[0019] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0020] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0021] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0022] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0023] Please see Figures 1 to 4 As shown, this invention provides a multimodal target detection and recognition method based on image processing, comprising: The image state is determined based on the richness of viewpoints and the stability of features, and the extraction strategy corresponding to the target video segment is adjusted from uniform extraction to compensated extraction based on the image state. When performing uniform sampling, the number of samples is adjusted based on the comparison between the viewpoint satisfaction and the preset viewpoint satisfaction. When performing compensated sampling, the compensated segment is determined based on the region parameter characterization value. Several reference images are obtained by extracting the target video segment according to the corresponding extraction parameters; For the extracted baseline image, the system determines whether to perform feature point optimization based on the feature anchor point representation value and spatial coverage. In feature point optimization, the system determines whether to adjust the point selection strategy from determining key points based on call frequency to determining key points based on feature evaluation value based on the evaluation balance. The sub-region category is determined as either a Class I or Class II sub-region based on the regional prominence. The distribution status is determined based on the Class I quantity and Class I clustering degree. Based on the distribution status, it is determined whether to adjust the area of the feature region.
[0024] The application scenario of this invention is to perform multimodal re-identification of target vehicles. Several video frames are uniformly extracted within the target video segment, and these video frames are recorded as reference images. Each reference image contains several feature points. The feature points are used to train a feature point detection network using the vehicle key point dataset CarFusion. Each reference image is input into the feature point detection network, which outputs several heatmaps. The peak position of each heatmap is taken as the coordinates of each feature point. Each feature point corresponds to a feature region. This invention performs image state analysis on the first uniformly extracted reference image. In this invention, the target video segment comes from a fixed camera used for road traffic monitoring.
[0025] The value of the number of feature points is easy to understand. The higher the user's requirements for the re-identification effect, the larger the value of the number of feature points. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In the embodiments of this invention, the value of the number of feature points is 10, with the unit being individual points. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth, so as to meet the user's requirements for the re-identification effect.
[0026] The reference working conditions record several historical multimodal re-identification processes that met user requirements, including viewpoint richness, feature stability, pixel comparison accuracy, viewpoint satisfaction, regional parameter representation values, duration, feature anchor point representation values, spatial coverage, evaluation balance, surrounding area complexity representation values, number of neighboring points, feature evaluation values, distance, structural similarity, regional prominence, call frequency, spatial parameter similarity, number of classes, and clustering degree of classes. Whether multimodal re-identification meets user requirements can be determined based on self-defined indicators (e.g., re-identification accuracy). For example, if the re-identification accuracy is greater than the re-identification accuracy required by the user, then the user requirements are met. This is content already known to those skilled in the art and is not limited here. Taking the preset viewpoint richness as an example, a preset value selection method is provided. The viewpoint richness of the reference working conditions is extracted, outliers are removed, and the average viewpoint richness after removing outliers is recorded as the preset viewpoint richness. Methods for removing outliers include, but are not limited to, the 3σ method.
[0027] The preset values in this invention include: preset viewpoint richness, preset feature stability, preset viewpoint satisfaction, preset region parameter characterization value, preset pixel comparison degree, duration, preset feature anchor point characterization value, preset spatial coverage, preset tolerance hole radius, preset evaluation balance, preset surrounding complexity characterization value, preset number of neighboring points, preset feature evaluation value, preset distance, preset length, preset structural similarity, preset region prominence, preset number of first-class, preset first-class aggregation degree, benchmark extraction quantity, number of feature points, number of angle intervals, number of sub-regions, preset spatial parameter similarity, and preset calling frequency.
[0028] Specifically, for image states where the view richness is greater than the preset view richness and the feature stability is greater than the preset feature stability, the extraction strategy is determined to be uniform extraction.
[0029] Within the target video segment, the monitoring range of the camera is evenly divided into several angle intervals, and the coverage angle interval of each reference image is obtained. The richness of the viewpoint = the number of coverage angle intervals / the total number of angle intervals. The coverage angle interval is the angle interval between the driving vector corresponding to the driving direction of the target vehicle in the target video segment and the optical axis corresponding to the optical axis of the camera. The feature stability = (maximum driving speed - average driving speed) / (average driving speed - minimum driving speed).
[0030] The target feature point is tracked using optical flow. In the reference image, the coordinates of the target feature point are (u1, v1). In the preceding frame of the reference image, which is adjacent to the reference image in time, the coordinates of the target feature point are (u2, v2). Therefore, the pixel displacement of the target feature point is (Δu, Δv), where Δu = u1 - u2 and Δv = v1 - v2. The pixel displacement (Δu, Δv) is converted into the actual displacement ΔS through inverse perspective transformation. The instantaneous velocity of the target feature point is V = ΔS / the time interval between two adjacent reference images. The instantaneous velocity of the target feature point is recorded as the instantaneous driving speed of the target vehicle. The maximum instantaneous driving speed of the target vehicle within the target video segment is recorded as the maximum driving speed, the minimum instantaneous driving speed of the target vehicle within the target video segment is recorded as the minimum driving speed, and the average instantaneous driving speed of the target vehicle within the target video segment is recorded as the average driving speed. The target feature point is the center point of the smallest recognition box that can completely cover the target vehicle.
[0031] The number of angle intervals is determined by the user's requirements for re-identification performance. The present invention uses the false alarm rate of re-identification to characterize the re-identification performance. In this embodiment, the number of angle intervals is 10, with the unit being individual intervals. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for re-identification performance.
[0032] The values of preset view richness and preset feature stability are easily understood. The higher the user's requirements for the re-identification effect, the larger the value of the preset view richness and the larger the value of the preset feature stability. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In this embodiment of the invention, the preset view richness is set to 0.8 and the preset feature stability is set to 2. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0033] Specifically, during the uniform sampling process, when the viewpoint satisfaction is less than or equal to the preset viewpoint satisfaction, the sampling quantity is adjusted based on the viewpoint satisfaction coefficient. The viewpoint satisfaction coefficient is determined based on the viewpoint satisfaction degree and the preset viewpoint satisfaction degree. The number of samples extracted is positively correlated with the viewpoint satisfaction coefficient.
[0034] When the viewpoint satisfaction level is greater than the preset viewpoint satisfaction level, the baseline sampling quantity is used.
[0035] Obtain the heading angle of the target vehicle within the reference image. The heading angle is the angle between the front of the target vehicle and the horizontal coordinate axis of the reference image. The difference in heading angle between two temporally adjacent reference images is calculated as follows: The =2, 3, ..., Q, where Q is the number of reference images, and the viewpoint satisfaction = 1 - The The maximum heading angle within the reference image. The minimum heading angle within the reference image. This represents the maximum value of the heading angle difference.
[0036] The number of samples extracted = baseline number of samples extracted × viewpoint satisfaction coefficient, where the viewpoint satisfaction coefficient = preset viewpoint satisfaction / viewpoint satisfaction. The value of the baseline number of samples is determined by the user's higher requirements for re-identification performance; in this invention, the baseline number of samples is 100, with units of frames. The value of the baseline number of samples is related to the duration of the target video segment; it is understood that the longer the target video segment, the larger the value of the baseline number of samples. In this invention, the duration of the target video is 20 minutes.
[0037] Regarding the value of the preset viewpoint satisfaction, the higher the user's requirements for the re-identification effect, the larger the value of the preset viewpoint satisfaction. This invention characterizes the re-identification effect through the false alarm rate of re-identification. In this embodiment of the invention, the value of the preset viewpoint satisfaction is set to 0.85. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth, so as to meet the user's requirements for the re-identification effect.
[0038] Specifically, for image states where the viewpoint richness is less than or equal to the preset viewpoint richness or the feature stability is less than or equal to the preset feature stability, the extraction strategy is to perform compensation extraction on the compensation segment. The compensation segment is a video segment whose regional parameter characterization value is greater than the preset regional parameter characterization value.
[0039] The absolute value of the difference between the maximum and minimum pixel values within the feature region corresponding to the feature point is recorded as the pixel comparison degree. Feature regions with a pixel comparison degree less than or equal to the preset pixel comparison degree are recorded as regions to be analyzed. The region parameter characterization value = the number of regions to be analyzed / the number of feature regions. When the region parameter characterization value is greater than the preset region parameter characterization value, the local contrast of the feature region corresponding to the feature point is poor. At this time, the number of reference images is increased by compensation extraction to obtain more data for subsequent key point determination, thereby improving the setting accuracy of key points and thus improving the multimodal vehicle re-identification effect.
[0040] The video segment with the longest duration that meets the extraction criteria is selected and designated as the compensation segment. Other video segments are uniformly extracted. The extraction criteria are that the regional parameter representation value of each video frame within the target video segment is less than or equal to a preset regional parameter representation value, and feature points exist in each video frame. For the compensation segment, the number of reference images is uniformly increased during extraction. The adjusted number of reference images = the number of extracted reference images for the compensation segment × the duration of the compensation segment / preset duration. In this invention, the duration of the compensation segment is greater than the preset duration. It is easy to understand that the higher the user's requirements for re-identification performance, the smaller the preset duration. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In this embodiment, the preset duration is 5 minutes. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for re-identification performance.
[0041] When the duration of the compensation segment is less than or equal to the preset duration, the present invention determines that no compensation extraction is required for the compensation segment.
[0042] The preset pixel comparison degree is understood to be set such that the higher the user's requirements for the re-identification effect, the larger the preset pixel comparison degree value. This invention characterizes the re-identification effect by the false alarm rate of re-identification. In this embodiment of the invention, the preset pixel comparison degree value is set to 30. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth, so as to meet the user's requirements for the re-identification effect.
[0043] The value of the preset region parameter is understood to be larger as the user's requirements for the re-identification effect are higher. This invention uses the false alarm rate of re-identification to represent the re-identification effect. In this embodiment of the invention, the value of the preset region parameter is set to 0.8. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth, so as to meet the user's requirements for the re-identification effect.
[0044] Specifically, for a reference image where the feature anchor point representation value is greater than the preset feature anchor point representation value and the spatial coverage is greater than the preset spatial coverage, feature point optimization is performed.
[0045] A number of points for feature analysis are uniformly set on the surface of the target vehicle, and these points are recorded as feature points. The feature points include target feature points. For each feature point, the minimum distances between the feature point and other feature points are obtained, and the maximum value among the minimum distances is recorded as the effective distance.
[0046] The number of feature points is denoted as the feature anchor point representation value, and the spatial coverage = -1; where, For effective distance, The preset tolerance hole radius is set to a value that varies depending on the user's requirements for re-identification performance. The higher the user's requirements for re-identification performance, the smaller the preset tolerance hole radius will be. This invention uses the false alarm rate of re-identification to characterize the re-identification performance. In this embodiment of the invention, the preset tolerance hole radius is set to 0.8 cm. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for re-identification performance.
[0047] The preset feature anchor value and preset spatial coverage are set in such a way that it is easy to understand that the higher the user's requirements for the re-identification effect, the larger the preset feature anchor value and the larger the preset spatial coverage value will be. This invention uses the false alarm rate of re-identification to represent the re-identification effect. In this embodiment of the invention, the preset feature anchor value is 10 and the preset spatial coverage is 0.85. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0048] Specifically, when the evaluation balance is greater than the preset evaluation balance, the selection strategy for the optimal location is to determine the key point based on the call frequency. When the evaluation balance is less than or equal to the preset evaluation balance, the selection strategy for the optimal location is to determine the key points based on the feature evaluation value.
[0049] Assess balance = ;in, For the first Feature evaluation value of each feature point for The average of the feature evaluation values of each feature point The number of feature points is used to determine the preset evaluation balance value. The higher the user's requirements for the re-identification effect, the larger the preset evaluation balance value will be. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In this embodiment of the invention, the preset evaluation balance value is 0.5. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0050] Specifically, feature evaluation values are determined based on the complexity of the surrounding domain and the number of neighboring points; The feature evaluation value is positively correlated with the surrounding complexity representation value and the number of neighboring points.
[0051] Feature evaluation value = γ1 × surrounding complex representation value / preset surrounding complex representation value + γ2 × number of neighboring points / preset number of neighboring points; where γ1 and γ2 are weight coefficients, and γ1 + γ2 = 1. The value of γ1 and γ2 can be determined by the user through historical user data using deep learning to obtain the influence of surrounding complex representation value and the number of neighboring points on the feature evaluation value, and then select the corresponding weight coefficient value. It can be understood that the greater the influence of surrounding complex representation value on feature evaluation value, the larger the value of γ1, and the greater the influence of the number of neighboring points on feature evaluation value, the larger the value of γ2. This invention provides a value of γ1 and γ2, in which γ1 is 0.5 and γ2 is 0.5.
[0052] Feature points with feature evaluation values greater than preset feature evaluation values are recorded as key points. The point selection strategy for determining key points based on feature evaluation values is recorded as the second selection strategy. For the preset feature evaluation value, the higher the user's requirements for the re-identification effect, the larger the preset feature evaluation value. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In this embodiment of the invention, the preset feature evaluation value is set to 0.78. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0053] For a single feature point, a circular region with the feature point as the center and a preset length as the radius is denoted as the feature region of the feature point. The method for confirming the complexity representation value of the region is to record the duration for which the pixel comparison degree corresponding to the feature point in the target video segment is greater than the preset pixel comparison degree as the region stability duration of the feature point. The complexity representation value of the region = region stability duration / duration of the target video segment. Feature points whose distance from the feature point is less than the preset distance are recorded as neighboring feature points, and the number of neighboring feature points is recorded as the number of neighboring points.
[0054] The preset values for the complexity representation value of the surrounding area and the preset number of neighboring points are easily understood to be such that the greater the user's tolerance for the influence of the complexity representation value of the surrounding area on the feature evaluation value, the larger the preset value of the complexity representation value of the surrounding area will be, and the greater the user's tolerance for the influence of the number of neighboring points on the feature evaluation value will be, the larger the preset number of neighboring points will be. This invention uses the false alarm rate of re-identification to represent the re-identification effect. In this embodiment of the invention, the preset value of the complexity representation value of the surrounding area is 0.75, and the preset number of neighboring points is 3, with the unit being individual points. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0055] Regarding the preset length and preset distance values, the higher the user's requirements for the re-identification effect, the smaller the preset length and preset distance values will be. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In this embodiment of the invention, the preset length is 3 in mm and the preset distance is 8 in mm. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0056] Specifically, when executing the first optimization strategy, reference models with structural similarity greater than the preset structural similarity in the historical records are selected, and points with spatial parameter similarity greater than the preset spatial parameter similarity in the reference models are recorded as reference points. The number of reference points is recorded as the call frequency, and feature points with a call frequency greater than the preset call frequency are recorded as key points. The first preferred strategy is to determine key points based on the call frequency.
[0057] The system acquires historical records of the 3D model of the target vehicle that meet the user's re-identification accuracy requirements. The historical records include vehicle information and key point sets. The vehicle information includes, but is not limited to, vehicle color and vehicle model. The reference model should include vehicle type, shooting angle, spatial location of key points, and semantic labels.
[0058] Target vehicle model Compared with the reference model Perform coordinate alignment, where... =1,2,...,p, where p is the number of reference models for the target vehicle model. With any reference model Construct a model of the target vehicle With any reference model The smallest bounding box that completely contains the target vehicle model is used, and this bounding box is divided into several voxels. The number of points in each voxel is obtained, and the voxels with points are recorded as occupied voxels. A three-dimensional binary array of the target vehicle model and the reference model is then used. and I = U = Structural similarity = I / U; where, the This refers to coordinates in three-dimensional space. The direction in which the front of the vehicle points towards the rear. The direction from the left to the right of the vehicle. The direction in which the vehicle's wheels point towards the roof.
[0059] The distance between feature points of the target vehicle and feature points of reference models with a structural similarity greater than a preset structural similarity is extracted. For a single feature point of the target vehicle, the spatial parameter similarity is calculated as 1 / the minimum distance between that feature point and feature points of each reference model. The preset structural similarity and preset spatial parameter similarity values are determined by the user's higher requirements for re-identification performance; the higher the preset structural similarity and preset spatial parameter similarity values, the larger these values become. This invention characterizes the re-identification performance through the false alarm rate. In this embodiment, the preset structural similarity is 0.9 and the preset spatial parameter similarity is 0.3, in mm. This setting aims to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for re-identification performance.
[0060] Feature points with a call frequency greater than a preset call frequency are recorded as key points. The value of the preset call frequency is larger when the user has higher requirements for the re-identification effect. This invention characterizes the re-identification effect by the false alarm rate of re-identification. In this embodiment of the invention, the preset call frequency is set to 100. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0061] Specifically, the feature region is evenly divided into several sub-regions. When the region prominence is greater than the preset region prominence, the sub-region is determined to be a type of sub-region. Regions other than the first-class sub-regions are denoted as second-class sub-regions.
[0062] The method for confirming the region prominence is to obtain the pixel values of the pixels in each sub-region. For a single sub-region, the region prominence = number to be analyzed / total number of pixels in the sub-region. The number to be analyzed is the number of pixels in the region whose pixel values are greater than the average pixel value. The average pixel value is the average of the pixel values of the pixels in the feature region.
[0063] The value of the preset region prominence is easy to understand. The higher the user's requirements for the re-identification effect, the larger the value of the preset region prominence. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In the embodiments of this invention, the value of the preset region prominence is 0.6. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth, so as to meet the user's requirements for the re-identification effect.
[0064] Regarding the value of the number of sub-regions, the higher the user's requirements for the re-identification effect, the larger the value of the number of sub-regions. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In the embodiments of this invention, the value of the number of sub-regions is 25, with the unit being individual regions. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth, so as to meet the user's requirements for the re-identification effect.
[0065] Specifically, for a distribution state where the quantity of a certain category is greater than a preset quantity of another category and the degree of clustering of that category is greater than a preset degree of clustering of another category, the area of the adjustment feature region is determined based on the degree of regional salience difference.
[0066] The number of sub-regions of a class is denoted as the class quantity. The class degree of clustering is equal to (maximum number of adjacent pairs + 1) / class quantity. Two sub-regions of a class that share a common edge are denoted as an adjacent pair.
[0067] Regional prominence difference = Class I clustering degree / Preset Class I clustering degree; Adjusted feature region area = Unadjusted feature region area × Regional prominence difference; This invention sets a maximum feature region area, and the adjusted feature region area is at most twice the unadjusted feature region area.
[0068] Regarding the preset values for the number of categories and the clustering degree, the higher the user's requirements for the re-identification effect, the larger the value of the preset number of categories and the larger the value of the preset clustering degree. This invention uses the false alarm rate of re-identification to characterize the re-identification effect. In this embodiment of the invention, the preset number of categories is 10, with the unit being individuals, and the preset clustering degree is 0.6. Under this setting, the aim is to control the false alarm rate of re-identification to within one-thousandth to meet the user's requirements for the re-identification effect.
[0069] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A multimodal target detection and recognition method based on image processing, characterized in that, include: The image state is determined based on the richness of viewpoints and the stability of features, and the extraction strategy corresponding to the target video segment is adjusted from uniform extraction to compensated extraction based on the image state. When performing uniform sampling, the number of samples is adjusted based on the comparison between the viewpoint satisfaction and the preset viewpoint satisfaction. When performing compensated sampling, the compensated segment is determined based on the region parameter characterization value. Several reference images are obtained by extracting the target video segment according to the corresponding extraction parameters; For the extracted baseline image, the system determines whether to perform feature point optimization based on the feature anchor point representation value and spatial coverage. In feature point optimization, the system determines whether to adjust the point selection strategy from determining key points based on call frequency to determining key points based on feature evaluation value based on the evaluation balance. The sub-region category is determined as either a Class I or Class II sub-region based on the regional prominence. The distribution status is determined based on the Class I quantity and Class I clustering degree. Based on the distribution status, it is determined whether to adjust the area of the feature region.
2. The multimodal target detection and recognition method based on image processing according to claim 1, characterized in that, For image states where the view richness is greater than the preset view richness and the feature stability is greater than the preset feature stability, the extraction strategy is determined to be uniform extraction.
3. The multimodal target detection and recognition method based on image processing according to claim 2, characterized in that, During the uniform sampling process, if the view satisfaction is less than or equal to the preset view satisfaction, the sampling quantity is adjusted based on the view satisfaction coefficient. The viewpoint satisfaction coefficient is determined based on the viewpoint satisfaction degree and the preset viewpoint satisfaction degree. The number of samples extracted is positively correlated with the viewpoint satisfaction coefficient.
4. The multimodal target detection and recognition method based on image processing according to claim 1, characterized in that, For image states where the viewpoint richness is less than or equal to the preset viewpoint richness or the feature stability is less than or equal to the preset feature stability, the extraction strategy is determined to be to perform compensation extraction on the compensation segment. The compensation segment is a video segment whose regional parameter characterization value is greater than the preset regional parameter characterization value.
5. The image processing-based multimodal target detection and recognition method according to claim 2 or 4, characterized in that, For a reference image where the feature anchor point representation value is greater than the preset feature anchor point representation value and the spatial coverage is greater than the preset spatial coverage, feature point optimization is performed.
6. The multimodal target detection and recognition method based on image processing according to claim 5, characterized in that, When the evaluation balance is greater than the preset evaluation balance, the selection strategy for key points is to determine the key points based on the call frequency. When the evaluation balance is less than or equal to the preset evaluation balance, the selection strategy for the optimal location is to determine the key points based on the feature evaluation value.
7. The multimodal target detection and recognition method based on image processing according to claim 6, characterized in that, The feature evaluation value is determined based on the complex characterization value of the surrounding domain and the number of neighboring points; The feature evaluation value is positively correlated with the surrounding complexity representation value and the number of neighboring points.
8. The multimodal target detection and recognition method based on image processing according to claim 5, characterized in that, When executing the first optimization strategy, reference models with structural similarity greater than the preset structural similarity in the historical records are selected, and points with spatial parameter similarity greater than the preset spatial parameter similarity in the reference models are recorded as reference points. The number of reference points is recorded as the call frequency, and feature points with a call frequency greater than the preset call frequency are recorded as key points. The first preferred strategy is to determine key points based on the call frequency.
9. The multimodal target detection and recognition method based on image processing according to claim 1, characterized in that, The feature region is evenly divided into several sub-regions. When the region prominence is greater than the preset region prominence, the sub-region is determined to be a type of sub-region. Regions other than the first-class sub-regions are denoted as second-class sub-regions.
10. The multimodal target detection and recognition method based on image processing according to claim 9, characterized in that, For a distribution state where the quantity of one type is greater than the preset quantity of another type and the degree of clustering of one type is greater than the preset degree of clustering of another type, the area of the adjustment feature region is increased based on the degree of regional salience difference.