Image archiving method, device, electronic device and storage medium
By performing feature extraction and motion analysis of the combined image, a reasonable distance threshold is determined, which solves the problems of spatial and temporal contradictions and distance differences in image clustering, and improves the accuracy of the combined image clustering.
Patent Information
- Application Number
- CN202210120962.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-02-09
AI Technical Summary
The prior art is prone to time and space contradictions and distance differences in the process of image collection, which affects the accuracy of image collection.
By extracting the two images to be gathered, the movement mode and movement speed of the target object are determined, and a reasonable distance threshold is determined based on this information, so as to analyze and judge whether the spatial span of the target object is reasonable.
The accuracy of image clustering is improved, the spatial and temporal correlation of clustering results is ensured, and the clustering results obtained are more credible.
Smart Images

Figure CN114549882B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image archiving method, device, electronic device and storage medium. Background Art
[0002] With the continuous development of modern information technology, a new situation of intelligent construction has emerged. For example, by building an image recognition system covering public areas, the target objects within a certain coverage area can be captured and identified, and then relevant security services can be carried out according to the identity characteristics of the identified target objects. It plays an immeasurable role in security defense, smart access control and other fields.
[0003] Specifically, in order to ensure the accuracy of identifying the identity features of the target object, it is usually necessary to adopt Image Clustering (IC) technology to cluster the massive target images captured in the area covered by the system, and then form a feature file for the specified target object based on the image features of each target image belonging to the same image file.
[0004] For example, in the related art, a similarity comparison method is usually used to perform corresponding image clustering on the target image. Specifically, the image recognition system determines whether the image to be clustered is the target image corresponding to the target object based on the feature similarity between the image features of the image to be clustered and the corresponding file features in the feature file of the target object. If the feature similarity is greater than the preset similarity threshold, the image recognition system can determine that the current image to be clustered is the target image corresponding to the target object, and further, the target image is assigned to the specified image file, so as to update the feature file of the target object according to the image features of the corresponding target images in the image file. However, the above method still has the following defects:
[0005] 1. There is a contradiction between time and space.
[0006] In the related art, the method of determining feature similarity may easily lead to time-space contradictions between corresponding target images in the image archive. For example, when the number of images to be clustered is large and the system covers a large area, target images with high feature similarities but with similar capture time and long capture distance may be classified into the same image archive, resulting in a time-space contradiction in the image archive that the target object arrives at two distant capture locations in a short period of time, affecting the accuracy of image clustering.
[0007] 2. There is a distance difference.
[0008] In the related art, in order to avoid time and space contradictions, the corresponding two target images are usually determined based on the capture distance between the two corresponding capture images at adjacent capture times. For example, if the capture distance is less than a preset distance threshold, the corresponding two capture images are determined to be target images. However, in actual situations, due to the different movement modes of the target objects, the corresponding capture distances also vary greatly. For example, when the target object moves on foot, the corresponding capture distance is smaller, and when the target object moves in a vehicle or other manner, within the same time range, the capture distance increases significantly. In this case, it is difficult to determine an accurate distance threshold, which affects the accuracy of image aggregation. Summary of the invention
[0009] The embodiments of the present application provide an image archiving method, device, electronic device and storage medium for improving the accuracy of image archiving.
[0010] In a first aspect, an embodiment of the present application provides an image aggregation method, comprising:
[0011] Two to-be-assembled archive images containing a specified target object and acquisition information of each of the two to-be-assembled archive images are acquired, wherein the acquisition information at least includes: image acquisition time and image acquisition position of the corresponding to-be-assembled archive images.
[0012] Using a preset target recognition model, feature extraction is performed on the two images to be gathered to obtain the image features of each of the two images to be gathered, and based on the obtained image features, the corresponding target motion mode of the target object between the two images to be gathered is determined.
[0013] From a preset motion speed set, a target motion speed corresponding to the target motion mode is obtained, and based on the target motion speed, a corresponding motion duration interval is determined.
[0014] When it is determined that the corresponding acquisition time interval of the two to-be-grouped images belongs to the motion duration interval, the two to-be-grouped images are clustered.
[0015] In a second aspect, an embodiment of the present application provides an image aggregation device, comprising:
[0016] The acquisition module is used to acquire two to-be-assembled archive images containing the specified target object and the acquisition information of each of the two to-be-assembled archive images, wherein the acquisition information at least includes: the image acquisition time and image acquisition position of the corresponding to-be-assembled archive images.
[0017] The recognition module is used to use a preset target recognition model to extract features of the two images to be gathered, obtain the image features of each of the two images to be gathered, and determine the corresponding target movement mode of the target object between the two images to be gathered based on the obtained image features.
[0018] The determination module is used to obtain a target motion speed corresponding to the target motion mode from a preset motion speed set, and determine a corresponding motion duration interval based on the target motion speed.
[0019] The clustering module is used to cluster the two to-be-clustered images when it is determined that the corresponding acquisition time interval of the two to-be-clustered images belongs to the motion duration interval.
[0020] In an optional embodiment, before acquiring two to-be-assembled images containing the specified target object, the acquisition module is further configured to:
[0021] A training sample set is obtained, wherein one training sample includes: input information determined for at least one object feature of a target object and an object feature label.
[0022] The preset image recognition model is trained for multiple rounds of iterations using the training samples in the training sample set, and the target recognition model is output when the preset convergence conditions are met; wherein, in one round of iterative training, the following operations are performed:
[0023] An image recognition model is used to obtain corresponding image recognition results based on input information in training samples, and parameters of the image recognition model are adjusted based on residual values between the image recognition results and corresponding object feature labels.
[0024] In an optional embodiment, when acquiring two to-be-aggregated images containing a specified target object, the acquisition module is specifically configured to:
[0025] A sample detection image containing a target object is obtained, and sample acquisition information of the sample detection image is obtained, wherein the sample acquisition information at least includes: an image acquisition time and an image acquisition position of the sample detection image.
[0026] Based on the sample collection information, each candidate detection image corresponding to the sample detection image is obtained from a preset candidate image library, wherein the candidate image library contains the sample detection image.
[0027] The similarities between the sample detection image and each candidate detection image are determined respectively, and based on the obtained similarities, two corresponding candidate detection images are selected from each candidate detection image as the images to be clustered for the target object.
[0028] In an optional embodiment, when acquiring each candidate detection image corresponding to the sample detection image from a preset candidate image library based on the sample collection information, the acquisition module is used to:
[0029] Based on the image acquisition time of the sample detection image in the sample acquisition information and the preset detection duration, the corresponding candidate time range is determined.
[0030] Based on the image acquisition position of the sample detection image in the sample acquisition information and the preset detection distance, the corresponding candidate position range is determined.
[0031] From the preset candidate image library, candidate images that meet the candidate time range and the candidate position range are selected as corresponding candidate detection images.
[0032] In an optional embodiment, when determining the similarity between the sample detection image and each candidate detection image, the acquisition module is specifically used to:
[0033] For each candidate detection image, perform the following operations:
[0034] A candidate feature value of a candidate detection image is obtained, and a sample feature value of a sample detection image is obtained, wherein the sample feature value is used to characterize the object feature of the target object in the sample detection image.
[0035] A similarity comparison is performed on the candidate feature values and the sample feature values, and based on the comparison result, the similarity between the sample detection image and a candidate detection image is determined.
[0036] In an optional embodiment, when determining that the corresponding acquisition time intervals of the two to-be-grouped images belong to the motion duration interval, when clustering the two to-be-grouped images, the clustering module is specifically used to:
[0037] When determining that the corresponding acquisition time interval of two to-be-gathered images belongs to the motion duration interval, the corresponding target similarity is determined based on the image feature values of each of the two to-be-gathered images, wherein the target similarity represents the corresponding object feature similarity of the target object in the two to-be-gathered images.
[0038] When it is determined that the target similarity is not less than a preset similarity threshold, the two images to be clustered are clustered.
[0039] In an optional embodiment, after clustering the two to-be-clustered images, the clustering module is further used to:
[0040] Based on the image acquisition time of each of the two images to be aggregated, the corresponding aggregation time range is determined.
[0041] Based on the image acquisition positions of the two images to be gathered, the corresponding gathering position range is determined.
[0042] From the preset acquisition image library, various acquisition images that meet the gathering time range and the gathering position range are selected, and the similarities between the two to-be-gathered images and various acquisition images are determined respectively.
[0043] When it is determined that the similarity between the two to-be-clustered file images and each collected image is not less than the similarity threshold, the two to-be-clustered file images and each collected image are clustered.
[0044] According to a third aspect, an electronic device is provided, the electronic device comprising:
[0045] Memory, used to store computer instructions.
[0046] The processor is used to read computer instructions and execute the image archiving method as described in the first aspect.
[0047] According to a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the image archiving method as described in the first aspect.
[0048] The embodiments of the present application propose an image clustering method, device, electronic device and storage medium, which perform feature extraction on two images to be clustered, respectively determine the image features of the two images to be clustered, and analyze the corresponding target motion mode of the target object between the two images to be clustered based on the determined image features, thereby determining the corresponding distance threshold within the corresponding acquisition time range of the two images to be clustered based on the target motion speed corresponding to the target motion mode, and thereby analyzing and judging whether the spatial span of the target object between the two images to be clustered is reasonable based on the determined distance threshold.
[0049] On the one hand, the above method further judges the two associated images to be clustered based on the corresponding acquisition time interval and the corresponding acquisition position distance between the two images to be clustered, thereby ensuring the spatiotemporal correlation between the corresponding clustering results and further improving the accuracy of image clustering; on the other hand, the above method makes the determined distance threshold more consistent with the movement distance of the corresponding target object in the actual scene, so that the corresponding clustering results obtained are more credible and the accuracy of image clustering is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of an image aggregation scenario is provided for an embodiment of the present application;
[0051] Figure 2 A schematic diagram of another image aggregation scenario provided in an embodiment of the present application;
[0052] Figure 3 An image archiving system architecture diagram provided in an embodiment of the present application;
[0053] Figure 4 A flow chart of an image aggregation method provided in an embodiment of the present application;
[0054] Figure 5 A schematic diagram of a scenario for screening candidate detection images provided in an embodiment of the present application;
[0055] Figure 6 A schematic diagram of a method for determining target similarity provided in an embodiment of the present application;
[0056] Figure 7 A schematic diagram of an architecture for fusion information provided in an embodiment of the present application;
[0057] Figure 8 A logical schematic diagram of an image aggregation method provided in an embodiment of the present application;
[0058] Fig. 9 A logical schematic diagram of another image aggregation method provided in an embodiment of the present application;
[0059] Fig.10 A schematic diagram of an image aggregation device provided in an embodiment of the present application;
[0060] Fig.11 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0062] It should be noted that the embodiments described in this application are only part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] The design ideas of the embodiments of this application are as follows:
[0064] In the related art, whether to cluster the image to be inspected and the known image archives is determined based on the feature similarity of the target object in the image to be inspected and the known image archives. However, in actual situations, in the image recognition system, each image acquisition device independently acquires the image to be inspected in the corresponding area. Therefore, this method is prone to classifying target images with high feature similarity that are acquired very close in time and at a long acquisition distance into the same image archive.
[0065] For example, see Figure 1 As shown, image acquisition devices A, B, and C acquire images to be inspected (image 1 to be inspected, image 2 to be inspected, and image 3 to be inspected) for the target object at 10:00, 10:01, and 10:02, respectively. In a possible situation, if the feature similarities between the corresponding objects to be inspected and the target object in the above-mentioned images to be inspected are greater than the preset similarity threshold, the image acquisition system further clusters the above-mentioned images to be inspected.
[0066] However, the above clustering results, combined with their respective acquisition times, have time-space contradictions in actual conditions. For example, in a possible scenario where the inspected image 1 and the inspected image 2 are presented, if it is assumed that the inspected object 1 and the inspected object 2 are the same object, it is considered that the object is located in the area where the image acquisition device A is located at 10:00, and is located in the area where the image acquisition device B is located at 10:01, that is, the object has a spatial span from the corresponding area A to the area B in a relatively short period of time, while in actual scenarios, the image acquisition device A and the image acquisition device B may be deployed at a long distance, such as a deployment distance L = 5km. Therefore, it is further considered that the inspected object 1 has achieved a spatial span of an extremely long distance within an extremely short time range, which cannot be achieved by any mobile method in the existing method, indicating that the above assumption is inaccurate and the above clustering results are also inaccurate.
[0067] On the other hand, in the related art, by setting a possible distance threshold within the corresponding acquisition time range, it is judged whether the spatial span of the relevant object is reasonable. However, this method has the problem of limited judgment accuracy because the movement mode of the relevant object is not fixed.
[0068] For example, see Figure 2 As shown, image acquisition devices D and E acquire respective images to be inspected (image to be inspected a, image to be inspected b) for the target object at 10:00 and 10:10 respectively, and in the above-mentioned images to be inspected a and images to be inspected b, the feature similarities between the corresponding objects to be inspected and the target objects are both greater than the preset similarity threshold.
[0069] Based on the above analysis process, assuming that the corresponding objects a and b in the images a and b are considered to be the same object, it is considered that the object is located in the area where the image acquisition device D is located at 10:00, and in the area where the image acquisition device E is located at 10:10, that is, the object has a spatial span from the corresponding area D to the area E within a short time range (10 minutes). Further, assuming that the deployment distance between the image acquisition device D and the image acquisition device E is 1 km, and the set fixed distance threshold is also 800 m, then in a possible scenario, the image recognition system determines that the spatial span is unreasonable and does not cluster the images a and b. However, in actual conditions, the object a may use a moving method such as taking a car to achieve the spatial span (1 km) from the corresponding area D to the area E. In this case, the image recognition system will miss or misjudge due to the unreasonable determined distance threshold, which shows that the clustering result determined based on the above fixed distance threshold is also inaccurate.
[0070] In order to improve the accuracy of image clustering, the embodiments of the present application propose an image clustering method, device, electronic device and storage medium, which perform feature extraction on two images to be clustered, respectively determine the image features of the two images to be clustered, and analyze the corresponding target motion mode of the target object between the two images to be clustered based on the determined image features, thereby determining the corresponding distance threshold within the corresponding acquisition time range of the two images to be clustered based on the target motion speed corresponding to the target motion mode, and thereby analyzing and judging whether the spatial span of the target object between the two images to be clustered is reasonable based on the determined distance threshold.
[0071] On the one hand, the above method further judges the two associated images to be clustered based on the corresponding acquisition time interval and the corresponding acquisition position distance between the two images to be clustered, thereby ensuring the spatiotemporal correlation between the corresponding clustering results and further improving the accuracy of image clustering; on the other hand, the above method makes the determined distance threshold more consistent with the movement distance of the corresponding target object in the actual scene, so that the corresponding clustering results obtained are more credible and the accuracy of image clustering is higher.
[0072] See also Figure 3 As shown, it is a schematic diagram of the architecture of the image archiving system provided in the embodiment of the present application, and the system includes an image acquisition device 301, a terminal device 302 and a server 303. There are corresponding data transmission channels between the image acquisition device 301 and the terminal device 302, and between the terminal device 302 and the server 303. Optionally, the above data transmission channels are established by wireless communication or wired communication.
[0073] In an optional embodiment, the server 303 can access the network through cellular mobile communication technology to communicate with one or more terminal devices 302. The cellular mobile communication technology, for example, includes the fifth generation mobile communication (5th Generation Mobile Networks, 5G) technology.
[0074] In an optional embodiment, the server 303 may access the network via a short-range wireless communication method, thereby communicating with the terminal device 302. The short-range wireless communication method, for example, includes Wireless Fidelity (Wi-Fi) technology.
[0075] It should be noted that the above-mentioned server can be connected to one or more terminal devices, and each terminal device can be connected to one or more image acquisition devices. The embodiment of the present application does not limit the number of servers and the above-mentioned other devices. For the sake of ease of description, the embodiment of the present application takes one server as an example.
[0076] Furthermore, the image acquisition device 301 is an electronic device used to acquire images or record images, including a handheld image acquisition device with a wireless connection function, a head-mounted image acquisition device, and a fixed image acquisition device.
[0077] For example, in an optional embodiment, the image acquisition device 301 can be a camera, a camcorder, a digital still camera (DSC), a single-lens reflex camera (SLRC), other image acquisition devices with a camera function (mobile phones, tablet computers, etc.), a video capture card or a bayonet device, etc.
[0078] It is worth noting that in the embodiment of the present application, the image gathering method proposed in the embodiment of the present application is described by taking the bayonet device as an example. Then in the image gathering system, each bayonet device is used to respectively capture corresponding candidate images.
[0079] Furthermore, the terminal device 302 is a device that can provide voice and / or data connectivity to the user, including a handheld terminal device with wireless connection function, a vehicle-mounted terminal device, etc. Optionally, the terminal device can be: a mobile phone, a tablet computer, a laptop computer, a PDA, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal device in industrial control, a wireless terminal device in unmanned driving, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, or a wireless terminal device in a smart home, etc., so that each connected card port device can collect the corresponding candidate image set and upload it to the corresponding server.
[0080] Furthermore, the server 303 is an electronic device with relevant storage and computing functions. In the embodiment of the present application, the server 303 is used to receive the candidate image sets sent by each terminal device 302, and execute the image aggregation method proposed in the embodiment of the present application for each candidate image obtained.
[0081] Furthermore, in an embodiment of the present application, a trained target recognition model is deployed in the server 303, and the target recognition model is used to recognize and analyze image features of related images.
[0082] For example, in an optional embodiment, based on the Resnet network in deep learning, an identity mapping method is used to construct a corresponding target recognition model, and the target recognition model is obtained by iterative training of a preset image recognition model. In each iterative training process, the relevant parameters of the model need to be adjusted according to the residual value between the image recognition result of the image recognition model for the current input image and the original feature label of the input image. The residual value in each training process is shown as follows:
[0083]
[0084] Among them, x l and F represent the input image of the Lth residual value and its corresponding residual function respectively. Based on the above steps, when the model meets the preset convergence conditions, the corresponding target recognition model is output to perform feature extraction on the relevant image based on the obtained target recognition model.
[0085] See also Figure 4 As shown, based on the above system architecture, the embodiment of the present application proposes an image aggregation method, including:
[0086] S401: Acquire two to-be-assembled archive images containing a specified target object and acquisition information of each of the two to-be-assembled archive images, wherein the acquisition information at least includes: image acquisition time and image acquisition position of the corresponding to-be-assembled archive images.
[0087] Specifically, in the embodiment of the present application, the image to be gathered can be any candidate image captured by the camera device and containing a specified target object. Furthermore, the image to be gathered can be obtained by image comparison.
[0088] For example, in an optional embodiment, if image retrieval is required for a specified target object, a sample detection image containing the target object can be obtained. The sample detection image can be collected by a camera device or retrieved from a preset candidate image library. Based on the object features of the target object in the sample detection image, corresponding feature comparisons are performed one by one on each candidate image in the preset candidate image library. Based on the comparison results, each candidate detection image containing the target object other than the sample detection image is further obtained from the candidate image library, thereby further selecting two images to be clustered containing the target object from each candidate detection image.
[0089] Optionally, in order to reduce the amount of image comparison calculations and improve the efficiency of searching for candidate detection images, each candidate image to be compared can be screened based on the image acquisition time and image acquisition position of the sample detection image in a time-space correlation manner, that is, based on the position information of the target object in the sample detection image (such as the latitude and longitude information of the corresponding image acquisition device), the possible movement range of the target object is delineated within the corresponding time range.
[0090] For example, in an optional embodiment, a sample detection image containing a specified target object is obtained from image acquisition device a, and the image acquisition time of the sample detection image is 10:00:30, and its corresponding image acquisition position is (104.73°E, 31.49°N).
[0091] Based on the above image acquisition time and the preset detection duration, the corresponding candidate time range is determined, and based on the above image acquisition position and the preset detection distance, the corresponding candidate position range is determined to define the movement interval of the target object within the candidate time range.
[0092] For example, see Figure 5As shown, in an optional embodiment, assuming that the preset detection time is 30s, the corresponding candidate time range is determined according to the detection time, which is [10:00:00, 10:01:00], indicating an interval period for the sample detection image acquisition time; further, assuming that the preset detection distance is 1km, the corresponding candidate position range is expressed as within 1km of the image acquisition position (104.73°E, 31.49°N), then based on the above information, in the candidate image library containing the sample detection image, each candidate image to be compared is screened to determine each candidate detection image (candidate detection image 1-candidate detection image 3) that meets the corresponding time-space association conditions.
[0093] Furthermore, in order to ensure the accuracy of the acquired images to be aggregated, image recognition is used to compare the object features of the target object contained in the sample detection image and the object features of the target object contained in each candidate detection image to determine the corresponding similarities.
[0094] For example, based on the trained target recognition model, the candidate feature values of each candidate detection image are determined, wherein each candidate feature value is used to respectively characterize the object features (such as head and face features, clothing features, etc.) of the target object in the corresponding candidate detection image. Then, each candidate feature value and the sample feature value of the sample detection image are further compared similarly to obtain the similarity between each candidate detection image and the sample detection image, assuming that it is shown in the following Table 1:
[0095] Table 1
[0096] Candidate detection images Similarity Candidate detection image 1 99.1% Candidate detection image 2 99.0% Candidate detection image 3 97.6%
[0097] Further, based on the above-mentioned similarities, two to-be-grouped file images for the target object are selected. Optionally, two candidate detection images with higher similarities are selected as corresponding to-be-grouped file images. Based on the similarities shown in Table 1 above, candidate detection image 1 and candidate detection image 2 are respectively used as corresponding to-be-grouped file images. It is worth noting that the above-mentioned method is a preferred implementation mode of the present application. In actual situations, two specified candidate images can also be selected as corresponding to-be-grouped file images.
[0098] Further, it is assumed that the image acquisition time and image acquisition position of the two to-be-assembled files (to-be-assembled file image a, to-be-assembled file image b) are as shown in Table 2 below:
[0099] Table 2
[0100] Image to be aggregated Image acquisition time Image acquisition position Image to be aggregated 10:00:30 (104.73°E,31.49°N) Image to be aggregated b 10:02:00 (104.73°E,31.50°N)
[0101] S402: extracting features from the two images to be gathered to obtain the image features of each of the two images to be gathered, and determining the corresponding target motion mode of the target object between the two images to be gathered based on the obtained image features.
[0102] Furthermore, feature extraction is performed on the two images to be gathered, and the corresponding target motion mode of the target object between the two images to be gathered is determined based on the obtained image features. For example, if there are relevant entity image features such as "bus station" and "car station" in the two images to be gathered, the target motion mode is determined to be "bus" and "car".
[0103] For example, in the embodiment of the present application, it is assumed that the determined target motion mode is "car".
[0104] S403: Obtain a target movement speed corresponding to the target movement mode from a preset movement speed set, and determine a corresponding movement duration interval based on the target movement speed.
[0105] Specifically, according to the determined target motion mode, the target motion speed corresponding to the target motion mode is determined. For example, assuming that the average speed of the car is 40km / h, which is the corresponding target motion speed, then based on the target motion speed and the corresponding acquisition distance of 1.1km between the two images to be gathered, the corresponding motion duration interval is determined to be [0,99] (unit: seconds).
[0106] S404: When it is determined that the corresponding acquisition time interval of the two to-be-clustered images belongs to the motion duration interval, the two to-be-clustered images are clustered.
[0107] Specifically, when determining the corresponding acquisition time interval of the two files to be clustered, and when it belongs to the motion duration interval, the two files to be clustered are clustered. For example, based on the corresponding acquisition time interval of 90s of the two files to be clustered shown in Table 2 above, it is determined that the acquisition time interval is within the calculation duration interval, which means that the spatial span of the target object between the two files to be clustered is credible, and the two files to be clustered are clustered.
[0108] For further information, see Figure 6 As shown, in order to further improve the accuracy of the current image clustering, the corresponding target similarity is determined based on the image features of the two to-be-clustered images. The target similarity represents the similarity of the corresponding object features of the target object in the two to-be-clustered images.
[0109] For example, through the target recognition model, the object features of the target object in the two images to be clustered are extracted respectively, and the corresponding target similarity is determined to be 96%. Assuming that the preset similarity threshold is 95%, it is determined that the target similarity is not less than the preset similarity threshold, and the two images to be clustered are clustered.
[0110] Furthermore, after the first clustering of the target object is achieved based on the above two images to be clustered, the corresponding clustering time range and the corresponding clustering position range are determined based on the respective acquisition information of the above two images to be clustered, so as to further cluster each preset acquisition image.
[0111] For example, based on the image acquisition time and image acquisition position of each of the two to-be-aggregated archive images, it is determined that within the corresponding time range (10:00:30, 10:02:00), at longitude 104.73°E, each acquired image acquired by each image acquisition device whose latitude range belongs to (31.49°N, 31.50°N) can be regarded as the corresponding to-be-aggregated archive images. Assuming that the corresponding to-be-aggregated archive images are respectively acquisition images 1 to acquisition images 3, the to-be-aggregated archive images are sorted based on the similarity between each to-be-aggregated archive image and any current to-be-aggregated archive image, as shown in Table 3 below:
[0112] Table 3
[0113] Acquiring images Similarity Collect image 1 99.2% Collect image 2 98.3% Collect image 3 90.0%
[0114] Furthermore, when it is determined that the similarity between any of the two to-be-clustered archive images and each of the collected images is not less than the similarity threshold, the two to-be-clustered archive images and each of the collected images are clustered. For example, if the similarity threshold is 95%, the collected image 1 and the collected image 2 are clustered with the to-be-clustered archive image 1 and the to-be-clustered archive image 2 to form a target profile for the target object, as shown in Table 4 below:
[0115] Table 4
[0116] Poly file image Image 1 to be gathered Image 2 to be gathered Collect image 1 Collect image 2
[0117] Furthermore, in actual scenarios, in order to quickly extract the corresponding target motion mode of the target object, and based on the target motion mode and the corresponding acquisition information, the corresponding acquisition images are quickly extracted from the massive images, and the heterogeneous data fusion method is adopted to construct the corresponding clustering model.
[0118] For example, see Figure 7 As shown in the figure, the ERNIE technology is used to fuse the object features of the target object and the corresponding acquisition time and space information, where w 1 、w 2 ,…,w nThey represent the object features of the target object (e.g., appearance features), e 1 、e 2 ,…,e m are the collection time and space information (collection time and collection location) of the target object, and the two are fused as shown in the following formula:
[0119] h j =σ(W t ·w j +W e ·e k )
[0120] w j =σ(W t ·h j );e k =σ(W e ·h j )
[0121] Among them, W t and W e are the matrix parameters to be updated for the two models; w j is the object feature of the target object; k is the corresponding collection time and space information; h j It is the feature information formed by the fusion of the above data.
[0122] In this technical framework, the preset knowledge graph is used to capture the target movement mode of the target object, which is usually expressed as a triple, such as <target object, riding, bus stop>, <target object, driving, license plate number>, etc. According to the two stacking modules contained in the ERNIE framework, the object features (head and face, clothing information) and the corresponding spatiotemporal information (acquisition time, acquisition location, target movement mode) of the target object can be obtained respectively, so as to determine whether the spatial span of the target object is reasonable based on the fused feature information.
[0123] See also Figure 8As shown, it is a logical schematic diagram of an image gathering method provided by an embodiment of the present application, wherein the image acquisition time corresponding to the to-be-gathered image 1 containing the target object is 10:02:00, and the corresponding image acquisition position is (104.73°E, 31.50°N); the image acquisition time corresponding to the to-be-gathered image 2 is 10:00:30, and the corresponding image acquisition position is (104.73°E, 31.49°N); specifically, feature extraction is performed on the to-be-gathered image 1 and the to-be-gathered image 2, and the target motion mode corresponding to the target object is determined to be "taking a car"; based on the determined target motion mode and the above-mentioned acquisition information, it can be determined that the target object corresponding to the to-be-gathered image 1 and the to-be-gathered image 2 is within (10:00:30-10:02:00), and 1.1 km can be achieved by "taking a car". The spatial span is reasonable, thereby ensuring the spatiotemporal correlation of the archive images 1 and 2 to be clustered and avoiding spatiotemporal contradictions; further, based on the image features of the two archive images to be clustered, that is, the object features of the corresponding target objects, the corresponding target similarity is determined to ensure that there is an object feature association in the two archive images to be clustered, thereby ensuring the accuracy of the image clustering.
[0124] See also Fig. 9 As shown, it is a logical schematic diagram of another image clustering method provided by the embodiment of the present application based on the clustering results of the above-mentioned archive images 1 and 2 to be clustered. In order to improve the efficiency of image clustering, after the clustering of the archive images 1 and 2 to be clustered is completed, based on the above-mentioned image acquisition time and image acquisition position, the corresponding clustering time range and clustering space range are determined, and the corresponding individual acquired images (acquired image 1 to acquired image n+m) are determined; the above-mentioned individual acquired images satisfy the corresponding spatiotemporal correlation with the archive images 1 and 2 to be clustered, then further, based on the similarity between each acquired image and the archive image to be clustered, each acquired image whose similarity reaches a preset similarity threshold is clustered with the above-mentioned archive image to form a target image file for the target object.
[0125] See also Fig.10 As shown, the embodiment of the present application also provides an image clustering device, including an acquisition module 1001, an identification module 1002, a determination module 1003 and a clustering module 1004, wherein:
[0126] The acquisition module 1001 is used to acquire two to-be-assembled archive images containing a specified target object, and acquisition information of each of the two to-be-assembled archive images, wherein the acquisition information at least includes: image acquisition time and image acquisition position of the corresponding to-be-assembled archive images.
[0127] The recognition module 1002 is used to use a preset target recognition model to perform feature extraction on the two images to be gathered, obtain the image features of each of the two images to be gathered, and determine the corresponding target motion mode of the target object between the two images to be gathered based on the obtained image features.
[0128] The determination module 1003 is used to obtain a target motion speed corresponding to the target motion mode from a preset motion speed set, and determine a corresponding motion duration interval based on the target motion speed.
[0129] The clustering module 1004 is used to cluster the two images to be clustered when it is determined that the corresponding acquisition time interval of the two images to be clustered belongs to the motion duration interval.
[0130] In an optional embodiment, before acquiring two to-be-assembled images containing the specified target object, the acquisition module 1001 is further configured to:
[0131] A training sample set is obtained, wherein one training sample includes: input information determined for at least one object feature of a target object and an object feature label.
[0132] The preset image recognition model is trained for multiple rounds of iterations using the training samples in the training sample set, and the target recognition model is output when the preset convergence conditions are met; wherein, in one round of iterative training, the following operations are performed:
[0133] An image recognition model is used to obtain corresponding image recognition results based on input information in training samples, and parameters of the image recognition model are adjusted based on residual values between the image recognition results and corresponding object feature labels.
[0134] In an optional embodiment, when acquiring two to-be-aggregated images containing a specified target object, the acquisition module 1001 is specifically used to:
[0135] A sample detection image containing a target object is obtained, and sample acquisition information of the sample detection image is obtained, wherein the sample acquisition information at least includes: an image acquisition time and an image acquisition position of the sample detection image.
[0136] Based on the sample collection information, each candidate detection image corresponding to the sample detection image is obtained from a preset candidate image library, wherein the candidate image library contains the sample detection image.
[0137] The similarities between the sample detection image and each candidate detection image are determined respectively, and based on the obtained similarities, two corresponding candidate detection images are selected from each candidate detection image as the images to be clustered for the target object.
[0138] In an optional embodiment, when acquiring each candidate detection image corresponding to the sample detection image from a preset candidate image library based on the sample collection information, the acquisition module 1001 is used to:
[0139] Based on the image acquisition time of the sample detection image in the sample acquisition information and the preset detection duration, the corresponding candidate time range is determined.
[0140] Based on the image acquisition position of the sample detection image in the sample acquisition information and the preset detection distance, the corresponding candidate position range is determined.
[0141] From the preset candidate image library, candidate images that meet the candidate time range and the candidate position range are selected as corresponding candidate detection images.
[0142] In an optional embodiment, when determining the similarity between the sample detection image and each candidate detection image, the acquisition module 1001 is specifically used to:
[0143] For each candidate detection image, perform the following operations:
[0144] A candidate feature value of a candidate detection image is obtained, and a sample feature value of a sample detection image is obtained, wherein the sample feature value is used to characterize the object feature of the target object in the sample detection image.
[0145] A similarity comparison is performed on the candidate feature values and the sample feature values, and based on the comparison result, the similarity between the sample detection image and a candidate detection image is determined.
[0146] In an optional embodiment, when determining that the corresponding acquisition time intervals of the two to-be-grouped images belong to the motion duration interval, when clustering the two to-be-grouped images, the clustering module 1004 is used to:
[0147] When determining that the corresponding acquisition time interval of two to-be-gathered images belongs to the motion duration interval, the corresponding target similarity is determined based on the image feature values of each of the two to-be-gathered images, wherein the target similarity represents the corresponding object feature similarity of the target object in the two to-be-gathered images.
[0148] When it is determined that the target similarity is not less than a preset similarity threshold, the two images to be clustered are clustered.
[0149] In an optional embodiment, after clustering the two to-be-clustered images, the clustering module 1004 is further used to:
[0150] Based on the image acquisition time of each of the two images to be aggregated, the corresponding aggregation time range is determined.
[0151] Based on the image acquisition positions of the two images to be gathered, the corresponding gathering position range is determined.
[0152] From the preset acquisition image library, various acquisition images that meet the gathering time range and the gathering position range are selected, and the similarities between the two to-be-gathered images and various acquisition images are determined respectively.
[0153] When it is determined that the similarity between the two to-be-clustered file images and each collected image is not less than the similarity threshold, the two to-be-clustered file images and each collected image are clustered.
[0154] Based on the same inventive concept as the above-mentioned application embodiment, the present application embodiment also provides an electronic device, which can be used for image clustering. In one embodiment, the electronic device can be a server, or a terminal device or other electronic device. In this embodiment, the structure of the electronic device can be as follows: Fig.11 As shown, it includes a memory 1101 , a communication interface 1103 and one or more processors 1102 .
[0155] The memory 1101 is used to store computer programs executed by the processor 1102. The memory 1101 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0156] The memory 1101 may be a volatile memory, such as a random-access memory (RAM); the memory 1101 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or the memory 1101 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1101 may be a combination of the above memories.
[0157] The processor 1102 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 1102 is used to implement the above-mentioned image clustering method when calling the computer program stored in the memory 1101 .
[0158] The communication interface 1103 is used to communicate with terminal devices and other servers.
[0159] The specific connection medium between the memory 1101, the communication interface 1103 and the processor 1102 is not limited in the embodiment of the present application. Fig.11 In the embodiment, the memory 1101 and the processor 1102 are connected via a bus 1104. The bus 1104 is connected to the processor 1102 via a bus 1104. Fig.11 The connections between other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The bus 1104 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Fig.11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0160] According to one aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs any of the image clustering methods in the above-mentioned embodiments. The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, - but not limited to - an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0161] According to one aspect of the present application, the present application further provides a computer program product, which, when called by a computer, enables the computer to execute the method as described in the first aspect.
[0162] The embodiments of the present application propose an image clustering method, device, electronic device and storage medium, which perform feature extraction on two images to be clustered, respectively determine the image features of the two images to be clustered, and analyze the corresponding target motion mode of the target object between the two images to be clustered based on the determined image features, thereby determining the corresponding distance threshold within the corresponding acquisition time range of the two images to be clustered based on the target motion speed corresponding to the target motion mode, and thereby analyzing and judging whether the spatial span of the target object between the two images to be clustered is reasonable based on the determined distance threshold.
[0163] On the one hand, the above method further judges the two associated images to be clustered based on the corresponding acquisition time interval and the corresponding acquisition position distance between the two images to be clustered, thereby ensuring the spatiotemporal correlation between the corresponding clustering results and further improving the accuracy of image clustering; on the other hand, the above method makes the determined distance threshold more consistent with the movement distance of the corresponding target object in the actual scene, so that the corresponding clustering results obtained are more credible and the accuracy of image clustering is higher.
[0164] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0165] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0167] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. An image aggregation method, characterized in that: include: Acquire two to-be-assembled archive images containing the specified target object, and acquisition information of each of the two to-be-assembled archive images, wherein the acquisition information at least includes: image acquisition time and image acquisition position of the corresponding to-be-assembled archive images; Using a preset target recognition model, feature extraction is performed on the two to-be-gathered images to obtain image features of each of the two to-be-gathered images, and according to each of the obtained image features, a corresponding target motion mode of the target object between the two to-be-gathered images is determined; Acquire a target motion speed corresponding to the target motion mode from a preset motion speed set, and determine a corresponding motion duration interval based on the target motion speed; When it is determined that the corresponding acquisition time intervals of the two to-be-grouped images belong to the motion duration interval, clustering the two to-be-grouped images; Determining a corresponding aggregation time range based on the image acquisition time of each of the two images to be aggregated; Determining a corresponding focusing position range based on respective image acquisition positions of the two images to be focused; Selecting each collected image that meets the gathering time range and the gathering position range from a preset collection image library, and determining the similarity between the two to-be-gathered images and each collected image; When it is determined that the similarity between the two to-be-clustered archive images and the various collected images is not less than a similarity threshold, the two to-be-clustered archive images and the various collected images are clustered.
2. The method according to claim 1, characterized in that Before acquiring two to-be-assembled images containing the specified target object, the method further includes: Acquire a training sample set, wherein one training sample includes: input information determined for at least one object feature of the target object and an object feature label; The training samples in the training sample set are used to perform multiple rounds of iterative training on the preset image recognition model, and when the preset convergence conditions are met, the target recognition model is output; wherein, during one round of iterative training, the following operations are performed: The image recognition model is used to obtain corresponding image recognition results based on input information in training samples, and the parameters of the image recognition model are adjusted based on the residual value between the image recognition result and the corresponding object feature label.
3. The method according to claim 1 or 2, characterized in that The step of acquiring two to-be-assembled images containing a specified target object comprises: Acquire a sample detection image containing the target object, and acquire sample acquisition information of the sample detection image, wherein the sample acquisition information at least includes: image acquisition time and image acquisition position of the sample detection image; Based on the sample collection information, acquiring each candidate detection image corresponding to the sample detection image from a preset candidate image library, wherein the candidate image library contains the sample detection image; The similarities between the sample detection image and each of the candidate detection images are respectively determined, and based on the obtained similarities, two corresponding candidate detection images are selected from the each of the candidate detection images as the images to be gathered for the target object.
4. The method according to claim 3, characterized in that The acquiring, based on the sample collection information, each candidate detection image corresponding to the sample detection image from a preset candidate image library comprises: Based on the image acquisition time of the sample detection image in the sample acquisition information and the preset detection duration, determining a corresponding candidate time range; Based on the image acquisition position of the sample detection image in the sample acquisition information and in combination with a preset detection distance, determine a corresponding candidate position range; From a preset candidate image library, candidate images that meet the candidate time range and the candidate position range are selected as corresponding candidate detection images.
5. The method according to claim 4, characterized in that The determining of the similarity between the sample detection image and each of the candidate detection images includes: For each candidate detection image, the following operations are performed respectively: Acquire a candidate feature value of a candidate detection image, and acquire a sample feature value of the sample detection image, wherein the sample feature value is used to characterize an object feature of the target object in the sample detection image; A similarity comparison is performed on the candidate feature value and the sample feature value, and based on the comparison result, a similarity between the sample detection image and the one candidate detection image is determined.
6. The method according to claim 1 or 2, characterized in that: When determining that the corresponding acquisition time intervals of the two to-be-grouped images belong to the motion duration interval, clustering the two to-be-grouped images comprises: When determining that the corresponding acquisition time intervals of the two to-be-aggregated archive images belong to the motion duration interval, determining the corresponding target similarity based on the image feature values of the two to-be-aggregated archive images, wherein the target similarity represents the corresponding object feature similarity of the target object in the two to-be-aggregated archive images respectively; When it is determined that the target similarity is not less than a preset similarity threshold, the two images to be clustered are clustered.
7. An image aggregation device, characterized in that: include: An acquisition module, used to acquire two to-be-assembled archive images containing a specified target object, and acquisition information of each of the two to-be-assembled archive images, wherein the acquisition information at least includes: image acquisition time and image acquisition position of the corresponding to-be-assembled archive images; A recognition module, used to extract features of the two to-be-gathered images using a preset target recognition model, obtain image features of each of the two to-be-gathered images, and determine a corresponding target motion mode of the target object between the two to-be-gathered images based on the obtained image features; A determination module, configured to obtain a target motion speed corresponding to the target motion mode from a preset motion speed set, and determine a corresponding motion duration interval based on the target motion speed; A clustering module is used to cluster the two images to be clustered when it is determined that the corresponding acquisition time intervals of the two images to be clustered belong to the motion duration interval; determine the corresponding clustering time range based on the image acquisition times of the two images to be clustered; determine the corresponding clustering position range based on the image acquisition positions of the two images to be clustered; select each acquisition image that meets the clustering time range and the clustering position range from a preset acquisition image library, and respectively determine the similarity between the two images to be clustered and each acquisition image; when it is determined that the similarity between the two images to be clustered and each acquisition image is not less than a similarity threshold, cluster the two images to be clustered with each acquisition image.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the image archiving method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
High-resolution remote sensing image classification method based on residual network and transfer learning
CN112836614A
Personnel archiving method and device and electronic equipment
CN113807127A