Target counting method and trolley management method

Through the methods of image segmentation and spectrum analysis, the problem of inaccurate trolley counting in complex backgrounds is solved, and higher counting accuracy and stability are achieved.

CN120047878AActive Publication Date: 2025-05-27ZHEJIANG AIRPORT DIGITAL TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510526237.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In complex contexts, it is difficult for the prior art to accurately count trolleys, especially in scenes where light changes and noise interference are large, it is easy to misidentify other objects as trolleys.

Method used

The image segmentation model is used to segment and identify the images to be processed, and the target area images are obtained, and then the spectral analysis is performed on these area images, and the spectral map is used to perform statistics on the target objects to improve the accuracy of the count.

Benefits of technology

Through spectrum analysis, the unique frequency characteristics of the target object are captured, the influence of light changes and noise interference is reduced, the accuracy of target counting is improved, and the counting needs in complex contexts are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047878A_ABST
    Figure CN120047878A_ABST
Patent Text Reader

Abstract

The invention provides a target counting method and a handcart management method, and the method comprises the steps: carrying out the image segmentation of a to-be-processed image through employing an image segmentation model, and obtaining a plurality of region images; for each region image in the plurality of region images, carrying out image identification on the region image to obtain a target region image with an identification result as a target object; performing spectral analysis on the target area image to obtain a spectrogram; and performing statistics on the target objects in the target area image according to the spectrogram to obtain the number of the target objects in the to-be-processed image. In the implementation process of the scheme, image segmentation and image recognition are performed on the to-be-processed image by using the model, and then spectral analysis is performed on the recognized target area image, so that the obtained spectrogram can capture the unique frequency characteristics of the target object; the probability that densely-arranged target objects are detected mistakenly can be effectively reduced, and therefore the accuracy of target counting is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer vision and image processing. Specifically, it relates to a target counting method and a trolley management method. Background Art

[0002] Currently, in an airport terminal, there is often a shortage of trolleys. Especially when a large aircraft arrives, the trolleys will be quickly used up. Usually, when counting trolleys in a video, visual image features such as shape, color, or texture are extracted from the video frames, and the number of trolleys is counted based on these visual image features. However, in some scenarios with complex backgrounds such as changes in lighting and noise interference from other objects, there may be objects in the same video frame that are similar in shape, color, or texture to the trolleys, such as baby carriages or suitcases, etc., resulting in an overcount of trolleys. Therefore, the current accuracy of target counting is relatively low and it is difficult to meet the requirements of scenarios with complex backgrounds. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a target counting method and a trolley management method to improve the problem of relatively low accuracy of target counting.

[0004] The embodiments of this application provide a target counting method, including: using an image segmentation model to perform image segmentation on the image to be processed to obtain multiple regional images; for each of the multiple regional images, performing image recognition on the regional image to obtain a target regional image whose recognition result is a target object; performing spectral analysis on the target regional image to obtain a spectrogram; and counting the target objects in the target regional image according to the spectrogram to obtain the number of target objects in the image to be processed. In the implementation process of the above solution, by first using a model to perform image segmentation and image recognition on the image to be processed, and then performing spectral analysis on the recognized target regional image, the obtained spectrogram can capture the unique frequency characteristics of the target object. These characteristics are often more abstract but more discriminative than traditional shape, color, or texture features. Therefore, using the spectrogram can understand the image content from the perspective of different frequency components. Compared with directly operating in the spatial domain, frequency domain information is usually more resistant to factors such as lighting changes and noise interference, providing a more stable and reliable feature representation. Further, spectral analysis can help filter out some high-frequency noises caused by environmental factors such as reflections and shadows. After removing these interfering factors of high-frequency noises, the probability of misdetection of densely arranged target objects can be effectively reduced. Therefore, the accuracy of target counting is effectively improved, thus meeting the requirements of scenarios with complex backgrounds.

[0005] Optionally, in the embodiments of the present application, the image segmentation model includes: a prompt encoder, an image encoder, and an image decoder; using the image segmentation model to perform image segmentation on the image to be processed includes: obtaining the segmentation prompt information of the image to be processed; encoding the segmentation prompt information using the prompt encoder to obtain a segmentation prompt vector; encoding the image to be processed using the image encoder to obtain an embedded image vector; decoding the segmentation prompt vector and the embedded image vector using the image decoder to obtain a contour mask map; using the contour mask map to perform region segmentation on the image to be processed to obtain multiple region images. In the implementation process of the above solution, the image to be processed is segmented through the segmentation prompt vector. Since the segmentation prompt vector provides additional context information, the contour mask map generated by the image decoder is more accurate, reducing the possibility of misidentifying the background or other similar objects as the target object. Therefore, each object can be effectively separated from the complex background, reducing the influence of background noise on subsequent processing steps. Even in the case of large changes in lighting, occlusion, or background interference, the accuracy of target counting can be effectively improved.

[0006] Optionally, in the embodiments of the present application, the image to be processed is the current video image frame, and the segmentation prompt information is the prompt detection points in the current video image frame; obtaining the segmentation prompt information of the image to be processed includes: determining the randomly generated detection points as the prompt detection points; or, determining the pre-set detection points as the prompt detection points; or, in the case where the target object in the current video image is occluded, determining the detection points used in the historical video image frame as the prompt detection points. In the implementation process of the above solution, by flexibly selecting the prompt points according to different scenario requirements (random generation, pre-setting, or using the detection points in the historical frame), the system can adapt to a variety of complex environments. When the environment changes suddenly (such as changes in lighting conditions, movement of background objects, etc.), the randomly generated detection points can help the system quickly adapt to the new situation and avoid failure caused by a fixed pattern, thereby enhancing the adaptability and flexibility of image segmentation. Further, when the target object is occluded, using the effective detection points in the previous historical frame as a prompt can maintain the continuous tracking and recognition of the target, reducing the problem of target loss caused by short-term occlusion. By combining multiple prompt point strategies, the system can select the optimal detection point combination in different situations, thereby enhancing the accuracy of image segmentation.

[0007] Optionally, in the embodiments of the present application, image recognition of the regional image includes: dividing the regional image into multiple image patches; inputting the multiple image patches into a vision transformer model so that the vision transformer model can recognize the category of the regional image; if the category of the regional image is the category to which the target object belongs, then determining the regional image as the target regional image. In the implementation process of the above solution, by dividing the regional image into multiple image patches, each image patch can be analyzed independently, which enables the model to better capture the local detail features of the target object. The vision transformer model can not only process the information of a single image patch, but also effectively fuse the global context information from different image patches through the self-attention mechanism. Even if some local features are not obvious, accurate classification can be achieved by combining the information of other parts, thereby improving the classification accuracy.

[0008] Optionally, in the embodiments of the present application, the vision transformer model includes: a fully connected network, a transformation encoder, and a multi-layer perceptron network; after inputting the multiple image patches into the vision transformer model, it further includes: using the fully connected network to perform embedding processing on the multiple image patches to obtain embedded patch vectors; using the transformation encoder to perform transformation processing on the embedded patch vectors to obtain transformed intermediate feature vectors; using the multi-layer perceptron to perform category recognition on the transformed intermediate feature vectors to obtain the category of the regional image. In the implementation process of the above solution, by performing embedding processing on each image patch through the fully connected network, the original pixel information can be transformed into a higher-level feature representation, enabling the model to capture the fine-grained features in the image, thereby increasing the accuracy of regional image category recognition. In addition, since the transformation encoder can effectively integrate global and local information, it can perform better when distinguishing similar objects. For example, in an airport environment, the subtle differences between a trolley and other similar items (such as a baby carriage or a suitcase) can be accurately captured in this way, thereby further increasing the accuracy of regional image category recognition.

[0009] Optionally, in the embodiments of the present application, the spectrogram includes: a frequency-domain contour map; performing spectral analysis on the target area image includes: performing grayscale processing on the target area image to obtain a two-dimensional grayscale image; performing a transform on the two-dimensional grayscale image to obtain a two-dimensional frequency-domain representation matrix; calculating a frequency-domain contour map based on the two-dimensional frequency-domain representation matrix. In the implementation process of the above solution, by performing spectral analysis on the target area image, the frequency feature differences between different target objects can be identified, which helps to more accurately classify and count similar but different objects (such as shopping carts and baby carriages). Compared with relying only on spatial-domain information, this method provides an additional dimension to describe the target features, thereby identifying subtle structures and patterns that are difficult to detect in the spatial domain and effectively improving the counting accuracy. In addition, since spectral analysis mainly focuses on the frequency components in the image rather than the brightness values, it is insensitive to illumination changes, which means that this method can maintain high accuracy even for images taken under different illumination conditions.

[0010] Optionally, in the embodiments of the present application, counting the target objects in the target area image according to the spectrogram includes: performing frequency-domain signal analysis on the spectrogram to obtain multiple signal frequency values; screening out the signal frequency value with the largest signal intensity from the multiple signal frequency values; determining the signal frequency value with the largest signal intensity as the quantity of the target object in the image to be processed. In the implementation process of the above solution, by performing frequency-domain signal analysis on the spectrogram and screening out the signal frequency value with the largest signal intensity from the obtained multiple signal frequency values, it is possible to discover that there is a certain correlation between some specific frequency components and the presence or quantity of the target object. For example, structures such as the wheels and frames of a shopping cart may generate specific frequency patterns. Therefore, this frequency-domain analysis method can explore features that are difficult to capture in the spatial domain. For example, shopping carts of different sizes may exhibit different frequency distribution characteristics in the spectrogram, thereby improving the counting accuracy in complex backgrounds or in the presence of other similar objects.

[0011] The embodiment of the present application also provides a trolley management method, including: using an image segmentation model to perform image segmentation on the image to be processed to obtain multiple regional images; for each of the multiple regional images, performing image recognition on the regional image to obtain a target regional image whose recognition result is a trolley; performing spectral analysis on the target regional image to obtain a spectrogram; counting the trolleys in the target regional image according to the spectrogram to obtain the number of trolleys in the image to be processed; if the number of trolleys in the image to be processed is less than a preset threshold, performing supplementary scheduling and / or quantity warning on the trolleys. In the implementation process of the above solution, when it is detected that the number of trolleys is lower than the preset threshold, the system can trigger an automated supplementary scheduling process to notify the staff to replenish the trolleys in time, avoiding service interruption or inconvenience to passengers caused by insufficient trolleys. Especially in key areas or during peak periods, the instant warning function can be used to remind the management to take actions, thereby effectively improving the operation efficiency of the airport terminal building.

[0012] The embodiment of the present application also provides a target counting device, including: an image region segmentation module, configured to use an image segmentation model to perform image segmentation on the image to be processed to obtain multiple regional images; a regional image recognition module, configured to perform image recognition on each of the multiple regional images to obtain a target regional image whose recognition result is a target object; an image spectral analysis module, configured to perform spectral analysis on the target regional image to obtain a spectrogram; a target object counting module, configured to count the target objects in the target regional image according to the spectrogram to obtain the number of target objects in the image to be processed.

[0013] Optionally, in the embodiment of the present application, the image segmentation model includes: a prompt encoder, an image encoder, and an image decoder; the image region segmentation module includes: a segmentation prompt information acquisition sub-module, configured to acquire the segmentation prompt information of the image to be processed; a segmentation prompt vector acquisition sub-module, configured to encode the segmentation prompt information using the prompt encoder to obtain a segmentation prompt vector; an embedded image vector encoding sub-module, configured to encode the image to be processed using the image encoder to obtain an embedded image vector; an image vector decoding sub-module, configured to decode the segmentation prompt vector and the embedded image vector using the image decoder to obtain a contour mask map; an image region segmentation sub-module, configured to perform region segmentation on the image to be processed using the contour mask map to obtain multiple regional images.

[0014] Optionally, in the embodiments of the present application, the image to be processed is the current video image frame, and the segmentation prompt information is the prompt detection point in the current video image frame; the prompt information acquisition sub-module includes: a first detection point determination unit for determining a randomly generated detection point as the prompt detection point; or, a second detection point determination unit for determining a pre-set detection point as the prompt detection point; or, a third detection point determination unit for determining the detection point used in the historical video image frame as the prompt detection point when the target object in the current video image is occluded.

[0015] Optionally, in the embodiments of the present application, the region image recognition module includes: a region image division sub-module for dividing the region image into a plurality of image blocks; an image category recognition sub-module for inputting the plurality of image blocks into a vision transformer model so that the vision transformer model recognizes the category of the region image; a target image recognition sub-module for determining the region image as the target region image if the category of the region image is the category to which the target object belongs.

[0016] Optionally, in the embodiments of the present application, the vision transformer model includes: a fully connected network, a transformation encoder, and a multi-layer perceptron network; the region image recognition module further includes: a patch vector embedding sub-module for performing embedding processing on the plurality of image blocks using the fully connected network to obtain embedded patch vectors; a vector transformation processing sub-module for performing transformation processing on the embedded patch vectors using the transformation encoder to obtain transformed intermediate feature vectors; a vector category recognition sub-module for performing category recognition on the transformed intermediate feature vectors using the multi-layer perceptron to obtain the category of the region image.

[0017] Optionally, in the embodiments of the present application, the spectrogram includes: a frequency domain contour map; the image spectrum analysis module includes: an image grayscale processing sub-module for performing grayscale processing on the target region image to obtain a two-dimensional grayscale image; an image transformation sub-module for performing transformation on the two-dimensional grayscale image to obtain a two-dimensional frequency domain representation matrix; a contour map obtaining sub-module for calculating the frequency domain contour map according to the two-dimensional frequency domain representation matrix.

[0018] Optionally, in the embodiments of the present application, the target object statistics module includes: a frequency domain signal analysis sub-module for performing frequency domain signal analysis on the spectrogram to obtain a plurality of signal frequency values; a signal frequency screening sub-module for screening out the signal frequency value with the maximum signal intensity from the plurality of signal frequency values; a target quantity determination sub-module for determining the signal frequency value with the maximum signal intensity as the quantity of the target object in the image to be processed.

[0019] The embodiment of the present application further provides a trolley management device, including: an image region segmentation module, configured to perform image segmentation on a to-be-processed image using an image segmentation model to obtain multiple region images; a region image recognition module, configured to perform image recognition on each of the multiple region images to obtain a target region image whose recognition result is a trolley; an image spectrum analysis module, configured to perform spectrum analysis on the target region image to obtain a spectrogram; a target object statistics module, configured to count the trolleys in the target region image according to the spectrogram to obtain the number of trolleys in the to-be-processed image; a trolley scheduling and warning module, configured to, if the number of trolleys in the to-be-processed image is less than a preset threshold, perform supplementary scheduling and / or quantity warning on the trolleys.

[0020] The embodiment of the present application further provides an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the machine-readable instructions are run by the processor, the above-described method is executed.

[0021] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the above-described method is executed.

[0022] The embodiment of the present application further provides a computer program product, including: a computer program or computer instructions, and when the computer program or computer instructions are run by a processor, the above-described method is executed. Description of the Drawings

[0023] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 A flowchart showing the target counting method provided by the embodiment of the present application; Figure 2 A schematic diagram showing the overall processing process of the to-be-processed image provided by the embodiment of the present application; Figure 3 A schematic diagram showing the network structure of the image segmentation model provided by the embodiment of the present application; Figure 4 A schematic diagram showing the network structure of the vision transformer model provided by the embodiment of the present application; Figure 5 One of the schematic diagrams showing the frequency domain contour map provided by the embodiment of the present application; Figure 6 Schematic flowchart of the trolley management method provided by the embodiments of the present application shown; Figure 7 Schematic flowchart of the target counting device provided by the embodiments of the present application shown. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the embodiments of the present application are only for the purpose of illustration and description, and are not used to limit the protection scope of the embodiments of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the embodiments of the present application illustrate the operations implemented according to some embodiments of the embodiments of the present application. It should be understood that the operations in the flowchart may not be implemented in sequence, and the steps without logical context may be reversed or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the embodiments of the present application.

[0026] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed embodiments of the present application, but merely represents the selected embodiments of the present application.

[0027] It can be understood that "first" and "second" in the embodiments of the present application are used to distinguish similar objects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily mean different. In the description of the embodiments of the present application, the term "and / or" is only an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. The term "plural" refers to two or more (including two). Similarly, "multiple groups" refers to two or more groups (including two groups).

[0028] It should be noted that the target counting method provided by the embodiments of the present application can be executed by an electronic device, where the electronic device refers to a device terminal or a server with the function of executing a computer program. The device terminal includes, for example, a smart phone, a personal computer, a tablet computer, a personal digital assistant, or a mobile Internet device, etc. The server refers to a device that provides computing services through a network. The server includes, for example, an x86 server and a non-x86 server, and the non-x86 server includes mainframes, minicomputers, and UNIX servers.

[0029] In the related art, the trolley counting is usually based on the visual image features extracted from video frames. For example, a convolutional neural network is used to extract visual image features such as shape, color, or texture, and the number of trolleys is counted based on these visual image features. However, in the specific practice process, it is found that this method has many limitations. The trolley needs to be placed in a place with good lighting conditions and field of view conditions. For example, the trolley is placed in an area close to the window and can be directly photographed by the camera, and the camera cannot be too far away from the trolley. Therefore, the solution for counting trolleys based on the visual image features extracted from video frames is only applicable to the case where the trolleys are placed in fixed positions and cannot handle the placement of trolleys at non-fixed points, which increases the management difficulty and reduces the flexibility of the system.

[0030] In order to achieve a better detection effect, the existing solutions require special planning and adjustment of the position and angle of the camera, and also require the management personnel to place the trolleys according to specific rules to ensure that they are within the detectable range, which puts additional requirements on the daily operation. However, the solution based on visual image features is very sensitive to light changes. Factors such as the light difference at different time periods or weather conditions, the placement angle of the trolley, and the distance between the camera and the trolley may directly lead to a decrease in the counting accuracy. Moreover, in some complex background environments, there may be other objects (such as baby carriages, suitcases) similar to the trolley, and these similar objects are likely to be misidentified as trolleys when they appear in the same video frame as the trolley, resulting in inaccurate statistical results. Therefore, these light change interferences and the noise interferences of other objects also reduce the accuracy of trolley counting from images.

[0031] For the above problems, please refer to Figure 1Flow schematic diagram of the target counting method provided by the embodiment of the present application shown; the main idea of the target counting method is to first perform image segmentation and image recognition through a model, and then perform spectral analysis on the recognized target region image. Since the obtained spectrogram can capture the unique frequency characteristics of the target object, these characteristics are often more abstract but more discriminative than traditional shape, color, or texture features. Therefore, using the spectrogram can understand the image content from the perspective of different frequency components. Compared with directly operating in the spatial domain, frequency domain information is usually more resistant to factors such as illumination changes and noise interference. In addition, spectral analysis can help filter out some high-frequency noises caused by environmental factors such as reflections and shadows. After removing these interfering factors of high-frequency noises, the probability of misdetection of densely arranged target objects can be effectively reduced, thus effectively improving the accuracy of target counting. The implementation manner of the above target counting method may include: Step S110: Use an image segmentation model to perform image segmentation on the image to be processed, and obtain multiple region images.

[0032] The image to be processed refers to the image for which target counting needs to be achieved. It can be the original video image frame extracted from the video stream captured by the surveillance camera as the image to be processed, or a single image obtained during the debugging and testing process as the above image to be processed. Among them, the above video stream can be transmitted through the Real Time Streaming Protocol (RTSP).

[0033] The image segmentation model (Image Segmentation Model) refers to a neural network model used to segment images. This model can divide a complete image into multiple small regions with different semantic meanings. Here, the image segmentation model can generally be an algorithm model based on machine learning or deep learning. For example: the Segment-Anything Model (SAM) or the Segment-Anything2 model. Both of these models can separate the trolley and other objects (such as the screen, luggage, or turntable, etc.) from the background, thereby generating a region image containing the trolley and region images of other objects. In addition, the above image segmentation model can also adopt models such as U-Net or Mask R-CNN, etc., so as to effectively separate the target object from the complex background and reduce the influence of background noise on subsequent processing steps. The above region image is a small-range image segment obtained through image segmentation. Each segment represents a part of the original image content and usually corresponds to a specific object.

[0034] Step S120: For each of the multiple region images, perform image recognition on the region image to obtain a target region image whose recognition result is a target object.

[0035] The target object refers to the specific item to be counted in the current application scenario. For example, in the scenario of an airport terminal, the target object can be a trolley.

[0036] The target area image is the area image that is confirmed to contain the target object (such as a trolley) after image recognition. For example, after image recognition confirmation, those area images that truly contain trolleys are separately selected for subsequent analysis.

[0037] Step S130: Perform spectral analysis on the target area image to obtain a spectrogram.

[0038] Spectral analysis is a transform analysis algorithm that converts an image from the spatial domain to the frequency domain to reveal the hidden periodic structures and texture features in the image. The analysis algorithms that can be used include Fourier transform or wavelet transform. It can be understood that, compared with simple image recognition, spectral analysis can capture more subtle structural information, which helps to more accurately identify and distinguish similar objects. Therefore, the above spectral analysis can reveal the hidden structures and periodic features in the image. For example, there may be significant frequency components in the wheels, frames, etc. of the trolley, and these features can be clearly shown through the spectrogram.

[0039] The spectrogram is an image representation generated by spectral analysis, showing the intensity distribution of the original image at different frequencies. The spectrogram here can be in the form of a two-dimensional spectrogram or a frequency domain contour map, etc.

[0040] Step S140: Count the target objects in the target area image according to the spectrogram to obtain the number of target objects in the image to be processed.

[0041] Please refer to Figure 2 The overall processing process schematic diagram of the image to be processed provided by the embodiment of the present application shown; it can be understood that the above solution uses an image segmentation model to perform image segmentation, image recognition, and spectral analysis on the image to be processed. Therefore, the cameras already installed and deployed in the airport terminal can be directly utilized, and there is no need to adjust the position and angle of the cameras to achieve better detection effects, thus making full use of the existing camera resources in the terminal. It is not only applicable to the placement of trolleys at fixed points but also can handle randomly distributed trolleys in a dynamic environment. The number of the above target objects in the image to be processed is the final output result of the target counting method, indicating how many target objects (such as trolleys) there are in the entire original input image.

[0042] In the implementation process of the above solution, the image to be processed is first segmented and recognized using a model, and then the spectrum analysis is performed on the recognized target region image. The obtained spectrogram can capture the unique frequency characteristics of the target object. These characteristics are often more abstract but more discriminative than traditional shape, color, or texture features. Therefore, using the spectrogram can understand the image content from the perspective of different frequency components. Compared with directly operating in the spatial domain, frequency domain information is usually more resistant to factors such as illumination changes and noise interference, providing a more stable and reliable feature representation. Further, spectrum analysis can help filter out some high-frequency noises caused by environmental factors such as reflections and shadows. After removing these interference factors of high-frequency noises, the probability of misdetection of densely arranged target objects can be effectively reduced. Therefore, the accuracy of target counting is effectively improved, thus meeting the scene requirements in complex backgrounds.

[0043] As an alternative implementation of the above step S110, the above image segmentation model may include: a prompt encoder, an image encoder, and an image decoder; the implementation of using the image segmentation model to segment the image to be processed may include: Step S111: Obtain the segmentation prompt information of the image to be processed.

[0044] Segmentation Prompt Information refers to the prompt information used as the image segmentation model. There are many types of such segmentation prompt information in theory, including: detection points (points), prompt boxes (box), text, and / or object masks (mask), etc.

[0045] Step S112: Encode the segmentation prompt information using the prompt encoder to obtain a segmentation prompt vector.

[0046] Please refer to Figure 3Schematic diagram of the network structure of the image segmentation model provided by the embodiment of the present application shown; The implementation manner of the above step S112 is, for example: object masks (masks), detection points (points), bounding boxes (boxes), and / or text, etc. can be used as segmentation prompt information. The above object masks, detection points, and bounding boxes can guide the image segmentation model to obtain information about the image regions where various objects need to be segmented. Although text cannot guide the image regions where objects are located, theoretically it can also be input to the image segmentation model as semantically related prompt information. Therefore, the segmentation prompt information can be input to the prompt encoder of the image segmentation model (such as the SAM model) so that the prompt encoder encodes the segmentation prompt information to obtain segmentation prompt vectors (Prompt tokens). It can be understood that if the image segmentation model adopts the SAM model, since the SAM model is a pre-trained model, it can be directly used without re-training the SAM model, which can effectively save the time of model training and improve the development efficiency of the entire target counting function.

[0047] Step S113: Use the image encoder to encode the image to be processed to obtain an embedded image vector.

[0048] The implementation manner of the above step S113 is, for example: input the image to be processed to the image encoder (Image Encoder) of the image segmentation model (such as the SAM model) so that the image encoder encodes the image to be processed, thereby obtaining the embedded image (Image Embedding) vector output by the image encoder. Here, the embedded image vector is actually a feature matrix vector that is dynamically updated during the calculation process.

[0049] Step S114: Use the image decoder to decode the segmentation prompt vector and the embedded image vector to obtain a contour mask image.

[0050] The implementation manner of the above step S114 is, for example: input the segmentation prompt vector and the embedded image vector together to the image encoder of the image segmentation model (such as the SAM model) so that the image decoder decodes the segmentation prompt vector and the embedded image vector to obtain a contour mask (Contour Mask) image. Among them, the contour mask image is a binary image or a grayscale image, where the region where the target object is located is marked as white (or high brightness), while the background and other non-target regions are marked as black (or low brightness).

[0051] Optionally, while the image decoder outputs a contour mask map after decoding the segmentation prompt vector and the embedded image vector, the Intersection over Union (IOU) confidence can also be output. That is to say, after the image decoder decodes the segmentation prompt vector and the embedded image vector, a contour mask map and the IOU confidence can be output simultaneously.

[0052] Step S115: Use the contour mask map to perform region segmentation on the image to be processed, and obtain multiple region images.

[0053] An implementation manner of the above step S115 is as follows: For example, if the image decoder outputs a contour mask map and the IOU confidence, the contour mask map can be used to perform region segmentation on the image to be processed only when the IOU confidence is greater than a preset confidence threshold. If the image decoder does not output the IOU confidence, the contour mask map is directly used to perform region segmentation on the image to be processed, that is, the contour mask map is used as a mask (Mask), and the contour mask map is multiplied pixel by pixel with the image to be processed. At the positions in the contour mask map marked as at least one target region, the pixels in the original image will be retained; the remaining parts are set to zero (i.e., black), thereby achieving the effect of region segmentation on the image to be processed, and finally obtaining multiple region images. It can be understood that if there are multiple target objects in the image, the contour mask map contains multiple non-connected regions, indicating that each region corresponds to an independent target object.

[0054] As an optional implementation manner of the above step S111, the above image to be processed can be the current video image frame, and the above segmentation prompt information can be the prompt detection points in the current video image frame; the implementation manner of obtaining the segmentation prompt information of the image to be processed can include: Step S111a: Determine the randomly generated detection points as the prompt detection points.

[0055] An implementation manner of the above step S111a is as follows: For example, when the image segmentation model has no prior knowledge of the image, such as video frames within a certain duration at the beginning of the video (e.g., the first video frame), an executable program written in a preset programming language can be used to randomly generate some detection points for each video image frame, and these randomly generated detection points are determined as the prompt detection points. The role of the above randomly generated detection points is to serve as the segmentation prompt information of the image segmentation model, so that the image segmentation model can segment each object in the image to be processed. In the specific practice process, in a scenario where the position of the trolley is not fixed and changes at any time, the randomly generated detection points can be evenly distributed on the entire image plane, or in a scenario where the position of the trolley does not change frequently, they can be distributed in the region of interest according to preset rules.

[0056] Alternatively, the above-described implementation manner of obtaining the segmentation prompt information of the image to be processed may include: Step S111b: Determine the preset detection points as the prompt detection points.

[0057] For example, the implementation manner of the above step S111b is as follows: In a scenario where the position of the trolley does not change frequently, a set of fixed detection points may also be preset. These points may be the positions that are most likely to provide useful information based on experience or analysis. These detection points can be determined during the algorithm design stage, and then the preset detection points are determined as the prompt detection points.

[0058] Alternatively, the above-described implementation manner of obtaining the segmentation prompt information of the image to be processed may include: Step S111c: When the target object in the current video image is occluded, determine the detection points used in the historical video image frames as the prompt detection points.

[0059] For example, the implementation manner of the above step S111c is as follows: If the target object in the current frame is partially occluded, the detection points used in the previous frame (historical frame) can be referred to as the prompt detection points, that is, the detection points used in the historical video image frames are determined as the prompt detection points. This method can utilize the continuity and similarity in the time series to continue the image segmentation of the image to be processed even when the target is occluded.

[0060] As an alternative implementation manner of the above step S120, the above-described implementation manner of performing image recognition on the regional image may include: Step S121: Divide the regional image into multiple image patches.

[0061] For example, the implementation manner of the above step S121 is as follows: The regional image can be divided into multiple image patches (Patch) through an executable program written in a preset programming language.

[0062] Step S122: Input the multiple image patches into the vision transformer model so that the vision transformer model can identify the category of the regional image.

[0063] It can be understood that since the network structure of the vision transformer (Vision Transformer, which can be abbreviated as ViT or VIT) model is relatively complex, the specific process of the vision transformer model processing the multiple image patches and identifying the category of the regional image will be described in detail below.

[0064] Optionally, in order to improve the image classification accuracy of the ViT model, the ViT model can also be trained before using the ViT model. For example, the ViT model can be trained using a dataset of 1000 images + image categories, thereby effectively improving the image classification accuracy of the ViT model.

[0065] Step S123: If the category of the regional image is the category to which the target object belongs, then determine the regional image as the target regional image.

[0066] An implementation manner of the above step S123 is as follows: Since the last layer of the ViT model is a Softmax classification network layer, the multi-layer perceptron can calculate multiple category probabilities of the regional image, and determine the category corresponding to the maximum probability value among the multiple category probabilities as the category of the regional image. For example, the multiple categories respectively include: 0 represents a trolley, 1 represents a person, and 2 represents luggage. Assume that the category to which the target object belongs is 0 trolley, and if one of the output results of the multi-layer perceptron is [0.9, 0.05, 0.05], then it means that the category subscript corresponding to the maximum probability value (i.e., 90%) among the multiple category probabilities is also 0, which exactly indicates that the category of the regional image is the category to which the target object (trolley) belongs. Therefore, the regional image can be determined as the target regional image.

[0067] Please refer to Figure 4 the schematic diagram of the network structure of the vision transformer model provided by the embodiment of the present application shown; As an alternative implementation manner of the above step S122, the vision transformer model includes: a fully connected network, a transformation encoder, and a multi-layer perceptron network; after inputting multiple image patches into the vision transformer model, the vision transformer model can also identify the category of the regional image, and this implementation manner may include: Step S122a: Use the fully connected network to perform embedding processing on multiple image patches to obtain embedded patch vectors.

[0068] An implementation manner of the above step S122a is as follows: Input multiple image patches into the fully connected network of the vision transformer model, so that the fully connected network in the vision transformer model performs embedding processing (Embedding) on the multiple image patches, thereby converting the multiple image patches into embedded patch vectors with a fixed dimension (i.e., Embedded PatchedVector).

[0069] Step S122b: Use the transformation encoder to perform transformation processing on the embedded patch vectors to obtain transformed intermediate feature vectors.

[0070] An implementation manner of the above step S122b is as follows: input the embedded tile vector into the Transformer Encoder of the vision transformer model, so that the Transformer Encoder in the vision transformer model performs transformation processing on the embedded tile vector, that is, uses the self-attention mechanism to capture the global dependencies in the image to generate a transformed intermediate feature vector.

[0071] Step S122c: Use a multi-layer perceptron to perform class recognition on the transformed intermediate feature vector to obtain the class of the regional image.

[0072] An implementation manner of the above step S122c is as follows: input the transformed intermediate feature vector into the multi-layer perceptron of the vision transformer model, so that the multi-layer perceptron of the vision transformer model performs class recognition on the transformed intermediate feature vector. Since the last layer of the multi-layer perceptron is a Softmax classification network layer, the multi-layer perceptron can calculate multiple class probabilities of the regional image, and determine the class corresponding to the maximum probability value among the multiple class probabilities as the class of the regional image. For example, the multiple classes include: 0 represents a trolley, 1 represents a person, and 2 represents luggage. Suppose one of the output results of the multi-layer perceptron is [0.9, 0.05, 0.05], then it means that the class corresponding to the maximum probability value (i.e., 90%) among the multiple class probabilities is a trolley.

[0073] As an alternative implementation manner of the above step S130, the above spectrogram may include: a frequency-domain contour map; the implementation manner of performing spectral analysis on the target regional image may include: Step S131: Perform grayscale processing on the target regional image to obtain a two-dimensional grayscale image.

[0074] An implementation manner of the above step S131 is as follows: the color information of each pixel point in the target regional image is simplified to its brightness value, usually represented by an integer between 0 and 255, and thus the grayscale processing of the target regional image can be completed, and finally a two-dimensional grayscale image is obtained. It can be understood that this grayscale image reduces the amount of data while retaining the main structural information of the image, facilitating subsequent processing such as spectral analysis.

[0075] Step S132: Perform a transformation on the two-dimensional grayscale image to obtain a two-dimensional frequency-domain representation matrix.

[0076] For example, the implementation of the above step S132 is as follows: performing a transformation such as a Fast Fourier Transform (FFT), a Discrete Cosine Transform (DCT), or a wavelet transform on the two-dimensional grayscale image, so as to convert the two-dimensional grayscale image from the spatial domain to the frequency domain space, thereby obtaining a two-dimensional frequency domain representation matrix. Among them, the above Fourier transform can adopt a fast Fourier transform or a double fast Fourier transform, so that the obtained two-dimensional frequency domain representation matrix has complex numbers for each element. This transformation converts a two-dimensional image signal into a two-dimensional signal, which helps to identify periodic patterns and structural features in the image. These periodic patterns and structural features are often difficult to discover solely by spatial domain processing. For example, the periodic pattern of the wheels of a wheelbarrow will be particularly prominent in the spectrogram, which helps to distinguish the wheelbarrow from other similar objects (such as a baby carriage or a suitcase).

[0077] Step S133: Calculate a frequency domain contour map according to the two-dimensional frequency domain representation matrix.

[0078] Please refer to Figure 5 One of the schematic diagrams of the frequency domain contour map provided by the embodiment of the present application shown; for example, the implementation of the above step S133 is as follows: assuming that the Fourier transform (FFT) is used to obtain the two-dimensional frequency domain representation matrix, since the two-dimensional frequency domain representation matrix is a complex matrix, the modulus (i.e., amplitude) of each element in the matrix can be calculated as the spectral value. Since the dynamic range of the spectral value may be relatively large, the spectral value can be logarithmically transformed first and then the absolute value is taken to enhance the contrast and highlight the details. Finally, the spectral value is scaled to a specific range (such as 0 to 255) to facilitate the generation of the frequency domain contour map, which is a method of visualizing spectral features. Among them, the spectral contour in the frequency domain contour map can be a two-dimensional array obtained after spectral conversion of the regional image of the wheelbarrow.

[0079] As an alternative implementation of the above step S140, the implementation of statistically analyzing the target object in the target region image according to the spectrogram may include: Step S141: Perform frequency domain signal analysis on the spectrogram to obtain multiple signal frequency values.

[0080] For example, the implementation of the above step S141 is as follows: assuming that the spectrogram uses a frequency domain contour map, that is, the frequency domain contour map is obtained through the above steps, then the frequency domain contour map can be analyzed to extract frequency components, such as calculating the intensity or amplitude of each frequency component.

[0081] Step S142: Screen out the signal frequency value with the largest signal intensity from the multiple signal frequency values.

[0082] Step S143: Determine the signal frequency value with the maximum signal strength as the quantity of the target object in the image to be processed.

[0083] For example, the implementation manners of the above Step S142 to Step S143 are as follows: Traverse all frequency components in the frequency-domain contour map, and find the signal frequency value with the highest strength among all frequency components from multiple signal frequency values. Then, directly determine the signal frequency value with the maximum signal strength as the quantity of the target object in the image to be processed.

[0084] Optionally, frequency-domain signal analysis can also be performed on the spectrogram to identify all local maxima (i.e., significant frequency peaks), that is, use a local maximum detection algorithm (such as non-maximum suppression NMS) to identify all significant peaks in the spectrogram. Then, cluster adjacent frequency peaks. Assume that each cluster corresponds to a target object. Specifically, a clustering algorithm (such as DBSCAN, K-means, etc.) can be used to group them according to the distance and similarity between frequency peaks. Finally, count the number of clusters as the quantity of the target object. That is, each cluster is regarded as an independent target object, so the total number of clusters is the quantity of the target object. In the implementation process of the above solution, by performing peak detection on the spectrogram and using a clustering algorithm to group adjacent peaks, the existence of each target object can be estimated more accurately, thereby improving the counting accuracy.

[0085] Please refer to Figure 6 the flowchart of the trolley management method provided by the embodiment of the present application shown in; The embodiment of the present application provides a trolley management method, including: Step S210: Use an image segmentation model to perform image segmentation on the image to be processed to obtain multiple regional images.

[0086] Step S220: For each of the multiple regional images, perform image recognition on the regional image to obtain a target regional image whose recognition result is a trolley.

[0087] Step S230: Perform spectral analysis on the target regional image to obtain a spectrogram.

[0088] Step S240: Count the trolleys in the target regional image according to the spectrogram to obtain the quantity of the trolleys in the image to be processed.

[0089] Among them, the implementation manners of the above Step S210 to Step S240 are similar to the implementation manners of the above Step S110 to Step S140. If there is anything unclear, reference can be made to the description of the implementation manners of the above Step S110 to Step S140, so it will not be elaborated here.

[0090] In the implementation process of the above solution, the images captured by the cameras already installed and deployed in the airport terminal can be used as the images to be processed, and the trolleys in the images to be processed can be counted by region, so as to be able to complete the real-time detection of the number of trolleys in the area captured by the cameras.

[0091] Step S250: If the number of trolleys in the image to be processed is less than the preset threshold, perform supplementary scheduling and / or quantity warning on the trolleys.

[0092] For example, the implementation method of the above step S250 is as follows: After obtaining the number of trolleys in the image to be processed, the number of trolleys can also be monitored, scheduled and warned through an executable program. For example, if the number of trolleys is less than the preset threshold (such as 10, 20 or 30, etc.), a supplementary reminder message can be sent to the administrator through a POST request, so that after seeing the reminder message, the administrator can perform supplementary scheduling on the trolleys according to the number of trolleys, and / or send a quantity warning message to a preset terminal, so that after seeing the quantity warning message, the terminal building operation and maintenance personnel can perform emergency scheduling on the trolleys in time, or the terminal building operation and maintenance personnel can forward the warning information in time, for example, forward the warning information to the administrator, so that the administrator can perform emergency scheduling on the trolleys.

[0093] In the implementation process of the above solution, when the detected number of trolleys is lower than the preset threshold, the system can trigger an automated supplementary scheduling process to notify the staff to replenish the trolleys in time, avoiding service interruptions or inconvenience to passengers caused by insufficient trolleys. Especially in key areas or peak periods, the instant warning function can be used to remind the management to take actions, thus effectively improving the operation efficiency of the airport terminal.

[0094] Please refer to Figure 7 the flowchart of the target counting device provided by the embodiment of the present application shown in; The embodiment of the present application provides a target counting device 300, including: An image region segmentation module 310, configured to perform image segmentation on the image to be processed using an image segmentation model to obtain multiple region images.

[0095] A region image recognition module 320, configured to perform image recognition on each of the multiple region images to obtain a target region image whose recognition result is a target object.

[0096] An image spectrum analysis module 330, configured to perform spectrum analysis on the target region image to obtain a spectrogram.

[0097] A target object statistics module 340, configured to count the target objects in the target region image according to the spectrogram to obtain the number of target objects in the image to be processed.

[0098] As an alternative implementation of the above device, the image segmentation model includes: a prompt encoder, an image encoder, and an image decoder; the image region segmentation module includes: A prompt information acquisition sub-module for acquiring the segmentation prompt information of the image to be processed.

[0099] A prompt vector obtaining sub-module for encoding the segmentation prompt information using the prompt encoder to obtain a segmentation prompt vector.

[0100] An image vector encoding sub-module for encoding the image to be processed using the image encoder to obtain an embedded image vector.

[0101] An image vector decoding sub-module for decoding the segmentation prompt vector and the embedded image vector using the image decoder to obtain a contour mask image.

[0102] An image region segmentation sub-module for performing region segmentation on the image to be processed using the contour mask image to obtain a plurality of region images.

[0103] As an alternative implementation of the above device, the image to be processed is the current video image frame, and the segmentation prompt information is the prompt detection points in the current video image frame; the prompt information acquisition sub-module includes: A first detection point determination unit for determining a randomly generated detection point as the prompt detection point.

[0104] Alternatively, a second detection point determination unit for determining a pre-set detection point as the prompt detection point.

[0105] Alternatively, a third detection point determination unit for determining the detection points used in the historical video image frame as the prompt detection points when the target object in the current video image is occluded.

[0106] As an alternative implementation of the above device, the region image recognition module includes: A region image division sub-module for dividing the region image into a plurality of image blocks.

[0107] An image category recognition sub-module for inputting the plurality of image blocks into a vision transformer model so that the vision transformer model recognizes the category of the region image.

[0108] A target image recognition sub-module for determining the region image as the target region image if the category of the region image is the category to which the target object belongs.

[0109] As an alternative embodiment of the above device, the vision transformer model includes: a fully connected network, a transformation encoder, and a multi-layer perceptron network; the region image recognition module further includes: a patch vector embedding sub-module for using the fully connected network to perform embedding processing on a plurality of image patches to obtain embedded patch vectors.

[0110] A vector transformation processing sub-module for using the transformation encoder to perform transformation processing on the embedded patch vectors to obtain transformed intermediate feature vectors.

[0111] A vector category recognition sub-module for using the multi-layer perceptron to perform category recognition on the transformed intermediate feature vectors to obtain the category of the region image.

[0112] As an alternative embodiment of the above device, the spectrogram includes: a frequency domain contour map; the image spectrum analysis module includes: An image grayscale processing sub-module for performing grayscale processing on the target region image to obtain a two-dimensional grayscale image.

[0113] A grayscale image transformation sub-module for transforming the two-dimensional grayscale image to obtain a two-dimensional frequency domain representation matrix.

[0114] A contour map obtaining sub-module for calculating a frequency domain contour map according to the two-dimensional frequency domain representation matrix.

[0115] As an alternative embodiment of the above device, the target object statistics module includes: A frequency domain signal analysis sub-module for performing frequency domain signal analysis on the spectrogram to obtain a plurality of signal frequency values.

[0116] A signal frequency screening sub-module for screening out the signal frequency value with the maximum signal intensity from the plurality of signal frequency values.

[0117] A target quantity determination sub-module for determining the signal frequency value with the maximum signal intensity as the quantity of the target object in the image to be processed.

[0118] The embodiment of the present application further provides a trolley management device, including: An image region segmentation module for using an image segmentation model to perform image segmentation on the image to be processed to obtain a plurality of region images.

[0119] A region image recognition module for performing image recognition on each of the plurality of region images to obtain a target region image with the recognition result being a trolley.

[0120] An image spectrum analysis module for performing spectrum analysis on the target region image to obtain a spectrogram.

[0121] The target object statistics module is used to count the carts in the target area image according to the frequency spectrum diagram to obtain the number of carts in the image to be processed.

[0122] The trolley scheduling and warning module is used to perform supplementary scheduling and / or quantity warning on the trolleys if the number of trolleys in the image to be processed is less than a preset threshold.

[0123] It should be understood that the device corresponds to the above-mentioned target counting method embodiment and can execute the various steps involved in the above-mentioned method embodiment. The specific functions of the device can be referred to in the above description, and the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in a memory in the form of software or firmware or solidified in the operating system (OS) of the device.

[0124] An embodiment of the present application also provides an electronic device, including: a processor and a memory, the memory storing machine-readable instructions executable by the processor, and the machine-readable instructions executing the above method when executed by the processor.

[0125] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to execute the above method. The computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.

[0126] The embodiment of the present application also provides a computer program product, including: a computer program or a computer instruction, and the computer program or the computer instruction executes the method described above when executed by a processor.

[0127] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the similarities between the various embodiments, reference can be made to each other. For device embodiments, since they are basically similar to method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0128] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code. A module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may also occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which mainly depends on the functions involved.

[0129] In addition, in each of the embodiments of the present application, the various functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. Furthermore, in the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0130] The above description is only an optional implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the embodiments of the present application, and all should be covered within the protection scope of the embodiments of the present application.

Claims

1. A target counting method, characterized in that: include: Use the image segmentation model to perform image segmentation on the image to be processed to obtain multiple region images; For each of the plurality of area images, performing image recognition on the area image to obtain a target area image whose recognition result is a target object; Performing spectrum analysis on the target area image to obtain a spectrum graph; Counting the target objects in the target area image according to the frequency spectrum to obtain the number of the target objects in the image to be processed.

2. The method according to claim 1, characterized in that The image segmentation model includes: a prompt encoder, an image encoder and an image decoder; the image segmentation model is used to perform image segmentation on the image to be processed, including: Obtaining segmentation prompt information of the image to be processed; Encoding the segmentation hint information using the hint encoder to obtain a segmentation hint vector; Encode the image to be processed using the image encoder to obtain an embedded image vector; Decoding the segmentation hint vector and the embedded image vector using the image decoder to obtain a contour mask image; The image to be processed is segmented into regions using the contour mask image to obtain the multiple region images.

3. The method according to claim 2, characterized in that The image to be processed is a current video image frame, and the segmentation prompt information is a prompt detection point in the current video image frame; The obtaining of segmentation prompt information of the image to be processed includes: Determine the randomly generated detection point as the prompt detection point; Alternatively, a pre-set detection point is determined as the prompt detection point; Alternatively, when the target object in the current video image is blocked, the detection point used by the historical video image frame is determined as the prompt detection point.

4. The method according to claim 1, characterized in that: The performing image recognition on the regional image includes: Dividing the regional image into a plurality of image blocks; Inputting the plurality of image blocks into a visual transformer model so that the visual transformer model recognizes the category of the image in the region; If the category of the region image is the category to which the target object belongs, the region image is determined as the target region image.

5. The method according to claim 4, characterized in that The visual converter model includes: a fully connected network, a conversion encoder and a multi-layer perceptron network; after the plurality of image blocks are input into the visual converter model, the model further includes: Using the fully connected network to embed the multiple image blocks to obtain embedded image block vectors; Using the conversion encoder to convert the embedded tile vector to obtain a converted intermediate feature vector; The multi-layer perceptron is used to perform category recognition on the converted intermediate feature vector to obtain the category of the regional image.

6. The method according to claim 1, characterized in that The spectrum diagram includes: a frequency domain contour map; the spectrum analysis of the target area image includes: Performing grayscale processing on the target area image to obtain a two-dimensional grayscale image; Transforming the two-dimensional grayscale image to obtain a two-dimensional frequency domain representation matrix; The frequency domain contour map is calculated according to the two-dimensional frequency domain representation matrix.

7. The method according to claim 1, characterized in that The performing statistics on the target objects in the target area image according to the frequency spectrum includes: Performing frequency domain signal analysis on the spectrum graph to obtain multiple signal frequency values; Filter out the signal frequency value with the largest signal strength from the multiple signal frequency values; The signal frequency value with the maximum signal intensity is determined as the number of the target objects in the image to be processed.

8. A trolley management method, characterized in that: include: Use the image segmentation model to perform image segmentation on the image to be processed to obtain multiple region images; For each of the plurality of area images, image recognition is performed on the area image to obtain a target area image whose recognition result is a cart; Performing spectrum analysis on the target area image to obtain a spectrum graph; Counting the carts in the target area image according to the frequency spectrum to obtain the number of the carts in the image to be processed; If the number of the trolleys in the image to be processed is less than a preset threshold, supplementary scheduling and / or quantity alarm are performed on the trolleys.

9. A target counting device, characterized in that: include: An image region segmentation module is used to perform image segmentation on the image to be processed using an image segmentation model to obtain multiple region images; A region image recognition module, configured to perform image recognition on each region image of the plurality of region images to obtain a target region image whose recognition result is a target object; An image spectrum analysis module, used for performing spectrum analysis on the target area image to obtain a spectrum diagram; The target object statistics module is used to count the target objects in the target area image according to the frequency spectrum to obtain the number of the target objects in the image to be processed.

10. An electronic device based on a target counting method or a trolley management method, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the machine-readable instructions are executed by the processor to perform any method according to claims 1 to 8.

11. A computer-readable storage medium based on a target counting method or a trolley management method, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is executed.

12. A computer program product based on a target counting method or a trolley management method, characterized in that: include: A computer program or a computer instruction, wherein when the computer program or the computer instruction is executed by a processor, the method according to any one of claims 1 to 8 is executed.

Citation Information

Patent Citations

  • Image and signal processing-based card identification and counting method

    CN107256384A

  • Method for counting number of targets in area, and method and device for training area identification model

    CN114842416A

  • Insulator chain disc segmentation positioning and state analysis method in infrared image

    CN116503470A

  • Image recognition method and device, equipment, storage medium and program product

    CN117036802A

  • Wire abnormity identification method and device, electronic equipment and storage medium

    CN118736223A