Detection method, device and equipment of left article, storage medium and program product

By processing and segmenting the surveillance video frames and combining them with timestamps to determine the length of time items have been left behind, the problem of low efficiency in handling left-behind items has been solved, and the customer can be quickly located, thereby improving the efficiency of handling left-behind items.

CN120673314APending Publication Date: 2025-09-19INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510781476.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies are unable to promptly detect left-behind items and quickly locate their owners, resulting in low efficiency in handling left-behind items.

Method used

By acquiring multiple image frames of surveillance video, mask processing is performed to remove moving objects and portrait areas, and segmentation processing is performed to determine the segmentation map of the left-behind items. The left-behind duration is determined by combining the timestamp, the customer is located, and a left-behind item prompt is generated.

Benefits of technology

It achieves the rapid location of left-behind items and the customers they belong to, improves the efficiency of handling left-behind items, reduces false detections and missed detections, and improves service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673314A_ABST
    Figure CN120673314A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a detection method and device for a left article, equipment, a storage medium and a program product, and relates to the technical field of image processing. The method comprises the following steps: acquiring a plurality of image frames corresponding to first video data, performing mask processing on an invalid region in each image frame to obtain a target still image, further performing segmentation processing on the plurality of target still images to obtain a left article segmentation image, and determining left article information based on the plurality of left article segmentation images and corresponding timestamps. And judging whether the leaving time length is greater than a preset time length, if so, determining an affiliation customer image of the left article from the second video data, and generating an article leaving prompt based on the affiliation customer image and the left article information, thereby realizing rapid positioning of the left article and the corresponding affiliation personnel. And an article leaving prompt is timely fed back to the affiliated personnel and related workers, so that the processing efficiency of the left article is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, device, storage medium and program product for detecting left-behind objects. Background Art

[0002] Business locations like office lobbies and bank branches are characterized by high turnover and complex business scenarios. Customers often leave personal belongings like documents and wallets behind in bank branches, resulting in financial losses. Furthermore, if banks handle these issues improperly, this can easily lead to customer dissatisfaction and negatively impact the quality of their service.

[0003] Currently, the primary means of discovering left-behind items and identifying their owners is through manual inspections and simple surveillance equipment. For example, bank staff patrol public areas like the bank's business hall to check for any left-behind items. If any are found, they review the bank's surveillance footage to identify the owner. However, manual inspections and review of surveillance footage are time-consuming and prone to false detections and missed detections.

[0004] Therefore, the existing technology is unable to discover the left-behind items in time and quickly locate the owner, which leads to the technical problem of low efficiency in handling the left-behind items. Summary of the Invention

[0005] The present application provides a method, device, equipment, storage medium and program product for detecting left-behind items, which are used to solve the technical problem that the existing technology cannot detect left-behind items in a timely manner and quickly locate the owner, resulting in low efficiency in handling left-behind items.

[0006] In a first aspect, the present application provides a method for detecting leftover items, comprising:

[0007] Acquire multiple image frames corresponding to the first video data, where the image frames carry timestamps;

[0008] Performing mask processing on the invalid area in each of the image frames to obtain a target still image, wherein the invalid area includes: an image area corresponding to the moving object and a portrait area;

[0009] Segmenting the plurality of target still images to obtain a plurality of left-behind item segmentation images;

[0010] Determining the left-behind item information based on the plurality of left-behind item segmentation graphs and corresponding timestamps, wherein the left-behind item information includes: a left-behind duration and a left-behind start time;

[0011] Determine whether the remaining time is greater than a preset time;

[0012] When the left-behind duration is greater than the preset duration, the customer image of the left-behind item is determined from the second video data, and an item left-behind prompt is generated based on the customer image and the left-behind item information. The second video data includes: video data within the monitoring area corresponding to the first preset duration, and the first preset duration includes: the left-behind duration, and / or the second preset duration before the start time of the left-behind.

[0013] In a second aspect, the present application provides a device for detecting left-behind objects, comprising:

[0014] An acquisition module, configured to acquire a plurality of image frames corresponding to the first video data, wherein the image frames carry a timestamp;

[0015] a processing module, configured to perform mask processing on an invalid area in each of the image frames to obtain a target still image, wherein the invalid area includes an image area corresponding to a moving object and a portrait area;

[0016] The processing module is further configured to segment the plurality of target still images to obtain a plurality of segmentation images of left-behind items;

[0017] A determination module, configured to determine information of a left-behind item based on the plurality of left-behind item segmentation graphs and corresponding timestamps, wherein the left-behind item information includes a left-behind duration and a left-behind start time;

[0018] A judging module, configured to judge whether the remaining time is greater than a preset time;

[0019] The determining module is further configured to determine, when the left-behind duration is greater than a preset duration, an image of the customer to which the left-behind item belongs from second video data, the second video data comprising: video data within a monitoring area corresponding to a first preset duration, the first preset duration comprising: the left-behind duration, and / or a second preset duration before the left-behind start time;

[0020] A generating module is used to generate an item left behind reminder based on the customer image and the item left behind information.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0022] The memory stores computer-executable instructions;

[0023] The processor executes the computer-executable instructions stored in the memory to implement the above first aspect and / or various possible implementations of the first aspect.

[0024] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementation methods of the first aspect.

[0025] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0026] The method for detecting left-behind items provided in the present application obtains multiple image frames corresponding to first video data, and performs mask processing on the invalid area in each image frame to obtain a target still image, and then performs segmentation processing on the multiple target still images to obtain multiple left-behind item segmentation maps. Based on the multiple left-behind item segmentation maps and the corresponding timestamps, the left-behind item information is determined, and it is judged whether the left-behind duration is greater than the preset duration. If the left-behind duration is greater than the preset duration, the customer image of the left-behind item is determined from the second video data, and based on the customer image and the above-mentioned left-behind item information, an item left-behind prompt is generated, thereby realizing rapid positioning of the left-behind item and the corresponding owner, and timely feeding back the item left-behind prompt to the owner and relevant staff, thereby improving the efficiency of handling left-behind items. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0028] Figure 1 Schematic diagram of the process of detecting leftover items provided in the embodiment of the present application Figure 1 ;

[0029] Figure 2 Schematic diagram of the process of detecting leftover items provided in the embodiment of the present application Figure 2 ;

[0030] Figure 3 A schematic diagram of a detection model for left-behind items provided in an embodiment of the present application;

[0031] Figure 4 Schematic diagram of the process of detecting leftover items provided in the embodiment of the present application Figure 3 ;

[0032] Figure 5 A schematic diagram of a homed customer detection model provided in an embodiment of the present application;

[0033] Figure 6A schematic diagram of a customer-attribution matching model provided in an embodiment of the present application;

[0034] Figure 7 A schematic diagram of the structure of the device for detecting leftover items provided in this application;

[0035] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application.

[0036] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0037] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0038] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0039] In addition, this application involves conducting big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.), and using artificial intelligence technology to make automated decisions, and making technical solutions that have a significant impact on personal rights and interests based on the results of automated decisions. The application provides users with corresponding operation entrances for them to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered.

[0040] It should be noted that the methods, devices, equipment, storage media and program products for detecting leftover items provided in this application can be used in the field of image processing technology, and can also be used in any field other than the field of image processing technology. The application fields of the methods, devices, equipment, storage media and program products for detecting leftover items in this application are not limited.

[0041] In business locations with high traffic and complex business processes, such as office halls and bank branches, customers are prone to leaving personal belongings in public places due to the multiple locations and long processing times required for business transactions. Furthermore, if the staff handles the left-behind items improperly, this can easily lead to customer dissatisfaction and reduce service quality.

[0042] Currently, the primary means of discovering left-behind items and identifying their owners relies on manual inspections and simple surveillance equipment. For example, bank staff patrol public areas, such as the bank's business hall, during designated hours to check for any left-behind items. If any are found, they replay the bank's surveillance footage to identify the owner. However, manual inspections and review of surveillance footage are time-consuming and prone to false detections and missed detections.

[0043] Therefore, the existing technology is unable to discover the left-behind items in time and quickly locate the owner, which leads to the technical problem of low efficiency in handling the left-behind items.

[0044] The method for detecting left-behind items provided in this application performs masking on multiple image frames corresponding to a surveillance video, removes invalid areas in the image frames, obtains a target still image, and then performs segmentation processing on the multiple target still images to determine multiple left-behind item segmentation maps. In addition, by identifying and analyzing the multiple left-behind item segmentation maps, information about the left-behind items is obtained. In the case of a long left-behind period, the image of the customer to whom the left-behind item belongs is determined, and based on the customer image and the left-behind item information, a left-behind item prompt is generated. This method can promptly detect and locate left-behind items and their corresponding customers, which is conducive to timely feedback of left-behind item prompts to the owner and relevant staff, thereby improving the efficiency of left-behind item processing.

[0045] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0046] Figure 1 Schematic diagram of the process of detecting leftover items provided in the embodiment of the present application Figure 1 .like Figure 1 As shown, the method includes:

[0047] S101: Acquire multiple image frames corresponding to first video data.

[0048] The first video data is surveillance video data of a target location, for example, surveillance video of a bank business hall.

[0049] Specifically, the first video data is subjected to frame extraction processing at a preset frame interval to obtain multiple image frames, each of which carries a timestamp. For example, in a bank business hall, a monitoring system uses monitoring equipment installed in the business hall to collect business scenes in real time to obtain first video data. The first video data is then subjected to frame extraction processing at a preset frame interval to obtain multiple image frames.

[0050] It should be noted that this method can be applied to bank monitoring systems, office lobby monitoring systems, etc., and can also be applied to left-behind item detection systems associated with monitoring systems. This application does not impose any restrictions on this.

[0051] S102: Perform mask processing on the invalid area in each image frame to obtain a target still image.

[0052] The invalid area includes: the image area corresponding to the moving object and the portrait area.

[0053] Specifically, a comparative analysis is performed on multiple image frames to determine the image area corresponding to the moving object. In addition, recognition analysis is performed on multiple image frames to determine the portrait area of ​​each image frame, and then mask processing is performed on the image area corresponding to the moving object and the portrait area to obtain the target still image.

[0054] Optionally, for each of the multiple image frames, feature extraction is performed on the image frame to obtain corresponding image features, and the image features of adjacent image frames are compared to obtain corresponding motion features. The motion features reflect the motion intensity and direction of multiple feature points corresponding to the previous image frame in the adjacent image frames. The motion features may be, for example, displacement features and velocity features.

[0055] Based on the motion features, a corresponding motion region is determined from the previous image frame and masked. Furthermore, the masked previous image frame is subjected to recognition analysis to extract portrait features from the previous image frame, and a portrait region corresponding to the portrait features is masked to obtain a target still image.

[0056] Before determining the left-behind items, this method removes the image area and portrait area corresponding to the moving object in each image frame, which is conducive to improving the accuracy and efficiency of determining the left-behind items.

[0057] It should be noted that this application does not restrict the timing of removing the image area corresponding to the moving object and removing the portrait area.

[0058] S103: Segment the multiple target still images to obtain multiple left-behind item segmentation maps.

[0059] Specifically, for each of the multiple target still images, feature extraction is performed on the target still image to obtain a corresponding feature map. Feature data corresponding to multiple pixels in the feature map are then classified, and multiple pixels belonging to the same category are determined as a segmented region, thereby obtaining a segmented image. This segmented image is then compared and analyzed with the target background segmented image corresponding to the target location in step S101 to obtain a segmentation map of the lost item.

[0060] This method effectively segments multiple different objects in the target still image through segmentation processing of the target still image. The segmentation map of the remaining items obtained based on the segmented target still image is more complete and accurate. When the segmentation map of the remaining items is subsequently identified and analyzed to determine the information of the remaining items, it is beneficial to improve the accuracy of the remaining item information.

[0061] S104: Determine the information of the left-behind items based on the multiple left-behind item segmentation graphs and the corresponding timestamps.

[0062] The information of the left-behind items includes: the duration of the left-behind and the starting time of the left-behind.

[0063] Multiple left-behind item segmentation images are identified and classified to obtain at least one left-behind item classification group. For each left-behind item classification group, target detection is performed on at least one left-behind item segmentation image within the left-behind item classification group to obtain a corresponding left-behind item category. Examples of left-behind item categories include a black wallet, a green umbrella, and a blue folder.

[0064] In addition, for each legacy item classification group, the legacy start time and end time of the legacy item are determined based on the timestamp corresponding to each legacy item segmentation map within the legacy item classification group, and the legacy duration of the legacy item is determined based on the legacy start time and end time.

[0065] The method determines the information of the left-behind items by analyzing and processing the segmentation map of the left-behind items, which is conducive to timely reporting the accurate information of the left-behind items to relevant personnel, thereby improving the efficiency of claiming the left-behind items.

[0066] S105. Determine whether the remaining duration is greater than the preset duration; if the remaining duration is greater than the preset duration, execute the following step S106; if the remaining duration is not greater than the preset duration, execute the following step S107.

[0067] A determination is made as to whether the left-behind time is greater than a preset time to determine whether the left-behind item was temporarily placed by the customer. If the left-behind time is greater than the preset time, it indicates that the corresponding left-behind item has been left in the target location for a long time, triggering the determination operation of the owner as shown in step S106 below. Alternatively, if the left-behind time is not greater than the preset time, it indicates that the corresponding left-behind item may have been temporarily placed by the customer, for example, a customer may have placed a school bag on a chair in the waiting area while conducting business.

[0068] It is understood that the above preset duration is determined based on the business processing time of the target scenario. For example, if the average business processing time in a bank business hall is 10 minutes, the preset duration is set to 13 minutes; if the average business processing time in an office hall is 20 minutes, the preset duration is set to 20 minutes.

[0069] This method filters out items that are temporarily placed by judging the length of time they have been left, avoiding alarm processing for items that customers temporarily place at the target location, thereby interfering with the customer's business processing progress and reducing customer satisfaction with the business processing location.

[0070] S106: Determine the customer image of the left-behind item from the second video data, and generate an item-left-behind prompt based on the customer image and the left-behind item information.

[0071] The second video data includes: video data in the monitoring area corresponding to a first preset time period, and the first preset time period includes: a legacy time period, and / or a second preset time period before the legacy start time period.

[0072] The information of the left-behind items includes: images of the left-behind items.

[0073] Specifically, if the judgment result of step S106 indicates that the length of time the item was left behind is greater than a preset length of time, the second video data is obtained, and feature extraction is performed on the image of the item left behind. Based on the extracted item features, the item left behind is located from the second video data, and multiple first position information corresponding to the item left behind is obtained. In addition, multiple customers are identified from the second video data, and second position information corresponding to the multiple customers is determined. Based on the overlap rate between the first position information and the second position information, the customer to whom the item was left behind is determined. For example, a customer whose overlap rate is greater than a preset overlap rate is determined as the customer to whom the item was left behind, and then the corresponding customer image is intercepted from the second video data.

[0074] After obtaining the customer image, the customer information is determined based on the customer image, and based on the customer information, the customer image and the information of the left-behind items, a left-behind item reminder is generated, and relevant information about the left-behind items is promptly fed back to the relevant staff and the customer, thereby improving the efficiency of handling left-behind items and enhancing the security management of target places such as bank branches.

[0075] S107: Use the third video data as the first video data and execute the above step S101.

[0076] Specifically, if the judgment result of step S106 indicates that the left-behind duration is not greater than the preset duration, third video data is acquired and it is determined whether there is any left-behind item in the third video data.

[0077] The method for detecting left-behind items provided in this embodiment obtains multiple image frames corresponding to the first video data, and performs mask processing on the invalid area in each image frame to obtain a target still image, and then performs segmentation processing on the multiple target still images to obtain multiple left-behind item segmentation maps. Based on the multiple left-behind item segmentation maps and the corresponding timestamps, the left-behind item information is determined, and it is judged whether the left-behind duration is greater than the preset duration. If the left-behind duration is greater than the preset duration, the customer image of the left-behind item is determined from the second video data, and based on the customer image and the above-mentioned left-behind item information, an item left-behind prompt is generated, thereby realizing rapid positioning of the left-behind item and the corresponding owner, and timely feeding back the item left-behind prompt to the owner and relevant staff, thereby improving the efficiency of handling left-behind items.

[0078] Figure 2 Schematic diagram of the process of detecting leftover items provided in the embodiment of the present application Figure 2 .like Figure 2 As shown, this embodiment Figure 1 Based on the embodiment, a possible method for detecting left-behind items is described in detail. The method includes:

[0079] S201: Acquire multiple image frames corresponding to first video data.

[0080] This step S201 is similar to the explanation of the above step S101 and will not be repeated here.

[0081] S202 , extracting features from multiple image frames to obtain feature data; and determining a motion intensity feature of the current image frame based on displacement features corresponding to the current image frame and the first image frame.

[0082] The feature data includes: a displacement feature, the current image frame is a previous image frame adjacent to the first image frame, and the motion intensity feature includes: a plurality of motion intensity values.

[0083] Specifically, feature extraction is performed on multiple image frames to obtain feature data corresponding to each image frame. For the adjacent current image frame and the first image frame, corresponding feature points are determined in the two frames. For example, a first feature point in the current image frame corresponds to a second feature point in the first image frame. Based on the displacement characteristics of the first and second feature points, a displacement offset vector is calculated. This displacement offset vector can be used to characterize the motion intensity feature, and the modulus of the displacement offset vector is the motion intensity value. For example, the displacement offset vector is (a, b).

[0084] This method quantifies the motion intensity of each feature point by comparing and analyzing multiple adjacent image frames, which is conducive to accurately constructing the motion intensity feature map corresponding to each image frame, so that in subsequent steps, the moving area and static area in each image frame can be more accurately determined.

[0085] S203 : Determine a sub-region in each current image frame whose motion intensity value is greater than a preset threshold as a motion region, and determine a sub-region whose motion intensity value is less than the preset threshold as a static region.

[0086] Specifically, for each of the multiple current image frames, a determination is made as to whether the motion intensity values ​​of multiple feature points in the current image frame are greater than a preset threshold. If the motion intensity value is greater than the preset threshold, the feature point corresponding to the motion intensity value is determined as a moving feature point. If the motion intensity value is not greater than the preset threshold, the feature point corresponding to the motion intensity value is determined as a stationary feature point. After determining the moving feature points and the stationary feature points, spatial clustering is performed on the multiple feature points, and areas with dense motion feature points are determined as moving areas, while areas with dense stationary feature points are determined as stationary areas.

[0087] Optionally, the motion region is marked. The marking may include, for example, determining a plurality of boundary points of each motion region, and constructing a corresponding motion region outline based on the plurality of boundary points, thereby achieving marking of the motion region.

[0088] It is understandable that the preset threshold value can be set according to needs, or can be an adaptive threshold value, which is not limited in this application.

[0089] S204: Perform mask processing on the motion area in each current image frame to obtain a first still image.

[0090] Specifically, a mask is generated based on the motion regions and applied to the current image frame to remove the motion regions in the current image frame while retaining the static regions, thereby obtaining a first static image. This method removes the motion regions in each image frame before identifying the left-behind object, reducing the computational effort required to identify the left-behind object and improving the efficiency and accuracy of the identification process.

[0091] Optionally, the masked image is subjected to processing such as dilation, erosion, and smoothing to eliminate noise and holes, thereby obtaining a first still image with higher quality.

[0092] S205 , performing recognition analysis on each first still image to obtain at least one corresponding portrait region; performing mask processing on the at least one portrait region to obtain a target still image corresponding to each first still image.

[0093] Specifically, feature extraction is performed on the first still image, and the extracted feature data is identified and analyzed to obtain a portrait region in the first still image. The portrait region is marked according to a preset rule. The preset rule may include, for example, marking the portrait region with a bounding box. Furthermore, after determining the portrait region corresponding to each first still image, a mask is performed on the portrait region to remove the portrait region from the first still image, thereby obtaining a target still image. For example, the mask corresponding to the portrait region is set to 0, and the masks corresponding to other regions are set to 1.

[0094] For example, Figure 3 This is a schematic diagram of a detection model for left-behind items provided in an embodiment of the present application. Figure 3 As shown, N image frames are input into the multi-frame comparison module corresponding to the above-mentioned detection model, where optical flow estimation is performed to obtain the motion region corresponding to each image frame. Furthermore, based on the feature data of the multiple image frames, the corresponding portrait region is determined. The motion region and the portrait region are then masked to obtain a first still image. The training set for the multi-frame comparison module includes: multiple image frames containing moving objects and people as input; the output for each image frame is: the first still image after removing the portrait region and the motion region.

[0095] This method removes the portrait area in the image frame before determining the left-behind items, reducing the recognition calculation amount and misjudgment rate of the left-behind items, thereby improving the detection efficiency and accuracy of the left-behind items.

[0096] S206 : For any target still image among the multiple target still images, perform feature extraction on the target still image to obtain a corresponding target feature map.

[0097] Specifically, feature extraction is performed on the target still image to obtain a corresponding target feature map, which includes target features corresponding to each feature point in the target still image, for example, pixel features corresponding to each feature point.

[0098] S207 , performing classification processing on the multiple target features, determining the target sub-region corresponding to at least one target feature belonging to the same category as a segmentation region, and obtaining a segmentation image.

[0099] Specifically, multiple target features in the target feature map are intensively processed, and subregions corresponding to feature points belonging to the same category are determined as a segmented region, thereby segmenting the target still image into at least one segmented region to obtain a segmented image. For example, based on a pixel map corresponding to the target still image, the category of each pixel in the pixel map is determined, and then the region corresponding to multiple pixels belonging to the same category is determined as a segmented region to obtain the segmented image corresponding to the target still image.

[0100] For example, Figure 3 As shown, the target still image is feature extracted through the transformer neural network and input into the image segmentation module to obtain segmentation results corresponding to multiple image frames, which include the above-mentioned segmentation map.

[0101] It should be noted that during the training process, Figure 3 The multi-frame comparison module is trained. After the training is completed, the first model parameters corresponding to the multi-frame comparison module are fixed, and then the image segmentation module is trained based on multiple image frames and corresponding segmented images to determine the second model parameters corresponding to the image segmentation module.

[0102] As can be understood, the multi-frame comparison module is pre-trained until convergence using the mean squared error (mse) loss as the loss function, thereby obtaining the corresponding first model parameters. Furthermore, the image segmentation module is trained until convergence using the pixel-by-pixel cross entropy loss as the loss function, thereby obtaining the corresponding second model parameters.

[0103] In the segmentation map obtained by this method, each object is accurately divided, providing clear boundary information for the subsequent identification of left-behind objects, which is conducive to improving the accuracy of left-behind object identification and can more accurately determine the size and shape of left-behind objects.

[0104] S208: Acquire a scene image corresponding to the first video data; and segment the scene image to obtain a scene segmentation image.

[0105] The scene image is an image in which no moving objects, human images or left-behind objects exist in the target location. The target location may be, for example, a public place such as a bank business hall, an office hall, a station or an airport.

[0106] Specifically, to accurately identify the item left behind in the segmented image obtained in step S207, a scene image corresponding to the first video data is obtained. This scene image may be, for example, a business hall devoid of people or moving objects. Furthermore, based on the feature data of the scene image, the scene image is segmented to obtain a scene segmentation map. For example, based on pixel features of the scene image, the scene image is segmented into at least one portion to obtain a scene segmentation map.

[0107] It is understood that the scene image is pre-stored in the left-behind item detection system, or the scene image is video data of the target location collected during a preset time period, which is not limited in this application. The preset time period is, for example, from 5:00 to 6:00 or from 1:00 to 2:00, which is determined based on the business practices of the target location, and is not limited in this application.

[0108] This method segments the scene image of the target place and obtains a scene segmentation image that only includes the background image of the target place. In the subsequent process of identifying left-behind objects, it is helpful to more quickly and accurately identify left-behind objects in the segmentation image.

[0109] S209: Compare the segmented image with the scene segmented image to obtain at least one segmented region in the segmented image that does not match the scene segmented image, and determine the at least one segmented region as a segmentation map of the left-behind items.

[0110] Specifically, after obtaining the scene segmentation image and the segmentation image, the segmentation image and the scene segmentation image are compared, and each segmentation area in the segmentation image is matched with multiple segmentation areas in the scene segmentation image to determine the segmentation area in the segmentation image that does not match the scene segmentation image. If the segmentation area includes the left-behind items, the segmentation area is determined as the left-behind item segmentation map.

[0111] For example, Figure 3 As shown, the fully connected layer of the image segmentation module transmits the feature data of the left-behind item segmentation map to different segmentation region extraction modules to obtain left-behind item segmentation information, which includes the above-mentioned left-behind item segmentation map.

[0112] This method is based on comparing the scene segmentation image with the segmentation image, and can more quickly and accurately determine the segmentation area corresponding to the left-behind object, thereby improving the detection efficiency of the left-behind object.

[0113] S210, performing feature extraction and classification processing on multiple left-behind item segmentation maps to obtain an item category and an item image of at least one left-behind item; and determining a left-behind start time and a left-behind duration of each left-behind item based on timestamps corresponding to the multiple left-behind item segmentation maps.

[0114] Specifically, feature extraction is performed on multiple left-behind item segmentation maps to obtain item features corresponding to each left-behind item segmentation map, and the item features include but are not limited to: color features, texture features, and shape features. Based on the similarity between the multiple item features, the left-behind items corresponding to the multiple left-behind item segmentation maps are classified. For example, the left-behind items corresponding to the item features whose similarity is less than a similarity threshold are determined as an item category. Target detection is performed on at least one left-behind item segmentation map in each item category to obtain the item category of the left-behind item, and an item image is output, which is the above-mentioned left-behind item segmentation map.

[0115] In addition, each left-behind item segmentation map corresponds to a timestamp. Based on the timestamp corresponding to at least one left-behind item segmentation map in each item category, the left-behind start time when the left-behind item first appears and the end time when the left-behind item last appears are determined, and then the left-behind duration is calculated based on the above-mentioned left-behind start time and end time.

[0116] Optionally, the left-behind location of the left-behind item is determined based on the position coordinates of the left-behind item in the image.

[0117] This method accurately identifies the category and time of left-behind items, provides a basis for determining whether the left-behind items are items temporarily stored by customers, and helps provide customers and relevant staff with specific left-behind item information, thereby improving the efficiency of left-behind item handling.

[0118] S211. Determine whether the remaining duration is greater than the preset duration; if the remaining duration is greater than the preset duration, execute the following step S212; if the remaining duration is not greater than the preset duration, execute the following step S213.

[0119] This step S211 is similar to the explanation of the above step S105 and will not be repeated here.

[0120] S212: Determine the customer image of the left-behind item from the second video data, and generate an item-left-behind prompt based on the customer image and the left-behind item information.

[0121] This step S212 is similar to the explanation of the above step S106 and will not be repeated here.

[0122] S213: Use the third video data as the first video data.

[0123] This step S213 is similar to the explanation of the above step S107 and will not be repeated here.

[0124] The method for detecting left-behind items provided in this embodiment obtains multiple image frames corresponding to first video data, performs feature extraction on the multiple image frames to obtain feature data, and determines the motion intensity feature of the current image frame based on the displacement features corresponding to the current image frame and the first image frame, and then determines the sub-region with a motion intensity value greater than a preset threshold in each current image frame as a motion region, and determines the sub-region with a motion intensity value less than a preset threshold as a static region, and performs mask processing on the motion region in each current image frame to obtain a first static image, and then performs recognition analysis on each first static image to obtain at least one corresponding portrait region, and performs mask processing on at least one portrait region to obtain a target static image corresponding to each first static image. The target static image removes the motion region and the portrait region, thereby preventing the moving object and the items carried by the customer from affecting the recognition of left-behind items, which is beneficial to improving the efficiency and accuracy of left-behind item detection. In addition, for any target still image among the multiple target still images, feature extraction is performed on the target still image to obtain a corresponding target feature map, and multiple target features are classified and processed, and the target sub-region corresponding to at least one target feature belonging to the same category is determined as a segmented region to obtain a segmented image, and then the scene image corresponding to the first video data is obtained, and the scene image is segmented to obtain a scene segmentation image; the segmented image is compared with the scene segmentation image to obtain at least one segmented region in the segmented image that does not match the scene segmentation image, and at least one segmented region is determined as a segmentation map of the left-behind items. This method not only obtains the boundary information of the left-behind items by segmenting the target still image, but also quickly and accurately locates the segmented region to which the left-behind items belong by comparing the segmented image with the scene segmentation image, which is conducive to improving the detection efficiency and accuracy of the left-behind items. Finally, the method extracts features and classifies multiple left-behind item segmentation images to obtain an item category and item image of at least one left-behind item. Based on the timestamps corresponding to the multiple left-behind item segmentation images, the method determines the starting time and duration of each left-behind item, and then determines whether the left-behind duration is greater than a preset duration. If the left-behind duration is greater than the preset duration, the method determines the customer image of the left-behind item from the second video data, and generates an item left-behind reminder based on the customer image and the left-behind item information, thereby promptly providing the left-behind item information to the customer and relevant staff, thereby improving the efficiency of processing the left-behind item information and accurately returning the left-behind item to the customer. Furthermore, if the left-behind duration is not greater than the preset duration, the method uses the third video data as the first video data to continue detecting whether there are left-behind items in the target location, so as to promptly detect left-behind items in the target location.

[0125] Figure 4 Schematic diagram of the process of detecting leftover items provided in the embodiment of the present application Figure 3 .like Figure 4 As shown, this embodiment Figure 1-3 Based on the embodiment, a method for determining the owner of a property and generating a reminder of a property left behind based on the owner of the property and the property left behind information is described in detail. The method includes:

[0126] S401: Perform frame extraction processing on the second video data to obtain multiple second image frames; perform recognition analysis on the multiple second image frames to obtain the movement trajectory of the left-behind object and the movement trajectories of multiple people.

[0127] The second image frame carries a timestamp.

[0128] Specifically, the second video data is subjected to frame extraction processing at preset time intervals to obtain a plurality of second image frames. Target recognition is performed on the plurality of second image frames, and at least one identified customer is labeled, with each customer corresponding to a customer identifier. Furthermore, target recognition is performed on the plurality of second image frames to obtain at least one left-behind item. The person position coordinates corresponding to each customer identifier and the item position coordinates corresponding to each left-behind item are determined from the plurality of second image frames. Furthermore, based on the timestamp, corresponding person position coordinates, and item position coordinates of each second image frame, a movement trajectory of the person corresponding to each customer identifier and a movement trajectory of the item corresponding to each left-behind item are constructed.

[0129] For example, Figure 5 This is a schematic diagram of a home customer detection model provided in an embodiment of the present application. Figure 5 As shown, the portrait image is an image of the customer entering the target location, triggered by a sensor at the door. The image frame is resized to capture multi-scale features corresponding to the image frame. These multi-scale features are then concatenated (CAT) and fed into a transformer module for feature extraction, yielding target features. These target features are then fed into a bounding box regression layer for classification. Bounding box labels are then assigned to at least one customer and at least one left-behind item, yielding at least one customer bounding box and at least one left-behind item bounding box. Bounding box information is then output, including the item location coordinates of the left-behind item and the item identifier corresponding to those location coordinates, as well as the customer's location coordinates and the customer identifier corresponding to those location coordinates. Furthermore, a similarity measurement is performed between the feature information corresponding to the portrait image and the target features to determine the customer image corresponding to the target feature, thereby improving the accuracy of customer identification.

[0130] After obtaining the item location coordinates of the item left behind by the person in each second image frame, the item identifier corresponding to the item location coordinates, and the customer location coordinates and the customer identifier corresponding to the customer location coordinates, the person location coordinates corresponding to each customer identifier are spliced ​​together based on the timestamp corresponding to each second image frame to delineate the person's movement trajectory corresponding to each customer identifier. Similarly, the item location coordinates corresponding to each item identifier are spliced ​​together based on the timestamp corresponding to each second image frame to delineate the item's movement trajectory corresponding to each item identifier.

[0131] S402: For any one of the plurality of human movement trajectories, determine an overlap value between the human movement trajectory and the object movement trajectory.

[0132] Specifically, for each person's movement trajectory and each object's movement trajectory, the person's movement trajectory is compared and analyzed with at least one object's movement trajectory to obtain the overlap value between the person's movement trajectory and each object's movement trajectory. For example, the time series of the person's movement trajectory and the object's movement trajectory are aligned, and the similarity between the two compared trajectories is calculated. This similarity can be used to represent the overlap value between the person's movement trajectory and the object's movement trajectory.

[0133] S403. Determine whether the overlap value is greater than the overlap threshold; if the overlap value is not greater than the overlap threshold, execute the above step S402; if the overlap value is greater than the overlap threshold, execute the following step S404.

[0134] Specifically, after obtaining the overlap value between the person's movement trajectory and the object's movement trajectory, it is determined whether the overlap value is greater than the overlap threshold, so as to timely lock the person's movement trajectory that is highly overlapped with the object's movement trajectory. Especially in complex business scenarios, this method helps to improve the efficiency and accuracy of identifying attributable customers.

[0135] It can be understood that the person's movement trajectory corresponds to the target customer, and the item's movement trajectory corresponds to the target left-behind item. If the overlap value is not greater than the overlap threshold, it means that the overlap value between the person's movement trajectory and the item's movement trajectory is low, and the target customer is determined not to be the original owner of the target left-behind item. In addition, if the overlap value is greater than the overlap threshold, it means that the item's movement trajectory and the person's movement trajectory are highly overlapped, and the target customer is determined to be the original owner of the target left-behind item.

[0136] It should be noted that if the overlap value is not greater than the overlap threshold, step S402 is executed to compare the overlap values ​​between the next person's trajectory and the multiple object's trajectory. This method outputs the corresponding customer image after identifying a person's trajectory that highly overlaps with an object's trajectory. This eliminates the need to compare and calculate the overlap values ​​between all person's and object's trajectories, reducing the computational effort involved in determining the customer, thereby improving the efficiency of determining the customer.

[0137] S404: Determine the person image corresponding to the person's movement trajectory as the customer image.

[0138] Specifically, if the result of step S403 indicates that the overlap value is greater than the overlap threshold, that is, if the item's movement trajectory and the person's movement trajectory highly overlap, the person image corresponding to the person's movement trajectory is determined as the image of the designated customer. This method quickly locates the designated customer when the overlap value is greater than the overlap threshold, improving the efficiency of designated customer detection and, consequently, the efficiency of handling left-behind items.

[0139] Optionally, if the overlap between the trajectory of multiple people and the trajectory of the item is no greater than a threshold, i.e., if the second video does not contain the corresponding customer of the left-behind item, a left-behind item reminder is generated based on the image of the customer and the left-behind item information. This left-behind item reminder can be displayed on a large screen at the target location and accompanied by a voice prompt to increase the likelihood of finding the customer and provide timely feedback to staff regarding the left-behind item.

[0140] S405 : For any customer image among the multiple customer images, match the attributed customer image with the customer image to obtain a matching value.

[0141] Specifically, the customer database stores multiple customer images. To determine the detailed information of the original customer, contact the original customer to return the left-behind items, and avoid losses to the customer, the original customer image is matched with the customer image for each of the multiple customer images to obtain a matching value. This matching value indicates the likelihood that the customer in the customer image is the original customer.

[0142] For example, Figure 6 This is a schematic diagram of a matching model for a customer provided in an embodiment of the present application. Figure 6 As shown, feature extraction is performed on any customer image in the database, and feature extraction is performed on the attributed customer image to obtain customer features and attributed customer features, and then the cosine value result between the customer features and the attributed customer features is calculated, and the cosine value result is the above-mentioned matching value.

[0143] S406. Determine whether the matching value is greater than the matching threshold; if the matching value is greater than the matching threshold, execute the following step S407; if the matching values ​​of multiple customer images and the attributed customer image are not greater than the matching threshold, execute the following step S408.

[0144] Specifically, the matching values ​​between the plurality of customer images in the customer database and the attributed customer image are sequentially determined, and a determination is made as to whether the matching values ​​are greater than a matching threshold. Once the matching value is determined to be greater than the matching threshold, step S407 is executed to determine the attributed customer information. If the matching value is not greater than the matching threshold, indicating that the customer corresponding to the customer image is not the attributed customer, the matching value between the next customer image and the attributed customer image is determined, and a determination is made as to whether the matching value is greater than the matching threshold. This process continues until the matching values ​​between the plurality of customer images and the attributed customer image are no greater than the matching threshold, indicating that the attributed customer does not exist in the second video data, and step S408 is executed.

[0145] This method quickly and accurately screens out customer images that match the owned customer images through matching thresholds, improves the efficiency of detecting owned customers, and is conducive to timely determining the owned customers of left-behind items and returning the left-behind items to the owned customers, thereby improving the efficiency of handling left-behind items.

[0146] S407: Determine the customer information corresponding to the customer image as the owned customer information; and generate an item left behind reminder based on the owned customer information and the left behind item information.

[0147] Specifically, if the result of step S406 indicates that the match value is greater than the match threshold, the customer corresponding to the customer image is determined to be the original customer, and the original customer information associated with the original customer is extracted. Based on this original customer information and the information about the item left behind, a left-behind notification is generated. For example, in a bank branch, based on the matched customer identity information, combined with the transaction time and counter information, the bank's internal transaction system searches for the customer's recent transaction records. These transaction records include information such as transaction time, transaction amount, and transaction counter to assist in confirming the ownership and time of loss of the item.

[0148] For example, the above-mentioned reminder of left-behind items may be, for example: Dear Mr. Zhang San: You left a black umbrella in waiting area A on January 1, 2025. Please contact the staff of the business hall to collect it.

[0149] Optionally, the left-behind item reminder is sent to a display screen at the target location for display processing, and / or the left-behind item reminder is sent to relevant staff, and / or the left-behind item reminder is sent to the customer based on the customer information.

[0150] S408: Generate an item left behind reminder based on the customer image and the item left behind information.

[0151] Specifically, when the judgment result of the above step S406 indicates that the matching values ​​of multiple customer images and the belonging customer image are not greater than the matching threshold, that is, the corresponding belonging customer cannot be queried from the customer database, therefore, an item left behind prompt is generated based on the belonging customer image and the left item information.

[0152] Optionally, the left-behind reminder is sent to a display screen at a target location for display processing, and / or the left-behind reminder is sent to relevant staff.

[0153] The method for detecting left-behind items provided in this embodiment obtains multiple second image frames by performing frame extraction processing on the second video data, and performs recognition and analysis on the multiple second image frames to obtain the object movement trajectory of the left-behind item and multiple person movement trajectories, and then determines the overlap value between the person movement trajectory and the object movement trajectory for any person movement trajectory among the multiple person movement trajectories, and judges whether the overlap value is greater than the overlap threshold. When the overlap value is greater than the overlap threshold, the person image corresponding to the person movement trajectory is determined as the image of the owned customer. According to the overlap value between the person movement trajectory and the item movement trajectory, the owned customer is quickly and accurately located, and the automatic identification of the owned customer is realized, which helps to improve the processing efficiency of left-behind items. Furthermore, after determining the attributable customer, the attributable customer image is matched against any of the multiple customer images to obtain a matching value. A determination is then made as to whether the matching value is greater than a matching threshold. If the matching value is greater than the matching threshold, the customer information corresponding to the customer image is determined as the attributable customer information. A left-behind item notification is then generated based on the attributable customer information and the left-behind item information. If the matching value is less than the matching threshold, a left-behind item notification is generated based on the attributable customer image and the left-behind item information. This method automatically filters and outputs attributable customer information from the customer database, facilitating timely notification of the attributable customer to collect the corresponding left-behind item, thereby improving the efficiency of left-behind item handling and enhancing the attributable customer's service satisfaction with the target location.

[0154] Figure 7 This is a schematic diagram of the structure of the detection device for leftover items provided in this application, as shown in Figure 7 As shown, the device 70 for detecting leftover items provided in this embodiment includes:

[0155] An acquisition module 701 is configured to acquire a plurality of image frames corresponding to first video data, wherein the image frames carry a timestamp;

[0156] The processing module 702 is configured to perform mask processing on the invalid area in each of the image frames to obtain a target still image, wherein the invalid area includes: an image area corresponding to a moving object and a portrait area;

[0157] The processing module 702 is further configured to segment the plurality of target still images to obtain a plurality of segmentation images of left-behind items;

[0158] A determination module 703 is configured to determine information about a left-behind item based on the plurality of left-behind item segmentation graphs and corresponding timestamps, wherein the left-behind item information includes a left-behind duration and a left-behind start time;

[0159] The judging module 704 is configured to judge whether the remaining time is greater than a preset time;

[0160] The determining module 703 is further configured to determine, when the left-behind duration is greater than a preset duration, the customer image to which the left-behind item belongs from second video data, the second video data comprising: video data within a monitoring area corresponding to a first preset duration, the first preset duration comprising: the left-behind duration and / or a second preset duration before the left-behind start time;

[0161] The generating module 705 is configured to generate an item left behind reminder based on the customer image and the item left behind information.

[0162] In a possible implementation, the processing module 702 is further configured to extract features from the plurality of image frames to obtain feature data, wherein the feature data includes: displacement features;

[0163] The determining module 703 is further configured to determine a motion intensity feature of the current image frame based on displacement features corresponding to the current image frame and the first image frame, wherein the current image frame is a previous image frame adjacent to the first image frame, and the motion intensity feature includes: a plurality of motion intensity values;

[0164] The determining module 703 is further configured to determine, in each current image frame, a subregion having a motion intensity value greater than a preset threshold as a motion region, and determine a subregion having a motion intensity value less than a preset threshold as a static region;

[0165] The processing module 702 is further configured to perform mask processing on the motion region in each current image frame to obtain a first still image;

[0166] The processing module 702 is further configured to perform recognition analysis on each of the first still images to obtain at least one corresponding portrait region;

[0167] The processing module 702 is further configured to perform mask processing on the at least one portrait region to obtain a target still image corresponding to each of the first still images.

[0168] In a possible implementation, the apparatus further includes: a comparison module 706;

[0169] The processing module 702 is further configured to extract features of any target still image among the plurality of target still images to obtain a corresponding target feature map, wherein the target feature map includes a plurality of target features;

[0170] The determining module 703 is further configured to classify the plurality of target features, determine a target sub-region corresponding to at least one target feature belonging to the same category as a segmented region, and obtain a segmented image;

[0171] The acquisition module 701 is configured to acquire a scene image corresponding to the first video data, wherein the scene image is an image without moving objects, human images, or left-behind objects;

[0172] The processing module 702 is further configured to segment the scene image to obtain a scene segmentation image;

[0173] The comparison module 706 is configured to compare the segmented image with the scene segmented image to obtain at least one segmented region in the segmented image that does not match the scene segmented image, and determine the at least one segmented region as a segmentation map of the left-behind items.

[0174] In a possible implementation, the processing module 702 is further configured to perform feature extraction and classification processing on the plurality of the left-behind item segmentation images to obtain an item category and an item image of at least one of the left-behind items;

[0175] The determining module 703 is further configured to determine the leaving start time and leaving duration of each of the left-behind items based on the timestamps corresponding to the plurality of left-behind item segmentation graphs.

[0176] In a possible implementation, the processing module 702 is further configured to perform frame extraction processing on the second video data to obtain a plurality of second image frames, where the second image frames carry a timestamp;

[0177] The processing module 702 is further configured to perform recognition analysis on the plurality of second image frames to obtain movement trajectories of the left-behind objects and movement trajectories of a plurality of people;

[0178] The determining module 703 is further configured to determine, for any one of the plurality of person movement trajectories, an overlap value between the person movement trajectory and the object movement trajectory;

[0179] The judging module 704 is further configured to judge whether the overlap value is greater than an overlap threshold;

[0180] The determining module 703 is further configured to determine the person image corresponding to the person movement trajectory as the belonging customer image when the overlap value is greater than the overlap threshold.

[0181] In a possible implementation, a plurality of customer images are stored in the customer database, and the apparatus further includes: a matching module 707;

[0182] The matching module 707 is configured to match the attribution customer image with any customer image among the multiple customer images to obtain a matching value;

[0183] The judging module 704 is further configured to judge whether the matching value is greater than a matching threshold;

[0184] The determining module 703 is further configured to determine the customer information corresponding to the customer image as the attributed customer information if the matching value is greater than a matching threshold;

[0185] The generating module 705 is further configured to generate an item left behind reminder based on the owned customer information and the left item information.

[0186] In a possible implementation, the generation module 705 is further configured to generate an item left behind prompt based on the customer image and the left item information when the overlap values ​​of the multiple person movement trajectories and the item movement trajectories are not greater than an overlap threshold.

[0187] In a possible implementation, the generation module 705 is further configured to generate an item left behind reminder based on the attributed customer image and the left-behind item information when the matching values ​​between the multiple customer images and the attributed customer image are not greater than a matching threshold.

[0188] The device for detecting left-behind items provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effects are similar, and are not described in detail in this embodiment.

[0189] Figure 8 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 8As shown, the electronic device 80 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the device 80 also includes a communication component 803. The processor 801, the memory 802 and the communication component 803 are connected via a bus 804.

[0190] During the specific implementation process, at least one processor 801 executes the computer-executable instructions stored in the memory 802, so that the at least one processor 801 performs the above method.

[0191] The specific implementation process of the processor 801 can be found in the above method embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.

[0192] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), etc. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in the present invention may be directly executed by a hardware processor or by a combination of hardware and software modules within the processor.

[0193] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage.

[0194] A bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be categorized as address buses, data buses, and control buses. For ease of illustration, the buses in the drawings of this application are not limited to just one bus or just one type of bus.

[0195] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0196] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0197] The readable storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0198] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0199] The division of units is merely a logical functional division; actual implementations may employ alternative divisions, such as combining or integrating multiple units or components into another system, or omitting or disabling certain features. Furthermore, any direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units, either through an interface, electrical, mechanical, or other means.

[0200] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0201] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0202] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0203] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented using hardware associated with program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0204] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.

[0205] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0206] It should be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0207] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.

[0208] If an integrated unit / module is implemented in hardware, the hardware may be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0209] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: U disk, read-only memory (ROM), random access memory (RAM), mobile hard disk, magnetic disk, or optical disk, etc., various media that can store program code.

[0210] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0211] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0212] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for detecting leftover items, characterized in that: include: Acquire multiple image frames corresponding to the first video data, where the image frames carry timestamps; Performing mask processing on the invalid area in each of the image frames to obtain a target still image, wherein the invalid area includes: an image area corresponding to the moving object and a portrait area; Segmenting the plurality of target still images to obtain a plurality of left-behind item segmentation images; Determining the left-behind item information based on the plurality of left-behind item segmentation graphs and corresponding timestamps, wherein the left-behind item information includes: a left-behind duration and a left-behind start time; Determine whether the remaining time is greater than a preset time; When the left-behind duration is greater than the preset duration, the customer image of the left-behind item is determined from the second video data, and an item left-behind prompt is generated based on the customer image and the left-behind item information. The second video data includes: video data within the monitoring area corresponding to the first preset duration, and the first preset duration includes: the left-behind duration, and / or the second preset duration before the start time of the left-behind.

2. The method according to claim 1, characterized in that The performing mask processing on the invalid area in each of the image frames to obtain a target still image includes: Extracting features from the plurality of image frames to obtain feature data, the feature data including: displacement features; Determining a motion intensity feature of the current image frame based on displacement features corresponding to the current image frame and the first image frame, wherein the current image frame is a previous image frame adjacent to the first image frame, and the motion intensity feature includes: a plurality of motion intensity values; Determining, in each current image frame, a subregion whose motion intensity value is greater than a preset threshold as a motion region, and determining a subregion whose motion intensity value is less than the preset threshold as a static region; Performing mask processing on the motion area in each current image frame to obtain a first still image; Performing recognition analysis on each of the first still images to obtain at least one corresponding portrait area; Mask processing is performed on the at least one portrait area to obtain a target still image corresponding to each of the first still images.

3. The method according to claim 1, characterized in that The step of segmenting the plurality of target still images to obtain a plurality of segmented image domains of left-behind objects includes: For any target still image among the plurality of target still images, extract features of the target still image to obtain a corresponding target feature map, wherein the target feature map includes a plurality of target features; Classifying the plurality of target features, determining a target subregion corresponding to at least one target feature belonging to the same category as a segmentation region, and obtaining a segmentation image; Acquire a scene image corresponding to the first video data, wherein the scene image is an image without moving objects, human images, or left-behind objects; Performing segmentation processing on the scene image to obtain a scene segmentation image; The segmented image is compared with the scene segmented image to obtain at least one segmented region in the segmented image that does not match the scene segmented image, and the at least one segmented region is determined as a segmentation map of the left-behind items.

4. The method according to claim 1, wherein The determining of the left-behind item information based on the plurality of left-behind item segmentation graphs and corresponding timestamps includes: Performing feature extraction and classification processing on the plurality of the left-behind item segmentation images to obtain an item category and an item image of at least one of the left-behind items; Based on the timestamps corresponding to the plurality of the left-behind item segmentation graphs, the leaving start time and the leaving duration of each left-behind item are determined.

5. The method according to claim 1, wherein Determining the customer image to which the left-behind item belongs from the second video data includes: Performing frame extraction processing on the second video data to obtain a plurality of second image frames, wherein the second image frames carry a timestamp; performing recognition analysis on the plurality of second image frames to obtain movement trajectories of the left-behind objects and movement trajectories of a plurality of people; For any one of the plurality of person movement trajectories, determining an overlap value between the person movement trajectory and the object movement trajectory; Determining whether the overlap value is greater than an overlap threshold; When the overlap value is greater than the overlap threshold, the person image corresponding to the person movement trajectory is determined as the belonging customer image.

6. The method according to claim 1, characterized in that A plurality of customer images are stored in the customer database. The generating of the item left behind reminder based on the customer image and the item left behind information includes: For any customer image among the multiple customer images, matching the attribution customer image with the customer image to obtain a matching value; Determining whether the matching value is greater than a matching threshold; In a case where the matching value is greater than a matching threshold, determining the customer information corresponding to the customer image as the attributed customer information; An item left behind reminder is generated based on the owned customer information and the left item information.

7. The method according to claim 5, characterized in that The method further comprises: When the overlap values ​​of the plurality of person movement trajectories and the object movement trajectory are not greater than an overlap threshold, an item left behind prompt is generated based on the attribution customer image and the left-behind item information.

8. The method according to claim 6, characterized in that The method further comprises: When the matching values ​​between the plurality of customer images and the attributed customer image are not greater than a matching threshold, a left-behind item prompt is generated based on the attributed customer image and the left-behind item information.

9. A device for detecting leftover items, characterized in that: include: An acquisition module, configured to acquire a plurality of image frames corresponding to the first video data, wherein the image frames carry a timestamp; a processing module, configured to perform mask processing on an invalid area in each of the image frames to obtain a target still image, wherein the invalid area includes an image area corresponding to a moving object and a portrait area; The processing module is further configured to segment the plurality of target still images to obtain a plurality of segmentation images of left-behind items; A determination module, configured to determine information of a left-behind item based on the plurality of left-behind item segmentation graphs and corresponding timestamps, wherein the left-behind item information includes a left-behind duration and a left-behind start time; A judging module, configured to judge whether the remaining time is greater than a preset time; The determining module is further configured to determine, when the left-behind duration is greater than a preset duration, an image of the customer to which the left-behind item belongs from second video data, the second video data comprising: video data within a monitoring area corresponding to a first preset duration, the first preset duration comprising: the left-behind duration, and / or a second preset duration before the left-behind start time; A generating module is used to generate an item left behind reminder based on the customer image and the item left behind information.

10. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.

12. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when the computer program is executed by a processor.