Article loss detection method and system

By combining a dual-background model architecture and a face recognition model, the problems of high false alarm rate and low recognition accuracy in lost item detection in bus station environments are solved, achieving efficient and reliable identification of lost items and their owners.

CN121170660APending Publication Date: 2025-12-19CCCC FOURTH HARBOR ENG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511043896.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing lost item detection systems have a high false alarm rate in crowded bus station environments, and existing facial recognition algorithms struggle to handle complex situations such as passenger movement and occlusion, resulting in low accuracy in identifying lost items.

Method used

A dual-background model architecture is adopted. The first background model retains the long-term stable features of the video scene and filters out temporary interference, while the second background model captures short-term dynamic changes. Combined with the face recognition model, the owner is identified within a preset time range, and lost and found information is generated.

Benefits of technology

It improves the accuracy of lost item identification, avoids invalid tracing caused by false detections, and ensures the reliability and accuracy of owner identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170660A_ABST
    Figure CN121170660A_ABST
Patent Text Reader

Abstract

The invention provides an article loss detection method and system, and relates to the technical field of intelligent video monitoring. The method comprises the following steps: modeling a video environment to obtain a first background model and a second background model; identifying a video stream of a monitoring device based on the first background model, determining whether a new target article exists in the video environment, and if the new target article exists in the video environment, determining whether the target article is a lost article based on the second background model; and under the condition that the target article is a lost article, calling video data in a preset time range in the video stream, identifying a loser from the video data based on a face judgment model, and generating lost article finding information. The method can improve the identification accuracy of lost articles.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent video monitoring, in particular to an article loss detection method and system. BACKGROUND

[0002] In the existing technical solutions, article loss detection mainly faces several prominent technical problems. Traditional monitoring systems mostly rely on manual review of video playback to find lost articles, which is inefficient and prone to missing key information. Although some intelligent monitoring systems can detect stationary objects, the false positive rate is extremely high in a crowded bus station environment, often incorrectly identifying articles temporarily placed by passengers as abandoned articles. Existing face association algorithms often only consider data matching at a single spatiotemporal point, making it difficult to cope with complex situations such as passengers moving around and being obstructed in the station, thus having the problem of low accuracy in identifying lost articles. SUMMARY

[0003] In a first aspect, the present application provides an article loss detection method, comprising: modeling a video environment to obtain a first background model and a second background model; wherein the video environment is an environment monitored by a monitoring device, the first background model is obtained by training an initial model with a set first learning rate, the second background model is obtained by training an initial model with a set second learning rate, and the numerical value of the first learning rate is less than that of the second learning rate; identifying a video stream of the monitoring device based on the first background model to determine whether there is a newly appearing target article in the video environment, and if there is a newly appearing target article in the video environment, determining whether the target article is a lost article based on the second background model; in the case that the target article is a lost article, retrieving video data of a preset time range in the video stream, identifying the owner from the video data based on a face determination model, and generating a lost article claim information.

[0004] According to one embodiment of the present application, in the case that the target article is a lost article, video data in a preset time range in the video stream is called, a lost owner is identified from the video data based on a face determination model, and lost article claim information is generated, including: calling video data in a first time range in the video stream, identifying face data within a preset radius range centered on the lost article in the video data in the first time range; if multiple face data within the preset radius range is identified, determining that first face data closest to the lost article is a pending owner, establishing a first data set associated with the lost article and the first face data, and adding a vote counting identifier to the first data set; calling video data in a second time range in the video stream, identifying face data within a preset radius range centered on the lost article in the video data in the second time range; if multiple face data within the preset radius range is identified, determining that second face data closest to the lost article is a pending owner; if the second face data identified this time is determined to represent the same pending owner as the first face data, adding a vote counting identifier to the first data set; if the second face data identified this time is determined to represent a different pending owner from the first face data, establishing a second data set associated with the lost article and the second face data, and adding a vote counting identifier to the second data set; in the case that a preset number of identification and determination operations are completed, determining that face data corresponding to a target data set with the most vote counting identifiers is the face data of the owner.

[0005] According to one embodiment of the present application, in the case that the target article is a lost article, video data in a preset time range in the video stream is called, a lost owner is identified from the video data based on a face determination model, and lost article claim information is generated, further including: in the case that first face data closest to the lost article is determined to be a pending owner, determining whether the first face data meets face recognition quality, and if so, adding a determination identifier to the first face data; in the case that second face data closest to the lost article is determined to be a pending owner, determining whether the second face data meets face recognition quality, and if so, determining whether it is the same as the first face data, and if so, adding a determination identifier to the first face data, otherwise adding a determination identifier to the second face data; in the case that face data corresponding to a data set with the most determination identifiers is determined to be the face data of the owner, extracting target face data with the determination identifier from the target data set, and generating lost property claim information based on the target face data.

[0006] According to one embodiment of the present application, the method for generating the lost property information based on the target face data comprises: uploading the target face data and the image of the lost property to a cloud platform; generating a two-dimensional code of the lost property information based on the cloud platform, and sending the two-dimensional code to one or more display devices for display.

[0007] According to one embodiment of the present application, the method further comprises: obtaining a monitoring video stream of another video environment, and identifying face data in the monitoring video stream to determine whether there is face data matching the face data of the owner; if so, sending a prompt message; receiving a lost property information query instruction sent by the owner, and providing the owner with the related information of the lost property.

[0008] In a second aspect, the present application further provides an article loss detection system, comprising: a training module configured to model a video environment to obtain a first background model and a second background model; wherein the video environment is an environment monitored by a monitoring device, the first background model is obtained by training an initial model with a set first learning rate, the second background model is obtained by training the initial model with a set second learning rate, and the numerical value of the first learning rate is smaller than that of the second learning rate; a first identification module configured to identify a video stream of the monitoring device based on the first background model to determine whether there is a newly appeared target article in the video environment, and if so, determine whether the target article is a lost article based on the second background model; and a second identification module configured to, in the case that the target article is a lost article, retrieve video data of a preset time range in the video stream, identify an owner from the video data based on a face determination model, and generate a lost property information.

[0009] According to one embodiment of the present application, the system further comprises: a camera group having a master-backup switching function, which switches to obtain a video stream of a video environment by another camera in the case that one camera in the camera group has a fault.

[0010] According to one embodiment of the present application, the system further comprises: a display device configured to display the related information of the lost property; a mobile terminal system configured to obtain the related information of the lost property or send a lost property information query instruction; and a cloud platform configured to save the target face data and the image of the lost property, and generate a two-dimensional code of the related information of the lost property.

[0011] In a third aspect, the present application further provides an intelligent bus station, wherein the intelligent bus station is provided with the article loss detection system of the above embodiments.

[0012] In a fourth aspect, the present application further provides a smart city system, which comprises one or more smart bus stations according to the above-mentioned embodiments.

[0013] Compared with the prior art, the present application has the beneficial effects that: the dual-background model architecture adopted is used to identify objects in a video scene, the first background model can retain long-term stable features of the video scene by virtue of a lower learning rate, and can effectively filter temporary interference; the second background model captures short-term dynamic changes by virtue of a higher learning rate, and the complementary verification of the two can improve the reliability of the left-behind object determination. This dual-verification mechanism guarantees from the source that only objects that truly meet the characteristics of the left-behind object will trigger the subsequent owner identification process, thereby improving the accuracy of the lost object identification and avoiding invalid retracing caused by false detection. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 A step schematic diagram of the lost object detection method provided by the embodiment of the present application.

[0015] Figure 2 A composition schematic diagram of the lost object detection system provided by the embodiment of the present application. DETAILED DESCRIPTION

[0016] The present application will be further described in detail below in combination with test examples and specific embodiments. However, it should not be understood that the scope of the above-mentioned subject matter of the present application is limited to the following embodiments only, and any technology realized based on the content of the present application falls within the scope of protection of the present application.

[0017] In the description of the specific embodiments of the present application, the terms indicating the orientation or positional relationship of “up”, “down”, “left”, “right”, “center”, “inner”, “outer”, “side” and the like appear without special indication, which are expressed based on the orientation or positional relationship shown in the drawings, or are the orientation or positional relationship used when the product / device / apparatus is placed. These terms of orientation or positional relationship are only for the convenience of describing the present application scheme or simplifying the description in the specific embodiments, so as to facilitate the quick understanding of the scheme by the technicians, and cannot be understood as indicating or implying that a specific device / component / element must have a specific orientation, or be constructed and operated in a specific positional relationship, and therefore cannot be understood as a limitation on the present application.

[0018] In the description of the embodiments of the present application, the technical terms “first”, “second” and the like only distinguish one entity or operation from another entity or operation, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. In the description of the embodiments of the present application, the meaning of “multiple” is two and more than two, unless otherwise specifically limited.

[0019] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a common set of embodiments, of the application, differing embodiments can be described.

[0020] Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a common set of embodiments, of the application, differing embodiments can be described. Figure 1 Figure 1 Reference to“an embodiment” herein means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase“in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all referring to a common set of embodiments, of the application, differing embodiments can be described.

[0021] S1, model the video environment to obtain a first background model and a second background model.

[0022] In the embodiments of the application, the video environment is an environment monitored by a monitoring device, and the video environment can be specifically a passenger gathering scene such as a bus station, a high-speed rail station, or a subway station.

[0023] The initial model can be a Gaussian Mixture Model (GMM). The first background model is obtained by training the initial model with a set first learning rate, and the second background model is obtained by training the initial model with a set second learning rate. The value of the first learning rate is smaller than the value of the second learning rate.

[0024] S2, identify the video stream of the monitoring device based on the first background model, determine whether there is a newly appearing target item in the video environment, and if there is a newly appearing target item in the video environment, determine whether the target item is a lost item based on the second background model.

[0025] Since the first background model is obtained by training with a small learning rate, it quickly adapts to new samples and can be used for identifying new appearing items in a video scene. Since the second background model is obtained by training with a large learning rate, it slowly adapts to new samples and can only be determined as part of the background after staying in the video scene for a long time. Therefore, it can be used to determine whether a newly appearing item is a lost item.

[0026] S3, in the case that the target item is a lost item, retrieve video data of a preset time range in the video stream, identify the owner from the video data based on a face determination model, and generate a lost item claim information.

[0027] For example, the specific way of S3 to identify the owner can include:

[0028] ​retrieve video data in a first time range in the video stream, and identify face data within a preset radius range centered on the lost item in the video data in the first time range;

[0029] If multiple face data are identified within the preset radius range, determine first face data closest to the lost item as a pending owner, establish a first data set associated with the lost item and the first face data, and add a voting identifier to the first data set;

[0030] retrieve video data in a second time range in the video stream, and identify face data within a preset radius range centered on the lost item in the video data in the second time range;

[0031] If multiple face data are identified within the preset radius range, determine second face data closest to the lost item as a pending owner, and if the second face data identified this time is determined to represent the same pending owner as the first face data, add a voting identifier to the first data set; if the second face data identified this time is determined to represent a different pending owner from the first face data, establish a second data set associated with the lost item and the second face data, and add a voting identifier to the second data set;

[0032] In a case where a preset number of identification and determination operations are completed, determine face data corresponding to a target data set with the most voting identifiers as face data of the owner.

[0033] In some other embodiments, the specific manner of S3 of identifying the owner can further include:

[0034] In a case where the first face data closest to the lost item is determined as a pending owner, determine whether the first face data meets face recognition quality, and if yes, add a determination identifier to the first face data;

[0035] In a case where the second face data closest to the lost item is determined as a pending owner, determine whether the second face data meets face recognition quality, and if yes, determine whether the second face data is the same as the first face data, and if yes, add a determination identifier to the first face data, otherwise, add a determination identifier to the second face data;

[0036] In a case where face data corresponding to a data set with the most determination identifiers is determined as face data of the owner, extract target face data with the determination identifier from the target data set, and generate lost property claim information based on the target face data.

[0037] For example, the application is applied to the lost article detection system of the intelligent bus station. When the passenger enters the bus station to wait for the bus, the camera of the bus station obtains the video stream and transmits the video stream to the video analysis server for detection. Once the article is identified as a lost article, the system will start a new thread to trace the owner of the lost article. Based on the large passenger flow density of the bus station, in order to eliminate misjudgment, the ticket counting evaluation method is used to trace the owner of the article:

[0038] When the system determines the lost article with the serial number NO. x, the determination time is defined as T0, the face determination model is loaded, and the video stream before T milliseconds (adjustable parameter) is called (T0-T time). The face data within the radius R range is identified around the current picture of the lost article. If multiple faces are detected in the set space, the closest face is temporarily determined as the owner of the article. The data set associated with the lost article and the face data is established. The face recognition quality evaluation algorithm is loaded. If the face recognition data meets the face recognition quality, it is identified as YES. If it is not qualified, it is identified as NO. Then, the video stream before 2T is called around the lost article, the face data is identified, and the face data is highly consistent with the face data at T time. The ticket of the face data is +1. If the face data is inconsistent, a new No. 2 person is associated with the lost article data set. The face data is YES or NO, and the record is recorded. The above process is repeated 10 times to find the person with the most votes as the owner of the article. The face data of the person is extracted to determine the YES face recognition picture. The lost article information is established and uploaded to the cloud platform system. The two-dimensional code of the lost article information is generated, and the information board of each bus station is displayed cyclically.

[0039] Further, the method provided by the application can further include the steps of collecting the lost article, specifically including:

[0040] Obtain the monitoring video stream of another video environment, and identify the face data in the monitoring video stream to determine whether there is face data matching the face data of the owner. If so, send a prompt message. Receive the lost article information query instruction sent by the owner, and provide the owner with the related information of the lost article.

[0041] For example, another video environment can be a bus stop of another site. When a passenger gets off the bus at the bus stop, the face data that matches the serial No. X lost property is captured, information is reminded through the speaker of the digital sign, and the lost property photo and information two-dimensional table are highlighted. The owner can obtain the relevant information of the lost property by scanning the two-dimensional code, including the lost time, associated pictures, bus stop location, camera link, etc. Through the camera link, the current real-time video stream of the bus stop is read to check whether the lost property is still there. If it has been taken away by others, it will be reported to the cloud management platform through one key. The management personnel will process the results (including the time of taking the property, the image at the time of taking the property, etc.) and send them to the email or mobile phone reserved by the reporter. If the owner is not captured by the bus station camera, the passenger can also scan the two-dimensional code or access the cloud platform for lost property inquiry and lost property claim by the management personnel for manual verification.

[0042] When the lost property is claimed, the face database is matched for recognition. If it is detected that it is the original owner, the taking time is automatically recorded, the taking pictures are taken from multiple angles, and are saved. If it is detected that the face data does not match the original owner, the digital sign will pop up a red box for warning and play a prompt sound "Please verify whether the personal belongings are correct" and the like, record the taking process, and transmit it to the cloud platform for manual verification by the administrator.

[0043] In summary, the dual background model architecture adopted in the scheme provided by the embodiment of the present application identifies the object in the video scene. The first background model can retain the long-term stable features of the video scene by virtue of a lower learning rate, and can effectively filter temporary interference. The second background model captures short-term dynamic changes through a higher learning rate. The complementation and verification of the two can improve the reliability of the left property determination. This dual verification mechanism guarantees from the source that only the target that truly meets the characteristics of the left property can trigger the subsequent owner identification process, thereby improving the accuracy of the lost property identification and avoiding the invalid tracing caused by false detection.

[0044] In the owner identification link, the system can establish a space-time correlation analysis framework through video data backtracking in a preset time range. The face judgment model is not used to analyze a single frame of image in isolation, but to intelligently analyze the continuous video stream in the period before and after the appearance of the target article. This time sequence correlation analysis method can effectively capture the actual interaction process between the owner and the article. The system particularly focuses on selecting the key time window before and after the article is left, which not only avoids irrelevant personnel interference caused by too early backtracking, but also prevents missing important information due to too late analysis. In addition, the owner identification link also implies a dynamic optimization face matching strategy. By analyzing the face data within a certain radius around the article, the system can intelligently filter the most possible associated objects, rather than relying on the simple principle of location proximity. The face quality evaluation link further ensures that the biological feature data used for comparison has sufficient identification value. This quality control mechanism greatly reduces the risk of false matching caused by factors such as image blur and obstruction, thereby improving the accuracy of owner identification.

[0045] Based on the same application concept, the embodiments of the present application also provide an article loss detection system, please refer to Figure 2 , Figure 2 The composition schematic diagram of the article loss detection system provided by the embodiments of the present application. The article loss detection system 20 can include:

[0046] The training module 21 is used to model the video environment to obtain a first background model and a second background model; wherein the video environment is an environment monitored by a monitoring device, the first background model is obtained by training an initial model with a set first learning rate, the second background model is obtained by training the initial model with a set second learning rate, and the numerical value of the first learning rate is less than the numerical value of the second learning rate;

[0047] The first identification module 22 is used to identify the video stream of the monitoring device based on the first background model, determine whether there is a newly appeared target article in the video environment, and if there is a newly appeared target article in the video environment, determine whether the target article is a lost article based on the second background model;

[0048] The second identification module 23 is used to retrieve the video data of a preset time range in the video stream in the case that the target article is a lost article, identify the owner from the video data based on a face judgment model, and generate a lost article claim information.

[0049] The camera group 24 has a main-backup switching function, and when one camera in the camera group fails, the other camera is switched to acquire the video stream of the video environment.

[0050] A display device 25 is configured to display the information related to the lost article. Specifically, the display device 25 can be an information digital signage, a display, or the like.

[0051] A mobile terminal system 26 is configured to acquire the information related to the lost article or send a query instruction of the lost article information.

[0052] A cloud platform 27 is configured to save the target face data and the image of the lost article, and generate a two-dimensional code of the information related to the lost article.

[0053] It should be understood that, in the above embodiments, the various modules of the system are provided when working, and only the division of the functional modules in the above description is exemplified. In actual application, the above functions can be distributed by different functional modules according to needs, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0054] Each functional module in the above embodiments can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for convenient distinction, and do not limit the protection scope of the embodiments of the present application.

[0055] Based on the same application concept, the embodiments of the present application also provide an intelligent bus station, wherein the intelligent bus station is provided with the article loss detection system in the above embodiments.

[0056] Based on the same application concept, the embodiments of the present application also provide a smart city system, wherein the smart city system comprises one or more intelligent bus stations in the above embodiments.

[0057] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An article loss detection method, characterized by, The application relates to a method for identifying a lost article in a video environment. The method comprises the following steps: modeling a video environment to obtain a first background model and a second background model; identifying a video stream of a monitoring device based on the first background model to determine whether a new target article exists in the video environment, and determining whether the target article is a lost article based on the second background model if the new target article exists in the video environment; 2. The method of claim 1, wherein, if the target article is a lost article, retrieving video data of a preset time range in the video stream, identifying a lost owner from the video data based on a face determination model, and generating a lost article claim information. The method for identifying a lost article in a video environment further comprises the following steps: retrieving video data of a first time range in the video stream, and identifying face data within a preset radius range around the lost article in the video data of the first time range; if multiple face data are identified within the preset radius range, determining first face data closest to the lost article as a pending lost owner, establishing a first data set associated with the lost article and the first face data, and adding a vote counting identifier to the first data set; retrieving video data of a second time range in the video stream, and identifying face data within a preset radius range around the lost article in the video data of the second time range; if multiple face data are identified within the preset radius range, determining second face data closest to the lost article as a pending lost owner, and adding a vote counting identifier to the first data set if the second face data identified this time is determined to represent the same pending lost owner as the first face data, or establishing a second data set associated with the lost article and the second face data and adding a vote counting identifier to the second data set if the second face data identified this time is determined to represent a different pending lost owner from the first face data; 3. The method of claim 2, wherein, after a preset number of identification and determination operations are completed, determining face data corresponding to a target data set with the most vote counting identifiers as face data of a lost owner. The method for identifying a lost article in a video environment further comprises the following steps: if the first face data closest to the lost article is determined to be a pending lost owner, determining whether the first face data meets face recognition quality requirements, and adding a determination identifier to the first face data if the first face data meets the face recognition quality requirements; if the second face data closest to the lost article is determined to be a pending lost owner, determining whether the second face data meets face recognition quality requirements, and adding a determination identifier to the second face data if the second face data meets the face recognition quality requirements. In a case where it is determined that the face data corresponding to the data set with the most determination identifiers of the vote is the face data of the owner of the lost item, target face data added with the determination identifier is extracted from the target data set, and lost and found information is generated based on the target face data.

4. The method of claim 3, wherein, The manner of generating the lost and found information based on the target face data includes: uploading the target face data and the image of the lost item to a cloud platform; generating a two-dimensional code of the information related to the lost item based on the cloud platform, and sending the two-dimensional code to one or more display devices for display.

5. The method of claim 1, wherein, The method further includes: obtaining a monitoring video stream of another video environment, and identifying face data in the monitoring video stream to determine whether there is face data matching the face data of the owner of the lost item; if so, sending a prompt message; receiving a lost item information query instruction sent by the owner, and providing the information related to the lost item to the owner.

6. An article loss detection system, characterized by, includes: a training module configured to model a video environment to obtain a first background model and a second background model; wherein the video environment is an environment monitored by a monitoring device, the first background model is obtained by training an initial model with a set first learning rate, and the second background model is obtained by training the initial model with a set second learning rate, and the numerical value of the first learning rate is less than the numerical value of the second learning rate; a first identification module configured to identify a video stream of the monitoring device based on the first background model to determine whether there is a newly appearing target item in the video environment, and if there is a newly appearing target item in the video environment, determine whether the target item is a lost item based on the second background model; a second identification module configured to, in a case where the target item is a lost item, retrieve video data of a preset time range in the video stream, identify the owner from the video data based on a face determination model, and generate lost item found information.

7. The system of claim 6, wherein, The system further includes: a camera group having a primary-backup switching function, when one camera in the camera group has a fault, the other camera is switched to obtain a video stream of a video environment.

8. The system of claim 6, wherein, The system further includes: a display device configured to display the information related to the lost item; a mobile sub-system configured to obtain the information related to the lost item or send a lost item information query instruction; a cloud platform configured to save the target face data and the image of the lost item, and generate a two-dimensional code of the information related to the lost item.

9. An intelligent bus stop, characterized in that, The smart bus stop is provided with the lost item detection system according to any one of claims 6 to 8.

10. A smart city system, characterized by, The smart city system includes one or more smart bus stops according to claim 9. The smart city system includes one or more smart bus stops according to claim 9.