Lost and found method and system based on intelligent bus platform
By using the trajectory matching scoring model on the smart bus platform and combining spatial trajectories and passenger behavior to perform structured modeling of lost property identification, the problem of low accuracy in lost property identification in existing technologies is solved, and higher identification precision and accuracy are achieved.
Patent Information
- Application Number
- CN202510861148.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing lost property recognition methods have the problem of low recognition accuracy on smart bus platforms, mainly because they ignore the correspondence between passenger riding behavior and object trajectory, resulting in excessive redundancy of candidate targets.
Through the trajectory matching scoring model, structural modeling is performed by combining the spatial trajectory of candidate objects in the car, the changes in static state and the passenger behavior trajectory. The correspondence between passenger behavior and object trajectory is clarified, a set of static objects is generated, and lost items are judged through semantic description and trajectory consistency analysis.
It improves the recognition accuracy of lost property identification and the accuracy of judgment results, reduces redundant data, and improves the accuracy of lost and found.
Smart Images

Figure CN120707888A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of smart platform technology, and in particular to a lost and found method and system based on a smart public transportation platform. Background Art
[0002] With the continuous development of intelligent transportation systems, public transportation platforms have gradually acquired the ability to integrate multi-source data, including onboard video capture, ticket information management, and in-car positioning. On this basis, efforts have been made to build lost and found modules to automatically identify and retrieve lost items. These modules use image recognition, semantic tagging, and time period analysis to assist passengers in retrieving their lost items.
[0003] There are deficiencies in the lost property identification technology in the prior art. Specifically, most of the existing lost property identification methods compare and judge lost property based on images or passenger descriptions. For example, the invention patent application with the publication date of August 5, 2022 and the publication number CN114863516A discloses a method and system for recovering lost property in public transportation based on image processing. The method obtains the first image information of the passenger entering the public transportation, and extracts the boarding data tuple based on the first image information; obtains the second image information of the passenger leaving the public transportation, and extracts the disembarkation data tuple based on the second image information; the boarding data tuple and the disembarkation data tuple both include the identity recognition feature data and carried item feature data of the corresponding passenger; then, the boarding data tuple of the passenger corresponding to the disembarkation data tuple is obtained based on the identity recognition feature data, and the carried item feature data in the boarding data tuple and the disembarkation data tuple are compared. If the carried item feature data in the boarding data tuple does not have a corresponding match in the disembarkation data tuple, it is determined that the passenger has lost the item when leaving the public transportation, and a reminder is given.
[0004] Although it can compare images of passengers getting on and off the bus, identify features that are lost when passengers get off the bus, and can use voice to remind passengers that they may have left behind items when they get off the bus, thereby increasing the chance of finding their items, it still has similar problems as other solutions in the existing technology. Specifically, it ignores the correspondence between riding behavior and object trajectory, resulting in a large amount of redundancy in candidate targets, and ultimately leads to low accuracy of judgment results, which further leads to low accuracy of lost and found. Summary of the Invention
[0005] Based on this, it is necessary to provide a lost and found method and system based on a smart bus platform that can structure the spatial trajectory, static state changes and passenger behavior trajectory of candidate objects in the car through a trajectory matching scoring model, clarify the correspondence between passenger behavior and object trajectory, improve the recognition accuracy of lost and found identification, and improve the accuracy of judgment results.
[0006] The technical solutions of the present invention are as follows: A lost and found method based on a smart public transportation platform, the method comprising: Obtaining target lost item information of the lost item, analyzing the target lost item, and generating a stationary object set; Generate semantic image descriptions based on a set of static objects, and generate a candidate lost object description set based on a preset semantic description generation function; Perform description similarity matching based on the candidate lost property description set to obtain a set of suspected targets; Based on a preset trajectory matching scoring model, trajectory consistency analysis is performed on the suspected target set to obtain a lost property judgment result.
[0007] Specifically, target lost item information of the lost item is obtained, the target lost item is analyzed, and a stationary object set is generated, including: Obtaining target lost property information of the lost item, and generating a target fragment set according to the target lost property information; Image feature analysis is performed based on the target segment set to obtain a stationary object set.
[0008] Specifically, image feature analysis is performed on the target segment set to obtain a stationary object set, including: Extracting spatial feature vectors based on the target segment set, and calculating spatial similarity of spatial feature vectors between frame images based on the spatial feature vectors; obtaining an image matrix based on the target segment set and generating a temporal stability score based on the image matrix; generating a spatiotemporal stability score of the lost object according to the spatial similarity and the temporal stability score; A set of stationary objects is obtained according to the spatiotemporal stability scores.
[0009] Specifically, semantic image descriptions are generated based on a set of stationary objects, and a candidate lost object description set is generated based on a preset semantic description generation function, including: Performing image feature extraction based on the target segment set to generate an image feature vector of a stationary object; Generate a matching weight of an anchor point according to the two-dimensional position coordinates of the stationary object, wherein the anchor point is pre-configured; Set the semantic label of the anchor point with the largest weight as the semantic reference of the stationary object; Based on the semantic description generation function, a sentence description is generated according to the semantic label of the image feature vector and the anchor point. A candidate lost property description set is generated according to the sentence description.
[0010] Specifically, description similarity matching is performed based on the candidate lost property description set to obtain a set of suspected targets, including: Obtain the descriptive information provided by the passenger when requesting lost property report; Based on a preset semantic encoder, description similarity matching is performed on the candidate lost property description set and the description information at the time of declaration to obtain a suspected target set.
[0011] Specifically, based on a preset trajectory matching scoring model, trajectory consistency analysis is performed on the set of suspected targets to obtain a lost property judgment result, including: generating a probability of object continuous stationary state, a stability score, and a spatial overlap degree according to the set of suspected targets; Based on a preset trajectory matching scoring model, a trajectory consistency score is generated according to the object's probability of continuous stationary state, the stability score, and the spatial overlap; The lost property judgment result is obtained based on the trajectory consistency score.
[0012] Specifically, the lost property judgment result is obtained based on the trajectory consistency score, including: Determining whether the trajectory consistency score is greater than a preset high confidence threshold; If the judgment is yes, the candidate object is judged to be a high-credible lost object, and the high-credible lost objects are aggregated to generate a lost object judgment result.
[0013] Specifically, a lost and found system based on a smart public transportation platform is also provided, the system comprising: A stationary object generation module is used to obtain target lost object information of lost items, analyze the target lost objects, and generate a stationary object set; A lost object description generation module is used to generate semantic image descriptions based on a set of stationary objects and generate a candidate lost object description set based on a preset semantic description generation function; A suspected target generation module is used to perform description similarity matching based on the candidate lost property description set to obtain a suspected target set; The judgment result generation module is used to perform trajectory consistency analysis on the suspected target set based on a preset trajectory matching scoring model to obtain a lost property judgment result.
[0014] Specifically, the stationary object generation module is further used to: obtain target lost object information of the lost item, and generate a target segment set according to the target lost object information; perform image feature analysis on the target segment set to obtain a stationary object set.
[0015] Specifically, the stationary object generation module is also used to: extract spatial feature vectors based on the target segment set, and calculate the spatial similarity of spatial feature vectors between frame images based on the spatial feature vectors; obtain an image matrix based on the target segment set, and generate a temporal stability score based on the image matrix; generate a spatiotemporal stability score of the lost object based on the spatial similarity and the temporal stability score; and obtain a stationary object set based on the spatiotemporal stability score.
[0016] Specifically, the lost property description generation module is also used to: perform image feature extraction based on the target fragment set to generate an image feature vector of a stationary object; generate a matching weight of an anchor point based on the two-dimensional position coordinates of the stationary object, wherein the anchor point is pre-configured; set the semantic label of the anchor point with the largest weight as the semantic reference of the stationary object; generate a sentence description based on the image feature vector and the semantic label of the anchor point based on a semantic description generation function, and generate a candidate lost property description set based on the sentence description.
[0017] Specifically, the suspected target generation module is further used to: obtain the declaration description information provided by the passenger when requesting lost property declaration; perform description similarity matching on the candidate lost property description set and the declaration description information based on a preset semantic encoder to obtain a suspected target set.
[0018] Specifically, the judgment result generation module is also used to: generate the probability of the object remaining still, the smoothness score and the spatial overlap based on the suspected target set; based on the preset trajectory matching scoring model, generate a trajectory consistency score according to the probability of the object remaining still, the smoothness score and the spatial overlap; and obtain the lost property judgment result according to the trajectory consistency score.
[0019] Specifically, the judgment result generation module is further used to: determine whether the trajectory consistency score is greater than a preset high confidence threshold; if so, determine that the candidate object is a high-confidence lost property, and summarize the high-confidence lost properties to generate a lost property judgment result.
[0020] Optionally, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps described in the above-mentioned lost and found method based on the smart bus platform are implemented.
[0021] Optionally, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the steps described in the lost and found method based on the smart bus platform are implemented.
[0022] The present invention relates to machine learning and deep learning technologies, and the technical effects achieved are as follows: The method obtains target lost item information of lost items, analyzes the target lost items, and generates a set of stationary objects; generates semantic image descriptions based on the stationary object set, and generates a candidate lost item description set based on a preset semantic description generation function; performs description similarity matching on the candidate lost item description set to obtain a suspected target set; performs trajectory consistency analysis on the suspected target set based on a preset trajectory matching scoring model to obtain a lost item judgment result; and then implements structured modeling of the spatial trajectory, stationary state change, and passenger riding behavior trajectory of the candidate objects in the car through the trajectory matching scoring model, clarifies the correspondence between the passenger riding behavior and the object trajectory, and judges the spatiotemporal correlation between the object and the passenger through the scoring mechanism, thereby reducing redundant data, improving the recognition accuracy of lost item recognition, and improving the accuracy of the judgment result. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 1 is a flow chart of a lost and found method based on a smart public transportation platform in one embodiment; Figure 2 This is a structural block diagram of a lost and found system based on a smart public transportation platform in one embodiment. DETAILED DESCRIPTION
[0024] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0025] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0026] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0027] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0028] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0030] In one embodiment, a terminal is provided, configured to: obtain target lost property information of a lost item, analyze the target lost property, and generate a set of stationary objects; generate a semantic image description based on the set of stationary objects, and generate a candidate lost property description set based on a preset semantic description generation function; perform description similarity matching on the candidate lost property description set to obtain a suspected target set; and perform trajectory consistency analysis on the suspected target set based on a preset trajectory matching scoring model to obtain a lost property judgment result.
[0031] The terminal may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices.
[0032] In one embodiment, Figure 1 As shown, a lost and found method based on a smart public transportation platform is provided, the method comprising: Step S100: obtaining target lost item information of the lost item, analyzing the target lost item, and generating a stationary object set; Step S200: generating semantic image descriptions based on the set of stationary objects, and generating a candidate lost property description set based on a preset semantic description generation function; Step S300: performing description similarity matching based on the candidate lost property description set to obtain a set of suspected targets; Step S400: performing trajectory consistency analysis on the set of suspected targets based on a preset trajectory matching scoring model to obtain a lost property judgment result.
[0033] The present application obtains target lost property information of lost items, analyzes the target lost properties, and generates a set of stationary objects; generates semantic image descriptions based on the stationary object set, and generates a candidate lost property description set based on a preset semantic description generation function; performs description similarity matching on the candidate lost property description set to obtain a suspected target set; performs trajectory consistency analysis on the suspected target set based on a preset trajectory matching scoring model to obtain a lost property judgment result; and then realizes structured modeling of the spatial trajectory, stationary state change, and passenger riding behavior trajectory of the candidate objects in the car through the trajectory matching scoring model, clarifies the correspondence between the passenger riding behavior and the object trajectory, and judges the spatiotemporal correlation between the object and the passenger through the scoring mechanism, thereby reducing redundant data, improving the recognition accuracy of lost property identification, and improving the accuracy of the judgment result.
[0034] In one embodiment, step S100: obtaining target lost item information of a lost item, analyzing the target lost item, and generating a set of stationary objects includes: Step S110: obtaining target lost property information of the lost item, and generating a target segment set according to the target lost property information; Step S120: performing image feature analysis based on the target segment set to obtain a stationary object set.
[0035] In this embodiment, in order to clarify the type of lost item and accurately limit the temporal and spatial range in which the lost item may appear, and to provide accurate boundary conditions and screening basis for the subsequent video frame segment extraction, image feature analysis and trajectory consistency analysis steps, the target lost item information of the lost item is obtained, and a target segment set is generated based on the target lost item information.
[0036] After losing an item, the passenger enters a description of the target lost item through the smart bus platform. The description includes the type of lost item, boarding time, disembarkation time, bus route number, boarding and alighting stops, and the location of the lost item. The type of lost item is the passenger's description of the type and status of the lost item, for example, a large black backpack containing something, and the location of the lost item can be described as the middle of the car or the back seat.
[0037] Based on the input bus route number, boarding time, alighting time, and boarding and disembarking stations, the platform queries the vehicles running on the route during the corresponding time period and determines the unique corresponding vehicle number, ensuring the uniqueness and accuracy of the video data extraction in the subsequent steps. After confirming the vehicle number, the platform retrieves the vehicle's running trajectory data and uses the GPS trajectory playback technology in the existing technology to obtain the complete driving path of the vehicle from the boarding time to the disembarking time, confirming the vehicle's stops. The driving path data includes a timestamp per second, the GPS coordinates at the corresponding time, the speed at the corresponding time, and the docking status at the corresponding time. Based on the number and installation position of the camera in the vehicle, a car compartment space structure diagram is established, and the location of the lost item described by the passenger is mapped to one or more specific camera viewing areas. The camera number set corresponding to the lost item location description in the car is confirmed, and the target lost item information is finally obtained, including the vehicle number, the complete driving path data during the time period from the boarding time to the disembarking time, and the camera number set corresponding to the lost item location description.
[0038] Next, based on the target lost item information, the system extracts onboard video frame segments to obtain a target segment set. This extraction process first determines the vehicle number and the frame index range of the onboard surveillance video corresponding to the time period based on the target lost item information. Then, from the onboard surveillance video stored in the vehicle's multi-channel video storage system, the system accurately extracts the video frame segments corresponding to the time period from the boarding time to the disembarkation time and the camera number. Specifically, this technology utilizes multi-channel time-synchronized index retrieval technology based on the vehicle's DVR system, which provides high stability and precise time alignment capabilities. First, multi-channel camera recognition and channel mapping are performed. The bus's internal video acquisition system consists of multiple fixed cameras, which are often connected to an onboard DVR to form a channel index table. Each camera channel corresponds to a fixed mounting point and is numbered and mapped to the vehicle's spatial model. Based on the obtained camera number, the video extraction scope is limited to the perspectives relevant to the passenger's description, avoiding redundant processing of irrelevant data, improving efficiency, and reducing the difficulty of subsequent recognition. Next, time segment extraction and frame index construction are performed. The onboard DVR system provides high-precision timestamps for each video channel. The system maps the corresponding time period into a frame index interval for each channel. Leveraging the DVR's synchronized index structure, the system achieves millisecond-level alignment, ensuring that the extracted video frames fully cover the window where the lost item may have appeared. Data segmentation and frame structure generation are then performed. The system extracts continuous frames from each camera channel at a fixed frame rate (e.g., 25 fps), forming a frame sequence within the corresponding time period and appending the camera number and capture time of each frame. Furthermore, information such as the average brightness between frames for abnormal occlusion detection and the inter-frame motion estimation to assist in determining stationary state can be added. Finally, a target segment set is output, including the camera number, corresponding time period, and the frame sequence within that time period, providing a high-precision, low-redundancy video foundation for image feature extraction. The target segment set includes multiple target segments, each of which includes multiple images. Each target segment has a corresponding time period, and each image corresponds to a camera number and a corresponding frame in the frame sequence within the corresponding time period. That is, each frame within the corresponding time period of each target segment corresponds to an image. Image feature analysis is then performed on the target segment set to obtain a set of stationary objects.
[0039] In one embodiment, step S120: performing image feature analysis on the target segment set to obtain a stationary object set includes: Step S121: extracting spatial feature vectors based on the target segment set, and calculating spatial similarity of spatial feature vectors between frame images based on the spatial feature vectors; Step S122: obtaining an image matrix according to the target segment set, and generating a temporal stability score according to the image matrix; Step S123: generating a spatiotemporal stability score of the lost object according to the spatial similarity and the temporal stability score; Step S124: obtaining a set of stationary objects according to the spatiotemporal stability scores.
[0040] In this embodiment, image feature analysis includes performing image feature analysis based on the target segment set using a spatiotemporal stability analysis model to identify stationary objects and obtain a stationary object set. The stationary object set contains relevant information about all stationary objects, including the object's image region feature vector, frame position information, position coordinate information, and associated camera number.
[0041] Specifically, the spatial feature vector of each frame image of the target segment set is calculated. Extraction, through the convolutional neural network CNN model, extract color, texture, shape and other features, and calculate the spatial similarity of spatial feature vectors between frame images ,The convolutional neural network CNN model can use ResNet-50, use fixed parameters, and do not perform training to ensure stability and reproducibility.
[0042] The spatial similarity is calculated as follows:
[0043] It is the spatial similarity between the spatial feature vectors extracted from the k-th frame and the k+1-th frame image. The higher the similarity, the more stable the object is in spatial position and it may be a stationary object. is the spatial feature vector extracted from the k-th frame image, is the spatial feature vector extracted from the k+1th frame image. is the Euclidean distance, i.e. the L2 norm.
[0044] Next, based on the temporal stability function, an image matrix is obtained according to the target segment set, and a temporal stability score is generated according to the image matrix. The temporal stability function is as follows:
[0045] in, It is the temporal stability score. The higher the score, the more stable the object is during the time period. It may be a stationary object. is the maximum step size for calculating the difference between the previous and next frames, is the image matrix of the i-th segment, j-th camera, and k+n-th frame, which represents the difference with the image matrix of the k-th frame in the time dimension. is the image matrix of the i-th segment, j-th camera, and k-th frame, It is the Euclidean distance of the grayscale value difference of all pixels between the kth frame and the k+nth frame image. It is obtained by squaring, summing and square rooting the grayscale difference of each pixel, and reflects the overall brightness difference between the two frames of images. is the Euclidean distance, i.e. the L2 norm.
[0046] Based on the spatiotemporal stability analysis model, the spatiotemporal stability score of the lost object is generated according to the spatial similarity and the temporal stability score. The spatiotemporal stability analysis model is as follows:
[0047] in, is the spatiotemporal stability score of the object, which is used to determine whether it is a stationary object. is the time stability function, is the image matrix of the i-th segment, j-th camera, and k-th frame, is the spatial feature vector extracted from the k-th frame image, is the spatial feature vector extracted from the k+1th frame image, Spatial similarity function between spatial feature vectors. It is a weight coefficient used to balance temporal stability and spatial similarity. The range of the weight coefficient is [0, 1]. It is set according to whether the monitoring image has occlusion, light interference, clarity and stability. If there is occlusion or window light reflection, the weight is reduced (such as 0.3 to 0.5) to reduce the sensitivity to brightness. If the picture clarity is poor and it is difficult to extract spatial features, the weight is increased (such as 0.6 to 0.8) to reduce the impact of spatial similarity.
[0048] In the existing technology, the object type is usually determined by relying on a single spatial feature, such as color, texture, shape, etc., while ignoring the temporal information in the video, especially the key factor of whether the object is stationary. In this embodiment, the temporal stationary and spatial characteristics of the object are simultaneously considered, and the spatial characteristics of each frame of the image are combined with the temporal stability of the object, that is, whether the object remains stationary in multiple frames, to perform multi-dimensional feature fusion of the image, thereby improving the detection accuracy of stationary objects and enhancing the accuracy and robustness of object recognition.
[0049] In one embodiment, step S200: generating semantic image descriptions based on a set of stationary objects and generating a candidate lost object description set based on a preset semantic description generation function includes: Step S210: extracting image features based on the target segment set to generate an image feature vector of the stationary object; Step S220: generating a matching weight of an anchor point according to the two-dimensional position coordinates of the stationary object, wherein the anchor point is pre-configured; Step S230: setting the semantic label of the anchor point with the largest weight as the semantic reference of the stationary object; Step S240: Generate a sentence description based on the image feature vector and the semantic label of the anchor point based on the semantic description generation function. Step S250: Generate a candidate lost property description set based on the statement description.
[0050] In this embodiment, when generating semantic image descriptions, the vehicle spatial structure anchor points are first constructed, and a descriptive semantic label is added to the spatial structure anchor points in the vehicle. Each anchor point contains the two-dimensional position coordinates of the anchor point and the semantic label of the anchor point (for example, "front door", "right side of the seat", "ground at the rear of the vehicle"). Then, the anchor point matching weight function is used to associate and match the stationary object with the nearest anchor point, so that the semantic label of the anchor point serves as the semantic reference of the stationary object. Finally, the semantic description generation function is used to jointly encode the image features of the stationary object and the semantic label of the anchor point to generate a semantic image description. All semantic image descriptions are grouped together to obtain a candidate lost property description set.
[0051] Specifically, when constructing the vehicle space structure anchor point, the vehicle space structure anchor point set is ,in, It is a spatial structure anchor point in the set, which contains the two-dimensional position coordinates of the anchor point and the semantic label information of the anchor point. Specifically, , in, is the two-dimensional position coordinate of the anchor point j, that is, the position of the center point of the object in the image coordinate system, is the semantic label of anchor point j, such as "front door", "right side of seat", "rear ground", which is used to describe the position of the anchor point in the vehicle spatial structure. j is the index of the spatial structure anchor point.
[0052] Furthermore, image feature extraction is performed on the target segment set to generate an image feature vector of the stationary object. After obtaining the target segment set, image feature extraction is performed on each frame image in the segment set through a convolutional neural network (CNN) model such as ResNet-50 to obtain an image feature vector of the stationary object. , and then generate the matching weight of the anchor point according to the two-dimensional position coordinates of the static object , where the anchor point is pre-configured. Matching weight Generated based on the following anchor matching weight function:
[0053] in, is the matching weight between the static object and the anchor point, is the two-dimensional position coordinate of the stationary object, It is the square of the Euclidean distance between the object and the anchor point. The smaller it is, the shorter the distance between the object and the anchor point is. The distance attenuation control factor is used to control the distance sensitivity when matching between a stationary object and an anchor point. It is usually a fixed value set according to the expected spatial tolerance range (in pixels) near the anchor point. For example, for high-definition images, a larger spatial drift can be tolerated, so the spatial tolerance range is set to 60, and the distance attenuation control factor is 1 / 60. 2 In small target detection, in order to improve accuracy, the spatial tolerance range is set to 20, and the distance attenuation control factor is 1 / 20. 2 By setting a negative sign in the exp function for calculation, the static object is associated with the nearest anchor point. The closer the distance, the greater the weight of the anchor point. Then the semantic label of the anchor point with the largest weight is set as the semantic reference of the static object; based on the semantic description generation function, a sentence description is generated according to the image feature vector and the semantic label of the anchor point. .
[0054] Among them, the function generates a statement description based on the semantic description :
[0055]
[0056] in, Is a statement description, is the image feature vector of the stationary object, is the semantic label of the anchor with the largest anchor matching weight, It is an image feature mapping function used to convert the image features of static objects into semantic embedding vectors, which is achieved by linear transformation plus nonlinear activation (such as ReLU). It is an anchor tag embedding function, which is used to convert the semantic label of the anchor into an embedding vector, which is implemented by word vector embedding (such as Word2Vec). is the operation representing vector concatenation, It is a language generation decoder used to generate the final description text.
[0057] Therefore, by and semantic labels of anchors After embedding, concatenation is performed and input into the language generation decoder Generate statement description ,For example, “a blue handbag is placed under the 4th row of seats”, a description is generated for each static object to obtain a candidate lost property description set.
[0058] In this embodiment, semantic descriptions are generated for stationary objects by associating and matching the stationary objects with the nearest anchor points, combined with the image features of the stationary objects. Existing image description generation methods, such as those based on CNN+LSTM or Transformer image caption generation models, often have problems in understanding the structural spatial location of the carriage (such as "next to the front door" or "under the seat") in specific closed environments such as train compartments, making it difficult to generate practical descriptions useful to passengers. They can only generate descriptions such as "a bag", lacking spatial adaptability and practicality. However, this application introduces a set of spatial structural anchor points, which has position perception capabilities, adapts to closed spaces, improves spatial adaptability, and generates descriptive information useful to passengers through semantic image description, such as descriptions such as "a black backpack near the back door", which improves practicality.
[0059] In one embodiment, step S300: performing description similarity matching based on the candidate lost property description set to obtain a set of suspected targets includes: Step S310: Obtain the description information provided by the passenger when requesting to report lost property; Step S320: Based on a preset semantic encoder, description similarity matching is performed on the candidate lost property description set and the description information at the time of declaration to obtain a suspected target set.
[0060] In this embodiment, each description in the candidate lost property description set corresponds to an identified stationary object. This description is then compared to the description provided by the passenger when initiating a lost property report on the smart public transportation platform using the semantic encoder SBERT (Sentence-BERT), combined with cosine similarity calculations. Based on the similarity matching results, all stationary object descriptions that meet the criteria (i.e., those with a similarity greater than a set threshold) are added to the set of suspected objects. If necessary, a maximum number of descriptions can be set, and the objects are sorted from high to low in terms of similarity to control computing resource consumption. The threshold is set empirically; for example, a similarity greater than 0.75 is generally considered a suspected object. The semantic encoder SBERT (Sentence-BERT) is a mature and stable natural language semantic representation method with good generalization capabilities and efficient online computing. By encoding the passenger description and the lost property description to obtain a semantic vector, cosine similarity is calculated to determine the similarity between the two descriptions. This method can capture contextual semantics, support complex description expressions, and adapt to different passenger language expressions. The model parameters are fixed and can run without complex tuning. It has good generalization ability and efficient online computing characteristics.
[0061] In one embodiment, step S400: performing trajectory consistency analysis on the set of suspected targets based on a preset trajectory matching scoring model to obtain a lost property judgment result includes: Step S410: generating an object's probability of continuous stationary state, a stability score, and a spatial overlap degree according to the set of suspected targets; Step S420: Based on a preset trajectory matching scoring model, a trajectory consistency score is generated according to the object's probability of continuous stationary state, the stability score, and the spatial overlap; Step S430: Obtaining a lost property judgment result based on the trajectory consistency score.
[0062] In this embodiment, each element in the set of suspected targets corresponds to a candidate identified as a stationary object. Because lost objects may experience multiple states of stillness, movement, and obstruction during bus operation, accuracy cannot be guaranteed solely through static image and semantic matching. Therefore, it is necessary to combine the object's time-series trajectory in the in-car video with the passenger's trajectory, perform trajectory consistency analysis using a trajectory matching scoring model, and perform a lost object determination for each target object in the set. A determination result is obtained, and the lost object is claimed based on the determination result.
[0063] Specifically, first generate the probability of the object being stationary continuously according to the set of suspected targets: , Stability score and spatial overlap .in, Indicates the probability that an object remained stationary during the time period in which the object was lost. A larger value indicates that the object was more stationary during the time period, reflecting the possibility that the lost item has not been taken away or moved. It can also be used to rule out situations where a passenger's carried object is mistakenly identified as lost property.
[0064] Stability score Indicates the stability score of the object's trajectory during the time period when the object was lost. The larger the value, the higher the stability score, which can effectively exclude non-stationary objects such as objects that were picked up after being placed by humans. Spatial overlap Indicates the spatial overlap between the passenger trajectory area and the object location. A larger value indicates that the passenger is likely to have left the object near the object location.
[0065] Among them, the probability of the object remaining still is generated include: ; in, is the probability that the object remains stationary, It is the spatial trajectory sequence of the object, including timestamp and two-dimensional position coordinates; is the mth frame image during the time period when the object is lost; is the time period during which the passenger described the object as lost; is the mode of the object's position (i.e., center coordinate) during the time period when the object was lost, that is, the position that appears most frequently; It is the fluctuation of the object's position during the time period when the object is lost. If the distance between the object's spatial position in each frame image and the mode of the object's position during the time period is less than the set threshold, that is, the position fluctuation tolerance range (such as 5 pixels), the value is 1, otherwise the value is 0; is the total number of frame images during the object loss period. The data representing the center coordinates of the object in the mth frame image during the time period when the item was lost belongs to , the mth frame image is During this time period.
[0066] Generating a stationary score as follows:
[0067] Among them, the stability control factor It is used to normalize the square of the calculated Euclidean distance between objects into a dimensionless number, and is set according to the tolerance for trajectory fluctuations. , The tolerance for trajectory fluctuations is in pixels. The higher the tolerance for trajectory fluctuations, the greater the tolerance. Under the same trajectory fluctuation, the higher the stability score. For example, if the tolerance for trajectory fluctuations is high, the tolerance is set to 20 pixels. If the tolerance for trajectory fluctuations is very sensitive and low, the tolerance is set to 5 pixels. : The two-dimensional position coordinates of the object in the previous frame image; : The square of the Euclidean distance between two frame images, reflecting the movement of the object.
[0068] This represents the square of the average distance an object moves between frames over that time period. If the object is stationary, the value is equal to or close to 0. If the object is constantly moving, the value varies significantly. The squared Euclidean distance is a geometric standard for two-dimensional translation. It is direction-independent, facilitating subsequent processing. The squared value avoids positive and negative offsets and more accurately reflects the intensity of movement.
[0069] The spatial overlap includes calculating the overlap ratio between the passenger boarding area (inferred from ticket information and boarding and alighting points) and the object appearance area. The specific calculation formula is as follows: ; in, It is the number of areas where the object appears and the passenger's possible activity area overlap (the compartment space can be divided into a 20*20 grid). It is obtained based on the number of areas where the object is photographed in the passenger's possible activity area. It is the number of possible passenger activity areas. Since each bus compartment is divided into multiple camera monitoring blocks, it is possible to find the number of times a passenger is captured by different cameras based on the passenger's riding time period and determine the number of possible passenger activity areas.
[0070] Finally, based on the preset trajectory matching scoring model, a trajectory consistency score is generated according to the object's probability of continuous stillness, the smoothness score, and the spatial overlap. The trajectory matching scoring model is as follows:
[0071] in, is the weight parameter representing the probability of the object remaining stationary. is the weight parameter representing the stationarity score; is a weight parameter representing the degree of spatial overlap. 、 and The sum of the weight parameters is 1, which is set according to the scene. For example, when the vehicle is stationary and the image is clear and stable, the static judgment is mainly used, and the weight parameters are 0.6, 0.3, and 0.1 respectively. During peak hours when there are dense crowds and the trajectories are mixed, the smoothness score is mainly used, and the weight parameters are 0.3, 0.5, and 0.2 respectively. When the image spatial position information is accurate, the spatial overlap is mainly used, and the weight parameters are 0.4, 0.2, and 0.4 respectively.
[0072] Most trajectory consistency analysis methods in the existing technology only focus on the object position or description results of a certain frame in the image. They are unable to model the dynamic trajectory behavior of the object over a period of time, and cannot determine whether the object is stationary or whether it is a passenger carrying an object for a short time. In addition, the object motion state analysis ability is weak, and it does not have the ability to analyze whether the object is stationary for a long time. In particular, it is unable to identify non-lost objects temporarily placed in the car, and the misrecognition rate is high. This embodiment combines the stationary probability, trajectory smoothness, and spatial overlap to score trajectory consistency. It supports the identification of real lost objects that have been stationary for a long time and have not been taken away, as well as objects that are carried by people or displaced midway, through the stationary probability and trajectory smoothness information, thereby improving the lost object recognition rate. In addition, the possible activity area of passengers in the car is inferred through ticketing data and boarding and alighting points, and the spatial overlap with the actual appearance area of the candidate target is calculated, which significantly improves the matching accuracy.
[0073] In one embodiment, step S430: obtaining a lost property determination result based on the trajectory consistency score includes: Step S431: determining whether the trajectory consistency score is greater than a preset high confidence threshold; Step S432: If the answer is yes, the candidate object is determined to be a highly credible lost property, and the highly credible lost properties are aggregated to generate a lost property judgment result. The candidate lost property image and description are pushed to the passenger, and the passenger is notified to claim it.
[0074] In this embodiment, if the judgment is no, that is, the trajectory consistency score is less than the preset high confidence threshold, it is further judged that it is greater than the low confidence threshold. At this time, it is judged as medium confidence lost property, and the passenger is notified to come and confirm whether it is lost property. Then, if the score is lower than the low confidence threshold, it is excluded from the lost property list.
[0075] In one embodiment, Figure 2 As shown, a lost and found system based on a smart bus platform is also provided, the system comprising: A stationary object generation module is used to obtain target lost object information of lost items, analyze the target lost objects, and generate a stationary object set; A lost object description generation module is used to generate semantic image descriptions based on a set of stationary objects and generate a candidate lost object description set based on a preset semantic description generation function; A suspected target generation module is used to perform description similarity matching based on the candidate lost property description set to obtain a suspected target set; The judgment result generation module is used to perform trajectory consistency analysis on the suspected target set based on a preset trajectory matching scoring model to obtain a lost property judgment result.
[0076] In one embodiment, the stationary object generation module is further configured to: obtain target lost item information of the lost item, and generate a target segment set based on the target lost item information; and perform image feature analysis based on the target segment set to obtain a stationary object set.
[0077] In one embodiment, the stationary object generation module is further used to: extract spatial feature vectors based on the target segment set, and calculate the spatial similarity of spatial feature vectors between frame images based on the spatial feature vectors; obtain an image matrix based on the target segment set, and generate a temporal stability score based on the image matrix; generate a spatiotemporal stability score of the lost object based on the spatial similarity and the temporal stability score; and obtain a stationary object set based on the spatiotemporal stability score.
[0078] In one embodiment, the lost property description generation module is further used to: perform image feature extraction based on the target fragment set to generate an image feature vector of a stationary object; generate a matching weight of an anchor point based on the two-dimensional position coordinates of the stationary object, wherein the anchor point is pre-configured; set the semantic label of the anchor point with the largest weight as the semantic reference of the stationary object; generate a sentence description based on the image feature vector and the semantic label of the anchor point based on a semantic description generation function, and generate a candidate lost property description set based on the sentence description.
[0079] In one embodiment, the suspected target generation module is further used to: obtain the declaration description information provided by the passenger when requesting lost property declaration; perform description similarity matching on the candidate lost property description set and the declaration description information based on a preset semantic encoder to obtain a suspected target set.
[0080] In one embodiment, the judgment result generation module is further used to: generate the probability of the object remaining still, the smoothness score and the spatial overlap based on the suspected target set; based on a preset trajectory matching scoring model, generate a trajectory consistency score based on the probability of the object remaining still, the smoothness score and the spatial overlap; and obtain a lost property judgment result based on the trajectory consistency score.
[0081] In one embodiment, the judgment result generation module is further used to: determine whether the trajectory consistency score is greater than a preset high confidence threshold; if so, determine that the candidate object is a high-confidence lost object, and summarize the high-confidence lost objects to generate a lost object judgment result.
[0082] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the lost and found method based on the smart bus platform are implemented.
[0083] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the lost and found method based on the smart bus platform are implemented.
[0084] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0085] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0086] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0087] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0088] An embodiment of the present application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps of any of the above-mentioned method embodiments when executing the computer program.
[0089] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0090] An embodiment of the present application provides a computer program product. When the computer program product is run on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned various method embodiments when executing the computer program product.
[0091] If the integrated unit is implemented as a software functional unit and sold or used as a standalone product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0092] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0093] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0094] In the embodiments provided in this application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0095] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0096] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
[0097] An embodiment of the present application also provides a computer device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any embodiment of the above method when executing the computer program.
[0098] The computer device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the above description is an example of a computer device and does not limit the computer device. The computer device may include more or fewer components than described above, or a combination of certain components, or different components. For example, the computer device may also include input / output devices, network access devices, etc.
[0099] The processor may be a central processing unit (CPU), and the processor 0 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0100] In some embodiments, the memory may be an internal storage unit of the computer device, such as a hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped with the computer device. Furthermore, the memory may include both an internal storage unit of the computer device and an external storage device. The memory is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory may also be used to temporarily store data that has been output or is about to be output.
[0101] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0102] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and such modifications and improvements are intended to fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A lost and found method based on a smart public transportation platform, characterized in that: The method comprises: Obtaining target lost item information of the lost item, analyzing the target lost item, and generating a stationary object set; Generate semantic image descriptions based on a set of static objects, and generate a candidate lost object description set based on a preset semantic description generation function; Perform description similarity matching based on the candidate lost property description set to obtain a set of suspected targets; Based on a preset trajectory matching scoring model, trajectory consistency analysis is performed on the suspected target set to obtain a lost property judgment result.
2. The lost and found method based on the smart public transportation platform according to claim 1 is characterized in that: Obtain target lost item information of the lost item, analyze the target lost item, and generate a stationary object set, including: Obtaining target lost property information of the lost item, and generating a target fragment set according to the target lost property information; Image feature analysis is performed based on the target segment set to obtain a stationary object set.
3. The lost and found method based on the smart public transportation platform according to claim 2 is characterized in that: Performing image feature analysis on the target segment set to obtain a stationary object set, including: Extracting spatial feature vectors based on the target segment set, and calculating spatial similarity of spatial feature vectors between frame images based on the spatial feature vectors; obtaining an image matrix based on the target segment set and generating a temporal stability score based on the image matrix; generating a spatiotemporal stability score of the lost object according to the spatial similarity and the temporal stability score; A set of stationary objects is obtained according to the spatiotemporal stability scores.
4. The lost and found method based on the smart public transportation platform according to claim 1 is characterized in that: Generate semantic image descriptions based on a set of static objects, and generate a candidate lost object description set based on a preset semantic description generation function, including: Performing image feature extraction based on the target segment set to generate an image feature vector of a stationary object; Generate a matching weight of an anchor point according to the two-dimensional position coordinates of the stationary object, wherein the anchor point is pre-configured; Set the semantic label of the anchor point with the largest weight as the semantic reference of the stationary object; Based on the semantic description generation function, a sentence description is generated according to the semantic label of the image feature vector and the anchor point. A candidate lost property description set is generated according to the sentence description.
5. The lost and found method based on the smart public transportation platform according to claim 1 is characterized in that: A description similarity matching is performed based on the candidate lost property description set to obtain a suspected target set, including: Obtain the descriptive information provided by the passenger when requesting lost property report; Based on a preset semantic encoder, description similarity matching is performed on the candidate lost property description set and the description information at the time of declaration to obtain a suspected target set.
6. The lost and found method based on the smart public transportation platform according to claim 1 is characterized in that: Based on the preset trajectory matching scoring model, the trajectory consistency analysis of the suspected target set is performed to obtain the lost property judgment result, including: generating a probability of object continuous stationary state, a stability score, and a spatial overlap degree according to the set of suspected targets; Based on a preset trajectory matching scoring model, a trajectory consistency score is generated according to the object's probability of continuous stationary state, the stability score, and the spatial overlap; The lost property judgment result is obtained based on the trajectory consistency score.
7. The lost and found method based on the smart public transportation platform according to claim 1 is characterized in that: The lost property judgment result is obtained based on the trajectory consistency score, including: Determining whether the trajectory consistency score is greater than a preset high confidence threshold; If the judgment is yes, the candidate object is judged to be a high-credible lost object, and the high-credible lost objects are aggregated to generate a lost object judgment result.
8. A lost and found system based on a smart public transportation platform, characterized in that: The system comprises: A stationary object generation module is used to obtain target lost object information of lost items, analyze the target lost objects, and generate a stationary object set; A lost object description generation module is used to generate semantic image descriptions based on a set of stationary objects and generate a candidate lost object description set based on a preset semantic description generation function; A suspected target generation module is used to perform description similarity matching based on the candidate lost property description set to obtain a suspected target set; The judgment result generation module is used to perform trajectory consistency analysis on the suspected target set based on a preset trajectory matching scoring model to obtain a lost property judgment result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Public transport lost article retrieval method and system based on image processing
CN114863516A
A real-time non-tracking monitoring video remnant detection method
CN109636795A
Lost article detection method and device, equipment and storage medium
CN114677605A
Article information transmission method and device
CN115982356A
Lost article retrieving method, system and equipment and storage medium
CN119380238A