Identity recognition method and device
By analyzing the object characteristics and behavioral characteristics in multiple video frames and generating identity recognition results in combination with deep learning models, the problem of poor identity recognition accuracy in the existing technology is solved, and higher recognition accuracy and flexibility are achieved.
Patent Information
- Application Number
- CN202510802012.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, identity identification based on facial information of static images is poor in the identification of personnel not entered into the system in advance, especially in an emergency situation, the identity of passengers and rescue personnel cannot be effectively identified.
By obtaining multiple video frames of the video to be identified, analyzing the object feature sequence, generating behavioral feature information of the target object, and combining object features and behavioral feature information to generate identity recognition results, using deep learning models such as YOLOv8-Pose, Transformer, Informer and other technologies, the object features and behavioral features in multiple video frames are analyzed, and the attribute information of the associated object is fused to improve the recognition accuracy.
It improves the accuracy and flexibility of identity recognition, can more comprehensively capture the static characteristics and dynamic behavior patterns of the target object, reduces misidentification, and supports object identity recognition that is not entered into the system.
Smart Images

Figure CN120339918A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to an identity recognition method and apparatus. Background Art
[0002] Identity recognition is an important research direction in the fields of computer vision and artificial intelligence, and is widely used in fields such as security monitoring, intelligent transportation, human-computer interaction, and smart home. Taking the rail transit field as an example, during the operation of a train, in case of an emergency that traps passengers, rapid and effective rescue operations are required. In such a situation, the automatic recognition of the identities of passengers and rescue personnel becomes particularly important to ensure that emergency commanders accurately perceive the specific situation and progress of the on-site rescue, provide support for rescue decision-making, and effectively monitor the implementation effect of rescue operations.
[0003] Currently, identity recognition mainly relies on face information in static images. However, the above-mentioned solution has a single basis of information, and there is a situation where it cannot recognize when dealing with the identity recognition of personnel who have not been pre-entered into the system, resulting in poor accuracy of identity recognition. Therefore, there is an urgent need for an identity recognition solution with high accuracy. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide an identity recognition method. One or more embodiments of this specification also relate to an identity recognition apparatus, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, an identity recognition method is provided, including: Obtain a video to be recognized for an identity recognition task, where the video to be recognized includes a plurality of video frames; Parse the plurality of video frames to obtain an object feature sequence of the video to be recognized, where the object feature sequence includes object feature information of a target object arranged based on time series, and the object feature information includes target attribute information of the target object and associated attribute information of an associated object, and the associated object is an object in the video frame that has a position association with the target object; Generate behavior feature information of the target object according to the object feature sequence; Generate an identity recognition result of the target object according to the object feature information and the behavior feature information.
[0006] According to the second aspect of the embodiments of this specification, an identity recognition apparatus is provided, including: An acquisition module, configured to obtain a video to be recognized for an identity recognition task, where the video to be recognized includes a plurality of video frames; A parsing module, configured to parse a plurality of video frames to obtain an object feature sequence of a video to be recognized, where the object feature sequence includes object feature information of a target object arranged based on time series, and the object feature information includes target attribute information of the target object and associated attribute information of an associated object, and the associated object is an object in the video frame that has a location association with the target object; A first generation module, configured to generate behavior feature information of the target object according to the object feature sequence; A second generation module, configured to generate an identity recognition result of the target object according to the object feature information and the behavior feature information.
[0007] According to a third aspect of the embodiments of the present specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method provided in the first aspect are implemented.
[0008] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method provided in the first aspect are implemented.
[0009] According to a fifth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method provided in the first aspect are implemented.
[0010] An identity recognition method provided by an embodiment of this specification includes: obtaining a video to be recognized for an identity recognition task, where the video to be recognized includes multiple video frames; parsing the multiple video frames to obtain an object feature sequence of the video to be recognized, where the object feature sequence includes object feature information of a target object arranged based on time sequence, and the object feature information includes target attribute information of the target object and associated attribute information of an associated object, and the associated object is an object in the video frame that has a positional association with the target object; generating behavior feature information of the target object according to the object feature sequence; and generating an identity recognition result of the target object according to the object feature information and the behavior feature information. Since the object feature information is arranged based on time sequence, the dynamic changes of the object features can be better understood, and the accuracy of the behavior feature information can be improved. By incorporating the associated attribute information of the associated object into the object feature information, the positional relationship and interaction behavior between the target object and the surrounding objects can be better understood, which helps to more accurately recognize the identity of the target object. By analyzing the object feature sequence and behavior feature information in multiple video frames, the static features (such as appearance, clothing, etc.) and dynamic behavior patterns (such as postures, gestures, etc.) of the target object can be more comprehensively captured, providing richer context information for the identity recognition process, reducing the misrecognition problem caused by a single recognition reference information, and at the same time supporting the identity recognition of objects not entered into the system, improving the accuracy and flexibility of identity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is an architecture diagram of an identity recognition system provided by an embodiment of this specification; Figure 2 is a flowchart of an identity recognition method provided by an embodiment of this specification; Figure 3 is a schematic diagram of data dimension conversion provided by an embodiment of this specification; Figure 4 is a schematic structural diagram of an identity recognition model provided by an embodiment of this specification; Figure 5 is a processing process flowchart of an identity recognition method provided by an embodiment of this specification; Figure 6 is a processing process flowchart of another identity recognition method provided by an embodiment of this specification; Figure 7 is a schematic structural diagram of an identity recognition device provided by an embodiment of this specification; Figure 8 is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] Numerous specific details are set forth in the following description to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0013] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items. The term "at least one" in one or more embodiments of this application refers to "one or more", and "a plurality" refers to "two or more". The term "comprising" is an open-ended description and should be understood as "including but not limited to", and other content may also be included on the basis of the described content.
[0014] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0015] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0016] First, the noun terms involved in one or more embodiments of this specification are explained.
[0017] YOLO (You Only Look Once): It is a real-time object recognition algorithm for object detection. Different from traditional object detection methods, YOLO processes the entire image at once to predict the location and category of objects, rather than detecting them region by region or using a sliding window. This end-to-end approach not only significantly improves the detection speed but also remarkably enhances the accuracy.
[0018] YOLOv8-Pose: It is a deep learning-based pose estimation technology that combines the YOLOv8 object detection algorithm and the pose estimation algorithm, enabling high-precision estimation of human poses in images or videos. The YOLOv8-Pose model is mainly used to identify and locate the key points of the human body, such as the head, shoulders, elbows, wrists, etc., thus achieving accurate estimation of human poses.
[0019] Deep self-attention (Transformer) model: It is a network structure based on the multi-head self-attention mechanism module, mainly used to process sequential data. The Transformer model consists of repeatable stacked encoding units (Encoders) and decoding units (Decoders). This design allows the Transformer to efficiently learn long-term dependencies and is applicable to various natural language processing tasks, including machine translation, text summarization, question-answering systems, etc.
[0020] Informer model: It is a deep learning model for long-sequence time series prediction. The Informer model improves the traditional Transformer architecture to handle long-term dependencies in time series data and significantly enhances the computational efficiency and memory usage efficiency, making it suitable for longer time series prediction tasks.
[0021] Tracking Algorithm: It refers to the technology used to identify, locate, and track one or more target objects in a series of consecutive frames (such as video streams or time series data). These algorithms are widely used in fields such as computer vision, robot navigation, autonomous driving, and surveillance systems. Their main purpose is to analyze the information in each frame, predict and update the state parameters of the target, such as position, speed, and direction, so as to achieve continuous tracking of the target. Tracking algorithms include but are not limited to the Simple Online and Realtime Tracking (SORT) algorithm, the Simple Online and Realtime Tracking with a Deep Association Metric (DeepSORT) algorithm, the Bag of Tricks for Efficient Multi-Object Tracking (BoT-SORT) algorithm, and the Simple Online and Realtime Tracking with Byte-Level Association (ByteTrack) algorithm.
[0022] Convolutional Neural Network (CNN): It is a deep learning model specifically designed to process data with a similar grid structure. By simulating the working principle of the human visual system, CNN can automatically extract hierarchical features from image data and perform tasks such as classification, detection, and segmentation.
[0023] Long Short-Term Memory (LSTM): It is a special type of recurrent neural network that can learn long-term dependencies. The LSTM is designed to address the difficulties encountered by traditional recurrent neural networks in dealing with long-term dependencies, namely the problem of vanishing gradients or exploding gradients. By introducing a structure called the "gating mechanism", the LSTM can effectively capture and store long-term dependency relationships in time series data.
[0024] Dropout: It is a technique used to prevent overfitting during the training of neural networks. It achieves this goal by randomly "dropping out" (i.e., setting to zero) a portion of neurons in each training iteration, so that the model does not overly rely on certain specific neurons, enhancing the generalization ability of the model.
[0025] Stochastic Gradient Descent (SGD): A commonly used optimization algorithm widely applied in machine learning and deep learning for minimizing loss functions. It gradually approaches the optimal solution by iteratively updating model parameters. Different from batch gradient descent, SGD updates parameters using only one training sample each time, hence the name "stochastic".
[0026] In this specification, an identity recognition method is provided. This specification also relates to an identity recognition device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0027] See Figure 1 , Figure 1 which shows the architecture diagram of an identity recognition system provided by an embodiment of this specification. The identity recognition system may include a client 100 and a server 200; The client 100 is configured to send a video to be recognized for an identity recognition task to the server 200, where the video to be recognized includes multiple video frames; The server 200 is configured to parse multiple video frames to obtain an object feature sequence of the video to be recognized, where the object feature sequence includes object feature information of a target object arranged based on time series. The object feature information includes target attribute information of the target object and associated attribute information of an associated object. The associated object is an object in the video frame that has a location association with the target object; generate behavior feature information of the target object according to the object feature sequence; generate an identity recognition result of the target object according to the object feature information and the behavior feature information; and send the identity recognition result to the client 100; The client 100 is further configured to receive the identity recognition result sent by the server 200.
[0028] Applying the solution of the embodiment of this specification, by analyzing the object feature sequence and behavior feature information in multiple video frames, it is possible to more comprehensively capture the static features and dynamic behavior patterns of the target object, provide richer context information for the identity recognition process, reduce the misrecognition problem caused by a single recognition reference information, and at the same time support the identity recognition of objects not entered into the system, improving the accuracy and flexibility of identity recognition.
[0029] In practical applications, the identity recognition system may include a server 200 and multiple clients 100. Communication connections can be established between the multiple clients 100 through the server 200. In the identity recognition scenario, the server 200 is used to provide identity recognition services between the multiple clients 100. The multiple clients 100 can serve as senders or receivers respectively and communicate through the server 200. Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the identity recognition scenario, it can be that the user publishes a data stream to the server 200 through the client 100, and the server 200 generates an identity recognition result based on the data stream and pushes the identity recognition result to other clients that have established communication.
[0030] Among them, a connection is established between the client 100 and the server 200 through a network. The network provides the medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the server 200.
[0031] The client 100 can be a browser, an application (APP, Application), or a web application such as a HyperText Markup Language 5 (H5) application, or a light application (also known as a mini-program, a lightweight application program), or a cloud application, etc. The client 100 can be developed based on the Software Development Kit (SDK) provided by the server 200 for the corresponding service, such as developed based on the Real Time Communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0032] The server 200 may include servers that provide various services. For example, a server that provides communication services for multiple clients, or a server for background training that supports models used on the client, or a server that processes data sent by the client, etc. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs, Content Delivery Networks), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0033] It is worth noting that the identity recognition method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client may also have a similar function to the server, so as to execute the identity recognition method provided in the embodiments of this specification. In other embodiments, the identity recognition method provided in the embodiments of this specification can also be jointly executed by the client and the server.
[0034] See Figure 2 , Figure 2 which shows a flowchart of an identity recognition method provided in an embodiment of this specification, specifically including the following steps: Step 202: Obtain the video to be recognized for the identity recognition task, where the video to be recognized includes multiple video frames.
[0035] It should be noted that the identity recognition task refers to the process of determining or verifying the identity of a specific object (such as a person, an animal, an item, etc.) by analyzing and processing data such as videos and images. The identity recognition task can be a task in different scenarios, and different scenarios include but are not limited to security monitoring and security protection scenarios, intelligent transportation management scenarios, and financial service scenarios. Taking the security monitoring and security protection scenario as an example, the identity recognition task can be a task of automatically identifying specific individuals or suspicious behaviors in public places (such as airports, railway stations, shopping malls, etc.) through video analysis technology, such as detecting whether there are blacklisted persons, or identifying abnormal behavior patterns to prevent crimes. Taking the rail transit management scenario as an example, the identity recognition task can be a task of automatically identifying the identities of passengers and rescue personnel in the rail transit scenario (such as ordinary operation scenarios, rescue drill scenarios, and actual emergency scenarios, etc.) through video analysis technology. The video to be recognized refers to a video file or video stream selected as the input source for identity recognition. The video to be recognized can be videos from different data sources, and different data sources include but are not limited to surveillance cameras, smartphones, or any other device capable of recording videos. A video frame refers to the basic unit that makes up the video to be recognized. The video to be recognized is actually a series of static images (i.e., video frames) that are displayed quickly and continuously. Each video frame is a complete image, representing the picture at a specific time point in the video to be recognized. By analyzing a single video frame, the image features of the video frame can be extracted. These image features of the video frames are arranged in chronological order and can form a continuous action sequence, which helps to understand the behavior pattern of the target object.
[0036] In practical applications, there are various ways to obtain the video frames to be recognized for the identity recognition task, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, the video to be recognized for the identity recognition task sent by the user through the client can be received. In another possible implementation manner of this specification, the video frames to be recognized for the identity recognition task can be read from other data acquisition devices or databases.
[0037] Step 204: Parse multiple video frames to obtain the object feature sequence of the video to be recognized, where the object feature sequence includes the object feature information of the target object arranged based on time sequence, and the object feature information includes the target attribute information of the target object and the associated attribute information of the associated object, and the associated object is the object in the video frame that has a location association with the target object.
[0038] It should be noted that the object feature sequence refers to a data set parsed from multiple video frames. The object feature sequence includes object feature information composed of the attribute information of the target object (target attribute information) and the attribute information of the object associated with the target object (associated attribute information). Since the target object or the behavior of the target object included in different video frames may be different, in order to ensure that the object feature sequence can reflect the behavior pattern of the target object, the object feature information of the target object can be arranged in the time order of multiple video frames to obtain the object feature sequence of the video to be recognized. For example, the first frame includes the target object A, and the second frame includes the target object B. The first frame is before the second frame. Therefore, the object feature sequence is {the object feature information of the target object A, the object feature information of the target object B}. Through the object feature sequence, a global view of the behavior and development of the target object based on the time dimension can be provided, so that identity recognition can be performed based on the behavior pattern of the target object. The target object refers to the object whose identity needs to be recognized in the video to be recognized, such as the people appearing in the video to be recognized, such as passengers and rescue workers. The associated object refers to a person, animal or item that has a location association with the target object in the same video frame, such as a rescue device. The number of associated objects can be one or more. The existence of a location association between the associated object and the target object can be that the location of the associated object coincides with that of the target object. For example, a rescue worker holds a rescue device in his hand, or the distance between the associated object and the target object is within a preset distance threshold. For example, there are rescue supplies and rescued people near the rescue worker. The target attribute information refers to the attribute information of the target object itself, including but not limited to the type, appearance characteristics (such as clothing, shape, color, etc.), and posture of the target object. The associated attribute information refers to the attribute information of the associated object itself, including but not limited to the type, appearance characteristics, and posture of the associated object. By using the associated attribute information of the associated object as part of the object feature information of the target object, rich context information can be provided for analyzing the behavior of the target object, assisting in understanding the environmental background and behavior motivation of the target object.
[0039] In practical applications, there are multiple ways to parse multiple video frames to obtain the object feature sequence of the video to be recognized, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In a possible implementation manner of this specification, multiple video frames can be input into a sequence generation model to obtain the object feature sequence of the object to be recognized output by the sequence generation model. Among them, the sequence generation model refers to a deep learning model trained based on the training object feature sequence and the corresponding multiple video frames trained with the training object features. In another possible implementation manner of this specification, the attribute information of each object can be recognized first, and then the associated objects of the target object can be determined from multiple objects. According to the respective attribute information of the target object and the associated objects, the object feature sequence of the video to be recognized can be generated.
[0040] In an optional embodiment of this specification, the video to be recognized includes multiple objects, and the multiple objects include a target object. The step of parsing multiple video frames to obtain the object feature sequence of the video to be recognized may include the following steps: Parse multiple video frames to obtain the attribute information of multiple objects, where the attribute information includes the position information of the object in the video frame; According to the position information, filter out the associated objects of the target object from the multiple objects; Generate the object feature sequence of the video to be recognized according to the target attribute information of the target object and the associated attribute information of the associated object.
[0041] It should be noted that the attribute information is used to describe various characteristics of the object, including position, type, and posture, etc., providing a deeper understanding of the object. The type is used to indicate which category the object belongs to, such as personnel, items, animals, etc. The type can be represented by a classification label, such as "person", "car", "dog", etc., or by a classification score, such as "person: 0.95". The type helps to distinguish different objects. The posture is used to describe the body structure and the current posture state of the object. For the human body, the posture usually involves the positions and angles of various parts of the body. The posture can be represented by a set of key points, and the key points include but are not limited to the head, shoulders, elbows, wrists, hips, knees, and ankles. Each key point is represented by a pair of (x, y) coordinates. The posture can also be represented by a skeleton diagram formed by connecting the key points, which is used to show the body structure and posture relationship of the object. For example, the line connecting the shoulder to the elbow and then to the wrist. The posture information helps to understand the action and behavior patterns of the object. The position information refers to the specific spatial coordinates or bounding box or center point of each object in the video frame. The position information is used to describe the specific position of the object in the video frame. The bounding box can represent a rectangular area with four values, namely the x coordinate, y coordinate, width, and height of the upper left corner, that is, [x, y, width, height]. It can also be represented by the center point coordinates and the width and height, that is, [center_x, center_y, width, height].
[0042] In practical applications, there are various ways to parse multiple video frames to obtain the attribute information of multiple objects, which are specifically selected according to the actual situation, and this specification embodiment does not make any limitations on this. In a possible implementation manner of this specification, multiple video frames can be input into an image parsing model to obtain the attribute information of multiple objects. In another possible implementation manner of this specification, the video frames can be matched with multiple video frame templates carrying object attribute information, and the object attribute information carried by the video frame template that matches the video frame successfully is determined as the object attribute information in the video frame.
[0043] Further, there are various ways to screen out the associated objects of the target object from multiple objects according to the position information, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In the first possible implementation manner of this specification, object tracking can be performed on multiple video frames based on a tracking algorithm according to the position information of multiple objects to obtain the tracking information of each object in each video frame, where the tracking information includes the tracking object identifier, category and confidence level, tracking box position information, human pose recognition information, etc.; according to the tracking information, the associated objects of the target object are screened out from the candidate objects. In the second possible implementation manner of this specification, the candidate objects whose candidate position information overlaps with the target position information of the target object can be determined as the associated objects. In the third possible implementation manner of this specification, the associated objects of the target object can be screened out from the candidate objects according to the degree of position association between the candidate position information of the candidate object and the target position information of the target object.
[0044] When generating the object feature sequence of the video to be recognized according to the target attribute information of the target object and the associated attribute information of the associated object, the associated attribute information can be concatenated after the target attribute information to obtain the object feature information of the target object, and then, the object feature information of different target objects is arranged according to the time sequence of multiple video frames to obtain the object feature sequence of the video to be recognized.
[0045] Applying the solution of the embodiments of this specification, by determining the position information and attribute information of each object in multiple video frames, and then determining the target object and the associated objects of the target object from multiple objects, and using the associated attribute information of the associated objects as the feature supplement of the target object, the richness of the object feature information is improved, and the accuracy of identity recognition is further improved.
[0046] In an optional embodiment of this specification, the above-mentioned screening out the associated objects of the target object from multiple objects according to the position information may include the following steps: According to the target position information of the target object and the candidate position information of the candidate object, a position association index is determined, where the position association index corresponds to the candidate object one by one, and the position association index is used to characterize the degree of position association between the candidate position information of the candidate object and the target position information, and the candidate object is an object other than the target object among multiple objects; According to the position association index, the associated objects of the target object are screened out from the candidate objects.
[0047] It should be noted that the candidate objects are the objects other than the target object among multiple objects, and the candidate objects can be people, animals, or items. The larger the position association index or the higher the level, the greater the position association degree between the candidate object and the target object; the smaller the position association index, the smaller the position association degree between the candidate object and the target object. The position association index can be a specific association value, such as 0.7, or an association level, such as very relevant, relatively relevant, irrelevant.
[0048] In practical applications, there are various ways to determine the position association index based on the target position information of the target object and the candidate position information of the candidate object. It is specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, the ratio of the intersection area to the union area of the bounding boxes of the target object and the candidate object can be calculated, and the calculation result is determined as the position association index. In another possible implementation manner of this specification, the Euclidean distance value between the center point position of the target object and the center point position of the candidate object can be calculated, and the Euclidean distance value is determined as the position association index. Among them, the Euclidean distance value can be calculated through the following formula (1), where represents the center point position of the target object, The coordinates of are represents the center point position of the candidate object, The coordinates of are represents the Euclidean distance value between the center point position of the target object and the center point position of the candidate object.
[0049] (1) Furthermore, there are various ways to screen out the associated objects of the target object from the candidate objects according to the position association index. It is specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one optional embodiment of this specification, the candidate objects can be sorted from largest to smallest according to the position association index, and the top N candidate objects corresponding to the position association indexes are determined as the associated objects, where N is a positive integer. In another possible implementation manner of this specification, the candidate objects with a position association index greater than the association index threshold can be determined as the associated objects, where the association index threshold is specifically set according to the actual situation.
[0050] Applying the solution of the embodiments of this specification to screen out the associated objects of the target object from the candidate objects by using the position association index takes into account the position association degree between the target object and each candidate object, and improves the accuracy of the associated objects.
[0051] In one optional embodiment of this specification, the above-mentioned parsing of multiple video frames to obtain the attribute information of multiple objects may include the following steps: Input multiple video frames into an image parsing model respectively to obtain the attribute information of multiple objects.
[0052] It should be noted that the image parsing model is a deep learning model with the ability to extract the attribute information (position, pose, type, appearance) of objects from images. The image parsing model includes but is not limited to the YOLOv8-pose model and the YOLOv10 series models with the ability to predict specific points. The image parsing model is trained based on multiple sample video frames carrying the sample attribute information of sample objects.
[0053] Applying the solution of the embodiments of this specification, using the image parsing model to determine the attribute information of multiple objects realizes efficient and accurate determination of attribute information.
[0054] In an optional embodiment of this specification, the training process of the image parsing model is described. That is, before inputting multiple video frames into the image parsing model respectively to obtain the attribute information of multiple objects, the following steps may also be included: Obtain multiple sample video frames of a sample video, where the sample video frames carry the sample attribute information of sample objects; Input multiple sample video frames into the image parsing model respectively to obtain the predicted attribute information of the sample objects; Train the image parsing model according to the sample attribute information and the predicted attribute information to obtain a trained image parsing model.
[0055] It should be noted that the sample video refers to a video file or video stream used to train the image parsing model. The sample video frame is the basic unit that makes up the sample video. The sample object refers to the object in the sample video, such as a person, an animal, or an item in the sample video. The sample attribute information refers to the attribute information of the sample object itself, including but not limited to the position information, type (item, person type), appearance features (such as clothing, shape, color, etc.), and pose (only including the pose of a person, and the pose key points of other objects are uniformly set to zero) of the sample object. The person types in the sample attribute information are such as passengers, station staff, security personnel, volunteers, police officers, medical staff, firefighters, etc. The sample attribute information is the parsing target of the image parsing model and is used to guide the training process of the image parsing model. The predicted attribute information is the result obtained by the image parsing model through parsing and predicting the sample video frame. By comparing the sample attribute information and the predicted attribute information, the parsing ability of the image parsing model can be determined. The sample attribute information can be obtained by manual annotation. Before manually annotating the sample attribute information of the sample object in the sample video, data cleaning can be performed on the multiple sample video frames included in the sample video to screen out the sample video frames that can be used for annotation, such as the sample video frames including objects, and then the sample attribute information is annotated for the sample video frames after data cleaning.
[0056] Exemplarily, taking the sample video as an example of a video in a rail transit scenario, sample videos containing the rescue behaviors of passengers and rescue personnel can be collected through cameras in various operating places of rail transit. The sample videos related to passengers and rescue personnel include three types of scenario video data: normal operation scenarios, rescue drill scenarios, and actual emergency scenarios. The sample objects include passengers and rescue personnel. The clothing of passengers and rescue personnel includes various dressings such as different upper and lower body colors, styles, whether wearing hats, glasses, holding different rescue tools, and the clothes being damaged or contaminated. Rescue behaviors include guiding evacuation, preliminary medical assessment, first aid treatment, using equipment such as fire extinguishers and fire hoses, using professional equipment (such as hydraulic tools) and techniques to rescue trapped passengers, transporting rescue supplies for medical rescue, and security patrol at the accident scene, etc.
[0057] In practical applications, when training an image parsing model using multiple sample video frames of a sample video, in one possible implementation manner of this specification, multiple sample video frames can be used to train an image parsing model to obtain a trained image parsing model. In another possible implementation manner of this specification, multiple sample video frames can be used to train multiple image parsing models to obtain multiple trained image parsing models, and then, from the multiple trained image parsing models, an image parsing model with very good comprehensive performance in object classification and pose recognition is selected. Hyperparameters such as the model architecture, number of training rounds, solver optimizer, step size, batch size, etc. of the multiple image parsing models can be set according to actual needs.
[0058] Exemplarily, multiple sample video frames carrying the sample attribute information of the sample objects can be divided into a training set, a validation set, and a test set according to a certain ratio (such as 6:2:2); the training set is used to train multiple image parsing models, the validation set is used to verify the image parsing ability of each image parsing model, and after obtaining multiple trained image parsing models upon completion of training, the test set is used to detect the training effects of each trained image parsing model, such as detecting indicators such as the accuracy, precision, recall, average precision (AP, Average Precision), and mean average precision (mAP, mean Average Precision) of each trained image parsing model, and an image parsing model that can ultimately be used for identity recognition is selected from the multiple trained image parsing models based on these indicators.
[0059] Furthermore, the training process of any image parsing model will be described: When training the image parsing model according to the sample attribute information and the predicted attribute information, the parsing loss value can be calculated based on the sample attribute information and the predicted attribute information, and the model parameters of the image parsing model can be adjusted according to the parsing loss value until the training process meets the preset stop condition, and then the training is stopped to obtain the trained image parsing model. Among them, there are many functions for calculating the parsing loss value, such as the cross-entropy loss function, L1 norm loss function, maximum loss function, mean square error loss function, logarithmic loss function, etc., which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. The preset stop condition includes but is not limited to that the parsing loss value is less than or equal to the preset threshold, or the number of iterations reaches the preset number of iterations. The preset threshold and the preset number of iterations are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard.
[0060] In a possible implementation manner of this specification, after calculating the parsing loss value, the parsing loss value can be compared with the preset threshold. Specifically, if the parsing loss value is greater than the preset threshold, it indicates that the difference between the sample attribute information and the predicted attribute information is relatively large, and the image parsing model has a poor processing ability for the sample video. At this time, the model parameters of the image parsing model can be adjusted until the parsing loss value is less than or equal to the preset threshold, indicating that the difference between the sample attribute information and the predicted attribute information is relatively small, reaching the preset stop condition, and obtaining the trained image parsing model.
[0061] In another possible implementation manner of this specification, in addition to comparing the magnitude relationship between the parsing loss value and the preset threshold, the number of iterations can also be combined to determine whether the current image parsing model is trained. Specifically, if the parsing loss value is greater than the preset threshold, the model parameters of the image parsing model are adjusted until the preset number of iterations is reached, and then the iteration is stopped to obtain the trained image parsing model.
[0062] Applying the solution of the embodiments of this specification, the image parsing model is trained according to the sample attribute information and the predicted attribute information. When the preset stop condition is not met, the image parsing model is continuously trained until the preset stop condition is reached, and the trained image parsing model is obtained after the training is completed. By continuously training the image parsing model, the performance of the image parsing model is improved.
[0063] Step 206: Generate the behavior feature information of the target object according to the object feature sequence.
[0064] It should be noted that the behavioral feature information refers to the information about the behavioral pattern of the target object extracted based on the object feature sequence of the target object. The behavioral feature information can describe the behavioral pattern of the target object (such as walking, running, using rescue tools, etc.), the movement trajectory (the movement path of the target object in consecutive video frames), the interaction mode (the interaction situation between the target object and other objects or the environment), and other contents.
[0065] In practical applications, there are various ways to generate the behavioral feature information of the target object according to the object feature sequence, and specific selection is made according to the actual situation. The embodiments of this specification do not make any limitation thereto. In one possible implementation manner of this specification, the object feature sequence can be directly input into the behavior recognition model to obtain the behavioral feature information of the target object. In another possible implementation manner of this specification, in order to better adapt to the input requirements of the behavior recognition model and improve the model performance of the behavior recognition model, the object feature sequence can be processed according to the model processing conditions of the behavior recognition model, and the processed object feature sequence is input into the behavior recognition model to obtain the behavioral feature information of the target object.
[0066] In an optional embodiment of this specification, the generation of the behavioral feature information of the target object according to the object feature sequence may include the following steps: Process the object feature sequence according to the model processing conditions of the behavior recognition model to obtain the processed object feature sequence, where the processed object feature sequence meets the model processing conditions; Input the processed object feature sequence into the behavior recognition model to obtain the behavioral feature information of the target object.
[0067] It should be noted that the behavior recognition model refers to a deep learning model with the ability to predict the time series to obtain the behavioral feature information, such as the Informer model and the LSTM model. The behavior recognition model is trained based on the sample video carrying the sample object feature sequence and the sample behavioral feature information. The model processing conditions refer to the requirements of the behavior recognition model for the specific format, dimension, range, batch size, etc. of the input data. There are various ways to obtain the model processing conditions of the behavior recognition model, and specific selection is made according to the actual situation. The embodiments of this specification do not make any limitation thereto. In one possible implementation manner of this specification, the model processing conditions can be extracted from the model description file of the behavior recognition model. In another possible implementation manner, the model processing conditions of the behavior recognition model sent by the user through the client can be received.
[0068] Exemplarily, when processing the object feature sequence according to the model processing conditions of the behavior recognition model, feature segmentation and dimension conversion can be performed on the object feature sequence. During feature segmentation, a certain number of object feature information can be obtained from the object feature sequence based on the model processing conditions. During dimension conversion, the object feature sequence can be converted into the dimensions commonly used in multi-dimensional time series classification. Refer to Figure 3 , Figure 3 FIG. Figure 3 shows a schematic diagram of data dimension conversion provided by an embodiment of this specification. The object feature sequence is a three-dimensional time series, denoted as , where X represents the object feature sequence, C represents the target attribute information and associated attribute information included in the object characteristic information, such as confidence level, relative values of the center and height-width of the bounding box, category parameters, and pose key point coordinates, etc. L represents the number of frames of the object feature sequence, and N represents the number of target objects in this video frame. When performing data dimension conversion, the three-dimensional time series can be flattened, and let M = C * N, that is, the object feature information of all target objects in this frame is sequentially connected together to form a vector, and finally a two-dimensional matrix can be obtained.
[0069] Applying the solution of the embodiment of this specification to process the object features using the model processing conditions of the behavior recognition model can ensure that the behavior recognition model correctly predicts the behavior feature information and improves the accuracy of the behavior feature information.
[0070] In an optional embodiment of this specification, the training method of the behavior recognition model is described, that is, before inputting the processed object feature sequence into the behavior recognition model to obtain the behavior feature information of the target object, the following steps may further be included: Obtain a sample video, where the sample video includes multiple sample video frames, and the sample video carries a sample object feature sequence and sample behavior feature information; According to the model processing conditions of the behavior recognition model, process the sample object feature sequence to obtain a processed sample object feature sequence, where the processed sample object feature sequence conforms to the model processing conditions; Input the processed sample object feature sequence into the behavior recognition model to obtain predicted behavior feature information; Train the behavior recognition model according to the sample behavior feature information and the predicted behavior feature information to obtain a trained behavior recognition model.
[0071] It should be noted that the sample object feature sequence refers to a sample data set parsed and processed from multiple sample video frames. The sample object feature sequence includes sample object feature information composed of sample target attribute information of the sample target object and sample associated attribute information of the object associated with the sample target object. The sample behavior feature information refers to the information about the behavior pattern of the sample target object extracted based on the sample object feature sequence of the sample target object. The sample behavior feature information can be obtained by manual annotation. The sample behavior feature information is the recognition target of the behavior recognition model and is used to guide the training process of the behavior recognition model. The predicted behavior feature information is the result obtained by the behavior recognition model through behavior recognition prediction on the processed sample object feature sequence. By comparing the sample behavior feature information and the predicted behavior feature information, the behavior classification ability of the behavior recognition model can be determined. The generation method of the sample object feature sequence can refer to the implementation method of "parsing multiple video frames to obtain the attribute information of multiple objects, where the attribute information includes the position information of the object in the video frame; screening the associated objects of the target object from multiple objects according to the position information; and generating the object feature sequence of the video to be recognized according to the target attribute information of the target object and the associated attribute information of the associated object". The implementation method of "processing the sample object feature sequence according to the model processing conditions of the behavior recognition model to obtain the processed sample object feature sequence" can refer to the above-mentioned implementation method of "processing the object feature sequence according to the model processing conditions of the behavior recognition model to obtain the processed object feature sequence", and this will not be elaborated in the embodiments of this specification.
[0072] In practical applications, when training a behavior recognition model using a sample video carrying the sample object feature sequence and the sample behavior feature information, in one possible implementation manner of this specification, a behavior recognition model can be trained using the sample video carrying the sample object feature sequence and the sample behavior feature information to obtain a trained behavior recognition model. In another possible implementation manner of this specification, multiple behavior recognition models can be trained using the sample video carrying the sample object feature sequence and the sample behavior feature information to obtain multiple trained behavior recognition models, and then, from the multiple trained behavior recognition models, a behavior recognition model with very good behavior recognition effect can be selected. Hyperparameters such as the model architecture, number of training rounds, solver optimizer, step size, batch_size, etc. of the multiple behavior recognition models can be set according to actual needs. Further, the training method for any behavior recognition model, "training the behavior recognition model according to the sample behavior feature information and the predicted behavior feature information to obtain a trained behavior recognition model", can refer to the implementation method of "training the image parsing model according to the sample attribute information and the predicted attribute information to obtain a trained image parsing model" mentioned above, and this will not be elaborated in the embodiments of this specification.
[0073] Exemplarily, the processed sample object feature sequence can be divided into a training set, a validation set, and a test set according to a certain ratio (such as 6:2:2); the training set is used to train multiple behavior recognition models, the validation set is used to verify the behavior recognition classification ability of each behavior recognition model, and after obtaining multiple trained behavior recognition models upon completion of training, the test set is used to detect the training effect of each trained behavior recognition model, such as detecting indicators such as the accuracy, precision, average accuracy, mean average precision, confusion matrix, receiver operating characteristic (ROC) curve, and area under the curve (AUC) value of each trained behavior recognition model, and a behavior recognition model that can ultimately be used for behavior recognition is selected from multiple trained behavior recognition models based on these indicators.
[0074] Applying the solution of the embodiments of this specification improves the behavior recognition performance of the behavior recognition model by continuously training the behavior recognition model.
[0075] Step 208: Generate an identity recognition result of the target object according to the object feature information and the behavior feature information.
[0076] It should be noted that the identity recognition result of the target object refers to the result obtained by performing identity recognition on the target object in the video to be recognized. Taking the rail transit scenario as an example, the identity recognition result of the target object can be a passenger, a rescue worker, and the rescue workers include station staff, security personnel, volunteers, police officers, medical staff, firefighters, temporary rescue workers, etc.
[0077] In practical applications, there are various ways to generate the identity recognition result of the target object according to the object feature information and the behavior feature information, which are specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. In one possible implementation manner of this specification, an identity recognition model can be used to process the object feature information and the behavior feature information to obtain the identity recognition result of the target object. In another possible implementation manner of this specification, a decision tree can be used to perform a series of conditional judgments on the object feature information and the behavior feature information to determine the identity recognition result of the target object. In an optional embodiment of this specification, after generating the identity recognition result of the target object, the specific information of the target object can be screened from the object information library according to the identity recognition result. For example, face recognition can be performed to monitor the identity of a specific target object, and for a target object that has been previously entered into the system, specific information such as name, gender, and age is given, and for a target object that has not been previously entered into the system, identity category information is given, so as to improve the fineness and accuracy of identity judgment and provide information support for the decision-making of disposal personnel in case of emergency.
[0078] Applying the solution of the embodiments of this specification, since the object feature information is arranged based on time series, the dynamic changes of object features can be better understood, and the accuracy of behavior feature information can be improved. Incorporating the associated attribute information of associated objects into the object feature information can better understand the positional relationship and interaction behavior between the target object and surrounding objects, which helps to more accurately identify the identity of the target object. By analyzing the object feature sequences and behavior feature information in multiple video frames, the static features and dynamic behavior patterns of the target object can be more comprehensively captured, providing richer context information for the identity recognition process, reducing the misrecognition problem caused by a single recognition reference information, and at the same time supporting the identity recognition of objects not entered into the system, improving the accuracy and flexibility of identity recognition.
[0079] In an optional embodiment of this specification, the above-mentioned generating an identity recognition result of a target object based on object feature information and behavior feature information may include the following steps: Input the object feature information and behavior feature information into an identity recognition model to obtain an identity recognition result of the target object.
[0080] It should be noted that the identity recognition model refers to a deep learning model with the ability to recognize identities based on object feature information and behavior feature information. Since the input of the identity recognition model includes object feature information and behavior feature information, the identity recognition model can be called an object feature and behavior feature fusion classification model. The identity recognition model is trained based on sample videos carrying sample object feature information, sample behavior feature information, and sample identity recognition results. The identity recognition model is usually an m-layer multi-layer fully connected network, where m is a positive integer and is specifically set according to actual situations. Refer to Figure 4 , Figure 4 shows a schematic structural diagram of an identity recognition model provided by an embodiment of this specification. The identity recognition model includes an m-layer multi-layer fully connected network for processing object feature information and behavior feature information. The number of neurons in each layer can be set to e1, e2,..., em. The input layer is used to input object feature information and behavior feature information, and the intermediate layers all contain the activation function ReLU. In addition, a Dropout layer is introduced in the identity recognition model to reduce overfitting. The output dimension of the last fully connected layer of the identity recognition model is the number of categories of the target object, and the output layer uses the softmax activation function for multi-classification to obtain the identity recognition result.
[0081] In practical applications, when inputting object feature information and behavior feature information into the identity recognition model, the object feature information and behavior feature information can be merged through concatenation, and the merged feature obtained by the merger is input into the identity recognition model to obtain an identity recognition result of the target object.
[0082] Applying the solution of the embodiments of this specification, by combining the object feature sequence and the behavior feature information for identity recognition, it is possible to more comprehensively capture the static features and dynamic behavior patterns of the target object, provide richer context information for the identity recognition process, reduce the misrecognition problem caused by the single reference information for recognition, and improve the accuracy of identity recognition.
[0083] In an optional embodiment of this specification, the training method of the identity recognition model is described, that is, before inputting the object feature information and the behavior feature information into the identity recognition model to obtain the identity recognition result of the target object, the following steps may further be included: Obtain a sample video, where the sample video carries the sample object feature information, the sample behavior feature information, and the sample identity recognition result of the sample object; Input the sample object feature information and the sample behavior feature information into the identity recognition model to obtain a predicted identity recognition result; Train the identity recognition model according to the sample identity recognition result and the predicted identity recognition result to obtain a trained identity recognition model.
[0084] It should be noted that the sample identity recognition result is the recognition target of the identity recognition model and is used to guide the training process of the identity recognition model. The predicted identity recognition result is the result obtained by the identity recognition model predicting the sample object feature information and the sample behavior feature information. By comparing the sample object feature information and the sample behavior feature information, the identity classification ability of the identity recognition model can be determined. The sample identity recognition result can be obtained by manual annotation. The sample identity recognition results are such as station staff, security personnel, volunteers, police officers, medical staff, firefighters, temporary rescue personnel, and so on.
[0085] In practical applications, when training an identity recognition model using a sample video carrying the sample object's feature information, sample behavior feature information, and sample identity recognition result, in one possible implementation of this specification, a sample video carrying the sample object's feature information, sample behavior feature information, and sample identity recognition result can be used to train an identity recognition model to obtain a trained identity recognition model. In another possible implementation of this specification, a sample video carrying the sample object's feature information, sample behavior feature information, and sample identity recognition result can be used to train multiple identity recognition models to obtain multiple trained identity recognition models, and then, from the multiple trained identity recognition models, select an identity recognition model with very good identity recognition effect. Hyperparameters such as the model architecture, number of training rounds, solver optimizer (SGD can be selected), step size, batch_size, etc. of the multiple identity recognition models can be set according to actual needs. Further, the method of training any identity recognition model, "training the identity recognition model based on the sample identity recognition result and the predicted identity recognition result to obtain a trained identity recognition model", can refer to the implementation method of "training the image parsing model based on the sample attribute information and the predicted attribute information to obtain a trained image parsing model" above, and this will not be elaborated in the embodiments of this specification.
[0086] Exemplarily, the sample object feature information and the sample behavior feature information can be divided into a training set, a validation set, and a test set according to a certain ratio (such as 6:2:2); use the training set to train multiple identity recognition models, use the validation set to verify the identity recognition classification ability of each identity recognition model, and after training is completed to obtain multiple trained identity recognition models, use the test set to detect the training effect of each trained identity recognition model, such as detecting indicators such as the accuracy rate, precision rate, recall rate, and F1 score of each trained identity recognition model, and select the final identity recognition model that can be used for identity recognition from the multiple trained identity recognition models according to these indicators.
[0087] Applying the solution of the embodiments of this specification improves the identity recognition performance of the identity recognition model by continuously training the identity recognition model.
[0088] Automatic identification of passengers and rescue workers in the rail transit scenario is a complex issue involving multiple aspects such as safety, technology, and emergency response. In the event of an emergency in the rail transit field, passengers may be trapped, and rapid and effective rescue operations are required. In such a situation, automatic identification of the categories of passengers and rescue workers becomes particularly important to ensure that emergency commanders accurately perceive the specific situation and progress of on-site rescue, provide support for rescue decision-making, and effectively monitor the implementation effect of rescue operations. The following is combined with the attached Figure 5, taking the application of the identity recognition method provided in this specification in the rail transit scenario as an example, the identity recognition method will be further described. Among them, Figure 5 Fig. Figure 5 shows a flowchart of the processing procedure of an identity recognition method provided by an embodiment of this specification, which specifically includes the following steps: Step 502: Obtain the video to be recognized in the rail transit scenario, where the video to be recognized includes multiple video frames.
[0089] Step 504: Input the multiple video frames into the image parsing model respectively to obtain the attribute information of multiple objects, where the attribute information includes the position information of the object in the video frame.
[0090] Step 506: Determine the position association index according to the target position information of the target object and the candidate position information of the candidate object, and screen out the associated object of the target object from the candidate objects according to the position association index.
[0091] Step 508: Generate the object feature information of the target object according to the target attribute information of the target object and the associated attribute information of the associated object.
[0092] Step 510: Arrange the object feature information of the target object according to the time sequence of the multiple video frames to obtain the object feature sequence of the video to be recognized.
[0093] Step 512: Process the object feature sequence according to the model processing conditions of the behavior recognition model to obtain the processed object feature sequence, and input the processed object feature sequence into the behavior recognition model to obtain the behavior feature information of the target object.
[0094] Step 514: Input the object feature information and the behavior feature information into the identity recognition model to obtain the identity recognition result of the target object.
[0095] By applying the solution of the embodiment of this specification, it is possible to handle the recognition of rescue personnel categories that have not been pre-entered into the system, and the application scalability is strong. Moreover, the identity categories of passengers and rescue personnel are recognized by integrating the whole-body image features and whole-body behavior features of the target object, and the recognition accuracy is high; and when recognizing the whole-body behavior of personnel, the information of associated objects (including items and other personnel) in the video frame is integrated, and the behavior classification is also more accurate. When dealing with the situation where the clothing of rescue personnel is severely damaged or temporary personnel act as rescue personnel to carry out rescue operations, it is possible to better classify the identities of passengers and rescue personnel, providing data support for the statistical feedback of the actual situation on the rescue site.
[0096] See Figure 6 , Figure 6The flowchart of the processing procedure of another identity recognition method provided by an embodiment of this specification is shown, which specifically includes: obtaining a sample video containing the rescue behaviors of passengers and rescue personnel in the rail transit scenario. Performing data processing on the sample video, including data cleaning and data annotation. When performing data annotation, the sample attribute information of the sample objects, the sample behavior feature information of each sample object, and the sample identity recognition results of the sample objects can be annotated. Training an image parsing model using the sample attribute information of the sample objects to obtain a trained image parsing model. Determining the sample associated objects that are relatively close in distance to the sample target objects in the same video frame, and sorting out the sample attribute information of the sample target objects and the sample attribute information of the sample associated objects in chronological order to obtain a sample object feature sequence. Performing time series feature segmentation and dimension conversion on the sample object feature sequence to obtain a processed sample object feature sequence. Training a behavior recognition model using the processed sample object feature sequence and the sample behavior feature information of each sample object to obtain a trained behavior recognition model. Training an identity recognition model using the sample identity recognition results of the sample objects, the sample attribute information of the sample objects, and the sample behavior feature information of each sample object to obtain a trained identity recognition model.
[0097] Obtain the video to be recognized in the rail transit scenario, input each video frame in the video to be recognized into the trained image parsing model, and output the attribute information of multiple objects in the video to be recognized. Among them, the attribute information includes the position information of the object in the video frame. Perform object tracking on multiple video frames according to the position information of multiple objects, and screen out the associated objects of the target object from the candidate objects. Generate an object feature sequence of the video to be recognized according to the target attribute information of the target object and the associated attribute information of the associated object. Among them, the object feature sequence includes the object feature information of the target object arranged based on time series, and the object feature information includes the target attribute information of the target object and the associated attribute information of the associated object. Perform feature segmentation and dimension conversion on the object feature sequence to obtain a processed object feature sequence. Input the processed object feature sequence into the trained behavior recognition model to obtain the behavior feature information (rescue behavior classification feature information) of the target object. Input the object feature information and behavior feature information of the target object into the trained identity recognition model to obtain the identity recognition result (passenger and rescue personnel category classification information) of the target object.
[0098] Corresponding to the above method embodiment, this specification also provides an embodiment of an identity recognition device. Figure 7 The structural schematic diagram of an identity recognition device provided by an embodiment of this specification is shown. As Figure 7 shown, this device includes: An acquisition module 702, configured to acquire a video to be recognized for an identity recognition task, where the video to be recognized includes multiple video frames; A parsing module 704, configured to parse a plurality of video frames to obtain an object feature sequence of a video to be recognized, where the object feature sequence includes object feature information of a target object arranged in time series, and the object feature information includes target attribute information of the target object and associated attribute information of an associated object, and the associated object is an object in the video frame that has a location association with the target object; A first generation module 706, configured to generate behavior feature information of the target object according to the object feature sequence; A second generation module 708, configured to generate an identity recognition result of the target object according to the object feature information and the behavior feature information.
[0099] Optionally, the video to be recognized includes a plurality of objects, and the plurality of objects include the target object; the parsing module 704 is further configured to parse the plurality of video frames to obtain attribute information of the plurality of objects, where the attribute information includes location information of the object in the video frame; filter out the associated object of the target object from the plurality of objects according to the location information; generate an object feature sequence of the video to be recognized according to the target attribute information of the target object and the associated attribute information of the associated object.
[0100] Optionally, the parsing module 704 is further configured to determine a location association index according to the target location information of the target object and the candidate location information of the candidate object, where the location association index corresponds to the candidate object one by one, and the location association index is used to represent the location association degree between the candidate location information of the candidate object and the target location information, and the candidate object is an object other than the target object among the plurality of objects; filter out the associated object of the target object from the candidate objects according to the location association index.
[0101] Optionally, the parsing module 704 is further configured to input the plurality of video frames into an image parsing model respectively to obtain attribute information of the plurality of objects.
[0102] Optionally, the apparatus further includes: a first adjustment module, configured to obtain a plurality of sample video frames of a sample video, where the sample video frames carry sample attribute information of a sample object; input the plurality of sample video frames into an image parsing model respectively to obtain predicted attribute information of the sample object; train the image parsing model according to the sample attribute information and the predicted attribute information to obtain a trained image parsing model.
[0103] Optionally, the first generation module 706 is further configured to process the object feature sequence according to the model processing conditions of a behavior recognition model to obtain a processed object feature sequence, where the processed object feature sequence conforms to the model processing conditions; input the processed object feature sequence into the behavior recognition model to obtain behavior feature information of the target object.
[0104] Optionally, the apparatus further includes: a second adjustment module configured to obtain a sample video, where the sample video includes a plurality of sample video frames, and the sample video carries a sample object feature sequence and sample behavior feature information; process the sample object feature sequence according to the model processing conditions of the behavior recognition model to obtain a processed sample object feature sequence, where the processed sample object feature sequence conforms to the model processing conditions; input the processed sample object feature sequence into the behavior recognition model to obtain predicted behavior feature information; and train the behavior recognition model according to the sample behavior feature information and the predicted behavior feature information to obtain a trained behavior recognition model.
[0105] Optionally, the second generation module 708 is further configured to input the object feature information and the behavior feature information into an identity recognition model to obtain an identity recognition result of the target object.
[0106] Optionally, the apparatus further includes: a third adjustment module configured to obtain a sample video, where the sample video carries sample object feature information, sample behavior feature information, and a sample identity recognition result of a sample object; input the sample object feature information and the sample behavior feature information into an identity recognition model to obtain a predicted identity recognition result; and train the identity recognition model according to the sample identity recognition result and the predicted identity recognition result to obtain a trained identity recognition model.
[0107] By applying the solution of the embodiment of this specification, since the object feature information is arranged based on time series, the dynamic changes of the object features can be better understood, and the accuracy of the behavior feature information can be improved. Incorporating the associated attribute information of the associated object into the object feature information can better understand the positional relationship and interaction behavior between the target object and the surrounding objects, which helps to more accurately identify the identity of the target object. By analyzing the object feature sequence and behavior feature information in multiple video frames, the static features and dynamic behavior patterns of the target object can be more comprehensively captured, providing richer context information for the identity recognition process, reducing the misrecognition problem caused by a single recognition reference information, and at the same time supporting the identity recognition of objects not entered into the system, improving the accuracy and flexibility of identity recognition.
[0108] The above is a schematic solution of an identity recognition apparatus according to this embodiment. It should be noted that the technical solution of this identity recognition apparatus and the technical solution of the above identity recognition method belong to the same concept. For the details not described in detail in the technical solution of the identity recognition apparatus, reference can be made to the description of the technical solution of the above identity recognition method.
[0109] Figure 8The block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0110] The computing device 800 further includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interfaces (e.g., a Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (WiMAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0111] In an embodiment of this specification, the above components of the computing device 800 and Figure 8 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 8 the shown block diagram of the computing device is for illustrative purposes only and is not a limitation on the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0112] The computing device 800 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a Personal Computer (PC). The computing device 800 can also be a mobile or stationary server.
[0113] Among them, the processor 820 is used to execute a computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above-mentioned identity recognition method are implemented.
[0114] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-mentioned identity recognition method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above-mentioned identity recognition method.
[0115] An embodiment of this specification also provides a computer-readable storage medium, which stores a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above-mentioned identity recognition method are implemented.
[0116] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-mentioned identity recognition method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above-mentioned identity recognition method.
[0117] An embodiment of this specification also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above-mentioned identity recognition method are implemented.
[0118] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned identity recognition method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above-mentioned identity recognition method.
[0119] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0120] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0121] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0122] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0123] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. An identity recognition method, characterized in that, Including: Obtain a video to be recognized for an identity recognition task, where the video to be recognized includes a plurality of video frames; Parse the plurality of video frames to obtain an object feature sequence of the video to be recognized, where the object feature sequence includes object feature information of a target object arranged based on time sequence, and the object feature information includes target attribute information of the target object and associated attribute information of an associated object, and the associated object is an object in the video frame that has a position association with the target object; Generate behavior feature information of the target object according to the object feature sequence; Generate an identity recognition result of the target object according to the object feature information and the behavior feature information.
2. The method according to claim 1, wherein The video to be recognized includes a plurality of objects, and the plurality of objects include the target object; The step of parsing the plurality of video frames to obtain an object feature sequence of the video to be recognized includes: Parse the plurality of video frames to obtain attribute information of the plurality of objects, where the attribute information includes position information of the object in the video frame; Screen out the associated objects of the target object from the plurality of objects according to the position information; Generate an object feature sequence of the video to be recognized according to the target attribute information of the target object and the associated attribute information of the associated objects.
3. The method according to claim 2, characterized in that The step of screening out the associated objects of the target object from the plurality of objects according to the position information includes: Determine a position association index according to the target position information of the target object and the candidate position information of a candidate object, where the position association index corresponds to the candidate object one by one, and the position association index is used to represent the position association degree between the candidate position information of the candidate object and the target position information, and the candidate object is an object other than the target object in the plurality of objects; Screen out the associated objects of the target object from the candidate objects according to the position association index.
4. The method according to claim 2, characterized in that, The step of parsing the plurality of video frames to obtain attribute information of the plurality of objects includes: Input the plurality of video frames into an image parsing model respectively to obtain attribute information of the plurality of objects.
5. The method according to claim 4, wherein Before inputting the plurality of video frames into the image parsing model respectively to obtain attribute information of the plurality of objects, it further includes: Obtain a plurality of sample video frames of a sample video, where the sample video frames carry sample attribute information of sample objects; Input the plurality of sample video frames into the image parsing model respectively to obtain predicted attribute information of the sample objects; Train the image parsing model according to the sample attribute information and the predicted attribute information to obtain a trained image parsing model.
6. The method according to claim 1, wherein The step of generating behavior feature information of the target object according to the object feature sequence includes: Process the object feature sequence according to the model processing conditions of a behavior recognition model to obtain a processed object feature sequence, where the processed object feature sequence meets the model processing conditions; Input the processed object feature sequence into the behavior recognition model to obtain behavior feature information of the target object.
7. The method according to claim 6, wherein Before inputting the processed object feature sequence into the behavior recognition model to obtain the behavior feature information of the target object, the following steps are further included: Obtain a sample video, where the sample video includes a plurality of sample video frames, and the sample video carries a sample object feature sequence and sample behavior feature information; Process the sample object feature sequence according to the model processing conditions of the behavior recognition model to obtain a processed sample object feature sequence, where the processed sample object feature sequence conforms to the model processing conditions; Input the processed sample object feature sequence into the behavior recognition model to obtain predicted behavior feature information; Train the behavior recognition model according to the sample behavior feature information and the predicted behavior feature information to obtain a trained behavior recognition model.
8. The method according to claim 1, characterized in that, The step of generating the identity recognition result of the target object according to the object feature information and the behavior feature information includes: Input the object feature information and the behavior feature information into an identity recognition model to obtain the identity recognition result of the target object.
9. The method according to claim 8, wherein Before inputting the object feature information and the behavior feature information into the identity recognition model to obtain the identity recognition result of the target object, the following steps are further included: Obtain a sample video, where the sample video carries sample object feature information, sample behavior feature information, and a sample identity recognition result of the sample object; Input the sample object feature information and the sample behavior feature information into the identity recognition model to obtain a predicted identity recognition result; Train the identity recognition model according to the sample identity recognition result and the predicted identity recognition result to obtain a trained identity recognition model.
10. An identity recognition device, characterized in that, It includes: An acquisition module configured to acquire a video to be recognized for an identity recognition task, where the video to be recognized includes a plurality of video frames; An analysis module configured to analyze the plurality of video frames to obtain an object feature sequence of the video to be recognized, where the object feature sequence includes object feature information of a target object arranged in time series, and the object feature information includes target attribute information of the target object and associated attribute information of an associated object, and the associated object is an object in the video frame that has a position association with the target object; A first generation module configured to generate behavior feature information of the target object according to the object feature sequence; A second generation module configured to generate an identity recognition result of the target object according to the object feature information and the behavior feature information.
11. A computing device, characterized in that, It includes: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium, characterized in that, It stores computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product, characterized in that, It includes computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Human body action recognition method and equipment
CN110135246A
Object behavior identification method and device
CN111325292A
Behavior recognition method and device, equipment and storage medium
CN113111839A
Video behavior recognition method and related device, electronic equipment and storage medium
CN118587759A
Motion recognition method and motion recognition system combining segmentation and depth estimation
CN119723657A