Video-based object recognition method, apparatus and electronic device

By extracting single-frame images from videos and generating target subtitles, and using deep learning models to identify object types, the problem of high labor costs caused by manual verification is solved, and efficient and accurate object type identification is achieved.

CN115294502BActive Publication Date: 2026-07-21INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2022-08-18
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies rely on manual verification of financial dual-recorded videos to identify the type of object, resulting in high labor costs.

Method used

By acquiring the target video, extracting single-frame images and performing feature extraction, generating target subtitles, and using a deep learning model to identify object types, the text conversion and accurate recognition of video content can be achieved without extracting features from each frame.

Benefits of technology

It reduces labor costs, improves recognition efficiency and accuracy, avoids the shortcomings of manual analysis, and achieves automated object type recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294502B_ABST
    Figure CN115294502B_ABST
Patent Text Reader

Abstract

The application discloses a kind of object identification method, device and electronic equipment based on video.It relates to artificial intelligence field, the method includes: obtaining target video, wherein, target video is used to record the handling process of first target object for second target object to handle pending business;Multiple single-frame images are extracted from target video, and the target feature sequence corresponding to each single-frame image is obtained by carrying out feature extraction on each single-frame image;Based on target feature sequence, determine a plurality of target subtitles for describing the video content of target video;Identify the object type of first target object based on multiple target subtitles, wherein, object type is used to indicate whether the handling operation of first target object to handle pending business conforms to target rule.The present application solves the technical problem of high labor cost caused by relying on artificial verification of business handling video to identify object type in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and more specifically, to a video-based object recognition method, apparatus, and electronic device. Background Technology

[0002] As an economic form driven by digital productivity, the digital economy plays an increasingly prominent role in the national economy and is bringing about disruptive changes to people's lifestyles, ways of thinking, and behaviors. The booming development of online life and service models has further enhanced the influence of digital technology and promoted the rapid advancement of audio and video technology services. In high-risk transaction scenarios such as changing important personal / corporate information, activating a transaction function, or opening an account, financial dual-recording video applications used to record the business transactions between customer service representatives and customers are very common.

[0003] Currently, to ensure the security of transactions in dual-recording scenarios, it is necessary to manually analyze financial dual-recording videos to identify whether the behavior of relevant customer service personnel is standardized, which has the problem of high labor costs.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This invention provides a video-based object recognition method, apparatus, and electronic device to at least solve the technical problem of high labor costs caused by relying on manual verification of business processing videos to identify object types in the prior art.

[0006] According to one aspect of the present invention, a video-based object recognition method is provided, comprising: acquiring a target video, wherein the target video is used to record the process of a first target object handling pending business for a second target object; extracting multiple single-frame images from the target video, performing feature extraction on each single-frame image to obtain a target feature sequence corresponding to each single-frame image; determining multiple target subtitles based on the target feature sequence to describe the video content of the target video; and identifying the object type of the first target object based on the multiple target subtitles, wherein the object type is used to characterize whether the handling operation of the pending business by the first target object conforms to target rules.

[0007] Furthermore, the video-based object recognition method also includes: extracting features from multiple single-frame images to obtain a feature sequence corresponding to each single-frame image; and extracting features from the feature sequence corresponding to each single-frame image based on a target step size to obtain a target feature sequence corresponding to each single-frame image, wherein the target step size corresponds to the single-frame image.

[0008] Furthermore, the video-based object recognition method also includes: identifying the video content of the target video to obtain the type of pending business corresponding to the video content; dividing the video duration of the target video into at least one time interval based on the type of pending business, and determining the importance level of each time interval, wherein the importance level at least characterizes the degree of correlation between the processing operation of the first target object for the pending business and the pending business within the current time interval; determining the number of frames to be extracted from the video content corresponding to each time interval based on the importance level of each time interval; and extracting multiple single-frame images from the target video based on the number of extracted frames.

[0009] Furthermore, the video-based object recognition method also includes: before extracting features from the feature sequence corresponding to each single frame image based on the target step size to obtain the target feature sequence corresponding to each single frame image, determining the target step size corresponding to the single frame image corresponding to each time interval based on the importance level of each time interval.

[0010] Furthermore, the video-based object recognition method also includes: converting the target feature sequence into a target vector of a preset dimension through an encoder in the target model; and decoding the target vector through a decoder in the target model to obtain multiple target subtitles used to describe the video content of the target video.

[0011] Furthermore, the video-based object recognition method also includes: determining the display time information corresponding to each target subtitle; determining multiple target single-frame images from the target video based on the display time information; combining multiple target subtitles into at least one target single-frame image corresponding to the target subtitle based on the time information to obtain the video to be inspected; and identifying the object type of the first target object based on the video to be inspected.

[0012] Furthermore, the video-based object recognition method also includes: determining the storage time of at least one video stored in the target storage area; and determining the target video from the at least one video based on the storage time.

[0013] According to another aspect of the present invention, a video-based object recognition device is also provided, comprising: an acquisition module for acquiring a target video, wherein the target video is used to record the process of a first target object handling pending business for a second target object; an extraction module for extracting multiple single-frame images from the target video, performing feature extraction on each single-frame image to obtain a target feature sequence corresponding to each single-frame image; a determination module for determining multiple target subtitles for describing the video content of the target video based on the target feature sequence; and a recognition module for recognizing the object type of the first target object based on the multiple target subtitles, wherein the object type is used to characterize whether the handling operation of the pending business by the first target object conforms to the target rules.

[0014] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described video-based object recognition method at runtime.

[0015] According to another aspect of the present invention, an electronic device is also provided, the electronic device including one or more processors; a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are configured to run the programs, wherein the programs are configured to execute the above-described video-based object recognition method at runtime.

[0016] In this embodiment of the invention, multiple target subtitles describing the video content of a target video are generated to identify the object type of a first target object. This is achieved by acquiring the target video, extracting multiple single-frame images from it, performing feature extraction on each single-frame image to obtain a target feature sequence corresponding to each single-frame image, and then determining multiple target subtitles describing the video content based on these target feature sequences. This allows for the identification of the object type of the first target object based on the multiple target subtitles. The target video records the process of the first target object handling pending business for a second target object, and the object type characterizes whether the first target object's handling of pending business conforms to target rules.

[0017] In the above process, multiple single-frame images are extracted from the target video, and the target feature sequence corresponding to each single-frame image is determined. This avoids the low work efficiency caused by extracting features from every single frame of the target video. Furthermore, based on the target feature sequence, target subtitles are determined, realizing the text conversion of the video content. This avoids the problems of high labor costs and easy omissions caused by manual analysis of video content. At the same time, it avoids the low recognition accuracy caused by the lack of contextual information when analyzing a single screenshot in the target video. By identifying the object type of the first target object based on multiple target subtitles corresponding to the target video, the low efficiency and high labor costs caused by the large workload when manually recognizing target subtitles are further avoided.

[0018] Therefore, the solution provided in this application achieves the purpose of generating multiple target subtitles to describe the video content of the target video in order to identify the object type of the first target object, thereby achieving the technical effect of reducing labor costs and solving the technical problem of high labor costs caused by relying on manual verification of business processing videos to identify object types in the prior art. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0020] Figure 1 This is a schematic diagram of an optional object recognition system according to an embodiment of the present invention;

[0021] Figure 2 This is a schematic diagram of an optional video-based object recognition method according to an embodiment of the present invention;

[0022] Figure 3 This is a schematic diagram illustrating the execution of an optional data preprocessing module according to an embodiment of the present invention;

[0023] Figure 4 This is an execution diagram of an optional video dense subtitle generation module according to an embodiment of the present invention;

[0024] Figure 5 This is an execution diagram of an optional target model according to an embodiment of the present invention;

[0025] Figure 6 This is an execution diagram of an optional result processing module according to an embodiment of the present invention;

[0026] Figure 7 This is a schematic diagram of an optional video-based object recognition device according to an embodiment of the present invention;

[0027] Figure 8 This is a schematic diagram of an optional electronic device according to an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties. For example, this system has an interface with the relevant user or organization. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving consent from the aforementioned user or organization.

[0031] Example 1

[0032] According to an embodiment of the present invention, an embodiment of a video-based object recognition method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0033] In this application, an optional object recognition system is used as the implementing entity, such as... Figure 1 As shown, the object recognition system may include a data preprocessing module, a video dense subtitle generation module, and a result processing module.

[0034] Figure 2 This is a schematic diagram of an optional video-based object recognition method according to an embodiment of the present invention, such as... Figure 2 As shown, the method includes the following steps:

[0035] Step S201: Obtain the target video, wherein the target video is used to record the process of the first target object handling pending business for the second target object.

[0036] Optionally, the target video can be acquired through electronic devices, servers, application systems, etc., as described in this application. Figure 3As shown, the target video is obtained through the data preprocessing module of the object recognition system. The target video can be generated by an image acquisition device installed in the target financial institution. The target video at least records the process of a customer service representative (i.e., the aforementioned first target object) handling pending business for a customer (i.e., the aforementioned second target object). In this application, the video containing both the customer service representative and the customer (i.e., the target video) is confirmed as a dual-recording video. Optionally, the preprocessing module can directly interact with the aforementioned image acquisition device to obtain the target video, and the preprocessing module can retrieve pre-stored target videos from a preset storage area.

[0037] Step S202: Extract multiple single-frame images from the target video, perform feature extraction on each single-frame image, and obtain the target feature sequence corresponding to each single-frame image.

[0038] In step S202, the data preprocessing module in the object recognition system can extract multiple single-frame images from the target video based on an existing video analysis model. The aforementioned video analysis model can be a C3D model. The object recognition system can determine the number of frames between each single-frame image based on the video content of the target video or a preset value. The number of frames between multiple single-frame images can be the same or different.

[0039] Furthermore, after extracting multiple single-frame images, the object recognition system can perform feature extraction on each extracted single-frame image based on the video analysis model, thereby obtaining the target feature sequence corresponding to each single-frame image.

[0040] It should be noted that by extracting multiple single-frame images from the target video, the problem of low work efficiency caused by performing feature extraction on every single frame of the target video is avoided. By extracting features from the extracted single-frame images, the target feature sequence can be obtained, which can effectively remove noise from the single-frame images and facilitate the processing of the single-frame images.

[0041] Step S203: Based on the target feature sequence, determine multiple target subtitles to describe the video content of the target video.

[0042] In step S203, the object recognition system can input the target feature sequence into a preset deep learning model to determine dense captioning used to describe the video content of the target video. The dense captioning consists of multiple target captions and is used to describe many local details in the image using natural language. The aforementioned deep learning model can be a transformer deep learning model using a global residual neural network.

[0043] It should be noted that by determining the target subtitles based on the target feature sequence, the text conversion of video content is realized, thereby avoiding the problems of high labor costs and easy omissions caused by manual analysis of video content. At the same time, it avoids the problem of low recognition accuracy caused by the lack of contextual information when analyzing a single screenshot in the target video.

[0044] Step S204: Identify the object type of the first target object based on multiple target subtitles, wherein the object type is used to characterize whether the processing operation of the pending business of the first target object conforms to the target rules.

[0045] In step S204, the object recognition system can determine the processing operation of the first target object when handling pending business based on the target subtitle, and can compare the processing operation with the processing operation specified in the target rule to identify the object type of the first target object.

[0046] It should be noted that by identifying the object type of the first target object based on multiple target subtitles corresponding to the target video, the inefficiency and high labor costs caused by the large workload of manual target subtitle identification are avoided.

[0047] Based on the scheme defined in steps S201 to S204 above, it can be understood that in this embodiment of the invention, multiple target subtitles are generated to describe the video content of the target video in order to identify the object type of the first target object. This is achieved by acquiring the target video, extracting multiple single-frame images from the target video, performing feature extraction on each single-frame image to obtain a target feature sequence corresponding to each single-frame image, and then determining multiple target subtitles to describe the video content of the target video based on the target feature sequence, thereby identifying the object type of the first target object based on the multiple target subtitles. The target video is used to record the process of the first target object handling pending business for the second target object, and the object type is used to characterize whether the handling operation of the pending business by the first target object conforms to the target rules.

[0048] It is noteworthy that in the above process, multiple single-frame images are extracted from the target video, and the target feature sequence corresponding to each single-frame image is determined. This avoids the low work efficiency caused by extracting features from every single frame of the target video. Furthermore, based on the target feature sequence, the target subtitles are determined, realizing the text conversion of the video content. This avoids the problems of high labor costs and easy omissions caused by manual analysis of video content. At the same time, it avoids the low recognition accuracy caused by the lack of contextual information when analyzing a single screenshot in the target video. By identifying the object type of the first target object based on multiple target subtitles corresponding to the target video, the low efficiency and high labor costs caused by the large workload when manually recognizing target subtitles are further avoided.

[0049] Therefore, the solution provided in this application achieves the purpose of generating multiple target subtitles to describe the video content of the target video in order to identify the object type of the first target object, thereby achieving the technical effect of reducing labor costs and solving the technical problem of high labor costs caused by relying on manual verification of business processing videos to identify object types in the prior art.

[0050] In one optional embodiment, during the process of extracting features from each single-frame image to obtain the target feature sequence corresponding to each single-frame image, such as... Figure 3 As shown, the data preprocessing module in the object recognition system can utilize a C3D video analysis model to first extract features from multiple single-frame images, obtaining the feature sequence corresponding to each single-frame image and outputting the feature sequence, then... Figure 4 As shown, the feature sequence is then obtained through the video dense subtitle generation module of the object recognition system, and feature extraction processing is performed on the feature sequence corresponding to each single frame image based on the target strides. The aforementioned feature sequence consists of multiple vectors, and the aforementioned target strides correspond to a single frame image, representing the number of vectors between the currently extracted vector and the next extracted vector in the current dimension of the feature sequence corresponding to the current single frame image. For example, when the target strides corresponding to a single frame image represent a stride of 2 in a certain dimension, the 1st, 4th, 7th, 10th... vectors in that dimension of the feature sequence corresponding to that single frame image are extracted and combined to form the target feature sequence. Optionally, the target strides corresponding to all single frames images can be exactly the same, completely different, or only partially the same.

[0051] It should be noted that by extracting features from the feature sequence corresponding to each single frame image based on the target step size, the number of features corresponding to each single frame image is reduced, thereby improving the working efficiency of the user recognition system, that is, improving the recognition efficiency.

[0052] In one optional embodiment, different methods for extracting single-frame images can be set for the target video based on whether the first target object in the target video is handling pending business for the second target object. Optionally, in the process of extracting multiple single-frame images from the target video, the data preprocessing module can identify the video content of the target video to obtain the pending business type corresponding to the video content, then divide the video duration of the target video into at least one time interval based on the pending business type, and determine the importance level of each time interval. Then, based on the importance level of each time interval, determine the number of frames to be extracted from the video content corresponding to each time interval, thereby extracting multiple single-frame images from the target video based on the number of extracted frames. The importance level at least characterizes the degree of correlation between the handling of pending business by the first target object and the pending business within the current time interval.

[0053] Specifically, the data preprocessing module can identify the video content in the target video to determine the type of pending business that the first target object in the target video is handling for the second target object. Further, the data preprocessing module can retrieve the splitting rules corresponding to the pending business type from a preset target storage area, and split the video duration of the target video into at least one time interval based on the splitting rules corresponding to the pending business type. For example, for pending business A, whose process is divided into 5 operation steps, the video duration of the target video can be split based on the time duration corresponding to these 5 operation steps.

[0054] Furthermore, after breaking down the video duration of the target video, the data preprocessing module can determine the importance level of each time interval based on the type of pending task. For example, for pending task A, if the first target object does not participate in the first operation step, but the first target object needs to participate in the second operation step, then the importance level of the time interval corresponding to the first operation step can be determined as level one, and the importance level of the time interval corresponding to the second operation step can be determined as level two, with level two being higher than level one. As another example, for pending task A, if the first target object's action in the second operation step is to print the materials provided by the second target object, and the first target object's action in the third operation step is to help the second target object fill in relevant information, then the correlation between the action performed by the first target object in the second operation step and the pending task can be determined as lower than the correlation between the action performed by the first target object in the third operation step and the pending task. In other words, the importance level of the time interval corresponding to the second operation step can be determined as level three, with level three being higher than level two.

[0055] Optionally, after determining the importance level of each time interval corresponding to the target video, the data preprocessing module can obtain target extraction information from the target storage area, and then determine the number of frames to be extracted from the video content corresponding to each time interval based on the target extraction information. The target extraction information is used to characterize the correspondence between each importance level and the number of extracted frames. Optionally, the target extraction information can be information used to limit the number of frames between adjacent single-frame images to be extracted. Then, the data preprocessing module can use a video analysis model to extract multiple single-frame images from the target video based on the number of extracted frames.

[0056] It should be noted that by setting different methods for extracting single-frame images from the target video based on the different tasks that the first target object performs for the second target object in the target video, the amount of features that need to be processed can be further reduced while ensuring the recognition effect, thereby effectively improving the recognition efficiency.

[0057] In one alternative embodiment, before extracting features from the feature sequence corresponding to each single frame image based on the target step size to obtain the target feature sequence corresponding to each single frame image, the object recognition system can determine the target step size corresponding to the single frame image corresponding to each time interval based on the importance level of each time interval.

[0058] Optionally, after determining the importance level of each time interval, the video dense subtitle generation module can obtain the target step size corresponding to each time interval from the target storage area. This allows for the processing of single-frame images in the time interval corresponding to the target step size based on slightly different target step sizes. This further reduces the amount of features that need to be processed while ensuring the recognition effect, thus effectively improving the recognition efficiency.

[0059] In an alternative embodiment, during the process of determining multiple target subtitles for describing the video content of a target video based on a target feature sequence, such as Figure 4 , Figure 5 As shown, the video dense caption generation module can input the target feature sequence into the transformer video information extraction model (i.e., the target model) of the decoder-encoder framework. After the target model obtains the target feature sequence, it can use the transformer learning network to convert the target feature sequence into a target vector of a preset dimension through the encoder in the model. Then, the decoder in the model decodes the target vector to obtain and output multiple target captions describing the video content of the target video. Each target caption has a corresponding timestamp. The aforementioned transformer video information extraction model can be a model using a global residual neural network.

[0060] It should be noted that by processing the target feature sequence using a target model based on a decoder-encoder framework, the accurate determination of multiple target subtitles is achieved, thereby improving the accuracy of object type recognition for the first target object.

[0061] In an optional embodiment, during the process of identifying the object type of a first target object based on multiple target subtitles, the object recognition system can determine the display time information corresponding to each target subtitle, then determine multiple target single-frame images from the target video based on the display time information, and then combine multiple target subtitles into at least one target single-frame image corresponding to the target subtitle based on the time information to obtain the video to be inspected, thereby identifying the object type of the first target object based on the video to be inspected.

[0062] Optional, such as Figure 6 As shown, after acquiring multiple target subtitles describing the target video, the result processing module can determine the display time information corresponding to the target subtitles based on the timestamps corresponding to the target subtitles. Then, it determines the single-frame image corresponding to each time point in the target video that corresponds to at least one time point represented by the display time information, thereby obtaining the target single-frame image. Further, the result processing module can combine the current target subtitle with all its corresponding target single-frame images. After combining all target subtitles according to the aforementioned method, it obtains the video to be inspected marked with the target subtitles, thereby enabling the identification of the object type of the first target object based on the video to be inspected.

[0063] It should be noted that by combining the target subtitles with the target video, the object recognition system can more quickly identify the object type of the first target object and the operational steps in which it performs non-standard operations, thereby further improving recognition efficiency.

[0064] In one alternative embodiment, during the acquisition of the target video, the object recognition system can determine the storage time of at least one video in the target storage area, thereby identifying the target video from at least one video based on the storage time.

[0065] Optionally, the object recognition system can retrieve and process the earliest video from the target storage area, according to the order of its storage time. Alternatively, the object recognition system can prioritize processing the most recently stored video.

[0066] It should be noted that by determining the target video based on the video's storage time, this application is more reasonable when processing multiple video files, thereby improving its applicability.

[0067] Therefore, the solution provided in this application achieves the purpose of generating multiple target subtitles to describe the video content of the target video in order to identify the object type of the first target object, thereby achieving the technical effect of reducing labor costs and solving the technical problem of high labor costs caused by relying on manual verification of business processing videos to identify object types in the prior art.

[0068] Example 2

[0069] According to an embodiment of the present invention, a video-based object recognition device is provided, wherein, Figure 7 This is a schematic diagram of an optional video-based object recognition device according to an embodiment of the present invention, such as... Figure 7 As shown, the device includes:

[0070] The acquisition module 701 is used to acquire the target video, wherein the target video is used to record the process of the first target object handling pending business for the second target object;

[0071] The extraction module 702 is used to extract multiple single-frame images from the target video, perform feature extraction on each single-frame image, and obtain a target feature sequence corresponding to each single-frame image;

[0072] The determination module 703 is used to determine multiple target subtitles for describing the video content of the target video based on the target feature sequence;

[0073] The identification module 704 is used to identify the object type of the first target object based on multiple target subtitles, wherein the object type is used to characterize whether the processing operation of the pending business of the first target object conforms to the target rules.

[0074] It should be noted that the above-mentioned acquisition module 701, extraction module 702, determination module 703 and identification module 704 correspond to steps S201 to S204 in the above embodiments. The four modules and the corresponding steps implement the same examples and application scenarios, but are not limited to the content disclosed in the above embodiment 1.

[0075] Optionally, the extraction module further includes: a first extraction submodule, used to extract features from multiple single-frame images to obtain a feature sequence corresponding to each single-frame image; and a second extraction submodule, used to extract features from the feature sequence corresponding to each single-frame image based on a target step size to obtain a target feature sequence corresponding to each single-frame image, wherein the target step size corresponds to the single-frame image.

[0076] Optionally, the extraction module further includes: a first identification submodule, used to identify the video content of the target video to obtain the type of pending business corresponding to the video content; a splitting module, used to split the video duration of the target video into at least one time interval based on the type of pending business, and determine the importance level of each time interval, wherein the importance level at least represents the degree of correlation between the processing operation of the pending business by the first target object and the pending business in the current time interval; a first determination submodule, used to determine the number of extracted frames of the video content corresponding to each time interval based on the importance level of each time interval; and a third extraction submodule, used to extract multiple single-frame images from the target video based on the number of extracted frames.

[0077] Optionally, the video-based object recognition device further includes: a second determining submodule, used to determine the target step size corresponding to the single frame image of each time interval based on the importance level of each time interval.

[0078] Optionally, the determining module further includes: a first processing module, used to convert the target feature sequence into a target vector of a preset dimension through an encoder in the target model; and a second processing module, used to decode the target vector through a decoder in the target model to obtain multiple target subtitles for describing the video content of the target video.

[0079] Optionally, the recognition module further includes: a third determining submodule, used to determine the display time information corresponding to each target subtitle; a fourth determining submodule, used to determine multiple target single-frame images from the target video based on the display time information; a combining module, used to combine multiple target subtitles into at least one target single-frame image corresponding to the target subtitle based on the time information to obtain the video to be inspected; and a second recognition submodule, used to identify the object type of the first target object based on the video to be inspected.

[0080] Optionally, the acquisition module further includes: a fifth determining submodule, used to determine the storage time of at least one video in the target storage area; and a sixth determining submodule, used to determine the target video from at least one video based on the storage time.

[0081] Example 3

[0082] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, wherein the computer program is configured to execute the above-described video-based object recognition method at runtime.

[0083] Example 4

[0084] According to another aspect of the present invention, an electronic device is also provided, wherein, Figure 8This is a schematic diagram of an optional electronic device according to an embodiment of the present invention, such as... Figure 8 As shown, the electronic device includes one or more processors; and a memory for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to run the programs, wherein the programs are configured to execute the aforementioned video-based object recognition method during runtime.

[0085] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0086] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0089] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0090] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0091] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A video-based object recognition method, characterized in that, include: Acquire a target video, wherein the target video is used to record the process of a first target object handling pending business for a second target object; Multiple single-frame images are extracted from the target video, and feature extraction is performed on each single-frame image to obtain a target feature sequence corresponding to each single-frame image. Based on the target feature sequence, multiple target subtitles are determined to describe the video content of the target video, and each of the multiple target subtitles has a corresponding timestamp; The object type of the first target object is identified based on the multiple target subtitles, wherein the object type is used to characterize whether the processing operation of the first target object for handling pending business conforms to the target rules; The step of identifying the object type of the first target object based on the plurality of target subtitles includes: determining the display time information corresponding to each target subtitle; determining a plurality of target single-frame images from the target video based on the display time information; combining the plurality of target subtitles into at least one target single-frame image corresponding to the target subtitle based on the time information to obtain the video to be inspected; and identifying the object type of the first target object based on the video to be inspected. The step of determining the display time information corresponding to each target subtitle includes: determining the display time information corresponding to each target subtitle based on the timestamp corresponding to each target subtitle.

2. The method according to claim 1, characterized in that, Feature extraction is performed on each single-frame image to obtain the target feature sequence corresponding to each single-frame image, including: Feature extraction is performed on the multiple single-frame images to obtain a feature sequence corresponding to each single-frame image; Based on the target step size, feature extraction is performed on the feature sequence corresponding to each single frame image to obtain the target feature sequence corresponding to each single frame image, wherein the target step size corresponds to the single frame image.

3. The method according to claim 2, characterized in that, Multiple single-frame images are extracted from the target video, including: The video content of the target video is identified to obtain the type of pending business corresponding to the video content; Based on the type of pending business, the video duration of the target video is divided into at least one time interval, and the importance level of each time interval is determined. The importance level at least represents the degree of correlation between the processing operation of the pending business by the first target object and the pending business within the current time interval. Based on the importance level of each time interval, determine the number of frames to be extracted from the video content corresponding to each time interval; Based on the number of extracted frames, multiple single-frame images are extracted from the target video.

4. The method according to claim 3, characterized in that, Before extracting features from the feature sequence corresponding to each single frame image based on the target step size to obtain the target feature sequence corresponding to each single frame image, the method further includes: Based on the importance level of each time interval, the target step size corresponding to the single frame image of each time interval is determined.

5. The method according to claim 1, characterized in that, Based on the target feature sequence, multiple target subtitles are determined to describe the video content of the target video, including: The target feature sequence is converted into a target vector of a preset dimension by the encoder in the target model; The target vector is decoded using the decoder in the target model to obtain multiple target subtitles that describe the video content of the target video.

6. The method according to claim 1, characterized in that, Obtain the target video, including: Determine the storage time for at least one video in the target storage area; Based on the storage time, a target video is determined from the at least one video.

7. A video-based object recognition device, characterized in that, include: The acquisition module is used to acquire the target video, wherein the target video is used to record the process of the first target object handling pending business for the second target object; The extraction module is used to extract multiple single-frame images from the target video, perform feature extraction on each single-frame image, and obtain a target feature sequence corresponding to each single-frame image; The determining module is used to determine multiple target subtitles for describing the video content of the target video based on the target feature sequence, wherein each of the multiple target subtitles has a corresponding timestamp; The identification module is used to identify the object type of the first target object based on the plurality of target subtitles, wherein the object type is used to characterize whether the processing operation of the first target object for handling pending business conforms to the target rules; The identification module includes: a third determining submodule for determining the display time information corresponding to each target subtitle; a fourth determining submodule for determining multiple target single-frame images from the target video based on the display time information; a combining module for combining the multiple target subtitles into at least one target single-frame image corresponding to the target subtitle based on the time information to obtain the video to be inspected; and a second identification submodule for identifying the object type of the first target object based on the video to be inspected. The third determining submodule is further configured to: determine the display time information corresponding to each target subtitle based on the timestamp corresponding to each target subtitle.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the video-based object recognition method according to any one of claims 1 to 6 when it is run.

9. An electronic device, characterized in that, The electronic device includes one or more processors; A memory for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to be configured to run the programs, wherein the programs are configured to execute the video-based object recognition method according to any one of claims 1 to 6.