Target object identification method and device, storage medium and electronic equipment
By processing the target video frames and text content, and combining the knowledge base and target model, we can identify whether specific objects on the street are violated, solving the problem of low recognition accuracy of single data, achieving higher recognition accuracy and complex scene processing capabilities.
Patent Information
- Application Number
- CN202411999126.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, it is illegal to identify whether a specific object on a street based on a single data, resulting in low identification accuracy.
By obtaining the target video, video processing is performed to obtain a set of video frames and text collections, knowledge retrieval is performed in the knowledge base based on these data, and the target big model is used to identify whether the target object is illegal.
The accuracy of identifying violations of target objects is improved, and the ability to identify complex scenarios is enhanced through the comprehensive utilization of multiple types of data.
Smart Images

Figure CN119942403A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and more specifically, to a method, device, storage medium and electronic device for identifying a target object. Background Art
[0002] With the acceleration of urbanization, the management of some specific objects on the streets (such as mobile merchants) has become a challenge for urban management. The illegal behaviors of some mobile merchants may involve food safety, traffic safety and other issues, affecting the development of citizens and the city.
[0003] At present, some recognition systems based on video surveillance are widely used, but most of them rely on a single data source, making it difficult to cope with various complex scenarios, affecting the accuracy of recognition.
[0004] Currently, no effective solution has been proposed to the problem that related technologies use single data to identify whether a specific object on the street is in violation of regulations, resulting in low recognition accuracy. Summary of the invention
[0005] The main purpose of the present application is to provide a method, device, storage medium and electronic device for identifying a target object, so as to solve the problem in the related art of identifying whether a specific object on the street is in violation of the regulations based on a single data, resulting in low recognition accuracy.
[0006] According to one aspect of an embodiment of the present invention, a method for identifying a target object is provided, comprising: acquiring a target video, wherein the target video is obtained by collecting images of a street, and the street includes at least one target object; performing video processing on the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes text content in the target video; performing knowledge retrieval in a knowledge base based on the video frame set and the text set to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to target object behaviors of different categories; identifying whether at least one target object is in violation of regulations based on at least one reference knowledge, the video frame set, and the text set through a target macro model to obtain an identification result.
[0007] Furthermore, the target object recognition method also includes: performing video frame extraction processing on the target video to obtain a video frame set; performing optical character recognition processing on the video content of the target video to obtain a text set.
[0008] Furthermore, the target object recognition method also includes: obtaining video features of a video frame set and text features of a text set; performing feature fusion on the video features and the text features to obtain multimodal features; and determining at least one reference knowledge from the knowledge in the knowledge base based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base.
[0009] Furthermore, the target object recognition method also includes: using a large visual model to perform feature extraction processing on a video frame set to obtain video features; using a large text model to perform feature extraction processing on a text set to obtain text features.
[0010] Furthermore, the target object identification method also includes: calculating the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base to obtain the similarity corresponding to each knowledge; sorting each knowledge in order from high to low according to the similarity to obtain sorted knowledge; determining the first N knowledge in the sorted knowledge as reference knowledge, where N is a positive integer.
[0011] Furthermore, the target object recognition method also includes: generating an indication statement based on at least one reference knowledge, a video frame set, and a text set, wherein the indication statement is used to guide the target large model to identify whether at least one target object is in violation according to at least one reference knowledge, a video frame set, and a text set; inputting the indication statement into the target large model, and determining the recognition result through the target large model according to at least one reference knowledge, a video frame set, and a text set.
[0012] Furthermore, the target object identification method also includes: obtaining a training sample set, wherein the training samples in the training sample set include a sample video frame set and a sample text set corresponding to the sample video, and the true label of the training sample is at least used to characterize whether the target object in the sample video is in violation; training an initial large model based on the training sample set to obtain a target large model.
[0013] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a device for identifying a target object is provided. The device includes: a first acquisition module, which is used to acquire a target video, wherein the target video is obtained by collecting images of a street, and the street includes at least one target object; a first processing module, which is used to process the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes the text content in the target video; a retrieval module, which is used to perform knowledge retrieval in a knowledge base based on the video frame set and the text set to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to different categories of target object behaviors; a second processing module, which is used to identify whether at least one target object is in violation of the rules based on the target macro model, the video frame set, and the text set, and obtain an identification result.
[0014] Furthermore, the first processing module also includes: a first processing submodule, used for performing video frame extraction processing on the target video to obtain a video frame set; and a second processing submodule, used for performing optical character recognition processing on the video content of the target video to obtain a text set.
[0015] Furthermore, the retrieval module also includes: an acquisition submodule, which is used to acquire video features of a video frame set and text features of a text set; a third processing submodule, which is used to perform feature fusion on the video features and the text features to obtain multimodal features; and a first determination submodule, which is used to determine at least one reference knowledge from the knowledge in the knowledge base based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base.
[0016] Furthermore, the acquisition submodule also includes: a first processing unit, which is used to use the visual large model to perform feature extraction processing on the video frame set to obtain video features; and a second processing unit, which is used to use the text large model to perform feature extraction processing on the text set to obtain text features.
[0017] Furthermore, the first determination submodule also includes: a calculation unit, which is used to calculate the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base, and obtain the similarity corresponding to each knowledge; a third processing unit, which is used to sort each knowledge in order from high to low according to the similarity, and obtain the sorted knowledge; a determination unit, which is used to determine the first N knowledge in the sorted knowledge as reference knowledge, where N is a positive integer.
[0018] Furthermore, the second processing module also includes: a generation submodule, used to generate an instruction statement based on at least one reference knowledge, a video frame set and a text set, wherein the instruction statement is used to guide the target large model to identify whether at least one target object is in violation according to at least one reference knowledge, a video frame set and a text set; a second determination submodule, used to input the instruction statement into the target large model, and determine the recognition result through the target large model according to at least one reference knowledge, a video frame set and a text set.
[0019] Furthermore, the target object recognition device also includes: a second acquisition module, used to acquire a training sample set, wherein the training samples in the training sample set include a sample video frame set and a sample text set corresponding to the sample video, and the true label of the training sample is at least used to characterize whether the target object in the sample video is in violation; a training module, used to train an initial large model based on the training sample set to obtain a target large model.
[0020] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a computer-readable storage medium is provided, and the computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned target object identification method.
[0021] In order to achieve the above-mentioned purpose, according to another aspect of the present application, an electronic device is provided, the electronic device comprising a memory storing an executable program; and a processor for running the program, wherein the above-mentioned target object recognition method is executed when the program is running.
[0022] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a computer program product is provided, including computer instructions, which implement the steps of the above-mentioned target object recognition method when executed by a processor.
[0023] In an embodiment of the present application, a method based on multi-type data is adopted to identify whether a specific object on the street is in violation of the rules. By processing the target video to obtain a video frame set and a text set, the multi-dimensional features of the target object to be identified are obtained. By determining matching reference knowledge from a knowledge base based on the video frame set and the text set, relevant knowledge for reference can be accurately found based on various data of the target object to be identified. Through a retrieval enhancement-based method, a target large model is used to identify whether the target object is in violation of the rules based on multi-dimensional data and reference knowledge, thereby effectively improving the accuracy of identifying the target object's violation of the rules.
[0024] It can be seen that the solution provided in the present application achieves the purpose of identifying whether a specific object on the street is in violation of the rules based on multiple types of data, thereby achieving the technical effect of improving the accuracy of identifying whether a specific object on the street is in violation of the rules, and further solving the problem of low recognition accuracy in the related technology of identifying whether a specific object on the street is in violation of the rules based on a single data. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0026] Figure 1 is a hardware structure block diagram of a computer terminal provided according to an embodiment of the present application;
[0027] Figure 2 is a schematic diagram of a target object identification method provided according to an embodiment of the present application;
[0028] Figure 3 is a flow chart of a target object identification method provided according to an embodiment of the present application;
[0029] Figure 4 is a schematic diagram of a target object identification device provided according to an embodiment of the present application;
[0030] Figure 5 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data are in compliance with relevant laws, regulations and standards, necessary confidentiality measures are taken, and public order and good customs are not violated, and corresponding operation entrances are provided for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.
[0034] Example 1
[0035] According to an embodiment of the present application, an embodiment of a method for identifying a target object is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0036] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for identifying a target object. Figure 1As shown, the computer terminal 10 (or mobile device) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.
[0037] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0038] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the target object recognition method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned target object recognition method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0039] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0040] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0041] Under the above operating environment, this application provides Figure 2 The target object recognition method shown. Figure 2 It is a schematic diagram of a target object recognition method provided according to an embodiment of the present application.
[0042] Step S201, obtaining a target video, wherein the target video is obtained by capturing images of a street, and the street includes at least one target object.
[0043] Optionally, electronic devices, application systems, servers and other devices may be used as the execution subject of the present application. In this embodiment, the target processing system is used as the execution subject to execute the above-mentioned target object identification method.
[0044] Optionally, cameras or monitoring equipment may be arranged on the streets to collect images of the streets, so that the obtained video data is determined as a target video, wherein there is at least one target object to be identified on the street in the target video, and the target object is a mobile merchant.
[0045] Step S202, performing video processing on the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes text content in the target video.
[0046] Optionally, the target video is processed using a time series analysis method to segment the video into a series of video frames arranged in chronological order to obtain a set of images, i.e., a video frame set, each of which may contain key information about the mobile merchant. Then, the text information in the video frame is extracted using OCR (Optical Character Recognition) technology to obtain a text set, which may contain advertising slogans, product prices, and other content of the mobile merchant.
[0047] In some embodiments, after obtaining the video frame set, the target processing system extracts text information in the video frame, recognizes that there are words "Fresh Fruit, 50% Discount" on the merchant's stall, and the system stores this information in the text set.
[0048] Optionally, after obtaining the video frame set and the text set, feature extraction processing can be performed on the video frame set and the text set using a large visual model and a large text model respectively to obtain video features and text features.
[0049] Step S203, performing knowledge retrieval in a knowledge base based on the video frame set and the text set to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to different categories of target object behaviors.
[0050] Optionally, the knowledge base contains multiple knowledge items, and different knowledge items include image data (such as example images of legal stalls, standard images of stall placement, etc.) and text descriptions (such as laws and regulations and behavioral norms) corresponding to different categories of target object behaviors (i.e., mobile merchant behaviors), and the image data comes from surveillance cameras in various cities. After the system determines the video frame set and text set from the target video, it performs knowledge retrieval in the knowledge base based on the video frame set and text set to find the knowledge items most relevant to the mobile merchant behaviors in the target video as reference knowledge (this is the first stage of RAG (Retrieval Augmented Generation) technology). Among them, the target object behaviors in the knowledge base can be illegal or compliant. In some embodiments, there are multiple pieces of knowledge representing compliant target object behaviors and multiple pieces of knowledge representing illegal target object behaviors in the knowledge base.
[0051] In some embodiments, the knowledge base may include multiple categories of mobile merchant behaviors, such as selling fruits, vegetables and various daily necessities, snacks, breakfast, fresh food, etc. on the street, selling clothes on the roadside, placing tables and chairs outside restaurants, and other goals related to mobile merchant operations. Each type of behavior is represented in detail through image data and text descriptions.
[0052] Optionally, for the video frame set and text set obtained from the target video, video features and text features are obtained therefrom, and then the text features and video features are fused to obtain multimodal features (a feature vector). Finally, based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base, some knowledge in the knowledge base with the highest similarity to the multimodal features is determined as reference knowledge.
[0053] Step S204, identifying whether at least one target object violates the rules based on at least one reference knowledge, a video frame set, and a text set through a target large model, and obtaining a recognition result.
[0054] Optionally, the target large model is a multimodal large model with enhanced retrieval capability based on RAG (Retrieval Augmented Generation). It can comprehensively analyze and identify whether the behavior of the target object complies with the specification based on the video frame set, text set and reference knowledge obtained from the knowledge base, and output the final recognition result. If the behavior of the target object violates the regulations, the specific type of violation will also be output, such as non-compliant placement, etc. This is the second stage of the RAG technology.
[0055] Optionally, the target processing system can first generate an instruction statement based on reference knowledge, a video frame set and a text set. The instruction statement is used to guide the target large model to identify the behavior of the target object and determine whether it is a violation. The instruction statement is then input into the target large model to obtain a recognition result.
[0056] In an embodiment of the present application, a method based on multi-type data is adopted to identify whether a specific object on the street is in violation of the rules. By processing the target video to obtain a video frame set and a text set, the multi-dimensional features of the target object to be identified are obtained. By determining matching reference knowledge from a knowledge base based on the video frame set and the text set, relevant knowledge for reference can be accurately found based on various data of the target object to be identified. Through a retrieval enhancement-based method, a target large model is used to identify whether the target object is in violation of the rules based on multi-dimensional data and reference knowledge, thereby effectively improving the accuracy of identifying the target object's violation of the rules.
[0057] It can be seen that the solution provided in the present application achieves the purpose of identifying whether a specific object on the street is in violation of the rules based on multiple types of data, thereby achieving the technical effect of improving the accuracy of identifying whether a specific object on the street is in violation of the rules, and further solving the problem of low recognition accuracy in the related technology of identifying whether a specific object on the street is in violation of the rules based on a single data.
[0058] In an optional embodiment, in the process of performing video processing on the target video to obtain a video frame set and a text set, the target processing system can perform video frame extraction processing on the target video to obtain a video frame set; and perform optical character recognition processing on the video content of the target video to obtain a text set.
[0059] Optionally, a timing analysis method may be used to decompose a continuous video stream (target video) into a series of static video frames to form a set of video frames, each of which may contain key information related to mobile merchants, such as their stall placement, stall appearance, etc.
[0060] Optionally, for each frame image in the video frame set, use OCR technology to identify and extract text information from each frame image, such as advertising slogans, price tags, promotional slogans, etc. of mobile vendors' stalls, and organize all extracted text information into a text set.
[0061] Optionally, when using OCR technology to extract text information from multiple video frames, all text information can be arranged in time sequence according to the corresponding video frames to obtain a text set with a temporal relationship; or a text set can be obtained according to the rule of recording repeated text only once.
[0062] In some embodiments, when extracting text information from a video frame using OCR technology, the following steps may be performed:
[0063] 1. Video preprocessing: Since the light in the video shooting environment may be unstable, it is necessary to adjust the brightness and contrast of the video, and then stabilize the video frame to reduce image blur caused by camera shake or target object movement. Finally, use image segmentation technology to identify areas in the video frame that may contain text, such as billboards, price tags, etc.
[0064] 2. Application of OCR technology: In the segmented text area, deep learning or machine vision algorithms are applied, such as those based on convolutional neural networks or text detectors, to accurately detect the locations of all text blocks, and then the detected text blocks are further segmented into individual characters or words to prepare for subsequent recognition. Finally, OCR technology is used to recognize the segmented characters or words and convert the characters in the image into an editable and retrievable text format. Among them, OCR technology will recognize characters of different fonts, sizes and directions based on the trained character model.
[0065] 3. Post-processing and correction: The recognized characters are spliced into complete words or sentences, and the text layout (such as line spacing, paragraph division, etc.) is adjusted to restore the context of the original text. Then, meaningless characters or symbols generated during the recognition process, such as image noise, non-text elements, etc., are removed. Finally, based on the language model and context understanding, spelling errors and grammatical anomalies in the recognition results are automatically corrected to improve the accuracy of the text information.
[0066] 4. Text information integration: The extracted text information is organized into a text set in chronological order or recognition order for subsequent integration with the video frame set. Each extracted text information is then assigned its corresponding frame in the video to ensure consistency between the text content and the video timing.
[0067] 5. Storage and application: The complete text set is stored in the database of the target processing system so that the mobile merchant’s behavior can be subsequently identified based on RAG technology to see if it violates regulations.
[0068] It should be noted that by processing the target video to obtain a set of video frames, the target processing system can capture the behavioral details of mobile merchants more meticulously, providing a reliable data basis for subsequent identification of violations. By extracting text information from video frames to obtain a text set, it is possible to identify violations based on multi-faceted data, thereby improving the accuracy of identification.
[0069] In an optional embodiment, in the process of performing knowledge retrieval in a knowledge base based on a video frame set and a text set to obtain at least one reference knowledge, the target processing system can obtain video features of the video frame set and text features of the text set; perform feature fusion on the video features and the text features to obtain multimodal features; and determine at least one reference knowledge from the knowledge in the knowledge base based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base.
[0070] Optionally, when extracting video features from a video frame set, a deep learning model (such as a convolutional neural network) can be used to process each frame in the video frame set to extract key visual features such as the appearance of the mobile merchants' stalls, the types of goods, the stall layout, and the behavior patterns. These features can help the system understand the specific behavior and environmental status of the mobile merchants in the video. When extracting text features from a text set, natural language processing technology, such as word vector models, can be used to extract semantic features of the text, such as the advertising slogans, price tags, and promotional slogans of the stalls. These features help the system understand the semantic background of the mobile merchants' behavior, such as whether the prices are clearly marked and whether the advertising is compliant.
[0071] Optionally, the video features and text features are fused to form comprehensive multimodal features. Feature fusion can be achieved through deep learning multimodal encoders, attention mechanisms and other technologies. The fused features contain both visual and semantic information, which can more comprehensively reflect the behavioral characteristics of mobile merchants. For example, combining the recognized "fresh fruit" text information with the image features of the fruit stall in the video can more accurately determine whether the merchant is legally selling fruit.
[0072] Optionally, after obtaining the multimodal features, compare the multimodal features with the knowledge features of each knowledge item in the knowledge base, and select some knowledge items that are most similar to the multimodal features as reference knowledge by calculating the similarity between them (such as cosine similarity). These reference knowledge can help the system determine whether the behavior of mobile merchants in the video is illegal, such as whether they occupy the sidewalk, whether they sell in a prohibited area, whether they are clearly marked with prices, etc.
[0073] It should be noted that by extracting and fusing features of video frame sets and text sets, the visual and text information of mobile merchants' stalls can be accurately captured, ensuring that the system can identify the behavior of mobile merchants based on multi-dimensional information, thereby improving the accuracy of recognition. By selecting reference knowledge from the knowledge base based on multimodal features, it provides a strong basis for identifying whether mobile merchants have violated regulations, thereby improving the accuracy of recognition.
[0074] In an optional embodiment, in the process of acquiring video features of a video frame set and text features of a text set, the target processing system can use a visual big model to perform feature extraction processing on the video frame set to obtain video features; and use a text big model to perform feature extraction processing on the text set to obtain text features.
[0075] Optionally, the visual big model and the text big model are used to perform feature extraction processing on the video frame set and the text set respectively to obtain video features and text features.
[0076] In some embodiments, the multimodal big model (i.e., the target big model mentioned above) structure is composed of a visual big model and a text big model, but the training methods are different, that is, the parameters are different.
[0077] Optionally, the large visual model can be trained by:
[0078] Select a pre-trained initial visual large model (such as a Transformer-based image recognition model, where the encoder part of the model encodes the pixels in the image to convert the image into a series of vectors, which contain the global and local features of the image, and then uses an attention mechanism (such as a temporal attention mechanism) to fuse different features in the image to obtain video features, which can be used for other subsequent operations. The initial visual large model may also include an object detection module for identifying different objects in the image, such as people, objects, etc.), and then use a large amount of street scene video data as a visual training sample set, input it into the initial visual large model for fine-tuning, so that the model can recognize common mobile merchant behaviors, item categories, and environmental features, etc., to obtain a visual large model. Among them, the visual training sample set can include positive samples (such as relevant video frames of legal merchants) and negative samples (such as relevant video frames of illegal merchants). The true label of the training sample can be a text description, such as a description of the merchant's actions, behaviors, and item placement.
[0079] Optionally, when the video frame set is subjected to feature extraction processing using the visual large model and the video features are obtained, the following steps may be performed:
[0080] 1. Video frame preprocessing: Preprocess each frame image in the video frame set of the target video, including size standardization, color enhancement, etc., to obtain a preprocessed video frame set.
[0081] 2. Extract video features: Input the preprocessed video frame set into the visual big model. The visual big model will automatically learn and extract the appearance features, product features, behavior features and environmental features of the stalls through multi-layer convolution, pooling and full connection operations, forming a vector representation of the video features. These feature vectors can capture the key information of the behavior of mobile merchants in the video.
[0082] Optionally, a large text model can be trained in the following ways:
[0083] Select a pre-trained initial text model (such as a Transformer model, where the model's tokenizer decomposes the text into vocabulary or subword units and converts them into word embedding vectors to capture the basic semantic information. Then, based on the Transformer architecture encoder, the multi-head attention mechanism and feedforward neural network are used to deeply encode the text to form text features. Finally, these features are aggregated into a comprehensive feature representation of the entire text, which can be used for subsequent operations). Then, the text set obtained from the street scene video is used as a text training sample set and input into the initial text model for fine-tuning so that the model can understand the semantic information and contextual connection of the text, and obtain the text model. Among them, the text training sample set can be the specific content of the text information in the video (such as "pancake", "5 yuan", etc.), and the true label of the training sample can be a sentence, corresponding to the text content in the training sample, which can more clearly express the information of the training sample (such as "a pancake fruit stand").
[0084] Optionally, the text collection is subjected to feature extraction processing using a large text model. When the text features are obtained, the following steps can be performed:
[0085] 1. Text information preprocessing: Preprocess the text information in the text set of the target video, including text cleaning, word segmentation, part-of-speech tagging, etc., to obtain the preprocessed text set.
[0086] 2. Extract text features: The preprocessed text set is input into the large text model. The model extracts the semantic features of the text through operations such as the embedding layer and the encoding layer to form a vector representation of the text features, including keywords, phrases, sentence structures, etc. These feature vectors can reflect the stall’s advertising information, price tags, and descriptive text content related to the behavior of mobile merchants.
[0087] It should be noted that by using the visual big model and the text big model to perform feature extraction processing on the video frame set and the text set respectively, video features and text features are obtained, ensuring that high-quality video features and text features are extracted from the video frame set and the text set respectively, providing a reliable data basis for the identification of illegal stalls, improving the system's ability to identify stall violations, and thus improving the accuracy of recognition.
[0088] In an optional embodiment, in the process of determining at least one reference knowledge from the knowledge in the knowledge base based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base, the target processing system can calculate the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base to obtain the similarity corresponding to each knowledge; sort each knowledge in order from high to low according to the similarity to obtain sorted knowledge; and determine the first N knowledge in the sorted knowledge as reference knowledge, where N is a positive integer.
[0089] Optionally, after obtaining the multimodal feature, the target processing system compares the multimodal feature with each knowledge feature in the knowledge base, calculates the similarity between the multimodal feature and each knowledge feature, and then sorts all similarities for the corresponding knowledge in descending order.
[0090] In some embodiments, the process of obtaining the knowledge features of the knowledge is the same as the process of obtaining the multimodal features of the target video, so it will not be repeated here.
[0091] In some embodiments, methods such as cosine similarity or Euclidean distance can be used to calculate the similarity between multimodal features and knowledge features. The calculation of similarity can help the system determine the degree of match between the behavior of mobile merchants in the target video and the behavioral norms or violations stored in the knowledge base.
[0092] In some embodiments, if the target processing system calculates that a fruit stall in the target video has the highest similarity to the knowledge features about "legal fruit stall" in the knowledge base, then the knowledge entry of "legal fruit stall" will be ranked first and become the priority reference knowledge.
[0093] Optionally, for the sorted knowledge, the system selects the top N knowledge items with similarity ranking as reference knowledge based on a preset threshold N. A larger N value can provide richer information but may increase the calculation time, while a smaller N value can respond quickly but may miss some relevant information.
[0094] It should be noted that by calculating the similarity between the multimodal features and each knowledge item, the matching degree between the merchant behavior characteristics in the quantified target video and the data characteristics in the knowledge base is achieved. By selecting the N pieces of knowledge with the highest similarity to the multimodal features as reference knowledge, it provides a favorable reference for identifying violations of mobile merchants, thereby improving the accuracy of the system in identifying whether a specific object in the target video has violated the regulations.
[0095] In an optional embodiment, in the process of obtaining an identification result by identifying whether at least one target object is in violation of the rules based on at least one reference knowledge, a video frame set, and a text set through a target macro model, the target processing system may generate an instruction statement based on at least one reference knowledge, a video frame set, and a text set, wherein the instruction statement is used to guide the target macro model to identify whether at least one target object is in violation of the rules based on at least one reference knowledge, a video frame set, and a text set; the instruction statement is input into the target macro model, and the identification result is determined by the target macro model based on at least one reference knowledge, a video frame set, and a text set.
[0096] Optionally, when using the target large model to identify whether the target object is in violation of the rules, the target processing system will first generate an instruction statement based on N reference knowledge, video frame sets, and text sets. The format of the instruction statement can be read and parsed by the target large model, and is used to guide the target large model to identify whether the target object is in violation of the rules based on the input information. For example, the target processing system can obtain an instruction statement template, and the instruction statement template can include the words "Based on the following reference knowledge, determine whether the mobile merchants appearing in the video frame set and text set are in violation of the rules." The target processing system can import the obtained reference knowledge, video frame set, and text set into the instruction statement template to obtain the instruction statement.
[0097] Optionally, after the instruction statement is input into the target large model, the model will identify whether the target object has any violation based on the guidance of the instruction statement, combined with reference knowledge, video frame set and text set.
[0098] In some embodiments, the reference knowledge that the system selects from the knowledge base that best matches the behavior of mobile merchants in the target video includes graphic examples of "stalls must not occupy pedestrian walkways, and must not display billboards exceeding a certain size", as well as graphic examples of "legal fruit stalls" and graphic examples of "illegal barbecue stalls". When the target large model is used to identify whether the target object is in violation, the target processing system first gives a specific instruction statement based on the selected reference knowledge and video frame set and text set, and inputs the instruction statement into the target large model. Then the target large model outputs the recognition result after comprehensive analysis: "Mobile merchant A set up a stall on the pedestrian walkway, which violates the stall management regulations", "The size of the billboard displayed by mobile merchant B exceeds the compliance range".
[0099] It should be noted that by generating instruction statements for guiding the target big model based on reference knowledge, video frame sets and text sets, the target big model is made more targeted and specific in identifying violations, thereby improving recognition efficiency. By identifying whether mobile merchants have violated regulations based on instruction statements based on the target big model, the target big model can flexibly respond to complex scenes in the target video, thereby improving the accuracy of model recognition, thereby helping to realize intelligent and standardized urban management.
[0100] In an optional embodiment, the target processing system may determine the target large model in the following manner: obtain a training sample set, wherein the training samples in the training sample set include a sample video frame set and a sample text set corresponding to the sample video, and the true label of the training sample is at least used to characterize whether the target object in the sample video violates the rules; train the initial large model based on the training sample set to obtain the target large model.
[0101] Optionally, a large number of street videos are collected from surveillance videos of multiple cities as sample videos, which contain scenes of compliance and violation of street mobile merchants. For the collected sample videos, each sample video is converted into a video frame set based on a time series analysis method to obtain a sample video frame set. Then, text information is extracted from the video frames through OCR technology to form a sample text set. The sample video frame set and the sample text set of the same sample video correspond to each other, i.e., sample image-text pairs, wherein the sample image-text pairs are divided into positive samples (sample data of legal stalls) and negative samples (sample data of illegal stalls). Finally, the sample video frame set, the sample text set and a question are determined as a training sample in the training sample set, wherein the question is to guide the target large model to approach the output target (e.g., is the stall of Merchant A legal?).
[0102] Optionally, for the sample video frame set and the sample text set, the staff can annotate the image-text pairs according to the questions in the sample training set to indicate whether the behavior of the mobile merchants in the sample video is illegal, as well as the specific types of violations (such as occupying the sidewalk, displaying illegal advertisements, etc.), and determine the annotated content as the true label of the training sample in the training sample set.
[0103] Optionally, a pre-trained multimodal large model is selected as the initial large model. After obtaining the training sample set, the initial large model is trained to obtain the target large model. In some embodiments, the structure of the initial large model consists of an initial visual large model and an initial textual large model, and the structure of the target large model consists of a visual large model and a textual large model, but due to different training methods, their model parameters are different. During training, the true labels of the training samples are used as supervisory signals to guide the model to learn how to accurately judge whether street mobile merchants have violated regulations, as well as the specific types of violations. When the preset iteration number threshold is reached, the model training ends and the target large model is obtained.
[0104] It should be noted that obtaining a training sample set based on videos of real scenes provides a real and reliable data basis for the training of the target large model, improves the adaptability of the target large model to different scenes, and thus improves the recognition accuracy. By training the target large model based on the training sample set, the target large model can learn a wider range of video features and rules, thereby improving the accuracy of the target large model in identifying whether mobile merchants have violated regulations, effectively supporting urban construction and efficient management.
[0105] In some embodiments, Figure 3 is a flow chart of a target object identification method provided in an embodiment of the present application. Figure 3 As shown in the figure, the process of the target processing system identifying the illegal behavior of mobile merchants is shown. First, the street is imaged to obtain the target video, and the target video is processed to obtain a video frame set and a text set. Then, based on the visual large model and the text large model, video features and text features are obtained from the video frame set and the text set respectively. The two features are fused to obtain a multimodal feature vector. Then, according to the multimodal feature vector, the N knowledge items with the highest similarity are found from the knowledge base as reference knowledge. The reference knowledge, the video frame set and the text set are input into the target large model together, and an instruction statement is automatically generated to guide the target large model for identification. After the instruction statement is generated, the target large model will identify the illegal behavior of the specific object according to the instruction statement and based on the input information, and output the identification result.
[0106] In an embodiment of the present application, a method based on multi-type data is adopted to identify whether a specific object on the street is in violation of the rules. By processing the target video to obtain a video frame set and a text set, the multi-dimensional features of the target object to be identified are obtained. By determining matching reference knowledge from a knowledge base based on the video frame set and the text set, relevant knowledge for reference can be accurately found based on various data of the target object to be identified. Through a retrieval enhancement-based method, a target large model is used to identify whether the target object is in violation of the rules based on multi-dimensional data and reference knowledge, thereby effectively improving the accuracy of identifying the target object's violation of the rules.
[0107] It can be seen that the solution provided in the present application achieves the purpose of identifying whether a specific object on the street is in violation of the rules based on multiple types of data, thereby achieving the technical effect of improving the accuracy of identifying whether a specific object on the street is in violation of the rules, and further solving the problem of low recognition accuracy in the related technology of identifying whether a specific object on the street is in violation of the rules based on a single data.
[0108] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0109] Example 2
[0110] The present application also provides a target object recognition device. It should be noted that the target object recognition device of the present application can be used to execute the target object recognition method provided by the present application. The target object recognition device provided by the present application is introduced below.
[0111] According to an embodiment of the present application, a device for implementing the above-mentioned target object identification method is also provided. Figure 4 As shown, the device comprises:
[0112] The first acquisition module 401 is used to acquire a target video, wherein the target video is obtained by collecting images of a street, and the street includes at least one target object;
[0113] A first processing module 402 is used to perform video processing on the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes text content in the target video;
[0114] A retrieval module 403 is used to perform knowledge retrieval in a knowledge base based on the video frame set and the text set to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to different categories of target object behaviors;
[0115] The second processing module 404 is used to identify whether at least one target object violates the rules according to at least one reference knowledge, a video frame set and a text set through a target large model to obtain a recognition result.
[0116] It should be noted that the above-mentioned first acquisition module 401, first processing module 402, retrieval module 403 and second processing module 404 correspond to steps S201 to S204 in the above-mentioned embodiment, and the examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned embodiment 1.
[0117] In an embodiment of the present application, a method based on multi-type data is adopted to identify whether a specific object on the street is in violation of the rules. By processing the target video to obtain a video frame set and a text set, the multi-dimensional features of the target object to be identified are obtained. By determining matching reference knowledge from a knowledge base based on the video frame set and the text set, relevant knowledge for reference can be accurately found based on various data of the target object to be identified. Through a retrieval enhancement-based method, a target large model is used to identify whether the target object is in violation of the rules based on multi-dimensional data and reference knowledge, thereby effectively improving the accuracy of identifying the target object's violation of the rules.
[0118] It can be seen that the solution provided in the present application achieves the purpose of identifying whether a specific object on the street is in violation of the rules based on multiple types of data, thereby achieving the technical effect of improving the accuracy of identifying whether a specific object on the street is in violation of the rules, and further solving the problem of low recognition accuracy in the related technology of identifying whether a specific object on the street is in violation of the rules based on a single data.
[0119] Optionally, in the target object identification device provided in the embodiment of the present application, the first processing module also includes: a first processing sub-module, used to perform video frame extraction processing on the target video to obtain a video frame set; and a second processing sub-module, used to perform optical character recognition processing on the video content of the target video to obtain a text set.
[0120] Optionally, in the target object identification device provided in the embodiment of the present application, the retrieval module also includes: an acquisition submodule, used to acquire video features of a video frame set and text features of a text set; a third processing submodule, used to perform feature fusion on the video features and the text features to obtain multimodal features; and a first determination submodule, used to determine at least one reference knowledge from the knowledge in the knowledge base based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base.
[0121] Optionally, in the target object recognition device provided in the embodiment of the present application, the acquisition submodule also includes: a first processing unit, used to perform feature extraction processing on a video frame set using a large visual model to obtain video features; a second processing unit, used to perform feature extraction processing on a text set using a large text model to obtain text features.
[0122] Optionally, in the target object identification device provided in the embodiment of the present application, the first determination submodule also includes: a calculation unit, used to calculate the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base, to obtain the similarity corresponding to each knowledge; a third processing unit, used to sort each knowledge in order from high to low according to the similarity, to obtain sorted knowledge; a determination unit, used to determine the first N knowledge in the sorted knowledge as reference knowledge, where N is a positive integer.
[0123] Optionally, in the target object identification device provided in the embodiment of the present application, the second processing module also includes: a generation submodule, used to generate an instruction statement based on at least one reference knowledge, a video frame set, and a text set, wherein the instruction statement is used to guide the target large model to identify whether at least one target object is in violation according to at least one reference knowledge, a video frame set, and a text set; a second determination submodule, used to input the instruction statement into the target large model, and determine the recognition result through the target large model according to at least one reference knowledge, a video frame set, and a text set.
[0124] Optionally, in the target object identification device provided in the embodiment of the present application, the target object identification device also includes: a second acquisition module, used to acquire a training sample set, wherein the training samples in the training sample set include a sample video frame set and a sample text set corresponding to the sample video, and the true label of the training sample is at least used to characterize whether the target object in the sample video is in violation; a training module, used to train an initial large model based on the training sample set to obtain a target large model.
[0125] It should be noted that the first acquisition module 401, the first processing module 402, the retrieval module 403 and the second processing module 404 correspond to steps S201 to S204 in the above embodiment, and the examples and application scenarios implemented by the four modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment 1. It should be noted that the above modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n), and the above modules may also be part of the device and may be run in the computer terminal 10 provided in the first embodiment.
[0126] Example 3
[0127] An embodiment of the present application may provide an electronic device, Figure 5 is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5(only one is shown) processor 1002, memory 1004, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0128] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned target object recognition method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0129] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a target video, wherein the target video is obtained by collecting images of a street, and the street includes at least one target object; perform video processing on the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes text content in the target video; perform knowledge retrieval in a knowledge base based on the video frame set and the text set to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to different categories of target object behaviors; identify whether at least one target object is in violation of the rules based on the target macro model and at least one reference knowledge, the video frame set, and the text set to obtain an identification result.
[0130] The processor can also call the information and application programs stored in the memory through the transmission device to execute the following steps: performing video frame extraction processing on the target video to obtain a video frame set; performing optical character recognition processing on the video content of the target video to obtain a text set.
[0131] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain video features of a video frame set and text features of a text set; perform feature fusion on the video features and text features to obtain multimodal features; and determine at least one reference knowledge from the knowledge in the knowledge base based on the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base.
[0132] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: use the visual big model to perform feature extraction processing on the video frame set to obtain video features; use the text big model to perform feature extraction processing on the text set to obtain text features.
[0133] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: calculate the similarity between the multimodal features and the knowledge features of each knowledge in the knowledge base to obtain the similarity corresponding to each knowledge; sort each knowledge in order from high to low according to the similarity to obtain the sorted knowledge; determine the first N knowledge in the sorted knowledge as reference knowledge, where N is a positive integer.
[0134] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: generate an instruction statement based on at least one reference knowledge, a video frame set, and a text set, wherein the instruction statement is used to guide the target large model to identify whether at least one target object is in violation based on at least one reference knowledge, a video frame set, and a text set; input the instruction statement into the target large model, and determine the recognition result through the target large model based on at least one reference knowledge, a video frame set, and a text set.
[0135] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: obtain a training sample set, wherein the training samples in the training sample set include a sample video frame set and a sample text set corresponding to the sample video, and the true label of the training sample is at least used to characterize whether the target object in the sample video is in violation; train an initial large model based on the training sample set to obtain a target large model.
[0136] In an embodiment of the present application, a method based on multi-type data is adopted to identify whether a specific object on the street is in violation of the rules. By processing the target video to obtain a video frame set and a text set, the multi-dimensional features of the target object to be identified are obtained. By determining matching reference knowledge from a knowledge base based on the video frame set and the text set, relevant knowledge for reference can be accurately found based on various data of the target object to be identified. Through a retrieval enhancement-based method, a target large model is used to identify whether the target object is in violation of the rules based on multi-dimensional data and reference knowledge, thereby effectively improving the accuracy of identifying the target object's violation of the rules.
[0137] It can be seen that the solution provided in the present application achieves the purpose of identifying whether a specific object on the street is in violation of the rules based on multiple types of data, thereby achieving the technical effect of improving the accuracy of identifying whether a specific object on the street is in violation of the rules, and further solving the problem of low recognition accuracy in the related technology of identifying whether a specific object on the street is in violation of the rules based on a single data.
[0138] It can be understood by those skilled in the art that Figure 5 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Figure 5 The structure of the electronic device is not limited. Figure 5 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Figure 5 Different configurations are shown.
[0139] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0140] Example 4
[0141] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the target object recognition method provided in the first embodiment.
[0142] Optionally, in this embodiment, the above-mentioned storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0143] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the steps of the target object identification method.
[0144] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0145] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0146] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0147] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk and other media that can store program codes.
[0150] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for identifying a target object, characterized in that: include: Acquire a target video, wherein the target video is obtained by collecting images of a street, and the street includes at least one target object; Performing video processing on the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes text content in the target video; Based on the video frame set and the text set, knowledge retrieval is performed in a knowledge base to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to target object behaviors of different categories; The target large model identifies whether the at least one target object violates the regulations according to the at least one reference knowledge, the video frame set and the text set, and obtains a recognition result.
2. The method according to claim 1, characterized in that: Performing video processing on the target video to obtain a video frame set and a text set, including: Performing video frame extraction processing on the target video to obtain the video frame set; Optical character recognition is performed on the video content of the target video to obtain the text set.
3. The method according to claim 1, characterized in that Based on the video frame set and the text set, knowledge retrieval is performed in a knowledge base to obtain at least one reference knowledge, including: Acquire video features of the video frame set and text features of the text set; Performing feature fusion on the video features and the text features to obtain multimodal features; The at least one reference knowledge is determined from the knowledge in the knowledge base according to the similarity between the multimodal feature and the knowledge feature of each knowledge in the knowledge base.
4. The method according to claim 3, characterized in that Acquiring video features of the video frame set and text features of the text set includes: The video frame set is subjected to feature extraction processing by using a visual macro model to obtain the video features; the text set is subjected to feature extraction processing by using a text macro model to obtain the text features.
5. The method according to claim 3, characterized in that: Determining the at least one reference knowledge from the knowledge in the knowledge base according to the similarity between the multimodal feature and the knowledge feature of each knowledge in the knowledge base comprises: Calculating the similarity between the multimodal feature and the knowledge feature of each piece of knowledge in the knowledge base to obtain the similarity corresponding to each piece of knowledge; Sorting the various pieces of knowledge according to the order of similarity from high to low to obtain sorted knowledge; The first N pieces of knowledge in the sorted knowledge are determined as the reference knowledge, where N is a positive integer.
6. The method according to claim 1, characterized in that Identifying whether the at least one target object violates a rule according to the at least one reference knowledge, the video frame set, and the text set through the target macro model, and obtaining an identification result, including: Generate an instruction statement based on the at least one reference knowledge, the video frame set, and the text set, wherein the instruction statement is used to guide the target macro model to identify whether the at least one target object violates the rules according to the at least one reference knowledge, the video frame set, and the text set; The instruction statement is input into the target large model, and the recognition result is determined by the target large model according to the at least one reference knowledge, the video frame set and the text set.
7. The method according to claim 1, characterized in that The target large model is obtained by: Acquire a training sample set, wherein the training samples in the training sample set include a sample video frame set and a sample text set corresponding to the sample video, and the real labels of the training samples are at least used to characterize whether the target object in the sample video violates the rules; An initial large model is trained based on the training sample set to obtain the target large model.
8. A target object recognition device, characterized in that: include: A first acquisition module is used to acquire a target video, wherein the target video is obtained by collecting images of a street, and the street includes at least one target object; A first processing module is used to perform video processing on the target video to obtain a video frame set and a text set, wherein the video frame set includes multiple video frames in the target video, and the text set includes text content in the target video; A retrieval module, configured to perform knowledge retrieval in a knowledge base based on the video frame set and the text set to obtain at least one reference knowledge, wherein different knowledge in the knowledge base includes images and text descriptions corresponding to target object behaviors of different categories; The second processing module is used to identify whether the at least one target object violates the regulations according to the at least one reference knowledge, the video frame set and the text set through the target large model to obtain a recognition result.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the computer-readable storage medium is located is controlled to execute the target object recognition method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: A memory storing an executable program; A processor is used to run the program, wherein the program executes the target object recognition method described in any one of claims 1 to 7 when running.
11. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the target object recognition method described in any one of claims 1 to 7 are implemented.