Video recognition method and device based on artificial intelligence and model training method and device based on artificial intelligence
By combining multimodal large models with retrieval enhancement and guidance information, the problem of concise recognition results in video recognition is solved, resulting in richer and more interpretable recognition results and improving user experience.
Patent Information
- Application Number
- CN202511834267.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies struggle to provide sophisticated identification results in video recognition, making it difficult for users to glean detailed information about specific violations or target objects from concise category results.
By employing a video recognition method based on a multimodal large model, this method utilizes retrieved enhancement and guidance information to identify the video to be recognized, outputting target output information including category results and explanatory information. The method combines the outputs of multiple models for noise reduction and information fusion, thereby improving the richness and interpretability of the recognition results.
It improves the richness and interpretability of video recognition results, enhances user experience, and reduces the frequency of illusion problems in multimodal large model recognition.
Smart Images

Figure CN121600445A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, particularly to the fields of computer vision, deep learning, and large-scale models. Specifically, it relates to AI-based video recognition methods, model training methods, devices, intelligent agents, electronic devices, storage media, and intelligent agents themselves. Background Technology
[0002] With the rapid development of computer technology and networks, the massive amounts of data sources and rich data layers have made it increasingly difficult to analyze and process video information manually. Computer vision technology offers enormous potential to liberate human labor. Computer vision is a science that studies how to "see" using electronic devices; that is, it utilizes computer technology to replace the human eye in identifying, tracking, and measuring targets. Computer vision technology provides significant assistance for applications in public safety, information security, and financial security. Summary of the Invention
[0003] This disclosure provides a video recognition method, model training method, device, intelligent agent, electronic device, storage medium, and program product based on artificial intelligence.
[0004] According to one aspect of this disclosure, an artificial intelligence-based video recognition method is provided, comprising: performing a retrieval based on a first category result representing the category of a video to be recognized, obtaining retrieval enhancement information, wherein the retrieval enhancement information is used to describe the features of content matching the initial category result; adding the first category result, the retrieval enhancement information, and the video to be recognized to a first prompt template to obtain first prompt information, wherein the first prompt template includes guidance information for describing the output content of a multimodal large model; and inputting the first prompt information into the multimodal large model to utilize the multimodal large model to recognize the video to be recognized based on the retrieval enhancement information and the first category result, obtaining target output information matching the guidance information; wherein the target output information includes a second category result representing the category of the video to be recognized and explanatory information describing the second category result, wherein the second category result and the first category result are semantically matched.
[0005] According to another aspect of this disclosure, a model training method is provided, comprising: training a model to be trained using training samples to obtain a trained model; wherein the training samples include sample videos and labels, and the labels are obtained by identifying the sample videos using the method described above.
[0006] According to another aspect of this disclosure, a video recognition method is provided, comprising: inputting a target video into a target recognition model to obtain a second target recognition result; wherein the target recognition model is trained by the method described above.
[0007] According to embodiments of this disclosure, a video recognition device is provided, comprising: a retrieval module, configured to perform a retrieval based on a first category result representing a category of a video to be recognized, to obtain retrieval enhancement information, wherein the retrieval enhancement information is used to describe features of content matching the initial category result; a first prompt determination module, configured to add the first category result, the retrieval enhancement information, and the video to be recognized to a first prompt template to obtain first prompt information, wherein the first prompt template includes guidance information describing the output content of a multimodal large model; and a first recognition module, configured to input the first prompt information into the multimodal large model, so as to use the multimodal large model to recognize the video to be recognized based on the retrieval enhancement information and the first category result, to obtain target output information matching the guidance information; wherein the target output information includes a second category result representing a category of the video to be recognized and explanatory information describing the second category result, wherein the second category result and the first category result are semantically matched.
[0008] According to an embodiment of this disclosure, a model training apparatus is provided, comprising: a training module for training a model to be trained using training samples to obtain a trained model; wherein the training samples include sample videos and labels, and the labels are obtained by recognizing the sample videos using the apparatus described above.
[0009] According to an embodiment of this disclosure, a video recognition device is provided, comprising: a second recognition module, configured to input a target video into a target recognition model to obtain a second target recognition result; wherein the target recognition model is trained by the device described above.
[0010] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method described above; and an output module for outputting the output information obtained by the processing module.
[0011] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0012] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described above.
[0013] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described above.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0016] Figure 1 This illustration schematically shows an exemplary system architecture for applying artificial intelligence-based video recognition methods and apparatus according to embodiments of the present disclosure;
[0017] Figure 2 A flowchart illustrating an artificial intelligence-based video recognition method according to an embodiment of the present disclosure is shown schematically.
[0018] Figure 3 This diagram illustrates the process of determining target interpretation information.
[0019] Figure 4 A schematic diagram illustrating a first prompt message according to an embodiment of the present disclosure is shown.
[0020] Figure 5 A schematic diagram illustrating the determination of a first category result according to an embodiment of the present disclosure is shown;
[0021] Figure 6 A schematic diagram illustrating the determination of retrieval enhancement information according to an embodiment of the present disclosure is shown.
[0022] Figure 7 A flowchart illustrating a model training method according to an embodiment of the present disclosure is shown schematically.
[0023] Figure 8The schematic diagram illustrates a flowchart of a model training method according to an embodiment of the present disclosure;
[0024] Figure 9 A flowchart illustrating a video recognition method according to an embodiment of the present disclosure is shown schematically.
[0025] Figure 10 A schematic flowchart of a video recognition method according to an embodiment of the present disclosure is shown.
[0026] Figure 11 A block diagram of an artificial intelligence-based video recognition device according to an embodiment of the present disclosure is shown schematically.
[0027] Figure 12 A block diagram of a model training apparatus according to an embodiment of the present disclosure is shown schematically;
[0028] Figure 13 A block diagram of a video recognition device according to an embodiment of the present disclosure is shown schematically;
[0029] Figure 14 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure is shown.
[0030] Figure 15 A block diagram of an electronic device suitable for implementing an artificial intelligence-based video recognition method according to an embodiment of the present disclosure is illustrated. Detailed Implementation
[0031] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0032] Large Language Models (LLMs) are large-scale text processing models that can be built using a Transformer (decoder-encoder) network structure as the main architecture. The large language model provided in this disclosure is mainly used for semantic fusion or rewriting of multiple explanatory information to obtain fused explanatory information.
[0033] Multimodal Large Language Models (MLLMs) are a class of models that combine the natural language processing capabilities of large language models with the ability to understand and generate data from other modalities (such as visual and audio). The MLLMs provided in this disclosure are mainly used to identify videos to be recognized, and output results including category results and explanatory information.
[0034] Figure 1 The illustration schematically shows an exemplary system architecture for applying artificial intelligence-based video recognition methods and apparatus according to embodiments of the present disclosure.
[0035] It is important to note that Figure 1 The examples shown are merely examples of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. They do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture for applying the AI-based video recognition method and apparatus may include a terminal device. However, the terminal device can implement the AI-based video recognition method and apparatus provided in the embodiments of this disclosure without interacting with a server.
[0036] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0037] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0038] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0039] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0040] It should be noted that the AI-based video recognition method provided in this disclosure can generally be executed by terminal devices 101, 102, or 103. Correspondingly, the AI-based video recognition device provided in this disclosure can also be disposed in terminal devices 101, 102, or 103.
[0041] Alternatively, the AI-based video recognition method provided in this disclosure can generally be executed by server 105. Correspondingly, the AI-based video recognition device provided in this disclosure can generally be located in server 105. The AI-based video recognition method provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the AI-based video recognition device provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0042] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0043] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.
[0044] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.
[0045] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.
[0046] Figure 2 A flowchart illustrating an artificial intelligence-based video recognition method according to an embodiment of the present disclosure is shown.
[0047] like Figure 2 As shown, the method includes operations S210~S230.
[0048] In operation S210, based on the first category result representing the category of the video to be identified, a search is performed to obtain search enhancement information.
[0049] In operation S220, the first category result, search enhancement information, and video to be identified are added to the first prompt template to obtain the first prompt information.
[0050] In operation S230, the first prompt information is input into the multimodal large model, so that the multimodal large model can be used to identify the video to be identified based on the retrieved enhanced information and the first category result, and obtain the target output information that matches the guidance information.
[0051] Deep learning models or multimodal large models can be used to identify the video to be identified, and the first category result representing the category of the video to be identified can be obtained.
[0052] Taking a review scenario as an example, the first category of results can indicate whether the video content to be identified violates regulations. Taking an object recognition scenario as an example, the first category of results can indicate whether a target object exists in the video to be identified.
[0053] However, while the first category of results is relatively concise, in human-computer interaction scenarios, providing users with more detailed information about the identified content makes it difficult for them to grasp the nuances of the identification process. For example, in a review scenario, it's difficult to determine which content in the video being identified violates regulations and the reasons for such violations based solely on the first category of results. Similarly, in a target recognition scenario, it's difficult to determine which video segment contains the target object, its posture and movements, and its background environment based solely on the first category of results.
[0054] Information retrieval can be performed using the first category results as search terms to obtain enhanced search information. Enhanced search information can be used to describe the characteristics of content matching the first category results. The type of enhanced search information is not limited; for example, it can include at least one of the following: examples of content matching the first category results, definitions used to define the first category results, and characteristic information of the content matching the first category results.
[0055] The first category result and retrieval enhancement information can be added to the first prompt template as reference information to obtain the first prompt information. The first prompt information is then input into the multimodal large model to utilize the multimodal large model to recognize the video to be recognized based on the first category result and retrieval enhancement information.
[0056] Guiding information can also be added to the first prompt template to include guidance that defines the output content of the multimodal large model. This standardizes the content type and format of the multimodal large model output.
[0057] For example, guiding information can be used to limit the target output information to include a second category result that characterizes the video category to be identified and explanatory information that describes the second category result, and semantic matching between the second category result and the first category result, but it is not limited to this, and the first category result and the second category result can also be the same.
[0058] Taking the review scenario as an example, the target output information may include Chinese content in the following sentence format: Based on the content to be identified, ..., it can be seen that ... displays .... This display method leads to ..., therefore, the identification result is: ....
[0059] According to embodiments of this disclosure, by obtaining a first category result through initial identification, retrieval enhancement information describing content features matching the first category result is retrieved, and combined with guidance information in a first prompt template, thereby guiding the multimodal large model to combine reference information to assist in outputting a second category result with the same semantics, while outputting explanatory information, so as to improve the richness and interpretability of the identification results, enrich the functionality of the multimodal large model, and thus improve the user experience.
[0060] According to embodiments of this disclosure, the video recognition method described above can be applied in inference scenarios such as video review and target tracking, and can also be applied in sample generation and labeling scenarios.
[0061] To improve the accuracy and effectiveness of the target output information, multiple multimodal large models can be set up to fuse the target output information output by each of the multiple multimodal large models, thereby making the explanatory information content of the text type standardized, effective and comprehensive.
[0062] According to embodiments of this disclosure, when performing such Figure 2 Simultaneously with the operation S230 shown, the video recognition method may further include: when there are multiple multimodal large models, inputting the first prompt information into each of the multiple multimodal large models to obtain multiple target output information. Based on the multiple second-category results, denoising is performed on the multiple target output information to obtain multiple candidate recognition results. Based on the interpretation information of each of the multiple candidate recognition results, information fusion is performed to obtain the first target recognition result.
[0063] Optionally, multiple multimodal large models can have the same model architecture but be trained using different training samples, resulting in different model functions. However, this is not a limitation. They can also have different model architectures and be trained using different training samples, resulting in different model functions. Thus, by utilizing multiple multimodal large models of different types, target output information from different dimensions or analytical perspectives can be obtained, thereby improving the richness and completeness of the fused first target recognition result.
[0064] Cross-validation can be performed using the outputs of multiple multimodal large models to ensure that the explanatory information in the first target recognition result is standardized, effective, and comprehensive, even when none of the multimodal large models can generate explanatory information to describe the second category result.
[0065] Optionally, information fusion can be performed directly based on the interpretation information of the output information of multiple targets to obtain the first target recognition result.
[0066] Compared to directly fusing all explanatory information, this method can also denoise multiple target outputs based on multiple second-category results, yielding multiple candidate recognition results. This removes noise and improves the effectiveness and accuracy of subsequent explanatory information fusion.
[0067] Figure 3 The schematic diagram illustrates the process of determining target interpretation information.
[0068] like Figure 3 As shown, the first prompt information is input into the first multimodal large model, the second multimodal large model, ..., the Nth multimodal large model, respectively, to obtain the first target output information, the second target output information, ..., the Nth target output information, where N is an integer greater than or equal to 2. Based on the second category results of the first target output information, the second target output information, ..., the Nth target output information is denoised to obtain the first candidate recognition result, the second candidate recognition result, ..., the Mth candidate recognition result, where M is an integer greater than or equal to 2 and less than N. The explanation information of the first candidate recognition result, the second candidate recognition result, ..., the Mth candidate recognition result is input into the large language model for information fusion to obtain the target explanation information.
[0069] Based on the target explanation information and the second category result in the candidate recognition results, the first target recognition result is obtained.
[0070] According to embodiments of this disclosure, for example, Figure 3The operation shown for determining candidate recognition results involves denoising multiple target output information based on multiple second-category results to obtain multiple candidate recognition results. This can include: determining the target second-category result with the largest number of identical categories from the multiple second-category results; and determining multiple candidate recognition results from the multiple target output information based on the target second-category result.
[0071] Taking the review scenario as an example, the second category results of the first target output information, the second target output information, ..., the Nth target output information can include "involving violent content", "involving false advertising terms", ... "involving violent content".
[0072] The number of identical categories in the second category results can be counted, and the target output information corresponding to the target second category result with the largest number of identical categories, such as "involving violent content", can be used as the candidate recognition result.
[0073] Therefore, noise information that causes erroneous output results due to hallucinations or other problems in large multimodal models is removed, improving the accuracy and effectiveness of subsequent interpretation information fusion. Furthermore, denoising is performed using rules for results of the same or similar categories in the second category, simplifying the denoising method while improving denoising efficiency and accuracy.
[0074] According to embodiments of this disclosure, for example, Figure 3 The information fusion operation shown, based on the explanatory information of multiple candidate recognition results, fuses information to obtain a first target recognition result. This can include: obtaining second prompt information by combining multiple explanatory information and a second prompt template; and inputting the second prompt information into a large language model to obtain the first target recognition result.
[0075] The second prompt template may include information describing the task of rewriting or polishing. For example, it may describe at least one task performed by the large language model: deduplication of semantically repetitive content or fusion of semantically different content.
[0076] For example, the explanatory information for each of the first, second, ..., and Mth candidate identification results could include: "This video shows the target object raising its hand and rushing forward... This display method is likely to cause discomfort to viewers and has a tendency to...", "This video shows the target object raising its foot... This display method is likely to cause discomfort to viewers and has a tendency to...", ..., "This video shows the target object raising its hand... This display method is likely to cause discomfort to viewers and has a tendency to...". The content such as "raising its hand and rushing forward" and "raising its hand" can be deduplicated, and semantic fusion can be performed on content such as "raising its hand and rushing forward" and "raising its foot" to obtain the target explanation information.
[0077] Therefore, a large language model is used to refine multiple explanatory information in plain text, so that the overall sentence of the target explanation information in the final target recognition result is fluent, comprehensive and effective.
[0078] The preceding text explained how to obtain the first target recognition result using the target output information from multiple multimodal large models. The following text will explain how to obtain the first prompt information used as input to the multimodal large model.
[0079] According to embodiments of this disclosure, for example, Figure 2 The operation S220 shown, which adds the first category result, search enhancement information and the video to be identified to the first prompt template to obtain the first prompt information, may include: adding the search enhancement information, the first category result, the video introduction information, the audio of the video to be identified and the video to be identified to the first prompt template to obtain the first prompt information.
[0080] Video description information can include a video cover and a video title. Taking short videos as an example, the video cover can be an image used to introduce the video to be identified, and can be any frame from the video to be identified, but is not limited to this. It can also be image information of relevant objects captured in the video to be identified, which is related to the content of the video to be identified but does not appear in the video content itself. The video title can be text used to introduce the video to be identified. Alternatively, it can be related to the content of the video to be identified but does not appear in the video content itself, but is part of the link address of the video to be identified, so that the theme of the video to be identified can be known through the video title.
[0081] By using video introduction information as the content to be identified, the identification process becomes effective and comprehensive.
[0082] According to an optional example of this disclosure, enhanced retrieval information and first category results can be added to the first field of the first prompt template as reference information. The first prompt information is obtained by adding the video to be identified and video description information describing the video to be identified, or by adding the video to be identified, audio matching the video to be identified, and video description information describing the video to be identified, to the second field of the first prompt template as content to be identified.
[0083] This allows for a richer and more comprehensive range of objects to be identified, while the first and second fields are used to distinguish between the identified objects and the reference objects, improving the prompting effect of the prompting information and thus improving the recognition accuracy of the multimodal large model.
[0084] Figure 4 A schematic diagram illustrating a first prompt message according to an embodiment of the present disclosure is shown.
[0085] like Figure 4As shown, the first prompt template may include field identification information for identifying "content to be identified", field identification information for identifying "reference information", and field identification information for identifying "guidance information".
[0086] like Figure 4 As shown, the video storage address of the video to be identified, the video cover storage address, the video title, and the audio text can be added to multiple second fields according to the field identifier information marked "Content to be Identified" in the first prompt template, such as... Figure 4 The “…” position under “Content to be identified” is shown.
[0087] like Figure 4 As shown, the first category results and search enhancement information are used as reference information. Following the field identifier information for "reference information" in the first prompt template, they are added to multiple first fields respectively, such as... Figure 4 The “…” position under “Reference Information” is shown.
[0088] like Figure 4 As shown, by combining the field identifier information that identifies "guidance information" with specific guidance content such as task requirements and target output information, the first prompt information is obtained.
[0089] Optionally, audio text can be obtained by converting audio that matches the video to be recognized. For example, a speech-to-text tool can be used to convert the audio that matches the video to be recognized into audio text.
[0090] According to embodiments of this disclosure, different types of content are added using different methods and according to field identification information to different field positions of the first prompt template. This makes the multimodal large model clearly and explicitly include reference information, content to be identified, and guidance information, thereby accurately and effectively performing tasks, improving the effectiveness and accuracy of target output information, and reducing the frequency of hallucination problems.
[0091] Optionally, for example Figure 2 The operation S230 shown involves inputting the first prompt information into the multimodal large model, so that the multimodal large model can identify the video to be identified based on the retrieval enhancement information and the first category result, and obtain target output information that matches the guidance information. This may include: the multimodal large model obtaining the video to be identified and the video cover from the corresponding storage space based on the video storage address and the cover storage address, so as to identify the video to be identified, the video cover, the video title, and the audio text based on the retrieval enhancement information and the first category result, and obtain target output information that matches the guidance information.
[0092] Optionally, the video to be recognized can be split into frames to obtain a video frame sequence, which is then stored in the video storage address so that the multimodal large model can recognize the video frame sequence, thereby reducing the recognition difficulty of the multimodal large model and ensuring recognition accuracy.
[0093] The preceding text explained how to obtain the initial prompt information and how to use it to obtain the target output information. The following text will explain how to determine the first category of results and search enhancement information.
[0094] According to embodiments of this disclosure, for example, Figure 2 Before the operation S210 shown, the video recognition method may further include: performing category recognition on the video to be recognized and the video description information used to describe the video to be recognized, and obtaining a first category result.
[0095] The video description information includes at least one of the following: video cover and video title. The video description information can be processed using a multimodal large model to identify the video, audio text, and video description information such as the video cover and video title, to obtain a first-category result. However, it is not limited to this. Recognition tools that match the type of content to be identified can also be used to identify the video, audio text, video cover, and video title separately, obtaining a first-category recognition result. Any method that can achieve category recognition is acceptable.
[0096] By utilizing an initial identification method, a first-class result representing the category of the video to be identified is obtained, reducing the difficulty of identification. This initial identification operation is then compared with... Figure 2 The operation S230 shown combines dual recognition to improve the level of fine recognition while reducing the difficulty of obtaining interpretable information.
[0097] Figure 5 A schematic diagram illustrating the determination of a first category result according to an embodiment of the present disclosure is shown.
[0098] like Figure 5 As shown, image recognition is performed on the video frames and video cover of the video to be identified, yielding a first recognition result. Semantic recognition is performed on the audio that matches the video to be identified, yielding a second recognition result. Text recognition is performed on the video title in the video introduction information, yielding a third recognition result. Based on the first, second, and third recognition results, a first category result is obtained.
[0099] Image recognition models can be used for image recognition, and these models can be trained using deep learning. Any model capable of image category recognition is acceptable. Similarly, audio can be converted to text, and text semantic recognition models can be used to recognize both audio text and video titles. The type of text semantic recognition model is not limited, as long as it can perform semantic recognition.
[0100] Depending on the identification scenario, different fusion strategies can be employed to fuse the first, second, and third identification results to obtain the first category result. For example, in a violation review scenario, if any one of the first, second, or third identification results represents a violation, the first category result represents a violation. Another example is in a target tracking scenario, where the first, second, and third identification results are weighted and summed to obtain the first category result. Alternatively, the identification result with the highest number of entries in the same category among the first, second, and third identification results can be used as the first category result.
[0101] According to embodiments of this disclosure, different recognition methods are used for targeted recognition of different types of content, thereby improving recognition accuracy. Furthermore, different result fusion strategies are employed for different recognition scenarios, thereby enhancing the targeting and effectiveness of the recognition.
[0102] The above section explained how to obtain the first category results. The following section will explain how to determine search enhancement information based on the first category results.
[0103] According to embodiments of this disclosure, for example, Figure 2 The operation S210 shown, based on the first category result representing the category of the video to be identified, performs a search to obtain search enhancement information, which may include: searching from the knowledge base according to the keyword matching method to obtain search enhancement information that has a mapping relationship with the first category result.
[0104] Figure 6 A schematic diagram illustrating the determination of enhanced retrieval information according to an embodiment of the present disclosure is shown.
[0105] like Figure 6 As shown, the knowledge base stores category information with mapping relationships and feature information describing content features that match the category information. For example, the knowledge base stores a list of mapping relationships. The mapping relationships are represented by the row relationships in the mapping relationship list. In the same row, there is a mapping relationship between the category information stored in column A and the feature information stored in column B.
[0106] like Figure 6 As shown, based on the results of the first category, the target category information can be determined from column A of the knowledge base according to either keyword matching or semantic matching. Figure 4 The category information shown*. Based on the mapping relationship, feature information that has a mapping relationship with the target category information is determined from the knowledge base, such as... Figure 4 The feature information shown* serves as retrieval enhancement information.
[0107] Optionally, the feature information stored in the knowledge base may include at least one of the following: definition, example, and recognition rules.
[0108] Feature information can be used as retrieval enhancement information, highlighting features or characteristics that distinguish the categories of videos to be identified in a fine-grained manner, thereby providing guidance information and improving the ability to analyze and output explanatory information.
[0109] The above section explained how to obtain the first target recognition result. The following section will explain its application in the training sample generation scenario.
[0110] Figure 7 A flowchart illustrating a model training method according to an embodiment of the present disclosure is shown schematically.
[0111] like Figure 7 As shown, the method includes operation S710 and operation S720.
[0112] Use the S710 to obtain training samples.
[0113] When operating the S720, the training samples are used to train the model to be trained, and the trained model is obtained.
[0114] The training samples include sample videos and labels. Labels are obtained by identifying the sample videos using the video recognition method provided in the above embodiments. Labels may include sample category results characterizing the category of the sample video and explanatory information describing the sample category results.
[0115] Optionally, such as Figure 7 The model training method shown includes operations S710 and S720, but it is not limited to these; it can also include only operation S720. The key is to ensure that the model can be trained using training samples.
[0116] By using the video recognition method provided in the above embodiments to identify sample videos, a sample category result including a characterizing the category of the sample video and a label for describing the explanatory information of the sample category result is obtained. The training model is then trained using training samples with the label, which can improve the fine-grained analysis of the recognized videos by the trained model, thereby improving the richness and interpretability of the recognition results.
[0117] Optionally, the model to be trained may include the large multimodal model provided in the embodiments of this disclosure. However, it is not limited to this. Any large model capable of processing multimodal data is acceptable.
[0118] According to embodiments of this disclosure, training a model to be trained using training samples to obtain a trained model includes: adding sample videos to a third prompt template to obtain third prompt information; inputting the third prompt information into the model to be trained to recognize the sample videos and obtain sample output information matching the prompt information; and adjusting the parameters of the model to be trained based on the sample output information and labels to obtain a trained model.
[0119] The third prompt template includes guiding information describing the output of the model to be trained.
[0120] Figure 8 The schematic diagram illustrates a flowchart of a model training method according to an embodiment of the present disclosure.
[0121] like Figure 8 As shown, the third prompt template may include field identification information that identifies the "content to be identified" and field identification information that identifies the "guidance information".
[0122] like Figure 8 As shown, the video storage address of the sample video, the cover storage address of the sample video, the title of the sample video, and the sample audio text can be added to multiple second fields according to the field identification information of the "content to be identified" in the third prompt template.
[0123] like Figure 8 As shown, by combining the field identifier information that identifies "guidance information" with specific guidance content such as task requirements and sample output information, the third prompt information is obtained.
[0124] like Figure 8 As shown, the third prompt information is input into the model to be trained to obtain sample output information. Based on the sample output information and labels, the parameters of the model to be trained are adjusted to obtain the trained model.
[0125] Optionally, cross-entropy loss can be calculated on the category results in the sample output information and the category results in the labels to obtain a first loss value. Semantic similarity can be calculated on the explanatory information in the sample output information and the explanatory information in the labels to obtain a second loss value. The first loss value and the second loss value are weighted and summed to obtain the target loss value. Based on the target loss value, the parameters of the model to be trained are adjusted until the target loss value converges or the training epoch is reached, resulting in the trained model.
[0126] According to embodiments of this disclosure, using the third prompt information generated by the aforementioned third prompt template for model training can guide the model to be trained to output explanatory information. Furthermore, the combination of the label with category results and explanatory information can serve as a reference result to train the model to be trained to learn the ability to output explanatory information, thereby improving the richness and interpretability of the output results of the trained model.
[0127] The preceding text explained how to obtain a trained model capable of performing video recognition, producing category results, and interpreting information. The following text will explain how to utilize this trained model for video recognition.
[0128] Figure 9 A flowchart illustrating a video recognition method according to an embodiment of the present disclosure is shown schematically.
[0129] like Figure 9 As shown, the method includes operations S910~S920.
[0130] Using the S910, acquire the target video.
[0131] When operating the S920, the target video is input into the target recognition model to obtain the second target recognition result.
[0132] The target recognition model is trained using the model training method provided in the above embodiments.
[0133] Optionally, such as Figure 9 The model training method shown includes operations S910 and S920, but it is not limited to these; it can also include only operation S920. This is as long as the target recognition model can be used to process the target video.
[0134] According to embodiments of this disclosure, the target recognition model trained using the above-described model training method can improve the richness and explanatoryness of the second target recognition result by utilizing the category results and explanatory information included in the second target recognition result.
[0135] Figure 10 A schematic flowchart of a video recognition method according to an embodiment of the present disclosure is shown.
[0136] like Figure 10 As shown, the fourth prompt template may include field identification information that identifies the "content to be identified" and field identification information that identifies the "guidance information".
[0137] like Figure 10 As shown, the target video's storage address, the target video cover's storage address, the target video title, and the target audio text can be added to multiple second fields according to the field identification information marked "Content to be Identified" in the fourth prompt template.
[0138] like Figure 10 As shown, by combining the field identifier information that identifies "guidance information" with specific guidance content such as task requirements and examples of second target recognition results, the fourth prompt information is obtained.
[0139] like Figure 10 As shown, the fourth prompt information is input into the target recognition model to obtain the second target recognition result.
[0140] According to embodiments of this disclosure, different types of content are added in different ways and according to field identification information to different field positions of the fourth prompt template. This makes the target recognition model clearly identify the content to be recognized and the guidance information, thereby accurately and effectively performing the task, improving the effectiveness and accuracy of the second target recognition result, and reducing the frequency of hallucination problems.
[0141] Figure 11 A block diagram of an artificial intelligence-based video recognition device according to an embodiment of the present disclosure is shown schematically.
[0142] like Figure 11 As shown, the video recognition device 1100 includes: a retrieval module 1110, a first prompt confirmation module 1120, and a first recognition module 1130.
[0143] The retrieval module 1110 is used to perform a retrieval based on the first category results representing the category of the video to be identified, and obtain retrieval enhancement information. The retrieval enhancement information is used to describe the features of the content that matches the initial category results.
[0144] The first prompt determination module 1120 is used to add the first category result, search enhancement information, and the video to be identified to the first prompt template to obtain the first prompt information. The first prompt template includes guiding information describing the output content of the multimodal large model.
[0145] The first recognition module 1130 is used to input the first prompt information into the multimodal large model, so as to use the multimodal large model to recognize the video to be recognized based on the retrieved enhanced information and the first category result, and obtain the target output information that matches the guidance information.
[0146] The target output information includes a second category result to characterize the video category to be identified and explanatory information to describe the second category result, with semantic matching between the second category result and the first category result.
[0147] According to embodiments of this disclosure, the video recognition device further includes: a multi-recognition module, a noise reduction module, and a fusion module.
[0148] The multi-recognition module is used to input the first prompt information into multiple multi-modal large models when there are multiple models, so as to obtain multiple target output information.
[0149] The noise reduction module is used to reduce noise on multiple target output information based on multiple second-class results, so as to obtain multiple candidate recognition results.
[0150] The fusion module is used to fuse information based on the explanatory information of multiple candidate recognition results to obtain the first target recognition result.
[0151] According to embodiments of this disclosure, the fusion module includes: an explanation prompt submodule and a fusion submodule.
[0152] The Explanation Prompt Submodule is used to obtain the second prompt information from multiple explanation messages and the second prompt template.
[0153] The fusion submodule is used to input the second prompt information into the large language model to obtain the first target recognition result.
[0154] The second prompt template includes at least one task performed by the large language model: deduplication of semantically repetitive content and fusion of semantically different content.
[0155] According to embodiments of this disclosure, the noise reduction module includes: a comparison submodule and a screening submodule.
[0156] The comparison submodule is used to determine the target second-category result with the maximum number of identical categories from multiple second-category results.
[0157] The filtering submodule is used to determine multiple candidate recognition results from multiple target output information based on the target's second category results.
[0158] According to embodiments of this disclosure, the first prompt determination module includes: a first adding submodule and a second adding submodule.
[0159] The first addition submodule is used to add the search enhancement information and the first category result to the first field of the first prompt template as reference information.
[0160] The second addition submodule is used to add the video to be identified and the video description information to be identified to the second field of the first prompt template, which is the content to be identified, to obtain the first prompt information.
[0161] According to embodiments of this disclosure, the video description information includes a video cover and a video title.
[0162] According to an embodiment of this disclosure, the second adding submodule includes: a first adding unit.
[0163] The first adding unit is used to add the video storage address of the video to be identified, the cover storage address of the video cover, the video title and the audio text to multiple second fields according to the field identification information in the first prompt template. The audio text is obtained based on the audio conversion that matches the video to be identified.
[0164] According to embodiments of this disclosure, the retrieval module includes a retrieval submodule.
[0165] The retrieval submodule is used to retrieve information from the knowledge base based on keyword matching to obtain retrieval enhancement information that has a mapping relationship with the results of the first category.
[0166] The knowledge base stores category information that is mapped to other categories and feature information that describes the content features that match the category information.
[0167] According to embodiments of this disclosure, the video recognition device further includes an initial recognition module.
[0168] The initial recognition module is used to classify the video to be recognized and the video description information used to describe the video to be recognized, and obtain the first category result.
[0169] According to embodiments of this disclosure, the video description information includes at least one of the following: video cover and video title.
[0170] According to embodiments of this disclosure, the initial identification module includes: a first identification submodule, a second identification submodule, a third identification submodule, and a fourth identification submodule.
[0171] The first recognition submodule is used to perform image recognition on the video frames and video cover in the video to be recognized, and obtain the first recognition result.
[0172] The second recognition submodule is used to perform semantic recognition on the speech that matches the video to be recognized, and obtain the second recognition result.
[0173] The third recognition submodule is used to perform text recognition on the video title in the video introduction information to obtain the third recognition result.
[0174] The fourth identification submodule is used to obtain the first category result based on the first identification result, the second identification result, and the third identification result.
[0175] Figure 12 A block diagram of a model training apparatus according to an embodiment of the present disclosure is shown schematically.
[0176] like Figure 12 As shown, the model training device 1200 includes: a first acquisition module 1210 and a training module 1220.
[0177] The first acquisition module 1210 is used to acquire training samples.
[0178] Training module 1220 is used to train the model to be trained using training samples to obtain the trained model;
[0179] The training samples include sample videos and labels, and the labels are obtained by recognizing the sample videos using the video recognition device provided in the above embodiments.
[0180] Optionally, such as Figure 12 The model training device shown includes a first acquisition module 1210 and a training module 1220, but is not limited to these; it may also include only the training module 1220.
[0181] According to embodiments of this disclosure, the training module includes: a third adding submodule, a fifth recognition submodule, and a training submodule.
[0182] The third submodule adds sample videos to the third prompt template to obtain third prompt information. The third prompt template includes guiding information describing the output of the model to be trained.
[0183] The fifth recognition submodule is used to input the third prompt information into the model to be trained, so that the model can be used to recognize the sample video and obtain sample output information that matches the guidance information.
[0184] The training submodule is used to adjust the parameters of the model to be trained based on the sample output information and labels, so as to obtain the trained model.
[0185] Figure 13 A block diagram of a video recognition device according to an embodiment of the present disclosure is shown schematically.
[0186] like Figure 13 As shown, the video recognition device 1300 includes: a second acquisition module 1310 and a second recognition module 1320.
[0187] The second acquisition module 1310 is used to acquire the target video.
[0188] The second recognition module 1320 is used to input the target video into the target recognition model to obtain the second target recognition result;
[0189] The target recognition model is trained using the model training device provided in the above embodiments.
[0190] Optionally, such as Figure 13 The video recognition device shown includes a second acquisition module 1310 and a second recognition module 1320, but is not limited to this and may include only the second recognition module 1320.
[0191] According to embodiments of this disclosure, the second identification module includes a fourth adding submodule and a sixth identification submodule.
[0192] The fourth submodule is used to add the target video to the fourth prompt template to obtain the fourth prompt information.
[0193] The sixth recognition submodule is used to input the fourth prompt information into the target recognition model to obtain the second target recognition result.
[0194] The fourth prompt template includes guiding information that describes the output of the target recognition model.
[0195] Figure 14 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.
[0196] In embodiments of this disclosure, such as Figure 14 As shown, the intelligent agent 1400 may include an input module 1410, a processing module 1420, and an output module 1430.
[0197] Input module 1410 is used to receive input information.
[0198] The processing module 1420 is used to determine the target task based on the input information received by the input module, determine the large model based on the target task, and obtain output information by calling the large model to execute the video recognition method based on artificial intelligence provided in the embodiments of this disclosure, or by calling the large model to execute the model training method provided in the embodiments of this disclosure.
[0199] Output module 1430 is used to output the output information obtained by the processing module.
[0200] According to embodiments of this disclosure, the input module 1410 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the intelligent agent 1400 can understand and process. The input module 1410 is the primary link for the intelligent agent 1400 to interact with the outside world, enabling the intelligent agent 1400 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0201] In the example, input module 1410 can input the video to be identified, training samples, target video, etc., as described above.
[0202] In the example, processing module 1420 is the core support for the ability of agent 1400 to handle complex tasks. Processing module 1420 can execute the AI-based video recognition method, model training method, or video recognition method described above.
[0203] In the example, the performance of processing module 1420 is closely related to the large model on which agent 1400 is based. To fully leverage the capabilities of the large model, the internal structure of processing module 1420 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.
[0204] In the example, after the agent 1400 acquires the video to be identified, the processing module 1420 can perform a retrieval based on the first category result representing the category of the video to be identified, obtain retrieval enhancement information, add the first category result, retrieval enhancement information and the video to be identified to the first prompt template to obtain the first prompt information, input the first prompt information into the multimodal large model, so that the multimodal large model can identify the video to be identified based on the retrieval enhancement information and the first category result, obtain the target output information that matches the guidance information, and pass the target output information to the output module 1430.
[0205] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. When Agent 1400 is given the ability to invoke tools, it can perform tasks such as using a calculator to perform mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.
[0206] In the example, the output module 1430 can output the target output information described above, the trained model, or the second target recognition result.
[0207] The intelligent agent 1400 according to the embodiments of this disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0208] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0209] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0210] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.
[0211] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.
[0212] Figure 15 A schematic block diagram of an example electronic device 1500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0213] like Figure 15 As shown, device 1500 includes a computing unit 1501, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1502 or a computer program loaded from storage unit 1508 into random access memory (RAM) 1503. The RAM 1503 may also store various programs and data required for the operation of device 1500. The computing unit 1501, ROM 1502, and RAM 1503 are interconnected via bus 1504. Input / output (I / O) interface 1505 is also connected to bus 1504.
[0214] Multiple components in device 1500 are connected to input / output (I / O) interface 1505, including: input unit 1506, such as a keyboard, mouse, etc.; output unit 1507, such as various types of displays, speakers, etc.; storage unit 1508, such as a disk, optical disk, etc.; and communication unit 1509, such as a network card, modem, wireless transceiver, etc. Communication unit 1509 allows device 1500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0215] The computing unit 1501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1501 performs the various methods and processes described above, such as video recognition methods and model training methods. For example, in some embodiments, the video recognition methods and model training methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1508. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1500 via ROM 1502 and / or communication unit 1509. When the computer program is loaded into RAM 1503 and executed by the computing unit 1501, one or more steps of the video recognition methods and model training methods described above can be performed. Alternatively, in other embodiments, the computing unit 1501 may be configured to perform video recognition methods or model training methods by any other suitable means (e.g., by means of firmware).
[0216] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0217] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0218] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0219] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0220] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0221] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.
[0222] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0223] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video recognition method based on artificial intelligence, comprising: Based on the first category result representing the category of the video to be identified, a search is performed to obtain search enhancement information, wherein the search enhancement information is used to describe the features of the content that matches the initial category result; The first category result, the search enhancement information, and the video to be identified are added to the first prompt template to obtain the first prompt information. The first prompt template includes guiding information describing the output content of the multimodal large model; and The first prompt information is input into the multimodal large model, so that the multimodal large model can identify the video to be identified based on the retrieval enhancement information and the first category result, and obtain target output information that matches the guidance information; The target output information includes a second category result for characterizing the category of the video to be identified and explanatory information for describing the second category result, wherein the second category result is semantically matched with the first category result.
2. The method according to claim 1, further comprising: When the multimodal large model includes multiple models, the first prompt information is input into each of the multiple multimodal large models to obtain multiple target output information; Based on multiple results of the second category, noise reduction is performed on multiple target output information to obtain multiple candidate recognition results; as well as Based on the interpretation information of each of the multiple candidate recognition results, information fusion is performed to obtain the first target recognition result.
3. The method according to claim 2, wherein, The step of fusing information based on the interpretation information of each of the multiple candidate recognition results to obtain a first target recognition result includes: The second prompt information is obtained from multiple explanation information and second prompt templates; and The second prompt information is input into the large language model to obtain the first target recognition result; The second prompt template includes at least one task performed by the large language model: deduplication of semantically repetitive content and fusion of semantically different content.
4. The method according to claim 2 or 3, wherein, Based on multiple results of the second category, noise reduction is performed on multiple target output information to obtain multiple candidate recognition results, including: Determine the target second category result with the maximum number of identical categories from multiple second category results; and Based on the second category result of the target, a plurality of candidate recognition results are determined from the output information of the plurality of targets.
5. The method according to any one of claims 1 to 4, wherein, The step of adding the first category result, the search enhancement information, and the video to be identified to the first prompt template to obtain the first prompt information includes: Add the enhanced search information and the first category result to the first field of the first prompt template as reference information; and The video to be identified and the video description information used to describe the video to be identified are added to the second field of the first prompt template as the content to be identified, so as to obtain the first prompt information.
6. The method according to claim 5, wherein, The video description information includes the video cover and video title; The step of adding the video to be identified and video description information to the second field of the first prompt template as the content to be identified, to obtain the first prompt information, includes: The video storage address of the video to be identified, the cover storage address of the video cover, the video title, and the audio text are added to multiple second fields according to the field identification information in the first prompt template. The audio text is obtained based on the audio conversion that matches the video to be identified.
7. The method according to any one of claims 1 to 6, wherein, The retrieval is performed based on the first category result representing the category of the video to be identified, to obtain retrieval enhancement information, including: According to the keyword matching method, a search is performed in the knowledge base to obtain the search enhancement information that has a mapping relationship with the results of the first category; The knowledge base stores category information that is mapped to other categories and feature information that describes the content features that match the category information.
8. The method according to any one of claims 1 to 7, further comprising: The video to be identified and the video description information used to describe the video to be identified are classified into categories to obtain the first category result.
9. The method according to claim 8, wherein, The video description information includes at least one of the following: video cover, video title; The step of classifying the video to be identified and the video description information used to describe the video to be identified to obtain the first category result includes: Image recognition is performed on the video frames and the video cover in the video to be identified to obtain a first recognition result; Semantic recognition is performed on the speech that matches the video to be identified to obtain a second recognition result; Text recognition is performed on the video title in the video description information to obtain a third recognition result; and Based on the first identification result, the second identification result, and the third identification result, the first category result is obtained.
10. A model training method, comprising: The model to be trained is trained using training samples to obtain the trained model; The training samples include sample videos and labels, wherein the labels are obtained by identifying the sample videos using the method described in any one of claims 1 to 9.
11. The method according to claim 10, wherein, The step of training the model to be trained using training samples to obtain the trained model includes: The sample video is added to the third prompt template to obtain the third prompt information, wherein the third prompt template includes guiding information for describing the output content of the model to be trained; The third prompt information is input into the model to be trained, so that the model can be used to identify the sample video and obtain sample output information that matches the prompt information; and Based on the sample output information and labels, the parameters of the model to be trained are adjusted to obtain the trained model.
12. A video recognition method, comprising: The target video is input into the target recognition model to obtain the second target recognition result; The target recognition model is trained using the method described in claim 10 or 11.
13. The method according to claim 12, wherein, The step of inputting the target video into the target recognition model to obtain the second target recognition result includes: Add the target video to the fourth prompt template to obtain the fourth prompt information; The fourth prompt information is input into the target recognition model to obtain the second target recognition result; The fourth prompt template includes guidance information describing the output content of the target recognition model.
14. A video recognition device, comprising: The retrieval module is used to perform a retrieval based on the first category result representing the category of the video to be identified, and obtain retrieval enhancement information, wherein the retrieval enhancement information is used to describe the features of the content that matches the initial category result; The first prompt determination module is used to add the first category result, the search enhancement information, and the video to be identified to a first prompt template to obtain first prompt information. The first prompt template includes guiding information describing the output content of the multimodal large model; and The first recognition module is used to input the first prompt information into the multimodal large model, so as to use the multimodal large model to recognize the video to be recognized based on the retrieval enhancement information and the first category result, and obtain target output information that matches the guidance information; The target output information includes a second category result for characterizing the category of the video to be identified and explanatory information for describing the second category result, wherein the second category result is semantically matched with the first category result.
15. A model training device, comprising: The training module is used to train the model to be trained using training samples to obtain the trained model. The training samples include sample videos and labels, wherein the labels are obtained by recognizing the sample videos using the apparatus as described in claim 14.
16. A video recognition device, comprising: The second recognition module is used to input the target video into the target recognition model to obtain the second target recognition result; The target recognition model is obtained by training the apparatus as described in claim 15.
17. An intelligent agent of artificial intelligence, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method of any one of claims 1 to 13 by calling the large model to obtain output information; An output module is used to output the output information obtained by the processing module.
18. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
19. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.
20. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-13.