Processing method for multimedia data and related equipment
By obtaining the characteristic information of multimedia data, using artificial intelligence and large language models to automatically filter the target objects, the problem of inefficient manual editing of multimedia data description information is solved, and efficient and accurate description information generation is achieved.
Patent Information
- Application Number
- CN202410027891.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, manual editing of video or picture description information is inefficient and it is difficult to cope with the large number of multimedia data processing needs in the era of data explosion.
By obtaining the characteristic information of multimedia data, using artificial intelligence and large language models, we automatically filter out the target objects and generate description information, reducing manual intervention.
It improves the efficiency and accuracy of multimedia data description information generation, reduces redundant information, and improves the accuracy and recall rate of object screening.
Smart Images

Figure CN120277228A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and particularly to a method for processing multimedia data and related devices. Background Art
[0002] Currently, with the development of Internet technologies, in order to better provide important information for tasks such as recommendation distribution, review, and search, it is necessary to generate description information for some data to be searched (such as videos, images, etc.). For example, an introduction to a certain movie can be generated, so that these description information can be used to better complete the tasks.
[0003] Currently, most of the solutions for generating description information of data such as videos and images are manually edited by staff or authors. For example, staff manually edit posters for movies, including editing information such as the leading actors of the movie. Manual editing can, to a certain extent, accurately generate description information for data such as videos and images, and based on these description information, subsequent tasks such as recommendation and retrieval can be well completed. However, in the current era of data explosion, a large number of videos and pictures, especially short videos, are generated. Manually editing description information for these videos or pictures is time-consuming, laborious, and inefficient. Summary of the Invention
[0004] Embodiments of this application provide a method for processing multimedia data, which can efficiently generate description information for input data such as pictures or videos.
[0005] On the one hand, embodiments of this application provide a method for processing multimedia data, the method comprising:
[0006] Obtain input data to be processed, the input data including an image or an image sequence composed of multiple images;
[0007] Extract information from the input data to obtain feature information of the input data, the feature information including: description information of N image objects in the input data, where N is a positive integer;
[0008] Fill the description information of the N image objects into a question-and-answer prompt template to obtain screening indication information;
[0009] According to the screening indication information and the input data, screen out M target objects from the N image objects, and generate description information of the input data according to the M target objects, where M is a positive integer less than or equal to N.
[0010] On the one hand, embodiments of this application provide a device for processing multimedia data, the device comprising:
[0011] Obtain input data to be processed, the input data including an image or an image sequence composed of multiple images;
[0012] Extract information from the input data to obtain the feature information of the input data. The feature information includes: description information of N image objects in the input data, where N is a positive integer;
[0013] Fill the description information of the N image objects into the Q&A prompt template to obtain the screening indication information;
[0014] According to the screening indication information and the input data, screen out M target objects from the N image objects, and generate the description information of the input data based on the M target objects, where M is a positive integer less than or equal to N.
[0015] On the one hand, an embodiment of the present application provides a computer device, which includes:
[0016] A processor, suitable for executing a computer program;
[0017] A computer-readable storage medium, in which a computer program is stored. When the computer program is executed by the processor, the processing method for multimedia data as described above is implemented.
[0018] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is loaded and executed by the processor to implement the processing method for multimedia data as described above.
[0019] On the one hand, an embodiment of the present application provides a computer program product, which includes a computer program or computer instructions. When the computer program or computer instructions are executed by the processor, the processing method for multimedia data as described above is implemented.
[0020] In the embodiment of the present application, by extracting information from the obtained input data to be processed, the feature information of the input data can be obtained. The input data includes an image or an image sequence composed of multiple images. The feature information includes: description information of N image objects in the input data, where N is a positive integer. The description information of the N image objects can provide a basis for subsequent screening of image objects. Then, fill the description information of the N image objects into the Q&A prompt template to obtain the screening indication information; finally, according to the screening indication information and the input data, screen out M target objects from the N image objects. It can be seen that by screening the N image objects, relatively reliable target objects can be provided for generating the description information of the subsequent input data. Further, generate the description information of the input data based on the M target objects, where M is a positive integer less than or equal to N. The description information of the input data can be automatically generated based on the M target objects, without manual editing of the description information, and the description information of input data such as pictures or videos can be generated efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a flowchart for the recommendation distribution of tags in a tag system provided by an embodiment of the present application;
[0023] Figure 2 It is an architecture diagram of a processing system for multimedia data provided by an embodiment of the present application;
[0024] Figure 3 It is a schematic diagram of a processing flow for multimedia data provided by an embodiment of the present application;
[0025] Figure 4 It is a schematic diagram of a processing method flow for multimedia data provided by an embodiment of the present application;
[0026] Figure 5a It is a schematic diagram of an information generation interface provided by an embodiment of the present application;
[0027] Figure 5b It is a schematic diagram of an object playback interface provided by an embodiment of the present application;
[0028] Figure 6 It is a schematic diagram of a processing flow for encoding a cover image provided by an embodiment of the present application;
[0029] Figure 7 It is a schematic diagram of a process for filling a Q&A prompt template provided by an embodiment of the present application;
[0030] Figure 8 It is a schematic diagram of another processing method flow for multimedia data provided by an embodiment of the present application;
[0031] Figure 9 It is an object screening flowchart provided by an embodiment of the present application;
[0032] Figure 10 It is a schematic diagram of the structure of a processing device for multimedia data provided by an embodiment of the present application;
[0033] Figure 11 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0035] The embodiment of the present application provides a processing solution for multimedia data. This solution can extract feature information from the input data to be processed (such as images or videos). The feature information may include cover image information in the input data, text information, position information of image objects in the input data, etc. Then, in the form of questions and answers, screening indication information is generated through the feature information of the input data, and M target objects are screened out from N image objects in the input data according to the screening indication information and the input data. N is a positive integer, and M is less than or equal to N.
[0036] Among them, in the process of extracting the position information of image objects in the input data, the embodiments of this application involve Artificial Intelligence (AI). AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. Among them, Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0037] Specifically, in the process of extracting the position information of image objects in the input data, object recognition technology is used to determine information such as the duration and position of each image object in the input data. Schematically, if the image object can include a person, then the object recognition technology is face recognition technology. Through face recognition technology, the duration, position information, etc. of each face can be determined from the input data. When extracting the cover image information, the cover image can be encoded to obtain the word segmentation vector of the cover image.
[0038] In one implementation, the feature information can be filled into the Q&A prompt template to obtain the screening indication information, and M target objects are screened out from the N image objects in the input data according to the screening indication information and the input data. In another implementation, the embodiments of the present application relate to large language models. The feature information can be filled into the Q&A prompt template to obtain the screening indication information, and the screening indication information and the input data are input into the trained large language model together to screen the N image objects, and M target objects output by the trained large language model are obtained. N is a positive integer, and M is a positive integer less than or equal to N. A large language model, also known as a pre-training model or a foundation model, is a natural language processing model constructed based on deep learning technology. A large language model refers to a deep neural network (DNN) with a large number of parameters, which is trained on a large amount of unlabeled data. The function approximation ability of the large-parameter DNN is used to enable the PTM to extract common features from the data. Through techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is applicable to downstream tasks. Therefore, a large language model can achieve ideal results in few-shot or zero-shot scenarios. A large language model can be widely applied to tasks such as language modeling, machine translation, text summarization, and sentiment analysis. Schematically, a large language model generates coherent text by predicting the probability of the next word or character. Because of its large number of parameters, the effect of a large language model is better than that of a general language model. Among them, large language models can be divided into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal large language models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, a multi-modal large language model refers to a model that establishes feature representations of two or more data modalities. Specifically, a multi-modal large language model is a model that, based on a large language model, inputs multi-modal information (videos, graphics, text, voice, etc.) after special transformation into the large language model for the purpose of improving the effect and capabilities.
[0039] Through the processing solution for multimedia data provided by the embodiments of the present application, rich text information in the input data such as title information, the tokenization vectors of the cover images (i.e., OCR text information), and information such as the duration and position of each image object in the input data can be fully utilized to implement the screening of N image objects in the input data. In addition, by inputting the feature information into the large language model for processing multimedia data, the powerful content first-order ability of the large language model can be utilized to quickly and accurately screen out the target objects in the input data, thereby achieving the purpose of improving the accuracy and recall rate of object screening.
[0040] Furthermore, based on M target objects, description information of the input data can be generated. Specifically, the identification information of the M target objects can be obtained, and based on the identification information of the M target objects, description information of the input data can be generated. The identification information here can include object names, object nicknames, etc. By way of illustration, if the target objects are people, the object names of each target object can be determined according to the M target objects, and the object names of each target object can be used as the description information of the input data. Of course, a description sentence can be generated from the object names of each target object as the description information of the input data. The embodiments of the present application do not make any limitations in this regard. It can be seen that by screening the N image objects, the information redundancy of the description information of the input data can be reduced to a certain extent, the purity of the description information of the input data can be improved, and thus description information can be efficiently generated for input data such as pictures or videos.
[0041] In summary, adopting the processing solution for multimedia data provided by the embodiments of the present application can have the following beneficial effects: 1) By fully utilizing the title information in the input data, the tokenization vectors of the cover images related to the input data, the position information of the image objects in the input data, etc., a judgment basis can be provided for the screening of the image objects in the input data. 2) During the screening process, by generating screening indication information and guiding the screening of image objects in the form of questions and answers, the accuracy of object screening can be improved to a certain extent. 3) During the screening of N image objects, the trained large language model can be used to fully utilize the powerful content understanding ability of the large language model, which can further improve the accuracy of object screening. In addition, if multimodal information such as the title information in the input data, the tokenization vectors of the cover images related to the input data, and the position information of the image objects in the input data are used during the screening of N image objects using the trained large language model, the ability of the trained large language model can be further improved, thereby being able to further improve the accuracy of object screening and efficiently generate description information for input data such as pictures or videos.
[0042] The processing solution for multimedia data provided by the embodiments of this application can be applied to various application scenarios. By way of illustration, the application scenarios include but are not limited to: short video scenarios, label recall scenarios, video recommendation scenarios, video search scenarios, and so on. Next, the processing solution for multimedia data provided by the embodiments of this application will be elaborated through two specific scenarios.
[0043] (1) Short video scenario
[0044] In the short video scenario, the input data can be a short video. Through the processing solution for multimedia data provided by the embodiments of this application, information extraction can be performed on the short video to obtain the feature information of the short video. The feature information of the short video includes the description information of N image objects (i.e., faces). The description information of the N image objects is filled into the Q&A prompt template to obtain the screening indication information. Based on the screening indication information and the short video, the key object (i.e., the target object) can be determined from the N image objects in the short video. The key object is the leading actor object in the short video. Based on the identification information of the leading actor object, the description information of the short video is generated, and the description information of the short video is displayed.
[0045] In the short video scenario, through the processing solution for multimedia data provided by the embodiments of this application, the key object can be screened out from the short video, and the description information of the short video can be generated based on the key object, which can efficiently generate the description information for the short video and help users quickly and intuitively understand the leading actor object of a certain short video.
[0046] (2) Label recall scenario
[0047] In the label recall scenario, the input data can be various types of videos. Through the processing solution for multimedia data provided by the embodiments of this application, information extraction can be performed on various types of videos to obtain the feature information of various types of videos. The feature information of various types of videos includes the description information of N image objects (i.e., faces). The description information of the N image objects is filled into the Q&A prompt template to obtain the screening indication information. Based on the screening indication information and various types of videos, the key object (i.e., the target object) in various types of videos can be determined from the N image objects in various types of videos. Based on the identification information of the key object, the label (i.e., the description information) of various types of videos is generated, and the video with the generated label is added to the label system. Among them, the label system refers to a system that can attach various rich labels to the input data. The labels are such as the names of people, the names of plays, the names of songs, items, scenes, etc. in the input data. The generated labels are used for downstream recommendation, search, distribution and other services. By way of illustration, as Figure 1 shown, it is the process of using the label in a label system provided by the embodiments of this application for recommendation and distribution. In Figure 1Among them, the input data is Video A. Through the processing solution for multimedia data provided by the embodiments of the present application, the key objects "Xiaoming" and "Zhang San" can be screened out from Video A. According to the identification information of the key object "Xiaoming" and the identification information of "Zhang San", tags for Video A are generated. For example, the tags of Video A include person name tags. Schematically, the tags of Video A are "Xiaoming" and "Zhang San". In addition, the videos in this tag system can include drama name tags, other tags (such as duration tags, video type tags, etc.).
[0048] After generating tags for various types of videos, it provides important information for downstream tasks such as recommendation distribution, review, and search. Schematically, when a user wants to search for videos starring a certain main actor object, they can quickly find the tags that match the input identification information of the main actor object from the tag system according to the identification information of the input main actor object, and obtain the videos corresponding to the found tags and return them to the user.
[0049] In summary, in the tag recall scenario, through the processing solution for multimedia data provided by the embodiments of the present application, not only can N image objects in the input data be screened to improve the accuracy of object screening, but also the problem that there are many and complex person names in a video, resulting in redundant information due to too many person name tags being marked, which affects downstream applications can be solved. That is, by screening the image objects, the purity of the tags can be improved, thereby improving the effect of downstream services and the tag recall rate.
[0050] It should be specifically noted that in the embodiments of the present application, relevant data in the process of processing multimedia data is involved, such as: the identification information of the target object (such as object name, nickname, region), etc. In the process of applying the solutions corresponding to the relevant embodiments of the present application, object permission or consent must be obtained, and the collection, use, and processing processes of relevant data must comply with relevant laws, regulations, and standards, and meet the principles of legality, propriety, and necessity, and do not involve obtaining data types prohibited or restricted by laws and regulations. In some optional embodiments, the relevant data involved in the embodiments of the present application is obtained after the object's separate authorization. In addition, when obtaining the object's separate authorization, the information such as the purpose of the relevant data involved is shown to the object, and relevant operations such as data acquisition and utilization will only be carried out after the object's authorization.
[0051] Next, the processing system for multimedia data provided by the embodiments of the present application will be elaborated.
[0052] Please refer to Figure 2 for the architecture diagram of a processing system for multimedia data provided by the embodiments of the present application. In Figure 2In this case, the multimedia data processing system includes at least one terminal device 101 and a server 102; the present application does not limit the number of terminal devices and servers. The terminal device 101 and the server 102 in the multimedia data processing system can be directly or indirectly connected through wired or wireless communication means. Among them:
[0053] The terminal device 101 is a device used by a business object (such as a user), and the terminal device 101 can be used to present input data. In some implementation manners, the terminal device 101 can also display an information generation interface, and the business object can import the input data to be processed in this information generation interface. In other implementation manners, the terminal device 101 can also display an object playback interface, and there is a playback object being played in this object playback interface, and the business object can use the playback object being played as the input data to be processed in this object playback interface. In addition, the terminal device 101 can also display the description information of the input data. The terminal device 101 can include but is not limited to: smart phones, tablet computers, smart wearable devices, intelligent voice interaction devices, smart home appliances, personal computers, vehicle-mounted terminals, smart cameras, virtual reality devices (such as AR (Augmented Reality) devices), etc., and the present application does not limit this.
[0054] The server 102 can store the input data. For example, the server 102 can store various types of videos and images. In addition, the server 102 can also screen N image objects in the input data to obtain M target objects, and generate the description information of the input data based on the M target objects. Among them, the server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0055] In one embodiment, taking the terminal device 101 and the server 102 as examples to describe the multimedia data processing flow, please refer to Figure 3 , which is a schematic diagram of a multimedia data processing flow provided by an embodiment of the present application. The multimedia data processing flow can be the following three parts: basic information extraction; information extraction from the basic information in the input data; object screening.
[0056] For basic information extraction, when it is necessary to generate description information of a certain input data, the terminal device 101 can obtain the input data to be processed from the server 102. The terminal device 101 can perform basic information extraction on the input data to be processed. The basic information extraction includes at least one of the following: extracting title information from the input data; obtaining a cover image related to the input data; extracting multiple images from the input data. Schematically, extracting multiple images from the input data may include: extracting multiple images from the input data at equal intervals, such as extracting one image from the input data every two images. For example Figure 3 In, the input data includes an image sequence composed of 8 images, and the serial numbers of these 8 images in the image sequence are 1-8 respectively. If one image is extracted from the input data every other image, then the images extracted from the input data are: the image with serial number 1 ( Figure 3 the image 1 in), the image with serial number 3 ( Figure 3 the image 3 in), the image with serial number 5 ( Figure 3 the image 5 in) and the image with serial number 7 ( Figure 3 the image 7 in).
[0057] For information extraction of the basic information in the input data, for the cover image in the basic information, image-to-token vector conversion can be performed. Specifically, the terminal device 101 can perform encoding processing on the cover image to obtain the token vector corresponding to the cover image. In one implementation, the terminal device 101 can use an image embedding model to perform encoding processing on the cover image to obtain the token vector corresponding to the cover image. Among them, the token vector corresponding to the cover image is a dense vector, and the token vector corresponding to the cover image is a vector with a fixed vector length.
[0058] For the cover image and the multiple images extracted, object recognition technology can be used to perform object recognition on the cover image to obtain the image objects in the cover image, and object recognition technology can be used to perform object recognition on the multiple images extracted to obtain the image objects in each image. Schematically, if object recognition includes face recognition, then face recognition technology can be used to perform face recognition on the cover image to obtain the faces in the cover image; face recognition technology can be used to perform face recognition on the multiple images extracted to obtain the faces in each image.
[0059] The terminal device 101 determines the serial numbers of the images where each image object is located in the image sequence and the number of images where each image object is located according to the image objects in the cover image and the image objects in each image, and generates the position information of each image object in the input data according to the serial numbers of the images where each image object is located in the image sequence and the number of images where each image object is located. Schematically, inFigure 3 Among them, according to the image objects in the cover image and the image objects in each image, it is determined that the images where the image object "Xiaoming" is located are the cover image, Image 1, and Image 3 respectively, the images where the image object "Zhang San" is located are Cover Image 1 and Image 5 respectively, and the image where Passerby A is located is Image 5.
[0060] The terminal device 101 can generate description information of N image objects in the input data according to the word segmentation vector corresponding to the cover image and the position information of each image in the input data.
[0061] In some implementation manners, the title information of the input data can also provide a basis for object screening. Therefore, the title information can be extracted from the input data, and the description information of N image objects in the input data can be generated according to the title information of the input data, the word segmentation vector corresponding to the cover image, and the position information of each image in the input data.
[0062] For object screening, after obtaining the description information of N image objects, M target objects can be obtained by adding appropriate screening indication information (prompt) to screen the N image objects. Specifically, the description information of the N image objects is filled into the question-and-answer prompt template to obtain the screening indication information, and then according to the screening indication information and the input data, M target objects are screened out from the N image objects. M is less than or equal to N. In some implementation manners, the N image objects can be screened by means of model question and answer. Specifically, the screening indication information and the input data can be input into the trained large language model to screen the N image objects to obtain M target objects. Through the screening indication information, it is possible to guide the trained large language model to screen out M target objects from the N image objects in the form of question and answer.
[0063] The terminal device 101 generates description information of the input data according to the M target objects and displays the description information of the input data. Specifically, the identification information of the M target objects can be obtained, and the identification information of the M target objects can be used as the description information of the input data. Schematically, the identification information includes the object name, object nickname, etc. For example, the object names of the M target objects can be used as the description information of the input data.
[0064] It should be noted that the above-mentioned interaction process for processing multimedia data is only for illustration, and does not limit the specific execution process of the terminal device and the server. In some implementation manners, the above-mentioned steps may be executed by the server 102, and the terminal device 101 may be used to display the description information of the input data. In another implementation manner, some of the above-mentioned steps may be executed by the server 102, and the terminal device 101 may obtain the screening result from the server 102, and generate the description information of the input data according to the M target objects in the screening result.
[0065] In summary, information extraction can be performed on the input data to obtain the feature information of the input data. According to the description information of the N image objects in the input data included in the feature information, screening indication information is generated. According to the screening indication information, the N image objects can be better screened, improving the screening accuracy of the image objects in the input data, and thus being able to efficiently generate description information for input data such as images or videos.
[0066] Next, a method for processing multimedia data provided in an embodiment of the present application will be described in detail.
[0067] Please refer to Figure 4 , which is a schematic flowchart of a method for processing multimedia data provided in an embodiment of the present application. The method for processing multimedia data can be executed by a computer device, which can be the above-mentioned terminal device 101 or the server 102, and the embodiment of the present application does not make any limitation thereto. The method for processing multimedia data provided in an embodiment of the present application may include the following steps S401-S404:
[0068] S401. Obtain input data to be processed, where the input data includes an image or an image sequence composed of multiple images. In an embodiment of the present application, the input data may include multimedia data. For example, the multimedia data may include an image or a video, and the image may be a static image, a dynamic image, etc. A video refers to an image sequence composed of multiple images. In an embodiment of the present application, the video may include a short video (such as a video with a playing duration not exceeding a preset duration) or a video with a playing duration exceeding the preset duration. The input data may include one or more image objects, and the image objects may include, but are not limited to: people, animals, virtual characters, plants, etc., and the embodiment of the present application does not make any limitation thereto.
[0069] In one implementation manner, obtaining the input data to be processed may include: the computer device displays an information generation interface, which includes: an import control for receiving an import operation of the input data, and a display area for displaying the description information of the input data. As Figure 5aAs shown in the figure, an embodiment of the present application provides a schematic diagram of an information generation interface. In Figure 5a the information generation interface 501 includes an import control 51 and a display area 52 for displaying description information for the input data. In response to an import operation on the import control 51 in the information generation interface 501, the object selected by the import operation can be used as the input data to be processed. Among them, the import operation can be an operation such as clicking or double-clicking on the import control. The object selected by the import operation can be one or more, so each object selected by the import operation can be used as the input data to be processed. By importing multiple input data at one time, it is possible to realize the subsequent batch generation of description information for the input data and improve the description information generation efficiency.
[0070] It should be understood that the import control 51 can be fixedly displayed in an area such as Figure 5a or in other areas of the information generation interface. Of course, the import control 51 can also be floatingly displayed in the information generation interface. The present application embodiment does not make any limitation on the display position of the import control 51.
[0071] In another implementation manner, obtaining the input data to be processed may include: the computer device displays an object playback interface, and the object playback interface includes: a playback object, a trigger control, a first display area for displaying description information for the input data, and a second display area for displaying M target objects. Schematically, as Figure 5b shown in the figure, an embodiment of the present application provides a schematic diagram of an object playback interface. In this object playback interface 502, there are a playback object 53, a trigger control 54, a first display area 55 for displaying description information for the input data, and a second display area 56 for displaying M target objects. The business object can click or double-click the trigger control. Correspondingly, the computer device can respond to the trigger operation on the trigger control 54 in the object playback interface, and the playback object can be used as the input data to be processed. The trigger operation here can include click operations, double-click operations, and so on.
[0072] It should be understood that the trigger control 54 can be fixedly displayed in an area such as Figure 5b or in other areas of the object playback interface. Of course, the trigger control 54 can be floatingly displayed in the object playback interface. The present application embodiment does not make any limitation on the display position of the trigger control 54. The first display area and the second display area can be the same display area or different display areas.
[0073] S402. Extract information from the input data to obtain the feature information of the input data. The feature information includes: description information of N image objects in the input data, where N is a positive integer. Among them, the description information may include at least one of the following: the token vector corresponding to the cover image, the position information of the image object, the title information, etc. Among them, the token vector corresponding to the cover image and the title information can provide information related to the image object. Schematically, for the title information "Xiaoming likes scenery", this title information can provide information related to the image object "Xiaoming". The cover image can refer to the image related to the input data before presenting the input data, or can refer to the special image generated related to the input data, or the cover image can also be any image in the image sequence. Schematically, the cover image can be the first image, the middle image, etc. in the image sequence.
[0074] The position information of the image object may include: the serial number of the image where the image object is located in the image sequence, etc. Schematically, the input data includes an image sequence composed of 3 images, and the image object A appears in the first image (i.e., the serial number in the image sequence is 1) and the second image (i.e., the serial number in the image sequence is 2). Correspondingly, the position information of the image object A in the input data includes: the serial numbers of the images where the image object A is located in the image sequence are 1 and 2 respectively. In addition, the description information may also include the number of images where the image object is located. For example, in the above example, the number of images where the image object A is located is 2.
[0075] It should be understood that when the description information of N image objects includes at least two of the token vector corresponding to the cover image, the position information of the image object, and the title information, this description information can also be called multimodal information. Using multimodal information can provide a judgment basis for screening N image objects in the input data.
[0076] In some implementation manners, the specific implementation manner of step S402 may include: obtaining the cover image related to the input data; performing encoding processing on the cover image to obtain the token vector corresponding to the cover image; performing object recognition on the input data to determine the image objects in the input data, and performing object recognition on the cover image to determine the image objects in the cover image; generating the description information of N image objects in the input data according to the token vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image.
[0077] Among them, encoding the cover image to obtain the token vector corresponding to the cover image may include: performing image segmentation on the cover image, encoding each image block determined by the segmentation and the position of the image block to obtain the feature vector of each image block, and performing length conversion processing on the feature vectors of each image block according to the correlation relationship between the feature vectors of each image block to obtain the token vector of the cover image. Among them, in the embodiment of the present application, an image embedding model may be used to encode the cover image, and the image embedding model includes a trained visual extraction model (such as a VIT model) and a trained feature processing model (such as a Q-former model). As Figure 6 , it is a schematic flowchart of a process for encoding a cover image provided by an embodiment of the present application. In Figure 6 , first, perform image segmentation on the cover image, input each image block determined by the segmentation and the position of the image block into the trained visual extraction model, encode each image block determined by the segmentation and the position of the image block through the trained visual extraction model to obtain the feature vector of each image block, and input the feature vectors of each image block into the trained feature processing model. Through the trained feature processing model, perform length conversion processing on the feature vectors of each image block according to the correlation relationship between the feature vectors of each image block to obtain the token vector of the cover image.
[0078] During the process of object recognition, the input data can be subjected to object recognition through an object recognition module. This object recognition mainly includes multiple processes such as object detection, object correction, object feature extraction, retrieval comparison, and policy card threshold. In some implementation manners, object recognition of the input data to determine the image objects in the input data may include the following steps: (1) Perform object detection on the input data to obtain the object detection result of the input data, and the object detection result includes the key point position information of each initial image object in the input data. Among them, the initial image objects may include people, animals, etc. Schematically, when the initial image object includes a person, the object detection is face detection, and the key points of the initial image object may be eyes, nose, etc. In one implementation, an object detection model can be used to perform object detection on the input data to obtain the object detection result of the input data. The object detection model may include but is not limited to retina-face, etc. (2) Based on the key point position information of each initial image object in the input data, correct the key points of each initial image object to obtain the correction result of each initial image object. The correction here includes: aligning the key points to a standard position, that is, through operations such as translation, rotation, and scaling, so that the position and size of the object features corresponding to the key points in the image object match the standard position. (3) Based on the correction results of each initial image object, perform feature extraction on each initial image object to obtain the feature information of each initial image object. In some implementation manners, a trained object feature extraction model can be used to perform feature extraction on each initial image object based on the correction results of each initial image object to obtain the feature information of each initial image object. Among them, the object feature extraction model may include but is not limited to: resnet34 (deep residual network), CNN, etc., and the embodiments of the present application do not limit this. (4) Retrieval comparison: Determine the similarity confidence between the candidate objects in the candidate set and each initial image object according to the feature information of each initial image object and the feature information of the candidate objects in the candidate set. Among them, the candidate set may be a set of candidate objects associated with the input data, that is, a candidate set related to the input data will be generated in advance, which can reduce the complexity of the retrieval comparison. (5) Determine the image object of the input data from the initial image objects according to the similarity confidence. Specifically, the initial image object corresponding to the similarity confidence greater than the confidence threshold can be determined as the image object of the input data. Determining the object of the input data from the initial image through the similarity confidence here can be considered as a screening of the finally recognized initial image object, thereby reducing the complexity of subsequent screening of N image objects.For example, if there is feature information of 100 initial image objects in the input data, the similarity confidence can be calculated based on the feature information of these 100 initial image objects and the feature information of the candidate objects in the candidate set. Then, perhaps only the similarity confidences corresponding to 80 initial image objects are greater than the confidence threshold. These 80 initial image objects are determined as the image objects in the input data. In this way, it is not necessary to identify all the initial image objects, and removing some irrelevant image objects can reduce the complexity of subsequent object screening. The confidence threshold can be set according to requirements, such as setting the confidence to 90, 95, etc. In some alternative implementation manners, the initial image object with the maximum similarity confidence and the candidate object are considered to be the same object to a certain extent. Then, the identification information of the corresponding image object can be determined according to the candidate object with the maximum similarity confidence. For example, the identification information of the candidate object with the maximum similarity confidence is determined as the identification information of the corresponding image object.
[0079] In some implementation manners, when the input data includes an image sequence composed of multiple images, at this time, multiple images of the input data can be extracted at equal intervals, and object recognition is performed on the extracted multiple images to obtain the image objects in these multiple images, and these multiple images are used as the image objects in the input data. Schematically, 60 images, 100 images in the input data can be extracted at equal intervals. The embodiments of the present application do not make any limitation in this regard.
[0080] It should be understood that for the specific implementation manner of performing object recognition on the cover image and determining the image objects in the cover image, reference can be made to the above specific implementation manner of performing object recognition on the input data and determining the image objects in the input data, which will not be elaborated here.
[0081] In some implementations, the input data is an image sequence composed of multiple images. According to the word segmentation vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image, the description information of N image objects in the input data may include: determining the position information of each image object in the input data according to the image objects in the input data and the image objects in the cover image. The position information includes the serial number of the image where the corresponding image object is located in the image sequence. In addition, if the image object also appears in the cover image, the position information of the image object further includes the cover image. Schematically, if image object A appears in image 1 (i.e., the serial number in the image sequence is 1) and image 4 (i.e., the serial number in the image sequence is 4), then the position information of image object A includes: the serial numbers 1 and 4 of the images where image object A is located in the image sequence; generating the description information of N image objects in the input data according to the word segmentation vector corresponding to the cover image and the position information of each image object in the input data. Among them, the N image objects may include N image objects in the input data, or the N image objects may include the image objects in the input data and the image objects in the cover image, which is not limited in the embodiments of the present application. In some optional implementations, the input data may further include title information. The title information may be extracted from the input data, and then the description information of N image objects in the input data is generated according to the title information, the word segmentation vector corresponding to the cover image, and the position information of each image object in the input data.
[0082] S403. Fill the description information of N image objects into the Q&A prompt template to obtain the screening indication information.
[0083] Among them, the Q&A prompt template can be set according to requirements. In some implementations, obtain the Q&A prompt template. The Q&A prompt template includes the input positions of the description information of N image objects. The input positions include the first position and the second position. Different input positions in the Q&A prompt template correspond to different prompt information. Schematically, as Figure 7 shown, it is a schematic flowchart of a process for filling the Q&A prompt template provided by the embodiments of the present application. In Figure 7In this case, different input positions in the Q&A prompt template 71 correspond to different prompt messages, that is, the Q&A prompt template 71 includes a first position 72 and a second position 73; in the first position 72, there is a corresponding prompt message "token vector of the cover image", which is used to prompt filling the token vector of the cover image into the first position 72; in the second position 73, there is a corresponding prompt message "the recognized image objects in the input data are", which is used to indicate filling the position information of N image objects into the second position 73. Optionally, the input position includes a third position 74, and the third position 74 corresponds to a prompt message "Please find the target object among the above image objects". Then, according to the prompt message corresponding to the first position, the token vector corresponding to the cover image can be filled into the first position, and according to the prompt message corresponding to the second position, the position information of each image object in the input data can be filled into the second position to obtain the screening indication information. Schematically, in Figure 7 In this case, the N image objects include A, B, and C. The description information of the N image objects includes: the token vector corresponding to the cover image is the cover image token, the position information of A, the position information of B, and the position information of C. Filling the description information of the N image objects into the Q&A prompt template, the obtained screening indication information can be: "You are a video editor. Here is a video. The description information of the N image objects includes: [token vector of the cover image]. The recognized image objects in the input data are: A, which appears in the cover image, the first image (i.e., image 1), the second image (i.e., image 2), and the fourth image (image 4); B, which appears in the cover image, the first image, the third image (image 3), and the fourth image; C, which appears in the fifth image (i.e., image 5), the sixth image (image 6), the seventh image (i.e., image 7), and the eighth image (image 8). Please find the target object among the above image objects that appear".
[0084] In some implementation manners, the description information of the N image objects further includes title information, and the input position further includes a fourth position. According to the prompt message corresponding to the first position, the token vector corresponding to the cover image is filled into the first position, and according to the prompt message corresponding to the second position, the position information of each image object in the input data is filled into the second position. The obtained screening indication information may further include: according to the prompt message corresponding to the first position, the token vector corresponding to the cover image is filled into the first position, according to the prompt message corresponding to the second position, the position information of each image object in the input data is filled into the second position, and according to the prompt message corresponding to the fourth position, the title information is filled into the fourth position to obtain the screening indication information.
[0085] S404. According to the screening indication information and the input data, screen out M target objects from N image objects, and generate description information of the input data based on the M target objects, where M is a positive integer less than or equal to N.
[0086] It should be understood that N target objects can be screened out from the N image objects, that is to say, all N image objects are target objects. Of course, some target objects can be screened out from the N image objects. Schematically, N = 6. When screening these 6 image objects, it is possible to screen out 6 target objects, or 5 target objects, or 3 target objects.
[0087] In one implementation, according to the screening indication information and the input data, the correlation degree between each image object in the N image objects and the input data can be determined, and M target objects are screened out from the N image objects according to the correlation degree between each image object and the input data. Schematically, the screening indication information includes the word segmentation vector of the cover image, which indicates the information of image object 1, and the screening indication information includes the position information of image object 1 in the input data, such as image object 1 appears in the image cover and multiple images in the image sequence. Then it can be considered that image object 1 has a high correlation with the input data, and the correlation degree between image object 1 and the input data can be set to a relatively large value. In one implementation, screening out M target objects from the N image objects according to the correlation degree between each image object and the input data may specifically include: taking the image object corresponding to the correlation degree greater than the correlation degree threshold as the target object, and the correlation degree threshold can be set according to requirements, and the embodiments of the present application do not make any limitation on this.
[0088] In another implementation, a large prediction model can be pre-trained, and then the trained large prediction model is called to screen out M target objects from the N image objects according to the screening indication information and the input data. This part of the method will be described in detail in the subsequent embodiments.
[0089] Among them, generating description information of the input data based on the M target objects may include: obtaining the identification information of the M target objects, and generating description information of the input data based on the identification information of the M target objects. Among them, the description information may include tags, and the identification information of the M target objects is used as the tags of the input data. Schematically, the M target objects include image object C and image object D, and the object name "Zhang San" of image object C and the object name "Li Si" of image object D can be obtained, and "Zhang San" and "Li Si" are used as the tags of the input data. In some implementations, the description information may include tags and titles. The identification information of the M target objects can be used as the tags of the input data, and the tags and titles of the input data are used to generate the description information of the input data.
[0090] After generating the description information of the input data, the description information of the input data can be displayed. In one implementation, as Figure 5a described in [reference], if the input data to be processed is the object selected for the import operation, then after generating the description information of the input data, the description information of the input data can be displayed in the display area 52 of the information generation interface. It should be understood that when the number of input data to be processed is multiple, the description information of multiple input data can be displayed in the display area 52 of the information generation interface.
[0091] In another implementation, as Figure 5b described in [reference], if the input data to be processed is a playback object, then after generating the description information of the input data, the description information of the input data is displayed in the first display area 55 of the object playback interface. Taking the input data as a video as an example, correspondingly, the object playback interface includes a video playback interface. The significance of displaying the description information of the input data in a certain display area of the object playback interface lies in that when a user watches a video, since the video is very long and it is not known whether it is good-looking, through the method provided by the embodiments of the present application, the target object (such as the leading actor object) of the input data can be quickly known. In addition, in addition to displaying the description information of the input data, the images where M target objects are located are displayed in the second display area 56. In one implementation, the images where M target objects are located can be images containing these M target objects. In another implementation, the images where M target objects are located can be images containing the object features of the complete target object (such as an image including the complete facial features of the leading actor object); in still another implementation, the images where M target objects are located can be images containing M target objects and capable of reflecting the characteristics of the input data. For example, if the input data includes a video and the video belongs to the type of funny videos, then the images where M target objects are located can be images containing M images and having funny characteristics. By presenting the description information of the input data and presenting the images where M target objects are located, the target object in the input data (such as knowing who the leading actor in the video is) can be quickly and intuitively understood, which can help the user understand the input data more quickly.
[0092] In an embodiment of the present application, input data to be processed is obtained. The input data includes an image or an image sequence composed of multiple images; information extraction is performed on the input data to obtain feature information of the input data. The feature information includes: description information of N image objects in the input data, where N is a positive integer; the description information of the N image objects is filled into a question-and-answer prompt template to obtain screening indication information; according to the screening indication information and the input data, M target objects are screened out from the N image objects. It can be seen that by filling the description information of the N image objects in the input data into the question-and-answer prompt template, object screening of the N image objects can be guided in the way of screening indication information, thereby effectively improving the screening accuracy of image objects. Further, description information of the input data is generated according to the M target objects, where M is a positive integer less than or equal to N. According to the target objects, the description information of the input data can be automatically generated without manual editing of the description information, and the description information of input data such as pictures or videos can be efficiently generated.
[0093] Please refer to Figure 8 , which is a schematic flowchart of another method for processing multimedia data provided by an embodiment of the present application. This method for processing multimedia data can be executed by a computer device, and the computer device can be the above-mentioned terminal device or the above-mentioned server. This method for processing multimedia data can include the following steps S801 - S804:
[0094] S801. Obtain input data to be processed, where the input data includes an image or an image sequence composed of multiple images.
[0095] S802. Perform information extraction on the input data to obtain feature information of the input data. The feature information includes: description information of N image objects in the input data, where N is a positive integer.
[0096] S803. Fill the description information of the N image objects into a question-and-answer prompt template to obtain screening indication information.
[0097] S804. Call the trained large language model, screen out M target objects from the N image objects according to the screening indication information and the input data, and generate description information of the input data according to the M target objects, where M is a positive integer less than or equal to N.
[0098] Among them, the trained large language model can include but is not limited to: LLaMA, ChatGLM, bloomz-3b, etc., and the embodiments of the present application do not make any limitations in this regard. Schematically, as Figure 9 shown, it is a flowchart of object screening provided by an embodiment of the present application. In Figure 9In [the above scenario], the screening indication information and input data can be input into the trained large prediction model, and the trained large language model can screen out M target objects from N image objects according to the screening indication information. For example, Figure 9 In [the above scenario], the finally screened M target objects include A and B.
[0099] In one implementation, since the direct use of open-source large language models has poor effects, the embodiments of this application can train large language models through a sample set to obtain a trained large prediction model with better object screening capabilities. Among them, the sample set can include multiple sample data, and the number of the sample data can be 100,000, 1,000, etc., which is not limited in the embodiments of this application. The sample data can include images or image sequences composed of multiple images. The training process of the large language model includes the following steps: obtaining a sample set, which includes multiple sample data and supervision labels configured for each sample data; extracting information from each sample data to obtain sample feature information of each sample data, and the sample feature information includes description information of N sample image objects in the corresponding sample data, where N is a positive integer; filling the description information of the N sample image objects in each sample data into a question-and-answer prompt template to obtain sample screening indication information corresponding to each sample data; calling the large language model, and screening out M target sample objects corresponding to each sample data from the N sample image objects included in each sample data according to the screening indication information corresponding to each sample data and each sample data; adjusting the parameters of the large language model according to the M target sample objects corresponding to each sample data and the configured supervision labels to obtain a trained large language model.
[0100] Among them, adjusting the parameters of the large language model according to the M target sample objects corresponding to each sample data and the configured supervision labels to obtain a trained large language model can include: adopting a cross-entropy loss function, determining the model loss of the large language model according to the M target sample objects corresponding to each sample data and the configured supervision labels, and then adjusting the parameters of the large language model in the direction of reducing the model loss of the large language model to obtain a trained large language model.
[0101] In the embodiments of the present application, input data to be processed is obtained. The input data includes an image or an image sequence composed of multiple images; information extraction is performed on the input data to obtain the feature information of the input data. The feature information includes: description information of N image objects in the input data, where N is a positive integer; the description information of the N image objects is filled into a question-and-answer prompt template to obtain screening indication information; the trained large language model is called, and according to the screening indication information and the input data, M target objects are screened out from the N image objects. It can be seen that the description information of the N image objects is filled into the question-and-answer prompt template, and the screening indication information is input into the trained large language model, so that the trained large language model can be guided in a question-and-answer manner to screen out M target objects from the N image objects, improving the screening accuracy of the image objects. Further, description information of the input data is generated according to the M target objects, where M is a positive integer less than or equal to N, and there is no need to manually edit the description information, and the description information of input data such as pictures or videos can be efficiently generated.
[0102] Next, the processing device for multimedia data provided in the embodiments of the present application will be described in detail.
[0103] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a processing device for multimedia data provided in the embodiments of the present application. The processing device for multimedia data may be a computer program (including program code) in a computer device. For example, the processing device for multimedia data may be an application software in a computer device; the processing device for multimedia data can be used to execute Figure 4 and Figure 8 part or all of the steps in the method embodiments shown. Please refer to Figure 10 , the processing device for multimedia data includes the following units:
[0104] An obtaining unit 1001, configured to obtain input data to be processed, where the input data includes an image or an image sequence composed of multiple images;
[0105] A processing unit 1002, configured to perform information extraction on the input data to obtain the feature information of the input data. The feature information includes: description information of N image objects in the input data, where N is a positive integer;
[0106] The processing unit 1002 is further configured to fill the description information of the N image objects into a question-and-answer prompt template to obtain screening indication information;
[0107] The processing unit 1002 is further configured to screen out M target objects from N image objects according to the screening indication information and the input data, and generate description information of the input data based on the M target objects, where M is a positive integer less than or equal to N.
[0108] In some implementation manners, the processing unit 1002 is specifically configured to:
[0109] Obtain a cover image related to the input data;
[0110] Perform encoding processing on the cover image to obtain a word segmentation vector corresponding to the cover image;
[0111] Perform object recognition on the input data to determine the image objects in the input data, and perform object recognition on the cover image to determine the image objects in the cover image;
[0112] Generate description information of the N image objects in the input data based on the word segmentation vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image.
[0113] In some implementation manners, the processing unit 1002 is specifically configured to:
[0114] Perform image segmentation on the cover image, and encode each image block determined by the segmentation and the position of the image block to obtain a feature vector of each image block;
[0115] Perform length conversion processing on the feature vectors of the image blocks according to the association relationship between the feature vectors of the image blocks to obtain a word segmentation vector of the cover image.
[0116] In some implementations, the input data is an image sequence composed of multiple images, and the processing unit 1002 is specifically configured to:
[0117] Determine the position information of each image object in the input data according to the image objects in the input data and the image objects in the cover image, where the position information includes the serial number of the image in the image sequence where the corresponding image object is located;
[0118] Generate description information of the N image objects in the input data based on the word segmentation vector corresponding to the cover image and the position information of each image object in the input data.
[0119] In some implementation manners, the processing unit 1002 is specifically configured to:
[0120] Perform object detection on the input data to obtain an object detection result of the input data, where the object detection result includes the key point position information of each initial image object in the input data;
[0121] Based on the key point position information of each initial image object in the input data, correct the key points of each initial image object to obtain the correction results of each initial image object;
[0122] Based on the correction results of each initial image object, extract features from each initial image object to obtain the feature information of each initial image object;
[0123] According to the feature information of each initial image object and the feature information of the candidate objects in the candidate set, determine the similarity confidence levels between the candidate objects in the candidate set and each initial image object;
[0124] Determine the image object of the input data from the initial image objects according to the similarity confidence levels, and determine the identification information of the corresponding image object according to the candidate object with the highest similarity confidence level.
[0125] Among them, the processing unit 1002 is specifically configured to:
[0126] Obtain a Q&A prompt template, where the Q&A prompt template includes input positions of description information of N image objects; the input positions include a first position and a second position; different input positions in the Q&A prompt template correspond to different prompt information;
[0127] According to the prompt information corresponding to the first position, input the word segmentation vector corresponding to the cover image into the first position, and input the position information of each image object in the input data into the second position according to the prompt information corresponding to the second position to obtain screening indication information.
[0128] Among them, the obtaining unit 1001 is specifically configured to:
[0129] Display an information generation interface; the information generation interface includes: an import control for receiving an import operation for the input data, and a display area for displaying the description information of the input data;
[0130] In response to an import operation of the import control in the information generation interface, use the object selected by the import operation as the input data to be processed;
[0131] After obtaining the description information of the input data, display the description information of the input data in the display area of the information generation interface.
[0132] In some implementation manners, the obtaining unit 1001 is specifically configured to:
[0133] Display an object playback interface, where the object playback interface includes: a playback object, a trigger control, a first display area for displaying the description information of the input data, and a second display area for displaying M target objects;
[0134] In response to a triggering operation of a control in the object playback interface, the playback object is used as input data to be processed;
[0135] After obtaining the description information of the input data, the description information of the input data is displayed in the first display area of the object playback interface, and the images where the M target objects are located are displayed in the second display area.
[0136] In some implementation manners, the M target objects are screened from N image objects by calling a trained large language model according to the screening indication information and the input data. The obtaining unit 1001 is further configured to obtain a sample set, where the sample set includes multiple sample data and supervision labels configured for each sample data;
[0137] The processing unit 1002 is further configured to perform information extraction on each sample data to obtain sample feature information of each sample data, where the sample feature information includes description information of N sample image objects in the corresponding sample data, and N is a positive integer;
[0138] The processing unit 1002 is further configured to fill the description information of the N sample image objects in each sample data into a question-and-answer prompt template to obtain sample screening indication information corresponding to each sample data;
[0139] The processing unit 1002 is further configured to call the large language model, and screen out M target sample objects corresponding to each sample data from the N sample image objects included in each sample data according to the screening indication information corresponding to each sample data and each sample data;
[0140] The processing unit 1002 is further configured to adjust the parameters of the large language model according to the M target sample objects corresponding to each sample data and the configured supervision labels to obtain a trained large language model.
[0141] In an embodiment of the present application, input data to be processed is obtained, where the input data includes an image or an image sequence composed of multiple images; information extraction is performed on the input data to obtain feature information of the input data, and the feature information includes: description information of N image objects in the input data, where N is a positive integer; the description information of the N image objects is filled into a question-and-answer prompt template to obtain screening indication information; a trained large language model is called, and according to the screening indication information and the input data, M target objects are screened out from the N image objects. It can be seen that by filling the description information of the N image objects into the question-and-answer prompt template and inputting the screening indication information into the trained large language model, the trained large language model can be guided in a question-and-answer manner to screen out M target objects from the N image objects, improving the screening accuracy of the image objects. Further, description information of the input data is generated according to the M target objects, where M is a positive integer less than or equal to N, and there is no need for manual editing of the description information, and description information can be efficiently generated for input data such as pictures or videos.
[0142] Next, the computer device provided by the embodiment of the present application will be described in detail.
[0143] Further, an embodiment of the present application also provides a schematic structural diagram of a computer device, and the schematic structural diagram of the computer device can be referred to Figure 11 ; the computer device may include: a processor 1101, an input device 1102, an output device 1103, and a memory 1104. The above-mentioned processor 1101, input device 1102, output device 1103, and memory 1104 are connected through a bus. The memory 1104 is used to store a computer program, and the computer program includes program instructions, and the processor 1101 is used to execute the program instructions stored in the memory 1104.
[0144] In one embodiment, the processor 1101 performs the following operations by running the program instructions in the memory 1104:
[0145] Obtain input data to be processed, where the input data includes an image or an image sequence composed of multiple images;
[0146] Perform information extraction on the input data to obtain feature information of the input data, and the feature information includes: description information of N image objects in the input data, where N is a positive integer;
[0147] Fill the description information of the N image objects into a question-and-answer prompt template to obtain screening indication information;
[0148] According to the screening indication information and the input data, screen out M target objects from the N image objects, and generate description information of the input data according to the M target objects, where M is a positive integer less than or equal to N.
[0149] In some implementations, when the processor 1101 extracts information from the input data to obtain the feature information of the input data, the following operations may be specifically performed:
[0150] Obtain a cover image related to the input data;
[0151] Perform encoding processing on the cover image to obtain a word segmentation vector corresponding to the cover image;
[0152] Perform object recognition on the input data to determine the image objects in the input data, and perform object recognition on the cover image to determine the image objects in the cover image;
[0153] Generate description information of N image objects in the input data according to the word segmentation vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image.
[0154] In some implementations, when the processor 1101 performs encoding processing on the cover image to obtain a word segmentation vector corresponding to the cover image, the following operations may be specifically performed:
[0155] Perform image segmentation on the cover image, and encode each image block determined by the segmentation and the position of the image block to obtain the feature vector of each image block;
[0156] Perform length conversion processing on the feature vectors of each image block according to the correlation relationship between the feature vectors of each image block to obtain the word segmentation vector of the cover image.
[0157] In some implementations, the input data is an image sequence composed of multiple images. When the processor 1101 generates description information of N image objects in the input data according to the word segmentation vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image, the following operations may be specifically performed:
[0158] Determine the position information of each image object in the input data according to the image objects in the input data and the image objects in the cover image, where the position information includes the serial number of the image in the image sequence where the corresponding image object is located;
[0159] Generate description information of N image objects in the input data according to the word segmentation vector corresponding to the cover image and the position information of each image object in the input data.
[0160] In some implementations, when the processor 1101 performs object recognition on the input data to determine the image objects in the input data, the following operations may be specifically performed:
[0161] Perform object detection on the input data to obtain the object detection result of the input data, where the object detection result includes the key point position information of each initial image object in the input data;
[0162] Based on the key point position information of each initial image object in the input data, correct the key points of each initial image object to obtain the correction result of each initial image object;
[0163] Based on the correction result of each initial image object, extract features from each initial image object to obtain the feature information of each initial image object;
[0164] According to the feature information of each initial image object and the feature information of the candidate objects in the candidate set, determine the similarity confidence between the candidate objects in the candidate set and each initial image object;
[0165] Determine the image object of the input data from the initial image objects according to the similarity confidence, and determine the identification information of the corresponding image object according to the candidate object with the highest similarity confidence.
[0166] In some implementation manners, when the processor 1101 fills the description information of N image objects into the Q&A prompt template to obtain the screening indication information, the following operations may be specifically performed:
[0167] Obtain the Q&A prompt template, where the Q&A prompt template includes the input positions of the description information of N image objects; the input positions include a first position and a second position; different input positions in the Q&A prompt template correspond to different prompt information;
[0168] According to the prompt information corresponding to the first position, input the token vector corresponding to the cover image to the first position, and according to the prompt information corresponding to the second position, input the position information of each image object in the input data to the second position to obtain the screening indication information.
[0169] In some implementation manners, when the processor 1101 obtains the input data to be processed, the following operations may be specifically performed:
[0170] Display an information generation interface; the information generation interface includes: an import control for receiving an import operation for the input data, and a display area for displaying the description information of the input data;
[0171] In response to the import operation of the import control in the information generation interface, use the object selected by the import operation as the input data to be processed;
[0172] The method further includes: after obtaining the description information of the input data, display the description information of the input data in the display area of the information generation interface.
[0173] In some implementations, when the processor 1101 obtains the input data to be processed, it can specifically perform the following operations:
[0174] Display an object playback interface, which includes: a playback object, a trigger control, a first display area for displaying description information about the input data, and a second display area for displaying M target objects;
[0175] In response to a trigger operation on the trigger control in the object playback interface, use the playback object as the input data to be processed;
[0176] The method further includes: after obtaining the description information of the input data, display the description information of the input data in the first display area of the object playback interface, and display the images where the M target objects are located in the second display area.
[0177] In some implementations, the M target objects are filtered from N image objects according to the filtering indication information and the input data by calling a trained large language model. The processor 1101 can also perform the following operations:
[0178] Obtain a sample set, which includes multiple sample data and supervision labels configured for each sample data;
[0179] Extract information from each sample data to obtain the sample feature information of each sample data. The sample feature information includes the description information of N sample image objects in the corresponding sample data, where N is a positive integer;
[0180] Fill the description information of the N sample image objects in each sample data into a question-and-answer prompt template to obtain the sample filtering indication information corresponding to each sample data;
[0181] Call the large language model, and filter out the M target sample objects corresponding to each sample data from the N sample image objects included in each sample data according to the filtering indication information and each sample data corresponding to each sample data;
[0182] According to the M target sample objects corresponding to each sample data and the configured supervision labels, adjust the parameters of the large language model to obtain a trained large language model.
[0183] In an embodiment of the present application, input data to be processed is obtained, and the input data includes an image or an image sequence composed of multiple images; information extraction is performed on the input data to obtain feature information of the input data, and the feature information includes: description information of N image objects in the input data, where N is a positive integer; the description information of the N image objects is filled into a question-and-answer prompt template to obtain screening indication information; the trained large language model is called, and M target objects are screened out from the N image objects according to the screening indication information and the input data. It can be seen that the description information of the N image objects is filled into the question-and-answer prompt template, and the screening indication information is input into the trained large language model, so that the trained large language model can be guided in a question-and-answer manner to screen out M target objects from the N image objects, improving the screening accuracy of the image objects. Further, description information of the input data is generated according to the M target objects, where M is a positive integer less than or equal to N, and there is no need to manually edit the description information, and description information can be efficiently generated for input data such as pictures or videos.
[0184] In addition, it should be noted here that: The embodiment of the present application also provides a computer-readable storage medium, and a computer program is stored in the computer-readable storage medium, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the methods in the corresponding embodiments described above Figure 4 and Figure 8 Therefore, details will not be described here again. For the technical details not disclosed in the embodiment of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiment of the present application. As an example, the program instructions can be deployed on one computer device, or executed on multiple computer devices located in one place, or, executed on multiple computer devices distributed in multiple places and interconnected through a communication network.
[0185] According to one aspect of the present application, a computer program product is provided, and the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device can execute the methods in the corresponding embodiments described above Figure 4 and Figure 8 Therefore, details will not be described here again.
[0186] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above various methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0187] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for processing multimedia data, characterized in that, Including: Obtain input data to be processed, where the input data includes an image or an image sequence composed of multiple images; Extract information from the input data to obtain feature information of the input data, where the feature information includes: description information of N image objects in the input data, and N is a positive integer; Fill the description information of the N image objects into a question-and-answer prompt template to obtain screening indication information; According to the screening indication information and the input data, screen out M target objects from the N image objects, and generate description information of the input data based on the M target objects, where M is a positive integer less than or equal to N.
2. The method according to claim 1, wherein The extracting information from the input data to obtain the feature information of the input data includes: Obtain a cover image related to the input data; Perform encoding processing on the cover image to obtain a word segmentation vector corresponding to the cover image; Perform object recognition on the input data to determine image objects in the input data, and perform object recognition on the cover image to determine image objects in the cover image; Generate description information of N image objects in the input data based on the word segmentation vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image.
3. The method according to claim 2, characterized in that The performing encoding processing on the cover image to obtain a word segmentation vector corresponding to the cover image includes: Perform image segmentation on the cover image, and encode each image block determined by the segmentation and the position of the image block to obtain feature vectors of the image blocks; Perform length conversion processing on the feature vectors of the image blocks according to the correlation relationship between the feature vectors of the image blocks to obtain the word segmentation vector of the cover image.
4. The method according to claim 2, wherein When the input data is an image sequence composed of multiple images, the generating description information of N image objects in the input data based on the word segmentation vector corresponding to the cover image, the image objects in the input data, and the image objects in the cover image includes: Determine the position information of each image object in the input data according to the image objects in the input data and the image objects in the cover image, where the position information includes the serial number of the image where the corresponding image object is located in the image sequence; Generate description information of N image objects in the input data based on the word segmentation vector corresponding to the cover image and the position information of each image object in the input data.
5. The method according to claim 2, characterized in that, The performing object recognition on the input data to determine image objects in the input data includes: Perform object detection on the input data to obtain an object detection result of the input data, where the object detection result includes key point position information of each initial image object in the input data; Based on the key point position information of each initial image object in the input data, correct the key points of each initial image object to obtain a correction result of each initial image object. Based on the correction results of the respective initial image objects, feature extraction is performed on the respective initial image objects to obtain the feature information of the respective initial image objects; According to the feature information of the respective initial image objects and the feature information of the candidate objects in the candidate set, determine the similarity confidence between the candidate objects in the candidate set and the respective initial image objects; Determine the image object of the input data from the initial image objects according to the similarity confidence, and determine the identification information of the corresponding image object according to the candidate object with the highest similarity confidence.
6. The method according to claim 4, characterized in that The filling the description information of the N image objects into the Q&A prompt template to obtain the screening indication information includes: Obtain a Q&A prompt template, where the Q&A prompt template includes input positions for the description information of the N image objects; the input positions include a first position and a second position; different input positions in the Q&A prompt template correspond to different prompt information; According to the prompt information corresponding to the first position, input the token vector corresponding to the cover image to the first position, and according to the prompt information corresponding to the second position, input the position information of the respective image objects in the input data to the second position to obtain the screening indication information.
7. The method according to claim 1, characterized in that, The obtaining the input data to be processed includes: Display an information generation interface; the information generation interface includes: an import control for receiving an import operation for the input data, and a display area for displaying the description information of the input data; In response to an import operation of the import control in the information generation interface, use the object selected by the import operation as the input data to be processed; The method further includes: after obtaining the description information of the input data, display the description information of the input data in the display area of the information generation interface.
8. The method according to claim 1, wherein The obtaining the input data to be processed includes: Display an object playback interface, where the object playback interface includes: a playback object, a trigger control, a first display area for displaying the description information of the input data, and a second display area for displaying M target objects; In response to a trigger operation of the trigger control in the object playback interface, use the playback object as the input data to be processed; The method further includes: after obtaining the description information of the input data, display the description information of the input data in the first display area of the object playback interface, and display the images where the M target objects are located in the second display area.
9. The method according to claim 1, characterized in that, The M target objects are screened from the N image objects according to the screening indication information and the input data by invoking a trained large language model, and the method further includes: Obtain a sample set, where the sample set includes multiple sample data and a supervision label configured for each sample data; Perform information extraction on each sample data to obtain the sample feature information of each sample data, where the sample feature information includes the description information of N sample image objects in the corresponding sample data, and N is a positive integer; Fill the description information of N sample image objects in each sample data into the Q&A prompt template to obtain the sample screening indication information corresponding to each sample data; Call a large language model, and screen out M target sample objects corresponding to each sample data from the N sample image objects included in each sample data according to the screening indication information corresponding to each sample data and each sample data; Adjust the parameters of the large language model according to the M target sample objects corresponding to each sample data and the configured supervision labels to obtain the trained large language model.
10. A processing device for multimedia data, characterized in that, Comprising: An acquisition unit for acquiring input data to be processed, where the input data includes an image or an image sequence composed of multiple images; A processing unit for extracting information from the input data to obtain feature information of the input data, where the feature information includes: description information of N image objects in the input data, and N is a positive integer; The processing unit is further configured to fill the description information of the N image objects into the Q&A prompt template to obtain screening indication information; The processing unit is further configured to screen out M target objects from the N image objects according to the screening indication information and the input data, and generate description information of the input data according to the M target objects, where M is a positive integer less than or equal to N.
11. A computer device, characterized in that, Comprising: A processor suitable for executing a computer program; A computer-readable storage medium storing a computer program, and when the computer program is executed by the processor, it executes the processing method for multimedia data according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer storage medium stores a computer program, and when the computer program is executed by the processor, it executes the processing method for multimedia data according to any one of claims 1-9.
13. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the processing method for multimedia data according to any one of claims 1-9.