Content description method and device, electronic equipment, storage medium and computer program product
Through the content description model of deep learning technology, combined with computer vision and natural language processing, the problem of low accuracy in video content description is solved, and the understanding of the relationship between video picture elements and accurate description of natural language forms is achieved.
Patent Information
- Application Number
- CN202510356475.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the accuracy of video content description is not high, and the tag summary lacks understanding of the relationship between picture elements, which leads to users need to speculate on the event content by themselves, which increases the understanding cost.
By obtaining the target picture and its corresponding prompt words, the target task is determined as a content recognition task, and the content description model of deep learning technology is used for content recognition, and the description information in natural language form is output, and cross-modal mapping is achieved by combining the computer vision module and the natural language processing module.
It improves the accuracy of video content description, can better understand the relationship between target content in different target pictures, and outputs more accurate and intuitive description information.
Smart Images

Figure CN120339903A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology. Specifically, this application relates to a content description method, apparatus, electronic device, storage medium, and computer program product. Background Art
[0002] With the development of deep learning technology, models such as convolutional neural networks (CNNs) have been widely applied in the field of video analysis. By detecting and recognizing objects in video frames, video summaries are generated.
[0003] Specifically, the CNN model can detect elements in the video frame and output the elements in the form of labels. However, this label-based video summary lacks an understanding of the relationships between the elements in the frame and is difficult to express the actual content information in the video. This results in users having to speculate on the event content themselves when viewing the summary, increasing the understanding cost and reducing the intuitiveness and practical application effect of the video summary.
[0004] As can be seen from the above, how to improve the accuracy of video description content description remains to be solved. Summary of the Invention
[0005] This application provides a content description method, apparatus, electronic device, and storage medium, which can solve the problem of low accuracy in video content description in the related art. The technical solutions are as follows:
[0006] According to one aspect of this application, a content description method includes: obtaining a target picture; determining a target task based on a prompt word corresponding to the target picture, where the prompt word is a word input by a user and used to indicate the task that the user expects to complete; if the target task is a content recognition task, performing content recognition on the target content in the target picture to obtain description information of the target picture, where the description information is information in the form of natural language used to describe the target content in the target picture.
[0007] According to one aspect of this application, a content description apparatus includes: a picture obtaining module for obtaining a target picture; a task determining module for determining a target task based on a prompt word corresponding to the target picture, where the prompt word is a word input by a user and used to indicate the task that the user expects to complete; a content recognition module for, if the target task is a content recognition task, performing content recognition on the target content in the target picture to obtain description information of the target picture, where the description information is information in the form of natural language used to describe the target content in the target picture.
[0008] In an exemplary embodiment, the content description device is further configured to obtain search indication information; the search indication information is information in the form of natural language used to describe the target content that the user expects to find; based on the search indication information, search for the screen where the target content exists to obtain the target screen.
[0009] In an exemplary embodiment, the content description device is further configured to obtain a target database, and the target database further includes multiple screens and their corresponding description vectors; perform a text vector conversion process on the search indication information to obtain a search indication vector; based on the search indication vector, perform data search in the target database. If there is a target description vector whose similarity with the search indication vector meets the set conditions, then use the screen corresponding to the target description vector as the target screen.
[0010] In an exemplary embodiment, the content description device is further configured to perform a text vector conversion process on the description information corresponding to the target screen to obtain a corresponding target description vector; store the target screen and its corresponding target description vector to generate the target database.
[0011] In an exemplary embodiment, the content description device is further configured to implement the content description method through a content description model; the content description model is a natural language model that has been trained and has the ability to identify the content of a screen.
[0012] In an exemplary embodiment, the content description device is further configured to obtain sample data; the sample data includes sample screens and their corresponding sample description information; the sample description information is information in the form of natural language used to describe the content in the sample screen; input the sample screen into an initial content description model for content recognition to obtain training description information; calculate a loss based on the sample description information and the training description information to obtain loss information; use the loss information to optimize the content description model and continue training until a trained content description model is obtained.
[0013] In an exemplary embodiment, the sample data further includes sample prompt words corresponding to the sample screen; the content description device is further configured to determine a target task based on the sample prompt words corresponding to the sample screen; in the case where the target task is content recognition, use the content description model to perform content recognition on the sample screen to obtain the training description information.
[0014] In an exemplary embodiment, the content description device is further configured to obtain a target video; use a set method to perform frame extraction on the target video to obtain the target screen.
[0015] According to one aspect of the present application, an electronic device includes at least one processor and at least one memory. Wherein, a computer program is stored on the memory, and when the computer program is executed by the processor, the content description method described above is implemented.
[0016] According to one aspect of the present application, a storage medium stores a computer program thereon, and when the computer program is executed by one or more processors, the content description method described above is implemented.
[0017] According to one aspect of the present application, a computer program product includes a computer program, and when the computer program is executed by one or more processors, the content description method described above is implemented.
[0018] The beneficial effects brought by the technical solution provided by the present application are as follows:
[0019] In the above technical solution, based on the prompt words corresponding to the target picture, the target task can be determined. When the target task is a content recognition task, the target picture is content-recognized, and description information is output, so as to present the target content in the target picture in the form of natural language, which is convenient for users to understand; in addition, the content recognition task supports static images (single frame / multiple photos) and dynamic videos (multiple frame extraction), and has the ability to understand the relationship between the target contents in different target pictures, so that the output description information can more accurately describe the target content, thereby effectively solving the problem of low accuracy of video content description in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative efforts.
[0021] Figure 1 is a schematic diagram of the implementation environment related to the present application;
[0022] Figure 2 is a hardware structure diagram of an electronic device shown according to an exemplary embodiment;
[0023] Figure 3 is a flowchart of a content description method shown according to an exemplary embodiment;
[0024] Figure 4 is a flowchart of the training process of a content description model shown according to an exemplary embodiment;
[0025] Figure 5 Yes Figure 2 It is a flowchart corresponding to the steps after step 350 in the embodiment;
[0026] Figure 6 Yes Figure 5 It is a flowchart of step 530 in an embodiment corresponding to the embodiment;
[0027] Figures 7 to 9 It is a schematic diagram of a specific implementation of a content description method in an application scenario;
[0028] Figure 10 It is a structural block diagram of a content description device shown according to an exemplary embodiment;
[0029] Figure 11 It is a structural block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0030] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.
[0031] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present disclosure means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0032] As mentioned above, the tag-based video summary lacks the understanding of the relationships between elements in the picture and is difficult to express the actual content information in the video.
[0033] The video summary detected by the CNN model only extracts the object information in the video and cannot understand the content information in the video. For example, the content information in the video is that a cat is eating cat food in a cat bowl on a table in the bedroom, and there is a dog beside it, and the dog is sleeping. And the tag-style summary is cat, dog, eating, bedroom, table, etc. Then, it is necessary for the user to understand the meaning of the tag-style summary by himself.
[0034] As can be seen from the above, there are still defects in the accuracy of video description content in the related technology.
[0035] Therefore, the content description method provided in this application can effectively improve the accuracy of content description. Correspondingly, this content description method is applicable to a content description device, and this content description device can be deployed on an electronic device, and this electronic device can be a computer device configured with a von Neumann architecture. For example, this computer device includes a desktop computer, a laptop computer, a server, etc.
[0036] To make the purpose, technical solution and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the drawings.
[0037] Figure 1 It is a schematic diagram of the implementation environment involved in a content description method. This implementation environment at least includes a user terminal 110, a smart device 130, a server side 170, and a network device. In Figure 1 it, the network device includes a gateway 150 and a router 190, and this is not a specific limitation here.
[0038] Among them, the user terminal 110, which can also be considered as the user side or the terminal, can deploy (also understood as install) the client associated with the smart device 130. This user terminal 110 can be an electronic device such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart control panel, and other devices with display and control functions, and this is not limited here.
[0039] Among them, the client, which is associated with the smart device 130, actually means that the user registers an account in the client and configures the smart device 130 in the client. For example, this configuration includes adding a device identifier to the smart device 130, etc., so that when the client runs in the user terminal 110, it can provide functions such as device display and device control for the user. This client can be in the form of an application program or in the form of a web page. Correspondingly, the interface for the client to display the device can be in the form of a program window or in the form of a web page, and this is not limited here either.
[0040] The intelligent device 130 is deployed in the gateway 150 and communicates with the gateway 150 through its configured communication module, and thus is controlled by the gateway 150. It should be understood that the intelligent device 130 generally refers to one of multiple intelligent devices 130. The embodiments of the present application only take the intelligent device 130 as an example. That is, the embodiments of the present application do not limit the number and device type of the intelligent devices deployed in the gateway 150. In an application scenario, the intelligent device 130 accesses the gateway 150 through a local area network, and thus is deployed in the gateway 150. The process of the intelligent device 130 accessing the gateway 150 through the local area network includes: the gateway 150 first establishes a local area network, and the intelligent device 130 joins the local area network established by the gateway 150 by connecting to the gateway 150. This local area network includes but is not limited to: ZIGBEE or Bluetooth. Among them, the intelligent device 130 can be an intelligent printer, an intelligent fax machine, an intelligent camera, an intelligent air conditioner, an intelligent door lock, an intelligent light, or a human body sensor, a door and window sensor, a temperature and humidity sensor, a water immersion sensor, a natural gas alarm, a smoke alarm, a wall switch, a wall socket, a wireless switch, a wireless wall sticker switch, a magic cube controller, a curtain motor, a millimeter wave radar, etc. configured with a communication module.
[0041] The interaction between the user terminal 110 and the intelligent device 130 can be realized through a local area network or through a wide area network. In an application scenario, the user terminal 110 establishes a wired or wireless communication connection with the gateway 150 through the router 190. For example, the wired or wireless method includes but is not limited to WIFI, etc., so that the user terminal 110 and the gateway 150 are deployed in the same local area network, and thus the user terminal 110 can realize the interaction with the intelligent device 130 through the local area network path. In another application scenario, the user terminal 110 establishes a wired or wireless communication connection with the gateway 150 through the server 170. For example, the wired or wireless method includes but is not limited to 2G, 3G, 4G, 5G, WIFI, etc., so that the user terminal 110 and the gateway 150 are deployed in the same wide area network, and thus the user terminal 110 can realize the interaction with the intelligent device 130 through the wide area network path.
[0042] Among them, the server 170 can also be considered as the cloud, the cloud platform, the platform side, the service side, etc. This server 170 can be a single server, or a server cluster composed of multiple servers, or a cloud computing center composed of multiple servers, so as to better provide background services for a large number of user terminals 110. For example, the background services include device control services.
[0043] In an application scenario, the user uses the intelligent device 130 to capture the scenario, and uploads the captured image or video to the server side 170. At the same time, the user uses the intelligent device 130 to input a prompt word, which corresponds to the target task that the user expects the server side 170 to perform on the image or video. For the server side 170, after receiving the video or image sent by the intelligent device 110 and the corresponding prompt word, it can process them to obtain the target picture and its corresponding prompt word, and determine the target task according to the prompt word. When the target task is a content recognition task, the server side 170 recognizes the target content of the target picture and obtains a description information, which is the information in the form of natural language corresponding to the target content in the target picture; then, the server side 170 sends the target information to the user terminal 110, and the user can view the description information through the user terminal 110.
[0044] It should be noted that the image / video acquisition task implemented by the above intelligent device 130 can also be implemented by the user terminal 110. For example, the user terminal 110 is a smart phone.
[0045] Please refer to Figure 2 , Figure 2 is a hardware structure diagram of an electronic device shown according to an exemplary embodiment. This electronic device is applicable to Figure 1 the server side 170 in the implementation environment shown.
[0046] It should be noted that this electronic device is only an example adapted to this application, and cannot be considered as providing any limitation to the scope of use of this application. This electronic device cannot be interpreted as needing to rely on or necessarily having Figure 2 one or more components in the exemplary electronic device 200 shown.
[0047] The hardware structure of the electronic device 200 may vary greatly due to different configurations or performances. As Figure 2 shown, the electronic device 200 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.
[0048] Specifically, the power supply 210 is used to provide working voltage for each hardware device on the electronic device 200.
[0049] The interface 230 includes at least one wired or wireless network interface 231, which is used to interact with external devices. For example, to perform Figure 1 the interaction between the intelligent device 130 and the server side 170 in the implementation environment shown.
[0050] Of course, in other examples adapted to this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, at least one USB interface 237, etc., as Figure 2 shown, and specific limitations are not imposed herein.
[0051] As a carrier for resource storage, the memory 250 can be a read-only memory, a random access memory, a magnetic disk, an optical disk, etc. The resources stored thereon include an operating system 251, application programs 253, data 255, etc., and the storage method can be transient storage or permanent storage.
[0052] Among them, the operating system 251 is used to manage and control each hardware device and application program 253 on the electronic device 200, so as to realize the operation and processing of the massive data 255 in the memory 250 by the central processing unit 270. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0053] The application program 253 is a computer program formed by computer-readable instructions that complete at least one specific task based on the operating system 251. It can include at least one module ( Figure 2 not shown), and each module can respectively contain corresponding computer-readable instructions. For example, the content description device can be regarded as an application program 253 deployed on the electronic device 200.
[0054] The data 255 can be photos, pictures, etc. stored in the magnetic disk, and can also be a target screen, which is stored in the memory 250.
[0055] The central processing unit 270 can include one or more processors, and is set to communicate with the memory 250 through at least one communication bus, so as to read the computer program stored in the memory 250, and then realize the operation and processing of the massive data 255 in the memory 250. For example, the content description method is completed in the form of reading the application program 253 stored in the memory 250 by the central processing unit 270.
[0056] In addition, this application can also be implemented by a hardware circuit or a combination of a hardware circuit and software. Therefore, the implementation of this application is not limited to any specific hardware circuit, software, and the combination of the two.
[0057] Please refer to Figure 3 , this application embodiment provides a content description method, which is applicable to an electronic device. For example, the electronic device can be Figure 1 the server side 170 in the shown implementation environment, and the hardware structure of the electronic device can be as Figure 2 shown.
[0058] In the following method embodiments, for the sake of convenience of description, the execution subject of each step of the method is taken as an example of an electronic device for illustration, but this is not a specific limitation thereto.
[0059] As Figure 3 shown, the method may include the following steps:
[0060] Step 310, obtain a target screen.
[0061] First of all, it should be noted that the target screen refers to the screen that the user expects to describe the content included therein.
[0062] The target screen can be obtained by shooting and collecting through an image acquisition device. Among them, the image acquisition device can be an electronic device with an image acquisition function, such as a camera, a smart phone equipped with a camera, and so on.
[0063] It can be understood that the shooting can be a single shooting or a continuous shooting. Then, for continuous shooting, a video can be obtained, and the target screen can be any frame in the video. For multiple shootings, multiple photos can be obtained, and the target screen can be any one of the multiple photos. In other words, the target screen in this embodiment can come from a dynamic image, such as multiple frames in a video or multiple photos, and can also come from a static image, such as any frame in a video or any one of multiple photos. Correspondingly, the content description in this embodiment is carried out in units of frames.
[0064] Regarding the acquisition of the target screen, the target screen can be sourced from the images captured and collected in real time by the image acquisition device, or can be the images captured and collected by the image acquisition device during a historical time period pre-stored in the electronic device. Then, for the electronic device, after the image acquisition device captures and collects the target screen, the target screen can be processed in real time, or can be pre-stored and then processed. For example, the target screen can be processed when the CPU of the electronic device is low, or processed according to the instructions of the staff. Therefore, the content description in this embodiment can be directed to the target screen obtained in real time or the target screen obtained during a historical time period, and no specific limitation is made here.
[0065] Step 330, determine a target task based on the prompt words corresponding to the target screen.
[0066] Among them, the prompt word is the word input by the user to indicate the task that the user expects to complete; the target task can refer to the task that the user expects to complete. For example, it can be a content recognition task. It should be understood that there may be an association between the target task and the prompt word. Specifically, if the user expects to complete a certain task, the corresponding prompt word can be input into the electronic device. When the electronic device is prompted to execute the content description method, the target task can be determined according to the prompt word corresponding to the target screen, so that the electronic product can achieve the target task based on the target screen.
[0067] Regarding the input of the prompt word, the user can type the prompt word using a physical keyboard, or write the prompt word by hand on a touch screen, or input the prompt word in the form of speech recognition through a voice module, etc. There is no specific limitation here.
[0068] In a possible implementation, semantic understanding processing can be performed on the prompt word to obtain the semantics represented by the prompt word, and then the target task corresponding to the prompt word can be obtained according to the semantics. Intent recognition processing can also be performed on the prompt word, or intent recognition processing can be performed on the basis of the semantics of the prompt word, to recognize the intent represented by the prompt word, and then the target task corresponding to the prompt word can be obtained according to the intent.
[0069] Furthermore, different target screens can correspond to different prompt words, or can correspond to the same prompt word. The correspondence between the target screen and the prompt word can depend on the task that the user expects to achieve.
[0070] For example, the user expects to perform a content recognition task on target screen A and expects to perform an image segmentation task on target screen B; then, for target screen A, the user can input "Please perform content recognition on it", and for target screen B, the user can input "Please perform image segmentation on it", etc. There is no specific limitation here.
[0071] Step 350, if the target task is a content recognition task, then perform content recognition on the target content in the target screen to obtain the description information of the target screen.
[0072] Among them, the description information is the information in the form of natural language used to describe the target content in the target screen, and the target content can refer to the content included in the target screen.
[0073] Regarding the content recognition task, it can refer to describing the target content in the target screen in natural language. For example, the content recognition task can include scene understanding, that is, parsing the semantics in the target content. For example, if the target screen includes a traffic sign (target content), then the content recognition task can parse this traffic sign and recognize the semantics of this traffic sign, such as no left turn, etc.
[0074] Further, the content recognition task may also include object detection, such as automatically identifying items in the picture; it may include behavior recognition, such as recognizing the behavior of the target object in the picture, etc., which will not be specifically limited here.
[0075] In a possible implementation, the content recognition task may be implemented by means of deep learning and convolutional neural network.
[0076] Through the above process, based on the prompt corresponding to the target picture, the target task can be determined. When the target task is a content recognition task, the content of the target picture is recognized, and the description information is output, so as to present the target content in the target picture in the form of natural language, which is convenient for users to understand; in addition, the content recognition task supports static images (single frame / multiple photos), dynamic videos (multiple frame extraction), and has the ability to understand the relationship between the target contents in different target pictures, so that the output description information can more accurately describe the target content.
[0077] In an exemplary embodiment, the above method may further include the following steps: implementing the content description method through a content description model; the content description model is a natural language model that has been trained and has the ability to recognize the content of the picture.
[0078] Among them, the content description model can convert visual information (target information) into a structured semantic description through deep learning technology.
[0079] Regarding the content description model, it may include a computer vision module VisionTransformer and a natural language processing module LLM. The fusion of the above two modules can achieve cross-modal mapping from the picture to the text.
[0080] For example, input the target picture into the content recognition model. The computer vision module of the content recognition model will recognize the target content in the target picture, and output the target content as description information through the natural language processing module.
[0081] It should be noted that different from the tag-style summary, the natural language processing module (LLM) can understand the association of each element in the picture and output it in the form of natural speech. That is to say, the description information can enable users to more easily and clearly understand the content in the picture.
[0082] For example, the tag-style summary: cat, dog, eating, bedroom, table, etc.; the description information in the form of natural language: A cat is eating cat food on the cat bowl on the table in the bedroom, and there is a dog beside it, and the dog is sleeping.
[0083] It should be added that if there are multiple target screens in the input content recognition model, then the content recognition model can extract the prompt words corresponding to each target screen, splice each target screen and the corresponding prompt words, and input them into the computer vision module; the computer vision module will determine the association of the target content between the target screens according to the target screen and the corresponding prompt words, and understand each target content, so as to output the corresponding description information through the natural language processing module. It can be understood that the description information after the above processing can reflect the association information of the target content between the target screens. If the target screens are obtained from the same video, then the description information can accurately describe the target content of the video, such as the story events of the video, etc.
[0084] Through the above process, using the content description model to recognize the content of the target screen can more accurately understand the target content, so that the output description information can more accurately and intuitively describe the target content, which is convenient for users to understand.
[0085] Please refer to Figure 4 , such as Figure 4 shown, in an exemplary embodiment, the training process of the content description model includes:
[0086] Step 410, obtain sample data.
[0087] Among them, the sample data includes sample screens and their corresponding sample description information; the sample description information is information in the form of natural language used to describe the content in the sample screen.
[0088] Regarding the acquisition of sample data, it can be obtained by manually understanding the content of the sample screen and then inputting the sample description information to obtain the sample description screen; it can also obtain sample data from a publicly available database, which is not specifically limited here.
[0089] Step 430, input the sample screen into the initial content description model for content recognition to obtain training description information.
[0090] Among them, the training description information is information in the form of natural language output by the initial content description model to describe the content in the sample screen.
[0091] It can be understood that the training description information can be regarded as the answer of the initial content description model. Since the initial content description model's ability to recognize the content of the screen is not yet perfect, there may be a deviation between the training description information and the sample description information.
[0092] For example, the content of the sample screen is "There are an old man and a dog sitting on the park bench", and the sample description information may be "There are people and animals on the chair".
[0093] Step 450: Calculate the loss based on the sample description information and the training description information to obtain loss information.
[0094] Among them, the loss information can quantify the deviation between the sample description information and the training description information. It can be understood that the smaller the above deviation, the stronger the recognition ability of the content description model in the training process for the content in the picture, that is, the closer it is to the state of training completion.
[0095] In a possible implementation, the loss information may include cross-entropy loss, contrastive learning loss, etc., which is not limited here.
[0096] Step 470: Optimize the content description model using the loss information and continue training until a trained content description model is obtained.
[0097] Specifically, the loss information can be used to optimize the model parameters of the content description model to improve the recognition ability of the content description model for the content in the picture.
[0098] In a possible implementation, the Adam optimizer can be used to optimize the content description model, which is not limited here.
[0099] Through the above process, the initial content description model is trained using sample data, the loss information is calculated based on the sample description information and the training description information, and the loss information is used to optimize the content description model, which can enable the content description model to gradually improve its recognition ability for the content in the picture. Thus, when the content description model is subsequently used to implement the content recognition task, the output description information can accurately present the target content.
[0100] In an exemplary embodiment, the sample data further includes sample prompt words corresponding to the sample pictures; step 430 may further include the following steps: determining the target task based on the sample prompt words corresponding to the sample pictures; in the case where the target task is content recognition, using the content description model to perform content recognition on the sample pictures to obtain training description information.
[0101] Among them, the sample prompt words are words corresponding to different characters. It can be understood that for the same target task, the sample prompt words can have different expression forms.
[0102] For example, when the target task is a content description task, at this time, the sample prompt word can be "Please describe the picture", or it can also be said "What is in the picture", etc.
[0103] Then, for different expression forms corresponding to the target task, training the content description model with different sample prompt words can enable it to accurately determine the same target task, that is, the content recognition task, through different sample prompt words.
[0104] Based on this, when it is determined that the target task is a content recognition task based on the sample prompt, the content description model can be used to perform content recognition on the sample picture, so as to obtain training description information.
[0105] Through the cooperation of the above embodiments, by using diverse sample prompts as the training process of the content description model, the understanding ability of the content description model for prompts in different expression forms can be improved, so as to accurately determine the content recognition task, thereby improving the generalization ability of the content description model.
[0106] Please refer to Figure 5 , in an exemplary embodiment, after step 350, the method may further include the following steps:
[0107] Step 510, obtain search indication information.
[0108] Among them, the search indication information is information in the form of natural language used to describe the target content that the user expects to find.
[0109] It should be noted that the user can enter the search prompt information using a physical keyboard, or can write the search prompt information on a touch screen, or can also input the search prompt information in the form of voice recognition through a voice module, etc., which is not specifically limited here.
[0110] It can be understood that when the user expects to find a certain target content, the corresponding search indication information can be input through the electronic device.
[0111] Step 530, search for the picture with the target content based on the search indication information to obtain the target picture.
[0112] It can be understood that the search indication information can describe the target content, so the corresponding target content can be found according to the search indication information.
[0113] In a possible implementation, the similarity between the search indication information and the description information corresponding to each picture is calculated. If the similarity between the description information corresponding to a picture and the search indication information exceeds 90%, it can be considered that this picture is the target picture.
[0114] Under the action of the above embodiments, using the search prompt message to find the picture with the target content to obtain the target picture can provide a way for the user to find the target picture using information in the form of natural language.
[0115] Please refer to Figure 6 , in an exemplary embodiment, step 530 may further include the following steps:
[0116] Step 531, obtain the target database.
[0117] Among them, the target database further includes multiple pictures and their corresponding description vectors; the description vector can be a vectorized representation of the description information corresponding to the picture.
[0118] Regarding the target database, a content description model can be used to identify the content of multiple pictures to obtain corresponding description information, and then the description information is subjected to a text vector conversion process to convert it into a corresponding description vector, and then each picture and its corresponding description vector are stored in the database to obtain the target database.
[0119] Step 533: Perform a text vector conversion process on the search indication information to obtain a search indication vector.
[0120] Regarding the text vector conversion process, through a text encoding model, for example, the BERT model (Bidirectional Encoder Representations from Transformers, a pre-trained encoding model based on transformers), the search hint information can be used to convert information in natural language form into vector form.
[0121] Step 535: Perform data search in the target database based on the search indication vector. If there is a target description vector whose similarity with the search indication vector meets the set conditions, the picture corresponding to the target description vector is used as the target picture.
[0122] Among them, the set condition is a condition used to determine whether the search indication vector is similar to the description vector. The set condition can be set according to the user's needs. For example, if the user expects the target description vector obtained from the data search to be more accurate, the set condition can be set more strictly, for example, the similarity reaches more than 95%.
[0123] Regarding the similarity calculation, the cosine similarity calculation method can be used. Specifically, refer to formula (1):
[0124]
[0125] Among them, A can be the search indication vector, and B can be the description vector in the target database.
[0126] Regarding the data search, specifically, it can refer to calculating the similarity between the search indication vector and each description vector in the target database, and determining the target description vector based on the similarity between the search indication vector and each description vector.
[0127] For example, if the similarity between the search indication vector and the description vector A is 90%, the similarity with the description vector B is 95%, and the similarity with the description vector C is 75%, then the description vector B is regarded as the target description vector.
[0128] It can be understood that the search indication vector can indicate the target content that the user expects to find, and the description vector can be used to indicate the content included in the corresponding screen. If the search indication vector is similar to the description vector, then the corresponding screen of the description vector includes the target content that the user expects to find, and this corresponding screen is the target screen. Based on this, the target screen requested by the user can be returned.
[0129] Under the action of the above embodiment, by searching for similar description vectors in the target database through the search indication vector to determine the target description vector, so as to find the target screen, the efficiency and accuracy of screen search can be improved.
[0130] In an exemplary embodiment, the above method may further include the following steps: performing a text vector conversion process on the description information corresponding to the target screen to obtain the corresponding target description vector; storing the target screen and its corresponding target description vector to generate a target database.
[0131] Specifically, after the content recognition of the target screen is completed to obtain the corresponding description information, in order to facilitate the user to search for the target screen later, the description information can be subjected to a text vector conversion process to obtain the corresponding target description vector.
[0132] Then, it can be stored to generate a target database. If the user expects to search for the target screen later, the search indication information in the form of natural language can be input, and the search indication information is subjected to a text vector conversion process to obtain a search indication vector, so as to calculate the similarity between the search indication vector and the target description vector, and thus find the target screen.
[0133] Through the above process, the description information obtained by the target content recognition of the target screen is converted into a description vector, and the target screen and its corresponding description vector are stored as a target database, which can facilitate the management of each target screen, and when the user expects to find the target screen, quickly find the target screen from the target database.
[0134] In addition, since the content recognition of the target screen is implemented by using the above content description method, the obtained description information more accurately reflects the target content. Then, when the user expresses the content expected to be searched through the search indication information in the form of natural language, the system can more accurately find the target content similar to the content (search indication vector) expected to be searched by the user.
[0135] In an exemplary embodiment, before step 310, the method may further include the following steps: obtaining a target video; performing frame extraction on the target video using a set method to obtain a target frame.
[0136] Among them, the target video may refer to a video that a user expects to perform content recognition on. It can be understood that the target video includes multiple frames of images, and each frame of image may correspond to the same content or different contents.
[0137] First of all, it should be noted that in the case of performing content recognition on the target video, if each frame of image is directly used for content recognition, if the target video is very long, frame-by-frame processing may be very resource-consuming.
[0138] Then, in order to save resources and improve the efficiency of content recognition, frame extraction can be performed on the target video to reduce the number of frames for content recognition.
[0139] It should be noted that the inventor found that the input of large LLM models (such as LLaMa, GPT-3, etc.) generally converts text information into vectors as input and cannot accept inputs such as pictures, text, or sounds.
[0140] So, how can the content of the target video be accurately described? The inventor thought that a video-to-vector module could be added to the input side of the original large LLM model as input. However, there are many image-to-vector models (such as VisionTransformer), but few video-to-vector models.
[0141] In order to achieve the technical effect of converting a video into a vector, the inventor also thought that since a video itself is composed of frames of images, key-position frames can be extracted by frame extraction, and the images can be converted into vectors through VisionTransformer as input.
[0142] Regarding the set method, it is a method for performing frame extraction on the target video. The frame difference method can be used to select the target frames for frame extraction. Specifically, key frames are extracted in video segments with large changes in frames, and frame skipping extraction is performed in this video segment to reduce the amount of calculation. For example, one frame is extracted every 1 second in a segment with large frame changes, which can reduce the amount of calculation while not losing the information of the target video.
[0143] It should be understood that for different types of target videos, there are also differences in the number of frames that need to be retained during frame extraction. For example, in a target video with a large frame change rate, more video frames need to be extracted and more target frames need to be retained to avoid loss of video content.
[0144] Furthermore, after content recognition of the target images obtained by frame extraction from the target video, description information corresponding to each target image can be obtained. Then, the description information can be processed by text vector conversion to obtain the corresponding description vectors, and thus the target images and their corresponding description vectors in the target video can be stored in the target database.
[0145] For example, assume that a total of 6 target images are extracted from the target video, and the prompt can be set as "Please add a summary description to the following video": <Video frame 1><Video frame 2><Video frame 3><Video frame 4><Video frame 5><Video frame 6>; the prompt part is converted into a vector through the encoding method of the LLM, and each target image is converted into a vector through VisionTransformer; after the target video and the corresponding prompt are both vectorized, the content description model can understand the target content of the entire target video and output the description information; and the description information is converted into a description vector through a text encoding model (such as OpenAI's text-embedding model) and stored on the disk. Assume there are a total of N videos, then the information on the disk can be as follows:
[0146] Video 1 ----> Vector 1; Video 2 ----> Vector 2;... Video n ----> Vector n
[0147] Then, when the user expects to find a certain video, a corresponding search indication vector can be generated to quickly find the target image in the target database, and the video corresponding to the target image is regarded as the target video, providing a way for the user to quickly locate the target video.
[0148] Similarly, when the user expects to find a certain image of the target video, a corresponding search indication vector can also be generated to quickly find the target image in the target database, providing a way for the user to quickly locate the target image.
[0149] In the above process, by setting a method to perform frame extraction on the target video, the amount of calculation can be reduced and resources can be saved without losing the picture content of the target video, thereby accelerating the efficiency of content recognition.
[0150] Figures 7 to 9 It is a schematic diagram of the specific implementation of a content description method in an application scenario.
[0151] Please refer to Figure 7 , Figure 7 which shows the training process of the content description model in this application scenario. Now, in combination with Figure 7 the training process will be described:
[0152] First, obtain the sample video and the sample prompt. Use the frame difference method to select the sample frames from the sample video. Specifically, extract key frames from the video segments with significant changes in the frames. In this video segment, reduce the computational complexity by skipping frames. For example, extract one frame every 1 second in the segment with significant frame changes.
[0153] Then, vectorize the sample frames and the sample prompt. Vectorization means converting the image prompt into a one-dimensional vector through a sampleable transformation matrix. Then, the image and the prompt can be concatenated. Suppose 3 key frames are extracted from a sample video, and the prompt is to describe the content in the video. The concatenation format is as follows: "Image vector 1" "Image vector 2" "Image vector 3" "Please describe the content in the video". Among them, the concatenation format during the training of the content description model is not fixed, but during the use process, the concatenation method should be consistent with the sample process.
[0154] Then, the sample frames can be input into the initial content description model for content recognition to obtain the training description information. The training description information can be regarded as the answer of the initial content description model. Since the initial content description model's ability to recognize the content of the frames is not yet perfect, there may be a deviation between the training description information and the sample description information.
[0155] Based on this, calculate the loss based on the sample description information and the training description information to obtain the loss information. The loss information can quantify the deviation between the sample description information and the training description information. It can be understood that the smaller the above deviation, the stronger the ability of the content description model in the training process to recognize the content of the frames, that is, the closer it is to the state of completed training.
[0156] Finally, use the loss information to optimize the parameters of the content description model and continue training until the content description model with completed training is obtained.
[0157] Please refer to Figure 8 , Figure 8 which shows the usage process of the content description model in this application scenario. Now, in combination with Figure 8 the usage process will be described as follows:
[0158] First, obtain the target video and the prompt input by the user. Use the frame difference method to select the target frames for extraction. Specifically, extract key frames from the video segments with significant changes in the frames. In this video segment, reduce the computational complexity by skipping frames. For example, extract one frame every 1 second in the segment with significant frame changes, which can reduce the computational complexity while not losing the information of the target video.
[0159] Furthermore, 6 target frames are extracted from the target video, and the prompt can be set as "Please add a summary description to the following video": <Video Frame 1><Video Frame 2><Video Frame 3><Video Frame 4><Video Frame 5><Video Frame 6>; The prompt part is converted into a vector through the encoding method of the LLM module of the content description model, and each target frame is converted into a vector through the VisionTransformer module of the content description model; After both the target frames and the corresponding prompts are vectorized, the content description model can understand the target content of the entire target video and output a description message; And the description message is converted into a description vector through a text encoding model (such as OpenAI's text-embedding model) and stored in the database to obtain a target database.
[0160] Please refer to the figure, Figure 9 which shows the search process of the target video in this application scenario. Now, in combination with Figure 9 the search process will be described:
[0161] Obtain search indication information, which is information in the form of natural language used to describe the target content that the user expects to find. Perform text vector conversion processing on the search indication information to obtain a search indication vector.
[0162] Furthermore, based on the search indication vector, data search is performed in the target database. The description vector with the highest similarity to the search indication vector is regarded as the target description vector. The frame corresponding to the target description vector is the target frame, and the video corresponding to the target frame is the target video.
[0163] Of course, it is also possible to sort according to the calculated similarity, and the one or more with the highest similarity are the target videos that the user expects to find.
[0164] In this application scenario, the content description model can greatly reduce the time of manual operation by video uploaders; fully consider the needs of video search users and provide users with description information that better fits their search intentions; use methods based on artificial intelligence deep learning technology to make the generated description information more accurately reflect the content of the video.
[0165] It can be understood that the trained content description model can achieve the association of picture elements in content recognition and output description information in the form of natural language, which is more intuitive than the traditional tag-based summary; Furthermore, it is also possible to search for videos through natural language instead of keyword matching. For example, to search for all videos with animals, all videos related to animals such as cats and dogs can be searched, and keyword matching cannot search through natural language.
[0166] It should be understood that although the steps in the flowchart of the accompanying drawings are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0167] The following is an apparatus embodiment of the present application, which can be used to execute the content description method involved in the present application. For the details not disclosed in the apparatus embodiment of the present application, please refer to the method embodiment of the content description method involved in the present application.
[0168] Please refer to Figure 10 , in an embodiment of the present application, a content description apparatus 900 is provided, including but not limited to: a screen acquisition module 910, a task determination module 930, and a content recognition module 950.
[0169] Among them, the screen acquisition module 910 is used to acquire a target screen.
[0170] The task determination module 930 is used to determine a target task based on a prompt word corresponding to the target screen; the prompt word is a word input by the user and used to indicate the task that the user expects to complete.
[0171] The content recognition module 950 is used to, if the target task is a content recognition task, perform content recognition on the target content in the target screen to obtain description information of the target screen; the description information is information in the form of natural language used to describe the target content in the target screen.
[0172] In an exemplary embodiment, the content description apparatus is further used to acquire search indication information; the search indication information is information in the form of natural language used to describe the target content that the user expects to find; based on the search indication information, search for a screen with the target content to obtain the target screen.
[0173] In an exemplary embodiment, the content description apparatus is further used to acquire a target database, and the target database further includes multiple screens and their corresponding description vectors; perform text vector conversion processing on the search indication information to obtain a search indication vector; based on the search indication vector, perform data search in the target database. If there is a target description vector whose similarity with the search indication vector meets the set conditions, then use the screen corresponding to the target description vector as the target screen.
[0174] In an exemplary embodiment, the content description device is further configured to perform text vector conversion processing on the description information corresponding to the target picture to obtain a corresponding target description vector; store the target picture and its corresponding target description vector to generate a target database.
[0175] In an exemplary embodiment, the content description device is further configured to implement the content description method through a content description model; the content description model is a natural language model obtained through training and having the ability to identify the content of a picture.
[0176] In an exemplary embodiment, the content description device is further configured to obtain sample data; the sample data includes a sample picture and its corresponding sample description information; the sample description information is information in the form of natural language used to describe the content in the sample picture; input the sample picture into the initial content description model for content recognition to obtain training description information; calculate a loss based on the sample description information and the training description information to obtain loss information; use the loss information to optimize the content description model and continue training until a trained content description model is obtained.
[0177] In an exemplary embodiment, the sample data further includes a sample prompt word corresponding to the sample picture; the content description device is further configured to determine a target task based on the sample prompt word corresponding to the sample picture; in the case where the target task is content recognition, use the content description model to perform content recognition on the sample picture to obtain training description information.
[0178] In an exemplary embodiment, the content description device is further configured to obtain a target video; perform frame extraction processing on the target video using a set method to obtain a target picture.
[0179] It should be noted that when the content description device provided in the above embodiment performs content description, only the above-mentioned division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the content description device will be divided into different functional modules to complete all or part of the functions described above.
[0180] In addition, the content description device provided in the above embodiment and the embodiment of the content description method belong to the same concept. The specific ways in which each module performs operations have been described in detail in the method embodiment and will not be repeated here.
[0181] Please refer to Figure 11 , in the embodiment of the present application, an electronic device 4000 is provided, and the electronic device 4000 may include: a desktop computer, a laptop computer, a server, etc.
[0182] In Figure 10Among them, the electronic device 4000 includes at least one processor 4001 and at least one memory 4003.
[0183] Among them, the data interaction between the processor 4001 and the memory 4003 can be realized through at least one communication bus 4002. The communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003. The communication bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0184] Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0185] The processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 4001 can also be a combination that realizes computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0186] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store computer programs in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited thereto.
[0187] A computer program is stored on the memory 4003, and the processor 4001 can read the computer program stored in the memory 4003 through the communication bus 4002.
[0188] The computer program is executed by one or more processors 4001 to implement the content description method in the above embodiments.
[0189] In addition, an embodiment of the present application provides a storage medium on which a computer program is stored, and the computer program is executed by one or more processors to implement the content description method as described above.
[0190] An embodiment of the present application provides a computer program product, including a computer program, and the computer program is executed by one or more processors to implement the content description method as described above.
[0191] Compared with the related art, the present solution can determine the target task based on the prompt word corresponding to the target picture. When the target task is a content recognition task, the target picture is content-recognized, and the description information is output, so as to present the target content in the target picture in the form of natural language, which is convenient for users to understand. In addition, the content recognition task supports static images (single frame / multiple photos) and dynamic videos (multiple frame extraction), and has the ability to understand the relationship between the target contents in different target pictures, so that the output description information can more accurately describe the target content.
[0192] In addition, by performing frame extraction processing on the target video through a setting method, the amount of calculation can be reduced and resources can be saved without losing the picture content of the target video, thereby accelerating the content recognition efficiency.
[0193] Furthermore, when the user expresses the content to be searched for in natural language by means of search indication information, the system can more accurately search for target content similar to the content to be searched for by the user (search indication vector).
[0194] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A content description method, characterized in that, It includes: Obtain a target screen; Determine a target task based on the prompt words corresponding to the target screen; the prompt words are words input by the user and used to indicate the task that the user expects to complete; If the target task is a content recognition task, perform content recognition on the target content in the target screen to obtain the description information of the target screen; the description information is information in the form of natural language used to describe the target content in the target screen.
2. The method according to claim 1, wherein, After performing content recognition on the target content in the target screen to obtain the description information of the target screen, the method further includes: Obtain search indication information; the search indication information is information in the form of natural language used to describe the target content that the user expects to find; Search for the screen where the target content exists based on the search indication information to obtain the target screen.
3. The method according to claim 2, wherein The searching for the screen where the target content exists based on the search indication information to obtain the target screen includes: Obtain a target database, and the target database further includes multiple screens and their corresponding description vectors; Perform text vector conversion processing on the search indication information to obtain a search indication vector; Perform data search in the target database based on the search indication vector. If there is a target description vector whose similarity with the search indication vector meets the set conditions, use the screen corresponding to the target description vector as the target screen.
4. The method according to claim 3, wherein, Before obtaining the target database, the method further includes: Perform text vector conversion processing on the description information corresponding to the target screen to obtain a corresponding target description vector; Store the target screen and its corresponding target description vector to generate the target database.
5. The method according to any one of claims 1 to 4, characterized in that, Implement the content description method through a content description model; the content description model is a natural language model obtained through training and having the ability to recognize the content of a screen.
6. The method according to claim 5, wherein The training process of the content description model includes: Obtain sample data; the sample data includes sample screens and their corresponding sample description information; the sample description information is information in the form of natural language used to describe the content in the sample screen; Input the sample screen into an initial content description model for content recognition to obtain training description information; Calculate a loss based on the sample description information and the training description information to obtain loss information; Use the loss information to optimize the content description model and continue training until a trained content description model is obtained.
7. The method according to claim 6, characterized in that, The sample data further includes sample prompt words corresponding to the sample screens; The inputting the sample screen into an initial content description model for content recognition to obtain training description information includes: Determine a target task based on the sample prompt words corresponding to the sample screen; In the case where the target task is content recognition, use the content description model to perform content recognition on the sample screen to obtain the training description information.
8. The method according to any one of claims 1 to 4, characterized in that Before obtaining the target screen, the method further includes: Obtain a target video; Use a set method to perform frame extraction processing on the target video to obtain the target screen.
9. A content description device, characterized in that, It includes: A screen acquisition module, configured to acquire a target screen; A task determination module, configured to determine a target task based on a prompt word corresponding to the target screen; the prompt word is a word input by a user and used to indicate a task that the user expects to complete; A content recognition module, configured to, if the target task is a content recognition task, perform content recognition on target content in the target screen to obtain description information of the target screen; the description information is information in a natural language form used to describe the target content in the target screen.
10. An electronic device includes at least one processor and at least one memory, wherein, A computer program is stored on the memory, characterized in that when the computer program is executed by the processor, the content description method according to any one of claims 1 to 8 is implemented.
11. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by one or more processors, the content description method according to any one of claims 1 to 8 is implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by one or more processors, the content description method according to any one of claims 1 to 8 is implemented.