Method, apparatus, storage medium and electronic device for identifying video scenes
By combining multiple modal information of vision, text and audio, using attribute recognition model and decision tree model, the problem of inaccurate recognition of dance scenes and singing scenes caused by insufficient visual information is solved, and higher recognition accuracy and personalized recommendations are achieved.
Patent Information
- Application Number
- CN202111295247.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-11-03
AI Technical Summary
In the prior art, it is difficult to accurately distinguish video scenes from dance scenes or singing scenes based on visual information alone, resulting in insufficient recognition accuracy.
Combining multiple modal information of vision, text and audio, the attribute recognition model and decision tree model are used to identify whether the video scene is a dance scene or a singing scene.
It improves the accuracy of video scene recognition, especially dance scenes and singing scenes, and enhances the effect of personalized recommendations.
Smart Images

Figure CN114005065B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of electronic information technology, and in particular, to a method, apparatus, storage medium, and electronic device for identifying video scenes. Background Art
[0002] With the development of the Internet, video data has been continuously growing. Efficiently and accurately understanding the content of videos is beneficial to providing users with more convenient services. Among them, accurately classifying the scenes of videos is the basis for understanding video content. Videos can be classified based on different video scenes, and corresponding services can be provided to users according to the classified videos. Therefore, how to accurately identify video scenes is crucial. Summary of the Invention
[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a method for identifying a video scene, including:
[0005] Obtaining a video to be identified;
[0006] Processing the video to be identified to obtain various modality information in the video to be identified, where the various modality information includes visual modality information, text modality information, and audio modality information;
[0007] Determining a target scene result of the video to be identified according to the various modality information, where the target scene result is used to indicate whether the video scene of the video to be identified is a target scene, and the target scene is a dance scene or a singing scene.
[0008] In a second aspect, the present disclosure provides an apparatus for identifying a video scene, including:
[0009] An obtaining module, configured to obtain a video to be identified;
[0010] A processing module, configured to process the video to be identified to obtain various modality information in the video to be identified, where the various modality information includes visual modality information, text modality information, and audio modality information;
[0011] A determining module, configured to determine a target scene result of the video to be identified according to the various modality information, where the target scene result is used to indicate whether the video scene of the video to be identified is a target scene, and the target scene is a dance scene or a singing scene.
[0012] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored. When the program is executed by a processing device, the steps of the method for identifying a video scene described in the first aspect are implemented.
[0013] In a fourth aspect, the present disclosure provides an electronic device, including:
[0014] a storage device, on which a computer program is stored;
[0015] a processing device, configured to execute the computer program in the storage device to implement the steps of the method for identifying a video scene described in the first aspect.
[0016] Through the above technical solutions, a video to be identified is obtained, the video to be identified is processed to obtain various modality information in the video to be identified, and the various modality information includes visual modality information, text modality information, and audio modality information; according to the various modality information, a target scene result of the video to be identified is determined, and the target scene result is used to indicate whether the video scene of the video to be identified is a target scene, and the target scene is a dance scene or a singing scene. By combining various modality information of vision, text, and audio to identify whether the scene of the video to be identified is a dance scene or a singing scene, the accuracy of video scene identification is improved.
[0017] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In combination with the drawings and with reference to the following specific implementation manners, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings:
[0019] Figure 1 is a flowchart of a method for identifying a video scene shown according to an exemplary embodiment.
[0020] Figure 2 is another flowchart of a method for identifying a video scene shown according to an exemplary embodiment.
[0021] Figure 3 is a schematic diagram of a decision-making process of a first decision tree model shown according to an exemplary embodiment.
[0022] Figure 4 is a schematic diagram of a decision-making process of a second decision tree model shown according to an exemplary embodiment.
[0023] Figure 5 is a block diagram of a device for identifying a video scene shown according to an exemplary embodiment.
[0024] Figure 6 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0026] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0027] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0028] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.
[0029] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless explicitly stated otherwise in the context, it should be understood as "one or more".
[0030] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0031] The video scene can be a video application scene, for example, sports, funny, games, singing, dancing, etc. Among many video scenes, the singing scene and the dancing scene often account for a relatively large proportion.
[0032] In the related art, for the recognition of video scenes, visual information (i.e., image information) is often used for recognition. However, for singing scenes or dancing scenes, relying solely on visual information cannot well distinguish between singing scenes and dancing scenes. For example, taking the judgment of a dancing scene as an example, there is visual feature information in the images extracted from the video where multiple people's hands are simply waving. A video with simply waving hands may not be judged as a dancing scene. However, if, on the basis of this visual feature information, other modal information besides vision is combined, there is a possibility of judging that the video is a dancing scene. Therefore, when a single modal information is insufficient to characterize the category of a video scene, the single modal information may not be able to accurately identify the video scene, and thus relying solely on visual information to identify the video scene is inaccurate.
[0033] In view of this, the embodiments of the present disclosure disclose a method, apparatus, electronic device, and storage medium for recognizing a video scene, which can recognize whether the scene of the video to be recognized is a dancing scene or a singing scene by combining multiple modal information of vision, text, and audio, thereby improving the accuracy of video scene recognition.
[0034] Figure 1 FIG. is a flowchart of a method for recognizing a video scene shown according to an exemplary embodiment. The method for recognizing a video scene can be applied to an electronic device, and the electronic device can be a device such as a server or a mobile terminal. As Figure 1 shown, the method for recognizing a video scene may include the following steps:
[0035] Step S101, obtain the video to be recognized.
[0036] It should be noted that the video to be recognized is a video with an unknown video scene.
[0037] In some embodiments, the video to be recognized may be a video crawled from a network database using web crawling technology. By recognizing the crawled video, it is beneficial to subsequently provide convenient services for users using the recognized video. For example, recommending videos of corresponding scenes to users who prefer singing scenes and dancing scenes.
[0038] In some embodiments, the video to be recognized may be a video stored locally.
[0039] In some embodiments, the video to be recognized may be a live video, a short video, a movie video, etc.
[0040] It should be noted that the above embodiments are only exemplary descriptions of the types and sources of the video to be recognized, and the present disclosure does not limit the types and sources of the video.
[0041] Step S102: Process the video to be recognized to obtain various modality information in the video to be recognized, where the various modality information includes visual modality information, text modality information, and audio modality information.
[0042] In some embodiments, the visual modality information may be image feature information obtained from the image frames of the video to be recognized. Image feature information can be obtained from the video to be recognized using image recognition technology. Exemplarily, the image feature information may be face feature information, action feature information, environmental feature information, etc. Among them, the environmental feature information can be understood as information characterizing the indoor environment or outdoor environment.
[0043] In some embodiments, the audio modality information may be audio information extracted from the video to be recognized. Exemplarily, the audio information corresponding to the video to be recognized can be obtained by converting the format of the video to be recognized into an audio format file.
[0044] In some embodiments, the text modality information may be text information recognized from the image frames of the video to be recognized using OCR (Optical Character Recognition); the text modality information may also be text information recognized from the audio information using ASR (Automatic Speech Recognition).
[0045] Step S103: Determine the target scene result of the video to be recognized according to the various modality information, where the target scene result is used to indicate whether the video scene of the video to be recognized is the target scene, and the target scene is a dance scene or a singing scene.
[0046] It should be noted that the various modality information uniquely determines a target scene result. The target scene result is used to indicate whether the video scene of the video to be recognized is the target scene. For example, when the video scene of the video to be recognized is a dance scene, the target scene result is used to indicate that the video scene of the video to be recognized is a dance scene.
[0047] Through the above technical solution, by combining various modality information of vision, text, and audio to identify whether the scene of the video to be recognized is the target scene, it is possible to avoid the situation where it is determined that the video to be recognized is not the target scene due to insufficient single modality information (such as visual modality information) while the video to be recognized is actually the target scene, improving the accuracy of recognizing singing scene videos and dance scene videos.
[0048] In some embodiments, Figure 1The shown step S103 may include: for each type of modality information, processing the modality information according to the attribute recognition model corresponding to the modality information to determine the attribute value of the attribute corresponding to the modality information; processing the attribute values of multiple modality information according to the pre-constructed decision tree model to determine the target scene result of the video to be recognized.
[0049] It should be noted that the attribute recognition model can be understood as a prediction model, which predicts the attribute value according to each modality information. The attribute recognition model can be obtained by training the neural network with sample data. It should be understood that the attribute recognition models corresponding to each modality information are different.
[0050] Among them, the attribute corresponding to the visual modality information is used to describe the picture information. Exemplarily, the attribute corresponding to the visual modality information may be an environmental attribute used to characterize indoor or outdoor, and the attribute value corresponding to the environmental attribute may be an indoor environment or an outdoor environment; the attribute corresponding to the visual modality information may be a number-of-people attribute used to characterize the number of people, and the attribute value corresponding to the number-of-people attribute may be multiple people or single person; the attribute corresponding to the visual modality information may be an action attribute used to characterize the action, and the attribute value corresponding to the action attribute may be continuous limb movements or non-continuous limb movements; the attribute corresponding to the visual modality information may be a microphone attribute used to characterize whether there is a microphone in the picture, and the attribute value corresponding to the microphone attribute may be there is a microphone or there is no microphone; the attribute corresponding to the visual modality information may be an instrument attribute used to characterize the instrument in the picture, and the attribute value corresponding to the instrument attribute may be the type of the instrument, the number of instrument types, etc.
[0051] The attribute corresponding to the text modality information is used to describe the text information. Exemplarily, the attribute corresponding to the text modality information may be a lyric attribute used to characterize whether it is a lyric, and the attribute value corresponding to the lyric attribute may be a lyric or not a lyric; the attribute corresponding to the text modality information may be a slogan attribute used to characterize the video scene. It should be understood that specific slogans of different video scenes can be used to characterize the category of the video scene. For example, if the slogan recognized in the video to be recognized is "Dance Competition", then this slogan can characterize the video scene of the video to be recognized as a dance scene.
[0052] The attribute corresponding to the audio modality information is used to describe the audio type. Exemplarily, the attribute corresponding to the audio modality information may be a human-voice attribute used to characterize whether the audio is a human voice, and the attribute value corresponding to the human-voice attribute may be a human voice or not a human voice; the attribute corresponding to the audio modality information may be a background-music attribute used to characterize whether the audio is a background music, and the attribute value corresponding to the background-music attribute may be a background music or not a background music.
[0053] Attributes can be used to characterize the categories of video scenes. For example, taking the vocal attribute corresponding to the audio modality information, the attribute value of the vocal attribute is that the vocal is more inclined to the singing scene. For another example, the attribute value of the microphone attribute is that the microphone is more inclined to the singing scene.
[0054] It should be noted that the decision tree model can be understood as a classification tree, which is a tree that can classify a data set. It requires that the possible values of each attribute in the data set are discrete. For the same data set, different partitioning attributes are selected, and the shape and depth of the resulting tree are also different.
[0055] An exemplary illustration is given of the decision tree model determining the target scene result of the video to be recognized based on some of the above-mentioned attributes. Taking the non-continuous limb movement of the action attribute and the slogan "Dance Competition" as an example, the first splitting node of the decision tree model cannot determine the type of video scene based on the non-continuous limb movement attribute. Further, the second splitting node of the decision tree model can determine that the video scene is a dance scene based on the slogan "Dance Competition".
[0056] In some embodiments, the attribute values output by different attribute recognition models can be concatenated, and the concatenation result is input into a pre-constructed decision tree model.
[0057] Refer to Figure 2 , the attribute recognition model corresponding to the visual modality information is the visual attribute recognition model. The visual model information is input into the visual attribute recognition model to obtain the corresponding video attribute value; the attribute recognition model corresponding to the audio modality information is the audio attribute recognition model. The audio model information is input into the audio attribute recognition model to obtain the corresponding audio attribute value; the attribute recognition model corresponding to the text modality information is the text attribute recognition model. The text model information is input into the text attribute recognition model to obtain the corresponding text attribute value. Finally, the video attribute value, audio attribute value, and text attribute value are concatenated and input into the decision tree model, so that the decision tree model classifies the video to be recognized according to the input, and the classification result of the decision tree model is the corresponding target scene result.
[0058] In some embodiments, the ID3 (Iterative Dichotomiser 3) algorithm can be used to construct a decision tree model. The ID3 algorithm determines the current splitting attribute by using the information gain of each attribute. Information gain represents the degree to which the classification uncertainty of the data set is reduced due to attribute A. A feature with a large information gain has a stronger classification ability. Therefore, in the decision tree model, it should be selected first, and the classification decision should be made preferentially based on this attribute. Specifically, the decision tree model can be constructed in the following way, including the following steps: Obtain a training set, where the training set includes multiple sample data, and each sample data includes sample attributes, sample attribute values, and sample scenario categories; Based on the training set, determine the information gain of the sample attributes; Based on the information gain of each sample attribute, construct a decision tree model.
[0059] First, it should be noted that there are three types of nodes in the decision tree model, the root node, internal nodes, and leaf nodes. Among them, the leaf node represents a class label, that is, a class label corresponds to a video scenario. The selection of the root node and internal nodes is determined based on the information gain of the sample attributes. When using the decision tree model to identify video scenarios, a path from the root node to the leaf node can be formed according to multiple attributes.
[0060] In some embodiments, the training set is pre-constructed. A sample data can be as shown in the table:
[0061]
[0062] When determining the sample attribute of the root node, calculate the information gain of all sample attributes, and use the sample attribute with the largest information gain as the root node sample attribute. Create different branches according to the values of the root node sample attribute. It can be understood that each value corresponds to a branch. If there is only one type of sample scenario category under a certain branch, the next node corresponding to this branch is a leaf node. If there are different sample scenario categories under a certain branch, the next node corresponding to this branch is an internal node, and continue to determine the sample attribute for this internal node. Among them, the sample attribute determined for this internal node is other sample attributes except the root node sample attribute, and the sample attribute of this internal node is also determined based on the information gain of the sample attributes until the preset condition is met to stop the determination of the node. Construct a decision tree model according to the determined root node and each determined internal node.
[0063] In some embodiments, the preset condition can be that any sample attribute among all sample attributes corresponds to a determined internal node.
[0064] In some embodiments, the preset condition can be that the information gain is less than the preset gain threshold.
[0065] In some embodiments, the preset condition may be that the branches under the last determined internal node belong to the same sample scenario category.
[0066] It can be understood that for different classification objectives (such as the classification objective of dance scenario or non-dance scenario and the classification objective of singing scenario or non-singing scenario), the influence degree of each attribute on the classification objective is different. Therefore, in order to improve the accuracy of song and dance scenario recognition, different decision tree models are set based on the classification of different target scenarios.
[0067] Specifically, the decision tree model includes a first decision tree model and a second decision tree model. Among them, the dance scenario corresponds to the first decision tree model. In this case, the step of processing the attribute values of multiple modal information according to the pre-constructed decision tree model to determine the target scenario result of the video to be recognized may include: processing the attribute values of multiple modal information according to the first decision tree model to determine the first target scenario result of the video to be recognized, and the first target scenario result is used to indicate whether the video scenario of the video to be recognized is a dance scenario; the singing scenario corresponds to the second decision tree model. In this case, the step of processing the attribute values of multiple modal information according to the pre-constructed decision tree model to determine the target scenario result of the video to be recognized may include: processing the attribute values of multiple modal information according to the second decision tree model to determine the second target scenario result of the video to be recognized, and the second target scenario result is used to indicate whether the video scenario of the video to be recognized is a singing scenario.
[0068] In one embodiment, multiple modal information of the video to be recognized can be input into the first decision model and the second decision model, so that the first decision model and the second decision model perform classification decisions on the video to be recognized to obtain the final video scenario, so as to avoid the situation that a single type of modal information (such as video modal information) cannot distinguish whether the video to be recognized is a singing scenario or a dance scenario.
[0069] The following combines Figure 3 and Figure 4 to further explain the decision-making processes of the first decision tree model and the second decision tree model.
[0070] Referring to Figure 3 and Figure 4 , Figure 3 is a schematic diagram of a decision-making process of the first decision tree model, Figure 4 is a schematic diagram of a decision-making process of the second decision tree model. In Figure 3 and Figure 4 , Figure 3 and Figure 4 the 1, 2, and 3 in Figure 3For example, the decision path is A1 - A2 - A3, and the decision result is a dance scene. For Figure 4 For another example, the decision path is B1 - B2 - B3, and the decision result is a singing scene.
[0071] It should be noted that when specifically constructing the first decision tree model and the second decision tree model, the specific attribute categories of 1, 2, and 3 in Figure 3 and Figure 4 can be determined according to the information gain of the attributes of different modalities. Continuing with the examples of the attributes corresponding to the above modality information, Figure 3 1 in Figure 3 can represent the action attribute corresponding to the video modality information, Figure 3 2 in Figure 4 can represent the background music attribute corresponding to the audio modality information, Figure 3 3 in Figure 3 can represent the lyrics attribute corresponding to the text modality information.
[0072] It should be noted that the above Figure 3 and Figure 4 The examples involved represent a tree structure of the first decision tree model and the second decision tree model. This embodiment does not limit the tree structure of the first decision tree model and the second decision tree model.
[0073] Through the above method, according to the classification objectives of different video scenes, the corresponding decision tree model is applied for decision-making, improving the accuracy of video scene classification and recognition.
[0074] In some embodiments, the video to be recognized can be labeled according to the target scene result of the video to be recognized, and personalized recommendations can be made for the video to be recognized according to the labeling result of the video to be recognized.
[0075] It should be noted that labeling refers to the process of adding labels to the video to be recognized.
[0076] In some embodiments, a label representing the target scene result in characters can be added to the video to be recognized. For example, the character can be a numerical character or other characters.
[0077] In some embodiments, personalized recommendations can be made for users according to the user profile. For example, the user profile can describe the user's video clicks, video views, video likes, video comments, video forwards, etc.
[0078] In the above manner, since the target scene result of the video to be recognized is determined based on multi-modal information, the label accuracy of the video to be recognized can be improved. On this basis, personalized recommendations are made for users according to the labels of the video to be recognized, improving the recommendation efficiency.
[0079] Figure 5 is a block diagram of a device for recognizing video scenes shown according to an exemplary embodiment. Referring to Figure 5 , the device 500 for recognizing video scenes includes:
[0080] An acquisition module 501, configured to acquire a video to be recognized;
[0081] A processing module 502, configured to process the video to be recognized to obtain various modal information in the video to be recognized, where the various modal information includes visual modal information, text modal information, and audio modal information;
[0082] A determination module 503, configured to determine a target scene result of the video to be recognized according to the various modal information, where the target scene result is used to indicate whether the video scene of the video to be recognized is a target scene, and the target scene is a dance scene or a singing scene.
[0083] In some embodiments, the determination module 503 includes:
[0084] A first determination sub-module, configured to, for each type of modal information, process the type of modal information according to an attribute recognition model corresponding to the type of modal information to determine an attribute value of an attribute corresponding to the type of modal information;
[0085] A second determination sub-module, configured to process the attribute values of the various modal information according to a pre-constructed decision tree model to determine the target scene result of the video to be recognized.
[0086] In some embodiments, the decision tree model includes a first decision tree model, the target scene is a dance scene, and the second determination sub-module includes:
[0087] A first determination unit, configured to process the attribute values of the various modal information according to the first decision tree model to determine a first target scene result of the video to be recognized, where the first target scene result is used to indicate whether the video scene of the video to be recognized is the dance scene.
[0088] In some embodiments, the decision tree model includes a second decision tree model, the target scene is a singing scene, and the second determination sub-module includes:
[0089] A second determination unit, configured to process the attribute values of the multiple-modal information according to the second decision tree model, and determine a second target scene result of the video to be recognized, where the second target scene result is used to indicate whether the video scene of the video to be recognized is the singing scene.
[0090] In some embodiments, the apparatus 500 further includes:
[0091] A training sample acquisition module, configured to acquire a training set, where the training set includes multiple sample data, and each sample data includes a sample attribute, a sample attribute value, and a sample scene category;
[0092] An information gain determination module, configured to determine the information gain of the sample attributes based on the training set;
[0093] A construction module, configured to construct a decision tree model based on the information gain of each sample attribute.
[0094] In some embodiments, the apparatus 500 further includes:
[0095] A labeling module, configured to label the video to be recognized according to the target scene result of the video to be recognized;
[0096] A recommendation module, configured to perform personalized recommendation on the video to be recognized according to the labeling result of the video to be recognized.
[0097] Next, referring to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0098] As Figure 6As shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 602 or a program loaded from the storage device 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0099] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 6 an electronic device 600 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.
[0100] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0101] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0102] In some embodiments, the electronic device can communicate using any currently known or future-developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0103] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.
[0104] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a video to be recognized; process the video to be recognized to obtain various modality information in the video to be recognized, where the various modality information includes visual modality information, text modality information, and audio modality information; determine a target scene result of the video to be recognized according to the various modality information, where the target scene result is used to indicate whether the video scene of the video to be recognized is a target scene, and the target scene is a dance scene or a singing scene.
[0105] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0107] The modules described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases. For example, the acquisition module may also be described as "the module for acquiring the video to be recognized".
[0108] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0110] According to one or more embodiments of this disclosure, Example 1 provides a method for identifying a video scene, including:
[0111] Obtain the video to be identified;
[0112] Process the video to be identified to obtain various modal information in the video to be identified, where the various modal information includes visual modal information, text modal information, and audio modal information;
[0113] Determine a target scene result of the video to be identified according to the various modal information, where the target scene result is used to indicate whether the video scene of the video to be identified is a target scene, and the target scene is a dance scene or a singing scene.
[0114] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, and the determining the target scene result of the video to be identified according to the various modal information includes:
[0115] For each type of modal information, process the modal information according to an attribute recognition model corresponding to the modal information to determine an attribute value of an attribute corresponding to the modal information;
[0116] Process the attribute values of the multiple modal information according to a pre-constructed decision tree model to determine the target scene result of the video to be recognized.
[0117] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2. The decision tree model includes a first decision tree model, and the target scene is a dance scene. The process of processing the attribute values of the multiple modal information according to the pre-constructed decision tree model to determine the target scene result of the video to be recognized includes:
[0118] Process the attribute values of the multiple modal information according to the first decision tree model to determine the first target scene result of the video to be recognized, where the first target scene result is used to indicate whether the video scene of the video to be recognized is the dance scene.
[0119] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 2. The decision tree model includes a second decision tree model, and the target scene is a singing scene. The process of processing the attribute values of the multiple modal information according to the pre-constructed decision tree model to determine the target scene result of the video to be recognized includes:
[0120] Process the attribute values of the multiple modal information according to the second decision tree model to determine the second target scene result of the video to be recognized, where the second target scene result is used to indicate whether the video scene of the video to be recognized is the singing scene.
[0121] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 2. The method further includes:
[0122] Obtain a training set, where the training set includes multiple sample data, and each sample data includes a sample attribute, a sample attribute value, and a sample scene category;
[0123] Based on the training set, determine the information gain of the sample attribute;
[0124] Based on the information gain of each sample attribute, construct a decision tree model.
[0125] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1. The method further includes:
[0126] Label the video to be recognized according to the target scene result of the video to be recognized;
[0127] Perform personalized recommendation on the video to be recognized according to the labeling result of the video to be recognized.
[0128] According to one or more embodiments of the present disclosure, Example 7 provides an apparatus for identifying a video scene, including:
[0129] An acquisition module, configured to acquire a video to be identified;
[0130] A processing module, configured to process the video to be identified to obtain various modality information in the video to be identified, where the various modality information includes visual modality information, text modality information, and audio modality information;
[0131] A determination module, configured to determine a target scene result of the video to be identified according to the various modality information, where the target scene result is used to indicate whether the video scene of the video to be identified is a target scene, and the target scene is a dance scene or a singing scene.
[0132] According to one or more embodiments of the present disclosure, Example 8 provides the apparatus of Example 7, where the determination module includes:
[0133] A first determination sub-module, configured to, for each type of the modality information, process the type of the modality information according to an attribute recognition model corresponding to the type of the modality information to determine an attribute value of an attribute corresponding to the type of the modality information;
[0134] A second determination sub-module, configured to process the attribute values of the various modality information according to a pre-constructed decision tree model to determine the target scene result of the video to be identified.
[0135] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in any one of Examples 1-6 are implemented.
[0136] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, including:
[0137] A storage device, on which a computer program is stored;
[0138] A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-6.
[0139] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0140] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0141] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. With regard to the apparatus in the foregoing embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated herein.
Claims
1. A method for identifying a video scene, characterized in that Including: Obtain the video to be recognized; Process the video to be recognized to obtain various modal information in the video to be recognized, where the various modal information includes visual modal information, text modal information, and audio modal information; Determine the target scene result of the video to be recognized according to the various modal information, where the target scene result is used to indicate whether the video scene of the video to be recognized is a target scene, and the target scene is a dance scene or a singing scene; The determining the target scene result of the video to be recognized according to the various modal information includes: for each type of modal information, process the modal information according to the attribute recognition model corresponding to the type of modal information to determine the attribute value of the attribute corresponding to the type of modal information; process the attribute values of the various modal information according to a pre-constructed decision tree model to determine the target scene result of the video to be recognized. The decision tree models for different classification targets are different, and the classification target is whether the video scene is a target scene. The attributes corresponding to the visual modal information include at least one of an environment attribute, a number of people attribute, an action attribute, a microphone attribute, and an instrument attribute. The attributes corresponding to the text modal information include at least one of a lyrics attribute and a slogan attribute. The attributes corresponding to the audio modal information include at least one of a human voice attribute and a background music attribute.
2. The method according to claim 1, characterized in that, The decision tree model includes a first decision tree model, and the target scene is a dance scene. The determining the target scene result of the video to be recognized by processing the attribute values of the various modal information according to the pre-constructed decision tree model includes: Process the attribute values of the various modal information according to the first decision tree model to determine the first target scene result of the video to be recognized, where the first target scene result is used to indicate whether the video scene of the video to be recognized is the dance scene.
3. The method according to claim 1, characterized in that The decision tree model includes a second decision tree model, and the target scene is a singing scene. The determining the target scene result of the video to be recognized by processing the attribute values of the various modal information according to the pre-constructed decision tree model includes: Process the attribute values of the various modal information according to the second decision tree model to determine the second target scene result of the video to be recognized, where the second target scene result is used to indicate whether the video scene of the video to be recognized is the singing scene.
4. The method according to claim 1, characterized in that, The method further includes: Obtain a training set, where the training set includes multiple sample data, and each sample data includes a sample attribute, a sample attribute value, and a sample scene category; Based on the training set, determine the information gain of the sample attribute; Construct a decision tree model based on the information gains of the sample attributes.
5. The method according to claim 1, wherein The method further includes: Label the video to be recognized according to the target scene result of the video to be recognized; Perform personalized recommendation on the video to be recognized according to the labeling result of the video to be recognized.
6. An apparatus for recognizing a video scene, characterized in that, Including: An obtaining module, configured to obtain the video to be recognized; A processing module for processing the video to be recognized to obtain various modal information in the video to be recognized, where the various modal information includes visual modal information, text modal information, and audio modal information; A determination module for determining a target scene result of the video to be recognized according to the various modal information, where the target scene result is used to indicate whether the video scene of the video to be recognized is a target scene, and the target scene is a dance scene or a singing scene; The determination module includes: A first determination sub-module for processing each type of modal information according to an attribute recognition model corresponding to the type of modal information to determine an attribute value of an attribute corresponding to the type of modal information; A second determination sub-module for processing the attribute values of the various modal information according to a pre-constructed decision tree model to determine a target scene result of the video to be recognized. The decision tree models for different classification targets are different, and the classification target is whether the video scene is a target scene. The attributes corresponding to the visual modal information include at least one of an environment attribute, a number of people attribute, an action attribute, a microphone attribute, and an instrument attribute. The attributes corresponding to the text modal information include at least one of a lyrics attribute and a slogan attribute. The attributes corresponding to the audio modal information include at least one of a human voice attribute and a background music attribute.
7. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processing device, it implements the steps of the method according to any one of claims 1-5.
8. An electronic device, characterized in that, It includes: A storage device storing a computer program thereon; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Video quality evaluation method and device
CN111212279A
Video file classification method and device, medium and electronic equipment
CN111488489A
Video classification method and device, equipment and storage medium
CN113159010A