Video tag generation method, device, medium and vehicle
By obtaining the facial expressions and playback status information of users when they experience the video and using machine learning algorithms to generate video tags, the problem of inaccurate video tags in the existing technology is solved, more accurate video tag generation and information recommendation are achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202310136565.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-02-10
AI Technical Summary
When generating video tags, the existing technology has difficulty in accurately representing video information, especially when there are many video elements, resulting in inaccurate tags.
By obtaining the facial expression information and playback status information of users when they experience the video, video tags are generated, and machine learning algorithms are used for facial expression recognition. Video clips are extracted in combination with the playback status information, and video tags are generated through analysis on the background server.
It improves the accuracy of video tags, enhances the success rate of information recommendation, improves user experience, and reduces the load on electronic devices.
Smart Images

Figure CN116095424B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a method, device, medium and vehicle for generating a video tag. Background Art
[0002] With the development of the information age, more and more videos are appearing on the internet. To facilitate video management, corresponding video tags can be generated for each video. Video tags are keywords identified by the video content or video copy. Video tags can quickly and accurately identify the information contained in the video, thereby recommending related products or other videos to users.
[0003] Currently, video tags are often generated by performing image recognition on the video being played. However, in some cases, such as when the video contains a large number of elements, image recognition results are inaccurate, making it difficult for the generated video tags to accurately represent the video information.
[0004] Accordingly, this field requires a new technical solution to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned defects, the present invention is proposed to solve or at least partially solve the problem of inaccurate video tags.
[0006] In a first aspect, a method for generating a video tag is provided, the method comprising:
[0007] Obtain and recognize facial expressions when users experience videos;
[0008] Get the video playback status information when the expression information is recognized;
[0009] Generate video tags for the video based on the expression information and playback status information.
[0010] In one technical solution of the above-mentioned method for generating a video tag, after the step of “generating a video tag for the video based on the expression information and the playback status information”, the method further includes:
[0011] Extract the video clip when the expression information is recognized according to the playback status information;
[0012] Analyze whether the emotion of the facial expression information is positive;
[0013] If so, information recommendation is made based on the video clip;
[0014] If not, no information recommendation is made.
[0015] In one technical solution of the above-mentioned method for generating video tags, the step of "extracting the video segment when the expression information is recognized based on the playback status information" specifically includes:
[0016] Determine the video playback progress when the emoticon information is recognized based on the playback status information;
[0017] Extract the video clips when the expression information is recognized according to the playback progress.
[0018] In one technical solution of the above-mentioned method for generating video tags, the step of "recommending information based on video clips" specifically includes:
[0019] Make product recommendations based on the products included in the video clips;
[0020] and / or,
[0021] Make video recommendations based on the video content of the video clips.
[0022] In one technical solution of the above-mentioned method for generating video tags, the step of "recommending products based on products included in the video clip" specifically includes:
[0023] Send the video clip to the backend server so that the backend server can output product recommendation information when analyzing that the video clip contains products;
[0024] Receive product recommendation information output by the backend server and make product recommendations.
[0025] In one technical solution of the above-mentioned method for generating video tags, the step of "obtaining and identifying the facial expression information of the user when experiencing the video" specifically includes:
[0026] Obtain facial images of users while experiencing the video;
[0027] Perform facial expression recognition on facial images to determine the user's facial expression information.
[0028] In one technical solution of the above-mentioned method for generating video tags, the step of "performing expression recognition on a facial image to determine the user's expression information" specifically includes:
[0029] Obtain a facial expression recognition model obtained using a machine learning algorithm;
[0030] A facial expression recognition model is used to perform expression recognition on facial images to determine the user's expression information.
[0031] In one technical solution of the above-mentioned method for generating video tags, the step of "generating a video tag for a video based on the expression information and the playback status information" specifically includes:
[0032] Determining information corresponding to the expression information and the video tag based on the recognized expression information;
[0033] Generate a video tag for the video based on the expression information, the information corresponding to the video tag, and the playback status information.
[0034] In one technical solution of the above-mentioned method for generating video tags, the step of "determining information corresponding to the expression information and the video tag based on the recognized expression information" specifically includes:
[0035] determining whether the emotion of the expression information is a positive emotion according to the expression information;
[0036] If so, it is determined that the information corresponding to the expression information and the video tag is positive information;
[0037] If not, it is determined that the information corresponding to the expression information and the video tag is negative information.
[0038] In one technical solution of the above-mentioned method for generating video tags, the step of "generating a video tag for a video based on the expression information and the playback status information" specifically includes:
[0039] The expression information and the playback status information are sent to the background server so that the background server can generate a video tag according to the expression information and the playback status information.
[0040] In a second aspect, a computer device is provided, comprising a processor and a storage device, wherein the storage device is adapted to store a plurality of program codes, and the program codes are adapted to be loaded and run by the processor to execute the method for generating a video tag as described in any one of the technical solutions of the method for generating a video tag.
[0041] In a third aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored in the computer-readable storage medium, wherein the program codes are suitable for being loaded and run by a processor to execute the method for generating a video tag as described in any one of the technical solutions of the method for generating a video tag.
[0042] In a fourth aspect, a vehicle is provided, comprising the computer device described in the above-mentioned computer device technical solution.
[0043] Solution 1. A method for generating a video tag, characterized in that the method comprises:
[0044] Obtain and recognize facial expressions when users experience videos;
[0045] Acquire playback status information of the video when the expression information is recognized;
[0046] A video tag of the video is generated according to the expression information and the playback status information.
[0047] Solution 2. The method for generating a video tag according to Solution 1, characterized in that after the step of "generating the video tag of the video based on the expression information and the playback status information", the method further comprises:
[0048] Extracting the video clip when the expression information is recognized according to the playback state information;
[0049] Analyzing whether the emotion of the facial expression information is positive;
[0050] If so, recommending information based on the video clip;
[0051] If not, the information recommendation is not performed.
[0052] Solution 3. The method for generating a video tag according to Solution 2 is characterized in that the step of "extracting the video segment when the expression information is recognized according to the playback state information" specifically includes:
[0053] Determining, according to the playback state information, the playback progress of the video when the expression information is recognized;
[0054] The video segment when the expression information is recognized is extracted according to the playback progress.
[0055] Solution 4. The method for generating video tags according to Solution 2, wherein the step of "recommending information based on the video clip" specifically includes:
[0056] Recommending products based on the products included in the video clip;
[0057] and / or,
[0058] Video recommendations are made based on the video content of the video clip.
[0059] Solution 5. The method for generating video tags according to Solution 4 is characterized in that the step of "recommending products based on the products included in the video clip" specifically includes:
[0060] Sending the video clip to a backend server so that the backend server can output recommendation information of the product when analyzing that the video clip contains the product;
[0061] Receive the recommendation information of the product output by the backend server and make product recommendations.
[0062] Solution 6. The method for generating a video tag according to Solution 1 is characterized in that the step of "obtaining and identifying facial expression information of a user when experiencing the video" specifically includes:
[0063] Obtain facial images of users while experiencing the video;
[0064] Perform expression recognition on the facial image to determine the user's expression information.
[0065] Solution 7. The method for generating video tags according to Solution 6 is characterized in that the step of "performing expression recognition on the facial image to determine the user's expression information" specifically includes:
[0066] Obtain a facial expression recognition model obtained using a machine learning algorithm;
[0067] The facial expression recognition model is used to perform expression recognition on the facial image to determine the user's expression information.
[0068] Solution 8. The method for generating a video tag according to Solution 1 is characterized in that the step of "generating a video tag for the video based on the expression information and the playback status information" specifically includes:
[0069] Determining information corresponding to the expression information and the video tag based on the recognized expression information;
[0070] A video tag for the video is generated according to information corresponding to the expression information and the video tag and the playback status information.
[0071] Solution 9. The method for generating a video tag according to Solution 8 is characterized in that the step of "determining information corresponding to the expression information and the video tag based on the recognized expression information" specifically includes:
[0072] determining whether the emotion of the expression information is a positive emotion according to the expression information;
[0073] If so, determining that the information corresponding to the expression information and the video tag is positive information;
[0074] If not, it is determined that the information corresponding to the expression information and the video tag is negative information.
[0075] Solution 10. The method for generating a video tag according to Solution 1, wherein the step of "generating a video tag for the video based on the expression information and the playback status information" specifically includes:
[0076] The expression information and the play status information are sent to a background server, so that the background server can generate the video tag according to the expression information and the play status information.
[0077] Solution 11. A computer device comprising a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, characterized in that the program codes are suitable for being loaded and run by the processor to execute the method for generating video tags according to any one of Solutions 1 to 10.
[0078] Solution 12. A computer-readable storage medium storing a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the method for generating a video tag according to any one of Solutions 1 to 10.
[0079] Option 13. A vehicle, characterized in that the vehicle includes the computer device described in Option 11.
[0080] The above one or more technical solutions of the present invention have at least one or more of the following beneficial effects:
[0081] In the technical solution implemented in the present invention, users will display different facial expressions when viewing different videos. By acquiring and identifying the facial expressions of users during video viewing, it is possible to infer the user's perception of the video. For example, if a user displays a laughing expression during a video viewing, it indicates that the user likely rated the video as having a good visual experience. Using the video's playback status information and the correspondence between the user's facial expression and the video tag when the facial expression is recognized as the video tag can improve the accuracy of the video tag. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The disclosure of the present invention will become more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Among them:
[0083] Figure 1 1 is a flow chart showing the main steps of a method for generating a video tag according to an embodiment of the present invention;
[0084] Figure 2 is a main structural block diagram of an information recommendation method according to an embodiment of the present invention;
[0085] Figure 3 It is a main structural diagram of a computer device embodiment according to the present invention. DETAILED DESCRIPTION
[0086] Some embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0087] In the description of the present invention, "processor" may include hardware, software or a combination of the two. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, hardware or a combination of the two. Non-transitory computer-readable storage media include any suitable media that can store program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, etc. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B or A and B.
[0088] The following describes an embodiment of a method for generating a video tag provided by the present invention.
[0089] See attached Figure 1 , Figure 1 FIG. 1 is a flow chart showing the main steps of a method for generating a video tag according to an embodiment of the present invention. Figure 1 As shown, the method for generating a video tag in the embodiment of the present invention mainly includes the following steps S101 to S103.
[0090] Step S101: Acquire and identify facial expression information of a user when experiencing a video.
[0091] In an embodiment of the present invention, a collection device can be used to obtain user information when a user experiences a video, and then the user information can be identified to obtain the user's facial expression information. The user experience of the video includes, but is not limited to, watching the video image and / or listening to the video sound.
[0092] For example, an image of the user can be captured by an image acquisition device such as a camera, and then the image can be recognized to obtain the user's facial expression information. Alternatively, a structured camera can be used to capture a facial image of the user, and then the facial image can be recognized to obtain the user's facial expression information. Of course, other methods can also be used to obtain the user's facial expression information. The embodiments of the present invention do not limit the method for obtaining the user's facial expression information. The facial expression information may include fear, anger, happiness, smile, laughter, shock, etc.
[0093] Step S102: Acquire the video playback status information when the expression information is recognized.
[0094] Video playback status information includes the video's playback progress and / or playback mode. The video's playback progress refers to the duration of the video. For example, if the video has been playing for 5 minutes when the user's facial expression is recognized, the video's playback progress is 5 minutes. Video playback modes include silent playback, looping playback, and so on.
[0095] Step S103: Generate a video tag for the video based on the expression information and the playback status information.
[0096] The user's facial expression information corresponds to the video's playback status information when the facial expression information is recognized. Therefore, the user's facial expression information and playback status information can be used as the video tag of the video. For example, if the user's video duration is 5 minutes, the facial expression information when the user is watching the video is recognized as crying, and the video is played in silent mode, then based on the recognized crying facial expression information, the video tag of the video can be determined as: playback progress 5 minutes, sad, silent playback.
[0097] When users experience a video, their facial expressions vary depending on the video content. For example, if a user's impression of a video is positive, they may display a happy, smiling, or laughing expression, and may even loop the video. However, if a user's impression of a video is negative, they may display a crying or fearful expression, and may even play the negative video on mute. Therefore, generating video tags based on facial expressions and playback status information allows video tags to more accurately reflect the user's impression of the video, improving the accuracy of video tags.
[0098] The above step S101 is further explained below.
[0099] In one possible implementation, the step of “obtaining and identifying the facial expression information of the user when experiencing the video (the aforementioned step S101)” specifically includes the following steps 11 and 12:
[0100] Step 11: Obtain the face image of the user while experiencing the video.
[0101] Among them, the facial image of the user when experiencing the video can be obtained through image acquisition devices such as cameras.
[0102] Step 12: Perform expression recognition on the face image to determine the user's expression information.
[0103] Expression recognition on facial images can be performed using a template matching-based expression recognition method, a probability model-based expression recognition method, or other methods. Of course, other methods can also be used to perform expression recognition on facial images. The embodiments of the present invention do not specifically limit the method of facial expression recognition.
[0104] Template matching-based facial expression recognition is a statistical recognition method. To recognize facial expressions in facial images, you first need to prepare different facial expressions as template images. By matching the facial image with the template images, you can determine the facial expression in the image.
[0105] Facial expression recognition methods based on probabilistic models require setting multiple feature points and a reference point for facial images with different expressions. By calculating the distance between the feature points and the reference point, the features of the faces with different expressions can be obtained. The method then selects the same feature points and the reference point for the face image to be recognized and measures the distance to obtain the features of the face image to be recognized. The probability that the features of the face image to be recognized match the features of faces with different expressions is calculated, and the expression with the highest probability of matching is selected as the expression of the face image to be recognized.
[0106] Based on the method described in steps 11 to 12 above, by obtaining facial images of users experiencing the video and performing facial expression recognition on the facial images, the facial expression information of the users experiencing the video can be accurately obtained.
[0107] In order to improve the accuracy and robustness of facial expression recognition, machine learning algorithms can be used to perform facial expression recognition on facial images.
[0108] In one possible implementation, the step of “performing expression recognition on the facial image to determine the user’s expression information (the aforementioned step 12)” specifically includes the following steps 21 to 22:
[0109] Step 21: Obtain a facial expression recognition model obtained using a machine learning algorithm.
[0110] Specifically, in this embodiment, expression image samples in different environments can be obtained, and a machine learning algorithm can be adopted and used to train a facial expression recognition model using these expression image samples, so that the trained facial expression recognition model has the ability to accurately recognize expression information from images in different environments. It should be noted that in this embodiment, a conventional model training method in the field of machine learning algorithm technology can be used to train the facial expression recognition model, and the embodiment of the present invention does not specifically limit this. For example, the expression image samples are input into the facial expression recognition model, the loss value of the model is calculated by forward propagation, the parameter gradient of the model parameters is calculated according to the loss value, and the model parameters are updated according to the parameter gradient back propagation until the facial expression recognition model meets the convergence condition and the training is stopped.
[0111] Step 22: Use a facial expression recognition model to perform expression recognition on the facial image to determine the user's expression information.
[0112] By inputting the facial image into the trained facial expression recognition model, the user's expression information can be obtained.
[0113] Based on the method described in steps 21 to 22 above, the facial expression recognition model can be trained using expression image samples in different environments. In this way, when the trained facial expression recognition model is used to recognize expression information, the expression information of users in different environments can be accurately obtained, thereby improving the robustness of expression information recognition.
[0114] The above is an explanation of step S101 , and the following is a further explanation of step S103 .
[0115] In one possible implementation, the step of “generating a video tag for the video based on the expression information and the playback status information (the aforementioned step S103)” specifically includes the following steps 31 to 32:
[0116] Step 31: Determine information corresponding to the expression information and the video tag based on the recognized expression information.
[0117] The information corresponding to the expression information and the video tag includes user emotion information determined based on the user's expression information, the user's preference for the video, and other information.
[0118] For example, if the recognized facial expression is smiling, the corresponding information of the facial expression and the video tag is determined to be happy, and the degree of liking the video is relatively favorable. If the recognized facial expression is laughing, the corresponding information of the facial expression and the video tag is determined to be happy, and the degree of liking the video is very favorable. If the recognized facial expression is crying, the corresponding information of the facial expression and the video tag is determined to be sad, and the degree of liking the video is relatively unfavorable.
[0119] Step 32: Generate a video tag for the video based on the expression information, the information corresponding to the video tag, and the playback status information.
[0120] The video tag of the video is generated by the method of steps 31 to 32 above. First, based on the facial expression information, the user's liking for the video, the user's emotions when watching the video, and other information corresponding to the video tag can be determined. Then, based on the facial expression information and the information corresponding to the video tag and the playback status information, the video tag of the video can be accurately generated.
[0121] In one possible implementation, the step of “determining information corresponding to the expression information and the video tag based on the recognized expression information (the aforementioned step 31)” specifically includes:
[0122] determining whether the emotion of the expression information is a positive emotion according to the expression information;
[0123] If so, it is determined that the information corresponding to the expression information and the video tag is positive information;
[0124] If not, it is determined that the information corresponding to the expression information and the video tag is negative information.
[0125] When the user's expression is smiling, laughing, etc., it means that the user's emotion is positive, that is, the user is in a good mood when watching the video, so it can be determined that the information corresponding to the video tag is positive information. Conversely, when the user's expression is crying, it means that the user's emotion is negative. Similarly, the user is in a depressed mood when watching the video, so it can be determined that the information corresponding to the video tag is negative information.
[0126] When users are watching videos, the performance of the electronic devices used to play the videos is limited. If tasks with too much workload are performed, the device may freeze or even crash, seriously affecting the user experience.
[0127] In a possible implementation, the step of “generating a video tag for the video based on the expression information and the playback status information (the aforementioned step S103)” specifically includes:
[0128] The expression information and the playback status information are sent to the background server. The background server can obtain the playback progress of the video according to the playback status information. The background server can generate video tags according to the expression information and the playback status information.
[0129] The backend server may be a cloud server. After receiving the user's expression information and the video playback progress, the backend server may generate a video tag including the expression information and the playback progress.
[0130] By sending expression information and playback status information to the background server, the background server generates video tags. Using a more powerful background server to complete the task of generating video tags can increase the speed of generating video tags and reduce the load on electronic devices used to play videos, so that the video can be played smoothly, ensuring the user experience when experiencing the video.
[0131] On the other hand, when users are in a more positive mood, they are more friendly to other information. For example, users are more willing to click on advertising links to purchase the recommended products or experience other recommended videos. In order to improve the user experience of the video, information recommendations can be made when the user is in a positive mood.
[0132] See attached Figure 2 After the step of “generating a video tag for the video based on the expression information and the playback status information (the aforementioned step S103)”, the method further includes the following steps S201 to S204 for information recommendation:
[0133] Step S201: extracting the video segment when the expression information is recognized according to the playback status information.
[0134] The video playback status information includes the video playback progress and the video playback mode. The video segment when the expression information is recognized can be determined based on the recognized video playback status information.
[0135] Step S202: Analyze whether the emotion of the facial expression information is a positive emotion. If so, execute step S203; if not, execute step S204.
[0136] Because video tags contain facial expressions, we can use this information to determine whether the user's emotion is positive. To determine whether a user's emotion is positive or negative, we can pre-categorize different facial expressions. For example, smiling and laughing are positive, while crying and fear are negative. Based on the label information in the video tags, we can determine the category to which the facial expression belongs, thereby determining whether the user's emotion is positive.
[0137] Step S203: Recommend information based on the video clip.
[0138] Playback progress refers to the duration of video playback, so video clips can be obtained based on the video's playback progress. Information recommendation refers to recommending other information, such as recommending other videos or advertising information.
[0139] Step S204: No information recommendation is performed.
[0140] Based on the method described in steps S201 to S204 above, the user's emotions can be determined based on the facial expression information contained in the video tags, and information can be recommended when the user's emotions are positive. This not only improves the success rate of information recommendation, but also ensures the user's video experience, and avoids further aggravation of the user's negative emotions due to recommended information when the user is in a negative mood.
[0141] In a possible implementation, the step of “extracting the video segment when the expression information is recognized according to the playback state information (the aforementioned step S201)” specifically includes:
[0142] The playback progress of the video when the expression information is recognized is determined based on the playback status information.
[0143] The video playback status information includes information such as the video playback progress and the video playback mode, so the video playback progress when the expression information is recognized can be determined based on the playback status information.
[0144] Extract the video clips when the expression information is recognized according to the playback progress.
[0145] After obtaining the playback progress when the facial expression information is recognized, the video segment at the time when the facial expression information is recognized can be extracted from the video according to the playback progress. For example, if a smiling facial expression is recognized between the 5th and 6th minutes of playback, the video segment between the 5th and 6th minutes of playback can be extracted from the playing video as the video segment at the time when the facial expression information is recognized.
[0146] When users are experiencing a video, they are more likely to be interested in the products in the video or information related to the video. In order to improve the success rate of recommendations, products or videos related to the content of the video can be recommended to users.
[0147] In a possible implementation, the step of “recommending information based on the video clip (the aforementioned step S203)” specifically includes:
[0148] Make product recommendations based on the products included in the video clips;
[0149] For example, if a certain brand of jacket appears in a video clip, then when making product recommendations, the jacket of this brand can be recommended to the user.
[0150] In another possible implementation, the step of “recommending information based on the video clip (the aforementioned step S203)” specifically includes:
[0151] Make video recommendations based on the video content of the video clips.
[0152] For example, if a certain star appears in a video clip, then when making video recommendations, other film and television works of the star can be recommended to the user.
[0153] By recommending related products based on the products contained in the video clip, the probability of users being interested in the products is increased, thereby improving the success rate of product recommendations. Similarly, by recommending videos based on the video content of the video clip, the probability of users being interested in highly relevant videos is greater, and the success rate of video recommendations is higher.
[0154] In a possible implementation, the step of “recommending products based on the products included in the video clip” specifically includes the following steps 31 and 32:
[0155] Step 31: Send the video clip to the backend server so that the backend server can output product recommendation information when analyzing that the video clip contains a product.
[0156] The method for the backend server to analyze whether a video clip contains goods can adopt traditional image recognition methods, such as the average hash algorithm, the perceptual hash algorithm, or can also adopt deep learning image recognition methods, such as using a convolutional neural network for image recognition or using a Bayesian probability generation network for image recognition. Of course, other network models can also be used for image recognition. The embodiment of the present invention does not specifically limit the method for analyzing the goods contained in the video clip.
[0157] Step 32: Receive the product recommendation information output by the backend server and make product recommendations.
[0158] Product recommendations can be made by popping up product images or playing product video ads on the video playback page. When the user clicks on the product image or clicks on the product video ad, the product details page can be opened to complete the product recommendation.
[0159] Based on the method described in steps 31 to 32 above, by sending the video clip to the background server, the background server completes the analysis of the goods contained in the video clip, thereby improving the efficiency of analyzing the goods contained in the video clip and reducing the load on the electronic device used to play the video.
[0160] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effects of the present invention, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present invention.
[0161] Those skilled in the art will appreciate that all or part of the processes in the method for implementing the above-mentioned embodiment of the present invention may also be accomplished by instructing the relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, it may implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal, and software distribution medium capable of carrying the computer program code. It should be noted that the content contained in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0162] Furthermore, the present invention also provides a computer device.
[0163] See attached Figure 3 , Figure 3 FIG. 1 is a schematic diagram of the main structure of a computer device according to an embodiment of the present invention. Figure 3 As shown, the computer device in the embodiment of the present invention mainly includes a storage device 31 and a processor 32. The storage device 31 can be configured to store a program for executing the method for generating video tags in the above-mentioned method embodiment, and the processor 32 can be configured to execute the program in the storage device, which includes but is not limited to the program for executing the method for generating video tags in the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present invention.
[0164] In the embodiment of the present invention, the computer device may be a control device device formed by various electronic devices. In some possible implementations, the computer device may include multiple storage devices 31 and multiple processors 31. The program for executing the method for generating video tags of the above method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute different steps of the method for generating video tags of the above method embodiment. Specifically, each subroutine can be stored in a different storage device 31, and each processor 32 can be configured to execute the program in one or more storage devices 31 to jointly implement the method for generating video tags of the above method embodiment, that is, each processor 32 executes different steps of the method for generating video tags of the above method embodiment to jointly implement the method for generating video tags of the above method embodiment.
[0165] The multiple processors 32 may be processors deployed on the same device. For example, the computer device may be a high-performance device composed of multiple processors, and the multiple processors 32 may be processors configured on the high-performance device. Furthermore, the multiple processors 32 may be processors deployed on different devices. For example, the computer device may be a server cluster, and the multiple processors 32 may be processors on different servers in the server cluster.
[0166] Furthermore, the present invention also provides a computer-readable storage medium.
[0167] In an embodiment of a computer-readable storage medium according to the present invention, the computer-readable storage medium can be configured to store a program for executing the method for generating a video tag of the above-mentioned method embodiment. The program can be loaded and executed by a processor to implement the above-mentioned method for generating a video tag. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present invention. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present invention is a non-transitory computer-readable storage medium.
[0168] Furthermore, the present invention also provides a vehicle.
[0169] In an embodiment of a vehicle according to the present invention, the vehicle may include the aforementioned computer device. The types of vehicles include, but are not limited to, internal combustion engine vehicles, new energy vehicles, and hybrid vehicles.
[0170] Thus far, the technical solution of the present invention has been described in conjunction with an embodiment shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A method for generating a video tag, characterized in that: The method comprises: Obtain and recognize facial expressions when users experience videos; Acquire playback status information of the video when the expression information is recognized; Generate a video tag for the video based on the expression information and the playback status information; the playback status information includes the playback progress of the video and the playback mode of the video, the playback progress of the video refers to the duration of video playback, and the playback mode of the video includes silent playback and loop playback; the expression information and the playback status information of the video when the expression information is recognized have a one-to-one correspondence, and generating the video tag for the video specifically includes using the expression information and the playback status information as the video tag of the video; Extracting the video clip when the expression information is recognized according to the playback state information; Analyzing whether the emotion of the facial expression information is positive; If so, recommending information based on the video clip; If not, the information recommendation is not performed.
2. The method for generating a video tag according to claim 1, wherein: The step of “extracting the video segment when the expression information is recognized according to the playback state information” specifically includes: Determining, according to the playback state information, the playback progress of the video when the expression information is recognized; The video segment when the expression information is recognized is extracted according to the playback progress.
3. The method for generating a video tag according to claim 1, wherein: The step of "recommending information based on the video clip" specifically includes: Recommending products based on the products included in the video clip; and / or, Video recommendations are made based on the video content of the video clip.
4. The method for generating a video tag according to claim 3, wherein: The step of "recommending products based on the products included in the video clip" specifically includes: Sending the video clip to a backend server so that the backend server can output recommendation information of the product when analyzing that the video clip contains the product; Receive the recommendation information of the product output by the backend server and make product recommendations.
5. The method for generating a video tag according to claim 1, wherein: The steps of "obtaining and identifying facial expressions when users experience a video" specifically include: Obtain facial images of users while experiencing the video; Perform expression recognition on the facial image to determine the user's expression information.
6. The method for generating a video tag according to claim 5, wherein: The step of “performing expression recognition on the facial image to determine the user’s expression information” specifically includes: Obtain a facial expression recognition model obtained using a machine learning algorithm; The facial expression recognition model is used to perform expression recognition on the facial image to determine the user's expression information.
7. The method for generating a video tag according to claim 1, wherein: The step of “generating a video tag for the video according to the expression information and the playback status information” specifically includes: Determining information corresponding to the expression information and the video tag based on the recognized expression information; A video tag for the video is generated according to information corresponding to the expression information and the video tag and the playback status information.
8. The method for generating a video tag according to claim 7, wherein: The step of “determining information corresponding to the expression information and the video tag based on the recognized expression information” specifically includes: determining whether the emotion of the expression information is a positive emotion according to the expression information; If so, determining that the information corresponding to the expression information and the video tag is positive information; If not, it is determined that the information corresponding to the expression information and the video tag is negative information.
9. The method for generating a video tag according to claim 1, wherein: The step of “generating a video tag for the video according to the expression information and the playback status information” specifically includes: The expression information and the play status information are sent to a background server, so that the background server can generate the video tag according to the expression information and the play status information.
10. A computer device comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, wherein: The program code is suitable for being loaded and run by the processor to execute the method for generating a video tag according to any one of claims 1 to 9.
11. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the method for generating a video tag according to any one of claims 1 to 9.
12. A vehicle, characterized in that: The vehicle includes the computer device of claim 10.
Citation Information
Patent Citations
Video processing method, device and system
CN104837059A
Video processing method and device
CN106792170A
Content recommendation method and device, model training method and device, equipment and storage medium
CN113515702A