Bullet screen generation method and device, electronic equipment and storage medium

By collecting user facial data in real time and analyzing emotional states to generate barrage content, the problems of barrage delay and inaccurate emotional expression in existing technologies are solved, and barrage generation with high real-time performance and high emotional expression is achieved.

CN120786128APending Publication Date: 2025-10-14SHANGHAI ZHONG YUAN NETWORK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510974869.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing barrage generation technology requires manual operation, which results in the barrage content being greatly affected by subjective factors such as the user's expressive ability, resulting in delays and inability to accurately express the user's emotional state.

Method used

The user's facial data is collected in real time through the camera equipment, the facial action unit combination is analyzed, the user's emotional state data is determined based on the facial action unit combination, and the barrage content that matches it is generated.

Benefits of technology

It achieves a high degree of matching between the barrage content and the user's emotional state, improves the real-time nature and emotional expression rate of the barrage, and conveys the user's true emotional state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120786128A_ABST
    Figure CN120786128A_ABST
Patent Text Reader

Abstract

The invention relates to a bullet screen generation method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting the current face data of a watching user in real time through camera equipment in a process of playing a target media content, analyzing a face action unit combination from the face data, and generating a bullet screen. And determining current emotional state data of the watching user according to the facial action unit combination, and generating and displaying bullet screen content in real time according to the emotional state data. The bullet screen content highly matched with the current emotional state of the user can be generated in real time according to the objective face data of the user, the physiological signals are dynamically converted into the bullet screen content, the bullet screen real-time performance and the emotional expression rate are improved, and the real emotional state of the user is transmitted.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and particularly relates to a method and device for generating bullet screen, electronic equipment and storage medium. BACKGROUND

[0002] With the continuous development of Internet media content, as a real-time interactive form, bullet screen has become a core carrier for users to express emotions and participate in content co-creation, which can not only make users feel the atmosphere of synchronous viewing and collective interaction, but also enhance the emotional involvement in the content. The existing bullet screen generation technology brings certain convenience to users in generating bullet screen.

[0003] However, the existing bullet screen generation technology needs manual operation, and the bullet screen content is greatly affected by subjective factors such as the expression ability of users, resulting in delay of bullet screen and inaccurate expression of the emotional state of users. SUMMARY

[0004] The present application provides a method and device for generating bullet screen, electronic equipment and storage medium, to solve the technical problem that the existing technology needs manual operation, and the bullet screen content is greatly affected by subjective factors such as the expression ability of users, resulting in delay of bullet screen and inaccurate expression of the emotional state of users.

[0005] In a first aspect, the present application provides a method for generating bullet screen, which comprises:

[0006] In the process of playing target media content, the facial data of a viewing user is collected by a camera device;

[0007] The facial action unit combination is parsed from the facial data;

[0008] According to the facial action unit combination, the emotional state data of the viewing user is determined;

[0009] According to the emotional state data, the bullet screen content is generated and displayed.

[0010] In a possible implementation, the facial action unit combination is parsed from the facial data, which comprises:

[0011] The action intensity parameter set of the target facial feature region is parsed from the facial data;

[0012] According to the preset facial action coding system rule, the action intensity parameter set is mapped into the facial action unit combination.

[0013] In a possible implementation, the emotional state data of the viewing user is determined according to the facial action unit combination, which comprises:

[0014] searching, from a preset emotion state model, emotion state data corresponding to the facial action unit, wherein the emotion state model comprises a mapping relationship between a facial action unit combination and emotion state data;

[0015] determining the searched emotion state data as current emotion state data of the watching user.

[0016] In a possible implementation, the generating and displaying of the barrage content according to the emotion state data comprises:

[0017] obtaining visual scene content of a currently played media content segment;

[0018] generating and displaying the barrage content based on the emotion state data and the visual scene content.

[0019] In a possible implementation, the generating and displaying of the barrage content based on the emotion state data and the visual scene content comprises:

[0020] extracting a scene keyword from the visual scene content, and obtaining an emotion state adjective corresponding to the emotion state data;

[0021] generating and displaying the barrage content according to the scene keyword and the emotion state adjective.

[0022] In a possible implementation, the generating and displaying of the barrage content according to the emotion state data comprises:

[0023] obtaining visual scene content of a currently played media content segment;

[0024] determining a confidence degree corresponding to the emotion state data according to the visual scene content;

[0025] generating and displaying the barrage content in a case where the confidence degree meets a set condition.

[0026] In a possible implementation, the determining of the confidence degree corresponding to the emotion state data according to the visual scene content comprises:

[0027] extracting a visual scene feature vector from the visual scene content, and generating an emotion state feature vector of the emotion state data;

[0028] performing similarity calculation on the visual scene feature vector and the emotion state feature vector, and determining the confidence degree corresponding to the emotion state data according to a calculated similarity value;

[0029] wherein the confidence degree corresponding to the emotion state data is positively correlated with the similarity value.

[0030] In a second aspect, the present application provides a barrage generation device, the device comprising:

[0031] a data collection module, configured to collect current facial data of a watching user through a camera device in a process of playing target media content;

[0032] a data analysis module, configured to analyze facial action unit combination from the facial data;

[0033] an emotional state determination module, configured to determine current emotional state data of the watching user according to the facial action unit combination;

[0034] a barrage generation module, configured to generate and display barrage content according to the emotional state data.

[0035] In a possible implementation, the data analysis module is specifically configured to:

[0036] analyze action intensity parameter set of target facial feature region from the facial data;

[0037] map the action intensity parameter set to facial action unit combination according to preset facial action coding system rule.

[0038] In a possible implementation, the emotional state determination module is specifically configured to:

[0039] find emotional state data corresponding to the facial action unit combination from preset emotional state model, wherein the emotional state model comprises mapping relationship between facial action unit combination and emotional state data;

[0040] determine the found emotional state data as the current emotional state data of the watching user.

[0041] In a possible implementation, the barrage generation module comprises:

[0042] a scene content acquisition unit, configured to acquire visual scene content of a currently played media content segment;

[0043] a barrage generation unit, configured to generate and display barrage content based on the emotional state data and the visual scene content.

[0044] In a possible implementation, the barrage generation unit is specifically configured to:

[0045] extract scene keywords from the visual scene content, and acquire emotional state adjectives corresponding to the emotional state data;

[0046] According to the scene keyword and the emotional state adjective, generate and display the barrage content.

[0047] In a possible implementation, the barrage generation module comprises:

[0048] a content acquisition unit configured to acquire visual scene content of a currently played media content segment;

[0049] a confidence determination unit configured to determine a confidence corresponding to the emotional state data according to the visual scene content;

[0050] a barrage generation unit configured to generate and display the barrage content in a case where the confidence meets a set condition.

[0051] In a possible implementation, the confidence determination unit is specifically configured to:

[0052] extract a visual scene feature vector from the visual scene content, and generate an emotional state feature vector of the emotional state data;

[0053] perform similarity calculation on the visual scene feature vector and the emotional state feature vector, and determine the confidence corresponding to the emotional state data according to a calculated similarity value;

[0054] wherein the confidence corresponding to the emotional state data is positively correlated with the similarity value.

[0055] In a third aspect, the present application provides an electronic device, comprising: a processor and a memory, the processor is used to execute the barrage generation program stored in the memory, to realize the barrage generation method in any one of the first aspect.

[0056] In a fourth aspect, the present application provides a storage medium, the storage medium stores one or more programs, the one or more programs can be executed by one or more processors to realize the barrage generation method in any one of the first aspect.

[0057] Compared with the prior art, the above technical solution provided by the embodiments of the present application has the following advantages: the method provided by the embodiments of the present application, in the process of playing the target media content, real-time collects the current facial data of the watching user through the camera device, parses the facial action unit combination from the facial data, determines the current emotional state data of the watching user according to the facial action unit combination, and then generates and displays the barrage content in real time according to the emotional state data. The barrage content generated in real time according to the objective facial data of the user is highly matched with the current emotional state of the user, which realizes the dynamic conversion of the physiological signal into the barrage content, improves the real-time performance of the barrage and the emotional expression rate, and transmits the real emotional state of the user. BRIEF DESCRIPTION OF DRAWINGS

[0058] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required by the embodiments or prior art description will be briefly introduced as follows. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without any creative effort.

[0060] One or more embodiments are illustrated by way of example in the drawings that are not intended to be limiting of the application, and the same or similar reference numerals designate similar or like elements throughout the several views of the drawings, as readily understood in the art. The drawings of the application are not to scale as required by the patent statutes, and in some instances, various drawings can have been manipulated to more clearly convey the principles of the application.

[0061] Figure 1 An embodiment flow chart of a barrage generation method provided by the embodiments of the present application;

[0062] Figure 2 An embodiment flow chart of another barrage generation method provided by the embodiments of the present application;

[0063] Figure 3 An embodiment flow chart of still another barrage generation method provided by the embodiments of the present application;

[0064] Figure 4 A structural block diagram of a barrage generation device provided by the embodiments of the present application;

[0065] Figure 5 A structural schematic diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0066] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort fall within the scope of protection of the present application.

[0067] The following disclosure provides many different embodiments, or examples, for implementing different structures of the present application. For the purpose of simplifying the present application, the components and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present application. In addition, the present application can repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not in itself indicate a relationship between the various embodiments and / or settings being discussed.

[0068] To solve the technical problems in the prior art that manual operation is required, the bullet screen content is greatly affected by subjective factors such as the expression ability of the user, the bullet screen is delayed, and the bullet screen content cannot accurately express the emotional state of the user, the present application provides a bullet screen generation method and device, electronic equipment and storage medium, which can generate bullet screen content highly matched with the current emotional state of the user according to the objective facial data of the user in real time, realize dynamic conversion of physiological signals into bullet screen content, improve the real-time performance of the bullet screen and the emotional expression rate, and make the generated bullet screen content can convey the real emotional state of the user.

[0069] In an embodiment, the execution subject of the embodiment of the present application is electronic equipment. The electronic equipment can be a hardware device or software that supports network connection to provide various network services. When the equipment is hardware, it can support various electronic equipment with a display screen, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, servers, etc. When the equipment is software, it can be installed in the above-mentioned listed electronic equipment.

[0070] Figure 1 An embodiment flowchart of a bullet screen generation method provided by the embodiment of the present application is shown in Figure 1 As shown in the figure, the method comprises the following steps:

[0071] Step 101, in the process of playing target media content, the current facial data of the watching user is collected by a camera device.

[0072] The target media content refers to the media content currently watched by the user, wherein the media content can be related content of films, television series, videos, etc. spread through various media channels.

[0073] In the embodiment of the present application, in order to obtain the facial expression of the user in real time, the electronic equipment connected with the camera device is used to play the target media content for the user to watch, and in the process of playing the target media content, the facial data sequence of the watching user is collected by the camera device. For example, the electronic equipment such as a smart TV or a tablet computer can be used to play the target media content, and the present application does not limit the playing way and playing device of the target media content.

[0074] The camera device refers to a tool for capturing dynamic images. In the embodiments of the present application, the camera device can be a camera module integrated on the electronic device playing the target media content, or a camera connected in communication with the execution subject of the embodiments of the present application. For example, the camera device in the embodiments of the present application can be a TureDepth camera, which can scan the user's face, capture facial motion information, and then collect facial expressions and motions, and obtain facial data of the user.

[0075] Further, the facial data obtained by the TureDepth camera refers to the facial mixed shape data obtained after the TureDepth camera captures the facial motion of the user in real time. For example, the facial mixed shape data can be regarded as a dictionary, which contains a series of key-value pairs, wherein the key refers to a representative motion of the face, and the value refers to a floating point number between 0 and 1, and the larger the value, the larger the motion amplitude of the face corresponding to the key. For example, the larger the value corresponding to the key of smiling mouth corner, the larger the amplitude of the user's smile.

[0076] In an embodiment, in the process of playing the target media content, the facial data of the watching user is collected by the camera device. Specifically, the current facial motion of the user is captured in real time by the camera device to obtain the real-time facial data of the watching user watching the media content segment.

[0077] Step 102, parse the facial motion unit combination from the facial data.

[0078] The facial motion unit is a standardized classification of human facial muscle movement based on FACS (Facial Action Coding System), which is a quantitative indicator corresponding to a specific facial muscle movement, and is used to accurately describe the basic constituent unit of facial expression.

[0079] For example, 52 basic AUs (Action Units) can be applied to emotional analysis of the facial motion of the user. Specifically, it includes AU1-AU52, a total of 52 facial motion units, each of which represents a different facial muscle movement, for example, AU1 represents "inner eyebrow up", AU12 represents "mouth corner up", etc. These facial motion units can be used alone or in combination to describe and analyze various facial expressions produced by the user when watching the media content.

[0080] In an embodiment, the corresponding facial action unit combination is parsed from the facial data, and the facial data includes the action amplitude of each feature point of the user's face. The facial feature point data that can be used to identify the user's emotional state is extracted from the facial data, and is mapped to the facial action unit in the facial action coding system, so as to determine the emotional state of the user according to the facial action unit subsequently.

[0081] The facial action unit can be a quantitative index of a specific facial muscle movement in the facial coding system, which is mapped to the facial mixed shape data collected by the camera. Exemplarily, the value of the facial action unit is a floating point number of 0.1-1.

[0082] In an embodiment, the specific implementation of parsing the corresponding facial action unit combination from the facial data is: parsing the action intensity parameter set of the target facial feature region from the facial data, and mapping the action intensity parameter set to the facial action unit combination according to the preset facial action coding system rule.

[0083] The target facial feature region refers to the facial feature region that is focused on in the facial action analysis and emotion calculation process of the embodiment. Exemplarily, the target facial feature can refer to the eyebrows, eyes, mouth, nose, and cheeks and chin. After the facial mixed shape data of these target facial feature regions is screened from the facial data, data support is provided for subsequent facial action analysis.

[0084] The action intensity parameter set refers to the facial mixed shape data of the target facial feature region in the facial data. These data cover the representative action of the target facial feature region, can comprehensively capture the facial muscle action related to emotion, and can accurately reflect the external performance of various emotions of the user, thereby providing rich raw data support for subsequent emotion analysis.

[0085] Exemplarily, the facial mixed shape data of the target facial feature region is first screened from the facial data collected by the camera, and then is parsed and mapped to the corresponding AU combination in the FACS according to the set rule in the facial action coding system, and each AU combination includes at least one AU. For example, the [AU1+AU3] combination includes two facial action units AU1 and AU3, wherein AU1 represents "inner eyebrow up", and AU3 represents "eyebrow lowerer and closer", the [AU5] combination includes one facial action unit AU5, and represents "levator muscle of upper eyelid movement causes upper eyelid to lift and pull back, thereby causing eyes to open wide", and the like.

[0086] For example, the facial action unit combination corresponding to the qualified facial mixed shape data "mouthSmileLeft" and "mouthSmileRight" in the facial data is [AU12] combination, wherein AU12 represents "smile".

[0087] Step 103, determining the current emotional state data of the watching user according to the facial action unit combination.

[0088] The emotional state data can be used to reflect the current emotion or mood of the user, and can also reflect the intensity of the current emotion or mood of the user. For example, the current emotional state data of the user is [happy: 0.9], which means that the emotional intensity of the current happy mood of the user is very high, that is, the user is very happy.

[0089] The facial action unit combination can reflect the muscle movement of the facial feature region of the user, and can also reflect the movement amplitude. Then, the current emotional state of the user can be analyzed according to the muscle movement of the facial feature region of the user and the movement amplitude. For example, the mapping relationship between the emotional state data and the facial action unit combination can also be pre-set, so as to determine the current emotional state of the user according to the current facial action unit combination of the user obtained.

[0090] In an embodiment, an example implementation of determining the current emotional state data of the watching user according to the facial action unit combination includes: searching for the emotional state data corresponding to the current obtained facial action unit combination of the user from the pre-set mapping relationship between the emotional state data and the facial action unit combination.

[0091] For example, it is assumed that the current obtained facial action unit combination is [AU6+AU12] combination, and the emotional state data corresponding to [AU6+AU12] combination is found to be [happy: 0.9] from the pre-set emotional state data. Then, [happy: 0.9] is determined as the current emotional state data of the watching user.

[0092] Step 104, generating and displaying the bullet screen content according to the emotional state data.

[0093] In an embodiment, the barrage content is generated and displayed according to the emotional state data. Specifically, the emotional state data of the user is obtained, the current emotion of the user is determined according to the emotional state data of the user, and then the barrage content capable of expressing the real emotion of the user is generated according to the current emotion of the user.

[0094] For example, the current emotional state data of the user is obtained: [happy: 0.9], which indicates that the user is very happy at present. At this time, a barrage content can be generated and displayed according to the emotional state data of the user: "I am so happy, ha ha ha". This is just an example. How to generate and display the barrage content according to the emotional state data will be described in the following embodiments.

[0095] In addition, in the embodiments of the present application, after the barrage content is generated according to the emotional state data, the face action (such as blinking, raising eyebrows) or head action of the user can be obtained and analyzed to trigger the sending of the barrage operation. In response to the face action of the user, the pre-generated barrage content is sent. This way can reduce the operation threshold and facilitate the participation of disabled users (such as hand disabled and visually impaired) in the barrage interaction.

[0096] The method provided by the embodiments of the present application can collect the current face data of the watching user in real time through the camera device in the process of playing the target media content, parse the face action unit combination from the face data, determine the current emotional state data of the watching user according to the face action unit combination, and then generate and display the barrage content in real time according to the emotional state data. The barrage content can be generated in real time according to the objective face data of the user, and the barrage content highly matched with the current emotional state of the user can be generated in real time. The physiological signal is dynamically converted into the barrage content, the real-time performance and the emotional expression rate of the barrage are improved, and the real emotional state of the user is transmitted.

[0097] The method provided by the embodiments of the present application can be applied not only to the media content watching scene, but also to other adaptive scenes, such as education and medical scenes, to help doctors or teachers better understand the emotional state of patients and students, especially for some disabled users, to realize barrier-free interaction.

[0098] Figure 2 Another embodiment flowchart of the barrage generation method provided by the embodiments of the present application is shown in Figure 2 The embodiment is based on Figure 1 The embodiment is based on

[0099] Step 201, in the process of playing the target media content, the current face data of the watching user is collected through the camera device.

[0100] Step 202, the face action unit combination is parsed from the face data.

[0101] For steps 201-202, see Figure 1 The description of the related embodiments.

[0102] Step 203, find the emotional state data corresponding to the facial action unit from the preset emotional state model, wherein the emotional state model includes a mapping relationship between facial action unit combination and emotional state data.

[0103] Step 204, determine the found emotional state data as the current emotional state data of the viewing user.

[0104] For steps 203-204, the following is a unified description:

[0105] The emotional state model can be a pre-constructed mapping system for establishing a corresponding relationship between facial action unit combination and human emotional state, and can also be other pre-trained large language models capable of identifying emotional state reflected by facial action unit, etc. The embodiments of the present application do not limit this.

[0106] The emotional state model is a bridge connecting user facial action and emotional content, and provides the emotional perception ability for the bullet screen generation system by establishing the mapping relationship between facial action unit and emotion. This model is not only limited to the video viewing scene, but also can be applied to emotional analysis in the fields of education, medical treatment, etc. Real emotions are expressed through bullet screens to enhance the user's interaction and expression experience.

[0107] In an embodiment, the emotional state data corresponding to the facial action unit is found from the preset emotional state model, and the found emotional state data is determined as the current emotional state data of the viewing user. Specifically, the above-mentioned emotional state model includes a mapping relationship between facial action unit combination and emotional state data. According to the current acquired facial action unit combination and the pre-set emotional state model, the corresponding emotion or mood is determined, and then the intensity of the current user's emotion is calculated using each facial action unit in the facial action unit combination. The obtained user's emotion and emotional intensity are combined to determine the current emotional state data of the user.

[0108] For example, the facial action unit combination of the current user is [AU6+AU12] combination, the emotional state model is found, the emotion corresponding to the [AU6+AU12] combination is happiness, then according to the specific floating point value 0.7 of AU6 and the specific floating point value 0.9 of AU12, the emotional intensity is calculated according to the set weighted summation mode, and the corresponding emotional intensity is 0.9. Finally, the emotional state data corresponding to the facial action unit combination [AU6+AU12] combination is [happy: 0.9].

[0109] Step 205, acquire visual scene content of the currently played media content segment.

[0110] Step 206, extract scene keywords from the visual scene content, and acquire emotion state adjectives corresponding to the emotion state data.

[0111] Step 207, generate and display the bullet screen content according to the scene keywords and the emotion state adjectives.

[0112] The following is a unified description of steps 205-207:

[0113] The visual scene content can refer to the sum of visual elements in the currently played media content segment (such as video frames, images), including but not limited to subject objects: people, animals, objects, etc., background environment: natural scene, indoor scene, weather, etc., visual style: color tone, lighting effect, picture composition, etc.

[0114] The visual scene content can be obtained by intercepting the target media content segment currently watched by the user, such as 2 frames of video before and after the current time point, to obtain a target media content segment containing 5 frames of video, and then identifying the person subject, scene content, and person action in the target media segment.

[0115] The scene keywords can be core semantic labels extracted from the visual scene content, which summarize the key information in the visual scene content, such as people, colors, scene types, etc. It can be matched with the current visual scene content by using a pre-set object, person, or action noun or verb dictionary, such as "skiing", "snow mountain", "sky", etc. Deep learning models can also be used to map visual scene content to corresponding text keywords, and the embodiments of the present application do not limit this.

[0116] The emotion state adjectives can be adjectives or phrases describing the user's current emotion state, which are usually converted from the emotion state data, for example, the emotion state is: happy, and the corresponding emotion state adjectives can be: happy, happy, excited, etc. The embodiments of the present application do not limit this.

[0117] In an embodiment, the visual scene content of the currently played media content segment is obtained, scene keywords are extracted from the visual scene content, and emotion state adjectives corresponding to the emotion state data are obtained, and the barrage content is generated and displayed according to the scene keywords and the emotion state adjectives. Specifically, the media content segment currently played by the user is obtained, the visual scene content information of the media content segment is recognized, scene keywords are extracted from the visual scene content information, and emotion state adjectives corresponding to the emotion state data are obtained by using a pre-set dictionary or a large language model, and the corresponding barrage content is generated by combining the emotion state adjectives and the scene keywords.

[0118] For example, the visual scene content of the currently played media content segment is obtained, the display content "a person is skiing in a snow mountain" of the visual scene content is recognized, scene keywords such as ["snow mountain", "skiing", "person (star)", "snow"] are extracted from the recognized scene, emotion state data [happy: 0.9] is obtained, emotion adjectives matching "happy" are obtained by using a pre-set dictionary or a large language model, and the barrage content "very happy to see a person (star) skiing" is finally generated.

[0119] Through Figure 2 According to the related description of the embodiment shown in the figure, the emotion state data corresponding to the user's face data is determined, the emotion state data is analyzed in combination with the visual scene content, and the barrage content is generated in real time, which improves the real-time nature of the barrage and enables the barrage to be more in line with the real feelings of the user and the video scene, significantly improving the user's participation and content interaction quality. Compared with the traditional barrage interaction mode, the intelligence of the barrage interaction is significantly improved, which not only optimizes the user's viewing and interaction experience, but also provides a practical solution for the fusion of emotion computing and multimedia interaction.

[0120] Figure 3 Another embodiment flowchart of the barrage generation method provided by the embodiment of the present application is shown in Figure 3 Based on the embodiment shown in the figure, the anti-misjudgment mechanism of barrage generation is mainly described, that is, how to determine whether to generate barrage content, including the following steps: Figure 1

[0121] Step 301, in the process of playing the target media content, the current face data of the watching user is collected by a camera device.

[0122] Step 302, the face action unit combination is parsed from the face data.

[0123] Step 303, the emotion state data of the watching user is determined according to the face action unit combination.

[0124] ​For the above steps 301-303, refer to the detailed description of the above related embodiments.

[0125] Step 304, acquiring visual scene content of the currently played media content segment.

[0126] Step 305, determining the confidence degree corresponding to the emotional state data according to the visual scene content.

[0127] Step 306, generating and displaying the barrage content in the case where the confidence degree meets the set condition.

[0128] The confidence degree corresponding to the emotional state data can refer to a quantitative index obtained by calculating the matching degree between the current emotional state of the user and the visual scene of the currently played media content, for measuring the reliable degree of the judgment that the emotional state of the user is caused by the currently watched media content.

[0129] Specifically, the high and low of the confidence degree directly reflects the strength of the association between the user's emotion and the media content scene. If the confidence degree is high, it means that the emotional state (such as happiness) of the user is highly matched with the visual scene (such as "happy picture") of the current media content, and it can be considered that the emotion of the user is indeed caused by the currently watched content, rather than other external factors (such as suddenly receiving a message and being happy). In the case of high confidence degree, the reliability of generating the barrage is high. If the confidence degree is low, it means that the emotion of the user is not matched with the current media content scene, for example, when a sad plot is being played, the user is recognized to show happiness, which means that a misjudgment has occurred, that is, the emotion of the user is caused by other events, and at this time, the barrage content can not be generated.

[0130] Accordingly, in step 305, the confidence degree corresponding to the emotional state data is determined according to the visual scene content, and further, in step 306, the barrage content is generated and displayed in the case where the confidence degree meets the set condition.

[0131] In an embodiment, an exemplary implementation of determining the confidence degree corresponding to the emotional state data according to the visual scene content includes: extracting a visual scene feature vector from the visual scene content, and generating an emotional state feature vector of the emotional state data, performing similarity calculation on the visual scene feature vector and the emotional state feature vector, and determining the confidence degree corresponding to the emotional state data according to the calculated similarity value.

[0132] The visual scene feature vector can be a numerical vector quantitatively representing visual scene content of the currently played target media content segment, wherein the visual scene feature vector contains rich visual information such as objects, scenes, colors, and textures in the visual image, and can reflect the emotional tendency conveyed by the visual image. For example, the visual scene content can be extracted by computer vision technology, such as a convolutional neural network (CNN) in a deep learning model, to convert key visual information in an image or a video frame into a numerical vector of a fixed length. In addition, the visual scene feature vector can also be obtained by other means, which are not limited in the embodiments of the present application.

[0133] The emotional state feature vector can be a numerical vector quantitatively representing the current emotional state data of the user, for accurately describing the type and intensity of the emotion. For example, the extracted emotional state can be converted into a numerical vector of a fixed length by natural language processing or emotion modeling technology, for example, the emotional state [happy: 0.9] can be mapped to a feature vector by a pre-trained emotion classification model.

[0134] As an optional implementation, the confidence of the emotional state data corresponding to the visual scene content can be determined by the following method: a pre-trained deep learning model is used to extract a visual scene feature vector from the visual scene content of the currently played media content segment, a pre-trained emotion classification model is used to obtain an emotional state feature vector corresponding to the emotional state data, a similarity value of the visual scene feature vector and the emotional state feature vector is calculated, for example, a cosine similarity is calculated, the similarity value is determined as the confidence of the emotional state data corresponding to the visual scene content, or the similarity value is further processed, for example, data normalization processing is performed, to obtain the final confidence of the emotional state data corresponding to the visual scene content, and the confidence of the emotional state data corresponding to the visual scene content is positively correlated with the similarity value.

[0135] The embodiments of the present application provide a false positive prevention mechanism for generating a barrage, which prevents the recognition of facial movements or emotional states caused by events other than the currently played media content segment, and further generates a barrage content according to the emotions caused by these other events, thereby improving the accuracy of generating the barrage content and avoiding the generation of barrage content that is not related or matched to the currently played media content.

[0136] Specifically, the confidence of the current emotional state data is determined by the recognized user emotional state data and the visual scene content of the currently played media content segment. The higher the confidence, the higher the matching degree of the current user's emotional state data and the visual scene content of the currently played media content, and it is not a false positive, and the barrage can be generated. On the contrary, the lower the confidence, the less reliable the currently recognized user's emotional state data, and the barrage content can not be generated. The confidence threshold can be set in advance to determine whether the currently recognized emotional state data is reliable, in addition, other ways can be used to determine whether the emotional state data is reliable, and the embodiments of the present application do not limit this.

[0137] For example, if the visual scene content of the currently played media content segment shows "a star skiing in a snow mountain", the corresponding visual scene feature vector is obtained by using computer vision technology, for example, [0.8, 0.3, 0.1, …, 0.6], and the user's emotional state data is [happy: 0.9], and the emotional state feature vector generated by the pre-trained large language model can be [0.95, 0.8, 0.02, …, 0.7], the cosine similarity of the two feature vectors is 0.9, and the cosine similarity value is close to the preset similarity threshold 1, which means that the user's current emotional state and the visual scene are highly matched, and the confidence of the emotional state data is high, which means that the user's "happiness" is suitable for the user to watch the media content segment, and the above confidence is reliable and not a false positive, that is, the user feels "happy" due to other things and not the content currently watched, and the barrage content can be generated. Similarly, the lower the confidence, the less reliable, which means that a false positive may occur, and the barrage content can not be generated.

[0138] In an embodiment, in a case where it is determined that the confidence meets the set condition, the barrage content is generated and displayed, and in a case where it is determined that the confidence does not meet the set condition, the barrage content can not be generated.

[0139] Specifically, the set condition can be used to determine the reliability of the current confidence. The above set condition can be a pre-set confidence threshold, and the confidence threshold can be dynamically set. In addition, the reliability of the confidence can also be determined by the large language model, and the embodiments of the present application do not limit this, and the confidence threshold is taken as an example in the following explanation.

[0140] As an optional implementation, in a case where it is determined that the confidence meets the set condition, the exemplary implementation of generating and displaying the barrage content includes: in a case where the confidence corresponding to the current obtained emotional state data is greater than or equal to a preset confidence threshold, the barrage content is generated and displayed.

[0141] For example, if the confidence corresponding to the current obtained emotional state data is 0.85, which is greater than the set confidence threshold 0.5, the barrage content (such as "too happy, ha ha ha") is generated.

[0142] Further, in the case where it is determined that the confidence meets the set condition, after the barrage content is generated, the sending operation of the barrage content can also be triggered according to the specified facial action set by the user. In particular, for some people with hand disabilities, the sending operation of the generated barrage content can be triggered by recognizing their facial actions, such as nodding or blinking, thereby lowering the threshold of the barrage interaction, realizing barrier-free interaction, and improving the user experience.

[0143] As another optional implementation, in the case where it is determined that the confidence does not meet the set condition, an exemplary implementation of generating and displaying the barrage content includes: in the case where the confidence corresponding to the current obtained emotional state data is less than the preset confidence threshold, the barrage content can not be generated.

[0144] For example, if the confidence corresponding to the current obtained emotional state data is 0.3, which is less than the preset confidence threshold 0.5, the barrage content is not generated.

[0145] In addition, in the case where it is determined that the confidence does not meet the set condition, i.e., the confidence corresponding to the current emotional state data is less than the preset confidence threshold, an option of whether to generate the barrage content can be displayed on the screen, and the user can decide whether to generate the barrage content in the case where the confidence is low. If the user determines to generate the barrage content, the system can be optimized according to the user's selection and the current obtained emotional state data, and the accuracy of generating the barrage content is further improved.

[0146] Through Figure 3 The anti-misjudgment mechanism based on visual scene content described in the embodiment shown in the figure can determine whether the current emotion of the user is caused by the target media content currently watched based on the confidence, and determine that it is caused by the target media content currently watched. This way can effectively filter the misjudgment caused by external factors and improve the accuracy of the barrage generation.

[0147] Figure 4 The structure block diagram of a barrage generation device provided by the embodiment of the present application is shown in FIG. 1. Figure 4 As shown in the figure, the device includes:

[0148] The data acquisition module 41 is configured to acquire the current facial data of the watching user through the camera device in the process of playing the target media content;

[0149] The data analysis module 42 is configured to analyze the facial action unit combination from the facial data.

[0150] an emotion state determining module 43, configured to determine emotion state data of the viewing user currently according to the facial action unit combination;

[0151] a barrage generating module 44, configured to generate and display barrage content according to the emotion state data.

[0152] In a possible implementation, the data analyzing module 42 is specifically configured to:

[0153] analyze the action intensity parameter set of the target facial feature region from the facial data;

[0154] map the action intensity parameter set into the facial action unit combination according to a preset facial action coding system rule.

[0155] In a possible implementation, the emotion state determining module 43 is specifically configured to:

[0156] find emotion state data corresponding to the facial action unit combination from a preset emotion state model, wherein the emotion state model comprises a mapping relationship between facial action unit combinations and emotion state data;

[0157] determine the found emotion state data as the emotion state data of the viewing user currently.

[0158] In a possible implementation, the barrage generating module 44 comprises:

[0159] a scene content obtaining unit, configured to obtain visual scene content of a currently played media content segment;

[0160] a barrage generating unit, configured to generate and display barrage content based on the emotion state data and the visual scene content.

[0161] In a possible implementation, the barrage generating unit is specifically configured to:

[0162] extract a scene keyword from the visual scene content, and obtain an emotion state adjective corresponding to the emotion state data;

[0163] generate and display barrage content according to the scene keyword and the emotion state adjective.

[0164] In a possible implementation, the barrage generating module 44 comprises:

[0165] a content obtaining unit, configured to obtain visual scene content of a currently played media content segment;

[0166] a confidence determination unit configured to determine a confidence corresponding to the emotional state data according to the visual scene content;

[0167] a barrage generation unit configured to generate and display a barrage content if the confidence meets a set condition.

[0168] In a possible implementation, the confidence determination unit is specifically configured to:

[0169] extract a visual scene feature vector from the visual scene content, and generate an emotional state feature vector of the emotional state data;

[0170] perform a similarity calculation on the visual scene feature vector and the emotional state feature vector, and determine the confidence corresponding to the emotional state data according to a similarity value obtained by the calculation;

[0171] The confidence corresponding to the emotional state data is positively correlated with the similarity value.

[0172] As shown in Figure 5 The embodiment of the present application provides an electronic device, which comprises a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112 and the memory 113 complete mutual communication through the communication bus 114,

[0173] The memory 113 is used for storing a computer program.

[0174] In an embodiment of the present application, the processor 111 is used for executing the program stored in the memory 113, and realizes the barrage generation method provided by any one of the preceding method embodiments, which comprises the following steps:

[0175] In the process of playing the target media content, the current facial data of a viewing user is collected through a camera device;

[0176] The facial action unit combination is parsed from the facial data;

[0177] The emotional state data of the viewing user is determined according to the facial action unit combination;

[0178] The barrage content is generated and displayed according to the emotional state data.

[0179] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the barrage generation method provided by any one of the preceding method embodiments.

[0180] The apparatus embodiments described above are only illustrative, and the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.

[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, and the computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the embodiments or some parts of the embodiments.

[0182] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes", "including" and "has", "having" are inclusive and therefore specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring their performance in the particular order in which they are described, unless specifically indicated as such. It is also to be understood that additional or alternative steps can be employed.

[0183] The above description is merely illustrative of the application and should not be taken as limiting. Numerous modifications and variations underlying the general principles of the applications can be made by those of ordinary skill in the art without departing from the spirit or scope of the application. Therefore, the application should not be limited to the embodiments shown herein but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating a barrage, characterized in that: The method comprises: During the playback of the target media content, the current facial data of the viewing user is collected through a camera device; Parsing facial action unit combinations from the facial data; determining the current emotional state data of the viewing user according to the facial action unit combination; The barrage content is generated and displayed according to the emotional state data.

2. The method according to claim 1, characterized in that The step of parsing facial action unit combinations from the facial data includes: Analyzing the action intensity parameter set of the target facial feature area from the facial data; According to a preset facial action coding system rule, the action intensity parameter set is mapped into a facial action unit combination.

3. The method according to claim 1, characterized in that Determining the current emotional state data of the viewing user according to the facial action unit combination includes: Searching for emotional state data corresponding to the facial action unit from a preset emotional state model, wherein the emotional state model includes a mapping relationship between facial action unit combinations and emotional state data; The found emotional state data is determined as the current emotional state data of the viewing user.

4. The method according to claim 1, wherein Generating and displaying barrage content according to the emotional state data includes: Get the visual scene content of the currently playing media content segment; Based on the emotional state data and the visual scene content, barrage content is generated and displayed.

5. The method according to claim 4, characterized in that The generating and displaying of bullet screen content based on the emotional state data and the visual scene content includes: extracting scene keywords from the visual scene content, and obtaining emotional state adjectives corresponding to the emotional state data; Based on the scene keywords and the emotional state adjectives, barrage content is generated and displayed.

6. The method according to claim 1, characterized in that Generating and displaying barrage content according to the emotional state data includes: Get the visual scene content of the currently playing media content segment; Determining, based on the visual scene content, a confidence level corresponding to the emotional state data; When it is determined that the confidence level meets the set conditions, the barrage content is generated and displayed.

7. The method according to claim 6, characterized in that Determining the confidence level corresponding to the emotional state data based on the visual scene content includes: extracting a visual scene feature vector from the visual scene content, and generating an emotional state feature vector for the emotional state data; Calculating similarity between the visual scene feature vector and the emotional state feature vector, and determining the confidence level corresponding to the emotional state data based on the calculated similarity value; The confidence level corresponding to the emotional state data is positively correlated with the similarity value.

8. A barrage generation device, characterized in that: The device comprises: A data acquisition module is used to collect the current facial data of the viewing user through a camera device during the playback of the target media content; A data parsing module, configured to parse facial action unit combinations from the facial data; an emotional state determination module, configured to determine the current emotional state data of the viewing user based on the facial action unit combination; The barrage generation module is used to generate and display barrage content based on the emotional state data.

9. An electronic device, characterized in that: include: A processor and a memory, wherein the processor is used to execute a barrage generation program stored in the memory to implement the barrage generation method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the barrage generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Bullet screen generation method, head-mounted display device and storage medium

    CN121462746A