Interaction processing method and device of live broadcast room, electronic equipment and storage medium
By displaying elements related to the live broadcast progress in the live broadcast room and automatically updating the comment editing area, the problem of time-consuming users' choice of emoticon packages is solved, and more efficient live broadcast room interaction is achieved.
Patent Information
- Application Number
- CN202410023878.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, it takes a long time for users to select emoticons related to live content in the live broadcast room, which affects the interaction efficiency.
It provides a live broadcast room interaction processing method. By displaying elements related to the live broadcast progress, the elements in the comment editing area are automatically updated. Users only need to select and generate interactive information matching the live broadcast progress.
It improves the user's interaction efficiency in the live broadcast room, reduces the time for users to find matching elements, and improves the real-time and efficiency of interaction.
Smart Images

Figure CN120281926A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and in particular, to an interactive processing method, apparatus, electronic device, and storage medium for a live broadcast room. Background Art
[0002] During a live broadcast, users can interact with the host or other users watching the live broadcast by sending text or emoticon pictures. However, in related technologies, the types of expressions in emoticons are rich and diverse. Users need to actively select emoticons related to the live broadcast content and send them to complete the interactive operation. It takes a long time for users to find suitable emoticons, which affects the interactive efficiency.
[0003] In related technologies, there is no good way to improve the interactive efficiency during a live broadcast. Summary of the Invention
[0004] Embodiments of this application provide an interactive processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for a live broadcast room, which can improve the interactive efficiency in the live broadcast room.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] Embodiments of this application provide an interactive processing method for a live broadcast room. The method includes:
[0007] Display the live broadcast room;
[0008] Display a comment editing area, where the comment editing area includes at least one first element, and the first element changes according to the current live broadcast progress of the live broadcast room;
[0009] In response to an element selection operation, display at least one first target element in a selected state, where the first target element is any one of the selected first elements;
[0010] In response to a sending operation, display first interactive information in the comment interaction area of the live broadcast room, where the first interactive information includes the first target element.
[0011] Embodiments of this application provide an interactive processing method for a live broadcast room. The method includes:
[0012] Display the live broadcast room;
[0013] Display a comment editing area, where the comment editing area includes at least one first element to be selected. The first element changes according to the current live broadcast progress of the live broadcast room. The first element is used to generate first interactive information, and the first interactive information is used to be sent to the live broadcast room.
[0014] An embodiment of the present application provides an interactive processing device for a live broadcast room, including:
[0015] A live broadcast room display, where the live broadcast room includes a comment editing entry;
[0016] In response to a trigger operation on the comment editing entry, a comment editing area is displayed, where the comment editing area includes at least one first element, the first element changes according to the current live broadcast progress of the live broadcast room, the first element is used to generate first interactive information, and the first interactive information is used to be sent into the live broadcast room.
[0017] An embodiment of the present application provides an interactive processing device for a live broadcast room, including:
[0018] A display module for displaying a live broadcast room, where the live broadcast room includes a comment editing entry;
[0019] The display module is further configured to display a comment editing area in response to a trigger operation on the comment editing entry, where the comment editing area includes at least one first element, the first element changes according to the current live broadcast progress of the live broadcast room, the first element is used to generate first interactive information, and the first interactive information is used to be sent into the live broadcast room.
[0020] An embodiment of the present application provides an electronic device, and the electronic device includes:
[0021] A memory for storing computer-executable instructions;
[0022] A processor, when executing the computer-executable instructions stored in the memory, implements the interactive processing method for a live broadcast room provided by an embodiment of the present application.
[0023] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, which, when being executed by a processor, implement the interactive processing method for a live broadcast room provided by an embodiment of the present application.
[0024] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions, which, when being executed by a processor, implement the interactive processing method for a live broadcast room provided by an embodiment of the present application.
[0025] The embodiments of the present application have the following beneficial effects:
[0026] Display at least one first element in the live broadcast room and the comment editing area, where the first element is related to the current live broadcast progress of the live broadcast room; by displaying elements related to the live broadcast progress, the user experience of the live broadcast progress is improved, and the user is promoted to perform interactive operations during the live broadcast. The first interactive information sent is generated based on the first target element, and the finally sent interactive information is related to the live broadcast progress. Compared with the existing interactive method of only displaying fixed elements, the elements displayed are changed in real time according to the live broadcast progress, and there is no need for the user to actively search for elements that match the live broadcast progress, which can improve the interactive efficiency of the user in the live broadcast room. Description of the Drawings
[0027] Figure 1 It is a schematic diagram of the application mode of the interactive processing method for the live broadcast room provided by the embodiment of the present application;
[0028] Figure 2 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present application;
[0029] Figures 3A to 3G It is a schematic diagram of the process of the interactive processing method for the live broadcast room provided by the embodiment of the present application;
[0030] Figure 4 It is a schematic diagram of the element synthesis principle provided by the embodiment of the present application;
[0031] Figures 5A to 5E It is a schematic diagram of the live broadcast room interface provided by the embodiment of the present application;
[0032] Figures 6A to 6D It is a schematic diagram of element synthesis provided by the embodiment of the present application;
[0033] Figures 7A to 7B It is a schematic diagram of the live broadcast replay video display method provided by the embodiment of the present application;
[0034] Figure 8 It is a preset element table provided by the embodiment of the present application;
[0035] Figure 9 It is a schematic diagram of the process of the interactive processing method for the live broadcast room provided by the embodiment of the present application. Detailed Embodiments
[0036] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0037] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0038] In the following description, the terms "first / second / third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0039] It should be noted that the collection and processing of relevant data in the present application (for example: the live content of the live broadcast room, the interactive information sent by users, etc.) should strictly comply with the requirements of relevant national laws and regulations during actual application, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.
[0040] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0042] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described, and the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0043] 1) Live broadcast, which is now often referred to as an online interactive live broadcast, is a social and commercial way. The anchor uses a computer or a mobile phone to synchronously broadcast what he / she is doing, and the audience can watch through a website or an application.
[0044] 2) Network live broadcast room, the interface displayed in the application for playing live content, where users can interact with the anchor or other users watching the live broadcast by sending interactive information. In the embodiments of the present application, the interface displayed in the application for playing live content is simply referred to as the live broadcast room.
[0045] 3) Text live broadcast, which is a live broadcast in text form, without pictures, and conveys information to the audience only through text. For example: a live broadcast of a sports game reporting the game situation and score through text.
[0046] 4) Element. The elements in the live broadcast room are a form of using pictures to represent emotions. The elements are stored in picture format, and the types of elements include but are not limited to image elements, expression elements, and text elements. The text element is a picture with text as the display content.
[0047] The embodiments of the present application provide an interactive processing method for a live broadcast room, an interactive processing device for a live broadcast room, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the interactive efficiency in the live broadcast room.
[0048] The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as a terminal device, such as a laptop computer, a tablet computer, a desktop computer, a set-top box, a smart TV, a mobile device (for example, a mobile phone, a portable music player, a personal digital assistant, a dedicated messaging device, a portable game device), a vehicle-mounted terminal, a virtual reality (VR) device, an augmented reality (AR) device, and other various types of user terminals, or can also be implemented as a server. Hereinafter, the exemplary applications when the electronic device is implemented as a terminal device or a server will be described.
[0049] Reference Figure 1 , Figure 1 is a schematic diagram of the application mode of the interactive processing method for the live broadcast room provided by the embodiments of the present application; for example, Figure 1 involves a server 200, a network 300, a terminal device 400, and a database 500. The terminal device 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0050] In some embodiments, the terminal device 400 stores an application program capable of playing live video. The server 200 can be a server of a live broadcast platform, and the database 500 stores content such as live replay videos of the live broadcast platform and interactive information of elements and objects.
[0051] For example, when the terminal device 400 starts the application program of the corresponding live video, the terminal device 400 sends a request carrying the live room identifier to the server 200. The server 200 returns the current live content of the corresponding live room and the first element associated with the live content to the terminal device 400. The user sends interactive information in the live room by clicking on the first element displayed in the terminal device 400. Compared with the prior art's interactive method of only displaying fixed elements, the elements displayed are changed in real time according to the live content, eliminating the need for the user to actively search for elements that match the live content, which can improve the interactive efficiency of the user in the live room.
[0052] In some embodiments, the interactive processing method of the live room in the embodiments of the present application can also be applied to the following application scenarios: video conferencing. During an online video conference, according to the live content played by the speaker, the corresponding text elements or expression elements are updated and displayed in real time. The users participating in the conference can send text elements and expression elements that are updated in real time following the video conference content to represent their views on the conference content, improving the efficiency of the online video conference.
[0053] The embodiments of the present application can be implemented through database technology. A database, in short, can be regarded as a place for storing electronic files in an electronic filing cabinet. Users can perform operations such as adding, querying, updating, and deleting data in the files. A so-called "database" is a data set stored together in a certain way, shared by multiple users, having the smallest possible redundancy, and independent of application programs.
[0054] A database management system (DBMS) is a computer software system designed to manage a database and generally has basic functions such as storage, interception, security guarantee, and backup. The database management system can be classified according to the database model it supports, such as relational, XML (Extensible Markup Language); or according to the type of computer it supports, such as server clusters, mobile phones; or according to the query language it uses, such as Structured Query Language (SQL), XQuery; or according to the performance measurement focus, such as maximum scale, highest running speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories. For example, they can support multiple query languages simultaneously.
[0055] Embodiments of this application can also be implemented through cloud technology. Cloud technology (Cloud Technology) is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, in the future, each item may have its own hash code identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various types of industry data require the support of a powerful system back-end, which can only be achieved through cloud computing.
[0056] In some embodiments, embodiments of this application can also be implemented through artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0057] Computer Vision Technology (Computer Vision, CV): Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, detection, and measurement on targets, and further perform graphic processing to make the computer process into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pretrained models in the visual field such as swin-transformer, ViT, V-MOE, MAE, etc. can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0058] In some embodiments, the server 200 may be composed of multiple servers, such as a live broadcast management server and an element acquisition server. Among them, the live broadcast management server is used to transmit the live broadcast content of the host to the terminal device of the audience, and the element acquisition server is used to determine the elements displayed in the live broadcast room of the terminal device of the audience according to the live broadcast content.
[0059] In some embodiments, the server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0060] In some embodiments, the terminal device 400 may implement the interactive processing method of the live broadcast room provided in the embodiments of the present application by running a computer program. For example, the computer program may be a native program or software module in the operating system; it may be a local (Native) application program (APP, APPlication), that is, a program that needs to be installed in the operating system to run, such as a live broadcast APP; it may also be a small program that can be embedded in any APP, that is, a program that only needs to be downloaded to the browser environment to run. In short, the above computer program may be any form of application program, module or plug-in.
[0061] See Figure 2 , Figure 2 is a schematic structural diagram of the electronic device provided in the embodiments of the present application. The electronic device may be Figure 1 the terminal device 400, Figure 2 The terminal device 400 shown in includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal device 400 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 440.
[0062] The processor 410 may be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0063] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.
[0064] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 optionally includes one or more storage devices that are physically located remotely from the processor 410.
[0065] The memory 450 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0066] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are illustratively described below.
[0067] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0068] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;
[0069] A presentation module 453 for enabling presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.);
[0070] An input processing module 454 for detecting one or more user inputs or interactions from one of one or more input devices 432 and translating the detected inputs or interactions.
[0071] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2 Shown is an interactive processing device 455 for a live broadcast room stored in the memory 450, which may be software in the form of a program and a plug-in, etc., including the following software modules: a display module 4551, an acquisition module 4552. These modules are logical, so they can be combined arbitrarily or further split according to the functions implemented. In Figure 2 For the convenience of expression, all the above modules are shown at once, but it should not be regarded as excluding the implementation where the interactive processing device 455 for the live broadcast room may only include the display module 4551. The functions of each module will be described below.
[0072] The interactive processing method for the live broadcast room provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the terminal device provided by the embodiments of the present application.
[0073] Next, the interactive processing method for the live broadcast room provided by the embodiments of the present application will be described. As mentioned above, the electronic device implementing the interactive processing method for the live broadcast room of the embodiments of the present application may be a terminal or a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.
[0074] It should be noted that in the examples of the interactive processing in the live broadcast room hereinafter, the live broadcast of a competition is taken as an example for description. Those skilled in the art can apply the interactive processing method for the live broadcast room provided by the embodiments of the present application to the processing including other types of live broadcasts according to the understanding of the following text, such as: product recommendation live broadcast, game live broadcast, daily life live broadcast or video conference live broadcast.
[0075] See Figure 3A , Figure 3A is a flowchart of the interactive processing method for the live broadcast room provided by the embodiments of the present application, and will be described in combination with the Figure 3A steps shown.
[0076] In step 301, display the live broadcast room.
[0077] Exemplarily, in the embodiments of the present application, a live broadcast room refers to the live broadcast room interface in an application for displaying live broadcast content, simply referred to as the live broadcast room. The live broadcast room interface is used to display live broadcast content, as well as audience comment content and interaction information between the anchor and the audience. Based on the above functions, the live broadcast room interface includes a comment editing entry (for triggering the comment editing area), a comment editing area (for displaying interaction information being edited), a comment interaction area (for displaying sent interaction information), and an area for displaying live broadcast content.
[0078] Exemplarily, referring to Figure 5A , Figure 5A is a schematic diagram of the live broadcast room interface provided by the embodiments of the present application. The live broadcast room interface 501A is an interface for sports text live broadcast, in which live broadcast content 502A and a comment interaction area 503A are displayed. If a user wants to join the interaction after viewing the interaction content or the live broadcast content in the comment interaction area 503A, they can trigger the comment editing area by clicking on the comment trigger entry or the blank area in the live broadcast room interface 501A.
[0079] In some embodiments, in response to the comment editing entry for the live broadcast room, the comment editing area is displayed, that is, the process of step 302 is entered.
[0080] In step 302, the comment editing area is displayed.
[0081] Here, the comment editing area includes at least one first element, and the first element changes according to the current live broadcast progress of the live broadcast room.
[0082] An element refers to a form in the live broadcast room that uses pictures to represent emotions, and the element is stored in picture format; the live broadcast progress refers to the progress of the entire live broadcast, including information such as all live broadcast stages and remaining time. The real-time progress refers to the progress of the currently ongoing live broadcast stage. In some cases, for some live broadcasts without a fixed progress bar, the live broadcast can be divided into progress according to the playing duration of the live broadcast; for some live broadcasts with a predetermined duration before the live broadcast, the live broadcast progress can be characterized by the progress of the live broadcast theme. For example, for a game live broadcast, the live broadcast progress can be characterized by the progress of the game; the first element changing with the current live broadcast progress of the live broadcast room means that the display and hiding of the first element are related to the live broadcast progress.
[0083] Exemplarily, in response to the trigger operation for the live broadcast room interface 501A, by Figure 5A enter Figure 5B interface. Referring to Figure 5B , Figure 5BIt is a schematic diagram of the live broadcast room interface provided by an embodiment of the present application; in the live broadcast room interface 501A, the live broadcast content 502A and the comment editing area 504A are displayed. The live broadcast theme is a basketball game. Among them, "LETS GO" and "DEFENCE" are elements updated according to the live broadcast progress, specifically text elements displayed according to the goal situation during the game. Multiple first elements 501B are displayed in the comment editing area 504A.
[0084] In some embodiments, the types of the first elements include:
[0085] Type 1, image element, where the image element is extracted from the video frame corresponding to the live broadcast content, or searched according to the live broadcast progress characteristics.
[0086] For example: In the current live basketball game broadcast, the picture of the video frame is an athlete of a certain team. Call the neural network model to perform segmentation processing on the video frame to obtain the team logo image on the athlete's clothes in the picture of the video frame, and use the team logo image as the image element. Or, the resolution of the team logo image is relatively high. Compress and redraw the team logo image so that the size and resolution of the team logo image are smaller, which is convenient for sending in the interactive information. That is, perform normalization processing on the original image of the extracted team logo image to obtain an image element with a smaller storage space occupancy compared to the original image.
[0087] Another example: Extract the image element from the preset element database based on the live broadcast progress characteristics of the live broadcast room.
[0088] Type 2, text element, where the text element is obtained by performing speech recognition on the audio data corresponding to the live broadcast content, or by performing text recognition on the video frame corresponding to the live broadcast content.
[0089] For example: Call the speech recognition model based on the audio data corresponding to the live broadcast content to obtain the text corresponding to the speech in the live broadcast content, and convert the text into a picture format to obtain the text element.
[0090] Another example: Call the neural network model to perform image segmentation processing on the video frame corresponding to the live broadcast content to obtain the image including text in the picture of the video frame, perform text recognition on the image including text by using the optical character recognition technology to obtain the text, and convert the text into a picture format to obtain the text element.
[0091] In some embodiments, the text element can be extracted from the preset element database based on the live broadcast progress characteristics of the live broadcast room. The scheme for extracting elements from the preset element database is applicable to any type of element.
[0092] Type 3, expression elements, where the expression elements are elements related to the emotional type represented by the live content, and the emotional type is obtained by performing emotional recognition on at least one of the video frames, audio data, and interaction situation corresponding to the live content.
[0093] Exemplarily, the expression elements are elements that can represent the emotions of the audience, and the forms of the expression elements are expressions of different emotional types of human or animated images, such as: laughing, crying, getting angry, etc.
[0094] The interaction situation refers to the interaction situation among users in the live broadcast room, including but not limited to the interaction among the audience, and the interaction between the audience and the anchor; based on the video frames, audio data, and interaction situation in the live content, a neural network model for emotion classification is called to perform emotion classification processing to obtain an emotion classification result. Emotion analysis can be trained and predicted using machine learning algorithms (such as Naive Bayes, Support Vector Machine, etc.) or deep learning models (such as Recurrent Neural Network, Convolutional Neural Network, etc.). The emotion classification result refers to the emotional type, and the obtained emotional type is matched with the stored expression elements to obtain the first element corresponding to the current live broadcast.
[0095] For example: Obtain the interaction situation among users, and the interaction situation can be represented by the interaction information of comments. The interaction information includes but not limited to images and texts. Based on the interaction information, a neural network model for emotion classification is called to obtain the target emotional type, and the expression element corresponding to the target emotional type can be used as the first element for display.
[0096] In some embodiments, the process of obtaining the first element can be frame-based, that is, updated in real time according to each frame; or updated according to or in a pre-configured period.
[0097] In step 303, in response to the element selection operation, at least one first target element in the selected state is displayed.
[0098] Here, the first target element is any one of the selected first elements.
[0099] Exemplarily, the element selection operation can be an operation such as clicking or long pressing, continue to refer to Figure 5B , among multiple first elements 501B, the first target element 502B is displayed in the selected state, and compared with other first elements, it has a circular mark including a tick.
[0100] In step 304, in response to the sending operation, the first interaction information is displayed in the comment interaction area of the live broadcast room.
[0101] Exemplarily, the first interaction information includes the first target element.
[0102] The content of the first interaction information includes the following types: 1. The content of the first interaction information only has the first target element; 2. The first interaction information includes the first target element and the text content edited by the user; 3. The first interaction information is synthesized by at least one first target element; 4. The first interaction information is synthesized by the first target element and other fixed elements, and the fixed elements do not change with the live broadcast progress during the live broadcast.
[0103] Reference Figure 5C , Figure 5C is a schematic diagram of the live broadcast room interface provided by the embodiment of the present application; if the user selects the first target element 502B in Figure 5B and sends the first target element 502B, then a first interaction information 506A is formed in the comment interaction area 505A of Figure 5C , and the content of the first interaction information 506A includes "LETS GO" of the first element.
[0104] In the embodiment of the present application, by displaying the first element that changes with the live broadcast progress, it is convenient for users to quickly find the expressions related to the live broadcast content to represent the emotions and ideas of watching the live broadcast, saving the time required for users to edit the interaction information and improving the interaction efficiency of users.
[0105] In some embodiments, when step 302 is executed, at least one second element to be selected is displayed in the comment editing area, where the second element is an element that remains unchanged during the live broadcast of the live broadcast room; when step 303 is executed, at least one second target element in a selected state is displayed in the comment editing area, where the second target element is any one of the selected second elements, and the second target element is used to combine with the first target element to generate the first interaction information.
[0106] Exemplarily, the type of the second element is interoperable with the type of the first element, and the second element can also be an image, text, and expression element. However, during the live broadcast, the second element will not be hidden or displayed with the change of the live broadcast progress. Reference Figure 5D , Figure 5D is a schematic diagram of the live broadcast room interface provided by the embodiment of the present application; in the comment editing area 504A, the user selects three different elements, and a third element 507A is generated based on the three elements. Among them, the comment interaction area 505A is Figure 5A the fully expanded comment interaction area 503A in
[0107] In the embodiments of the present application, a channel for users to self-set the elements to be sent is provided, which improves the freedom of editing interactive information, can more comprehensively represent the emotional types required by users through the combination of elements, and improves the efficiency of sending interactive information. Compared with the solution of storing a large number of preset elements, by allowing users to create combined elements themselves, the storage space required for storing elements can be saved.
[0108] In some embodiments, the first element is used to represent the current live progress characteristics of the live broadcast room, and the live progress characteristics include at least one of the following:
[0109] Type 1, the interaction progress among multiple objects in the live broadcast room.
[0110] For example, the object can be the interaction between the live audience and the host, the interaction between users in the live content, or the interaction between the audience. The interaction includes forms such as competitions, conversations, and comments. The interaction progress includes start, climax, gentle, and end. The classification model can be called based on the multi-modal characteristics of the live content for interactive progress classification processing to obtain the interactive progress.
[0111] Type 2, the keywords that appear in the live broadcast room.
[0112] For example, the keywords can be obtained by performing text recognition on the video frames of the live content, or by performing speech recognition on the audio data of the live broadcast. The content represented by the keywords includes but is not limited to: the excitement level of the live content, the emotional type of the live content (negative, positive, neutral), the theme of the live content, etc.
[0113] Type 3, the objects that appear in the live broadcast room.
[0114] For example, the objects that appear in the live broadcast room include but are not limited to people and objects. For example, if the live content is a product recommendation live broadcast, the host and the product that appear in the live broadcast room can both be used as live progress characteristics.
[0115] In some embodiments, after step 301, the live progress characteristics of the current live content of the live broadcast room are obtained. The method for determining at least one first element to be selected in the comment editing area includes any one of the following:
[0116] Method 1, when the live progress characteristic is the interaction situation among multiple objects in the live broadcast room, perform sentiment analysis processing on the interaction situation by calling a neural network model to obtain the current emotional type of the live broadcast room, and use the relevant elements of the current emotional type as at least one first element to be selected.
[0117] Exemplarily, the method for sentiment analysis processing can be a machine learning algorithm (such as Naive Bayes, Support Vector Machine, etc.), and the neural network model includes but is not limited to deep learning models (such as Recurrent Neural Network, Convolutional Neural Network, etc.). There is a pre-set mapping relationship between the emotion type and the relevant elements of the emotion type. When the current emotion type is determined, the relevant elements corresponding to the current emotion type are queried from the mapping relationship as the first element.
[0118] Method 2: When the live broadcast progress feature is an object appearing in the live broadcast room, perform graphic analysis processing on the object appearing in the live broadcast room, and use the obtained image elements as at least one candidate first element.
[0119] Exemplarily, the graphic analysis processing can be implemented through a neural network model. Perform image segmentation processing on the object appearing in the live broadcast room, and directly use the segmented image as the image element; or, perform compression and conversion processing on the segmented image, and use the processed image as the image element.
[0120] Method 3: When the live broadcast progress feature is a keyword appearing in the live broadcast room, generate text elements based on the keyword.
[0121] Exemplarily, if the text element is presented in a picture format, the keyword can be directly converted into text in a preset font, and the screenshot of the text is used as the text element. Or, call a generative pre-trained model based on the keyword to obtain a picture generated based on the keyword, and use this picture as the text element.
[0122] Method 4: Perform a search process based on the live broadcast progress feature to obtain search content, and perform graphic analysis processing on the search content to obtain at least one candidate first element.
[0123] Exemplarily, search for an image or text in the network based on the current live broadcast progress feature of the live broadcast room, and convert the image or text into a picture format that can be displayed in the live broadcast room based on image analysis processing to obtain the first element.
[0124] In the embodiments of the present application, generating the first element through multiple different methods can improve the efficiency of obtaining the first element, as well as improve the richness of the first element, and enhance the user experience of watching the live broadcast and performing interactions.
[0125] In some embodiments, the first element is used to represent the playback unit where the current live broadcast progress of the live broadcast room is located, where the playback unit is obtained by dividing the live broadcast content of the live broadcast room.
[0126] Exemplarily, the update of the first element is in units of the playback unit. The types of the playback unit include:
[0127] Type 1: A time unit obtained by periodically dividing according to a preset duration. For example, in practical applications, according to the estimated total duration of the live broadcast in the live broadcast room, a preset duration is set, and the live broadcast progress is divided periodically according to the preset duration to obtain multiple time units.
[0128] Type 2: A sub-shot unit obtained by dividing according to different camera shots in the live broadcast room.
[0129] For example, each operation of switching the camera is counted as a node, and the content captured by the same camera between two adjacent nodes is a time interval. If the live broadcast is a single-shot live broadcast, other types of playback unit division methods can be used.
[0130] Type 3: A plot unit obtained by dividing according to different plots in the live broadcast room.
[0131] For example: In a sports live broadcast, each score can be used as a division point, and the time interval between two scores is a plot unit; in a product promotion live broadcast, there are multiple different live broadcast products, and the promotion part corresponding to each product can be used as a plot unit. In a game live broadcast, for a story game, the plot unit can be divided according to the chapters in the game, and for a multiplayer battle game, each battle can be used as a node to divide the plot unit.
[0132] In the embodiments of the present application, the current live broadcast progress of the live broadcast room is characterized by a first element. Furthermore, the first element can update following the live broadcast progress with the playback unit as the unit, making the update of the first element match the live broadcast progress. Compared with the frequently updated scheme, it saves the computing resources required to update the first element and improves the user's live broadcast viewing experience and interaction efficiency.
[0133] In some embodiments, after step 301, refer to Figure 3B , Figure 3B is a schematic flowchart of the interactive processing method for the live broadcast room provided by the embodiments of the present application. Execute Figure 3B steps 3011 and 3012, which are specifically described below.
[0134] In step 3011, obtain the current live broadcast progress feature of the live broadcast room.
[0135] For example, the extraction method of the live broadcast progress feature has been explained in the different types of live broadcast progress features above and will not be elaborated here.
[0136] In step 3012, query at least one first element to be selected from the candidate element set based on the live broadcast progress feature.
[0137] For example, a large number of candidate elements are stored in the candidate element set, and the candidate elements can be used as the first element when being extracted. There is a pre-set mapping relationship between the elements in the candidate element set and the live broadcast progress feature. Based on the current live broadcast progress feature, the corresponding element can be found in the mapping relationship, and the element extracted from the candidate element set is used as the first element.
[0138] In some embodiments, before step 3012, refer to Figure 3C , Figure 3C FIG. Figure 3C is a schematic flowchart of the interactive processing method for the live broadcast room provided by the embodiments of the present application. The steps 30121 to 30126 of Figure 3C will be specifically described below.
[0139] In step 30121, obtain the preview information of the live broadcast room.
[0140] For example, the preview information is used to be displayed before the start of the live broadcast in the live broadcast room;
[0141] Taking a sports game as an example, the preview information is, for example: the teams participating in the game, team emblems, athlete names, etc. Taking a live product recommendation as an example, the preview information is, for example: the name of the anchor, information about the product to be recommended, etc.
[0142] In step 30122, perform feature extraction processing on the preview information to obtain attribute features.
[0143] For example, the attribute features include at least one of the following: image features, text features, and expression features. The feature extraction processing includes but is not limited to: image feature extraction processing, speech recognition, text feature extraction, emotion type analysis processing, and classification processing.
[0144] In step 30123, perform graphic analysis processing based on the image features to obtain image elements.
[0145] Here, the image elements are used as candidate elements.
[0146] Candidate elements are elements stored in the candidate element set and used as the first element when being displayed. For example: for a sports game, extract the image features of the team emblem from the preview information, and call a generative pre-trained model for image generation processing based on the image features of the team emblem to obtain image elements.
[0147] In step 30124, perform image conversion processing based on the text features to obtain text elements.
[0148] Here, the text elements are used as candidate elements.
[0149] Image conversion processing is to perform graphic processing on the text features of text to obtain a vector graphic, and use the vector image as a text element. For example, write the text in a preset font, take a screenshot of the text, and store the screenshot in the format of a vector graphic to obtain a text element. Or, call an image processing model based on the text to obtain a text element in the format of a vector graphic.
[0150] In step 30125, the expression element corresponding to the expression feature is used as a candidate element.
[0151] Exemplarily, the expression feature has a corresponding emotion type, and the relevant expression elements of the emotion type corresponding to the expression feature are used as candidate elements.
[0152] In step 30126, each candidate element is combined to form a candidate element set.
[0153] Exemplarily, each candidate element is used as a first element to be selected when queried based on the live broadcast progress feature. During the live broadcast, query in the candidate element set based on the live broadcast progress feature of the current live broadcast, and the query result can be displayed as the first element in the comment editing area.
[0154] In the embodiments of the present application, by constructing a candidate element set, the efficiency of obtaining the first element during the live broadcast can be improved. Compared with the solution of generating the first element in real time, the computing resources required to obtain the first element can be saved.
[0155] In some embodiments, when there are multiple selected first target elements, the multiple first target elements are carried in the first interaction information in a synthesized manner; when step 302 is executed, a synthesized element is displayed, where the synthesized element is obtained by fusing the multiple first target elements, and the synthesized element is used to be carried in the first interaction information.
[0156] Reference Figure 5D , Figure 5D is a schematic diagram of the live broadcast room interface provided by the embodiments of the present application; in the comment editing area 504A, the user selects three different elements, and a third element 507A is generated based on the three elements. Among them, the comment interaction area 505A is Figure 5A the fully expanded comment interaction area 503A in Figure 5E , Figure 5E is a schematic diagram of the live broadcast room interface provided by the embodiments of the present application; in response to a send operation on the third element 507A, a first interaction information 508A including the third element is displayed in the comment interaction area 505A.
[0157] In some embodiments, after the composite element is displayed, in response to a cancel operation on any first target element, the canceled first target element is switched from a selected state to an unselected state; and the composite element is updated based on the current first target element.
[0158] For ease of understanding, the following is described in conjunction with the accompanying drawings. Figure 6B , Figure 6B 604B is a schematic diagram of the element synthesis provided by the embodiment of the present application. The arrows between the comment editing area 601B, the comment editing area 602B, and the comment editing area 603B represent the order. When the user selects multiple elements in the comment editing area 601B, the selected element 604B in the comment editing area 602B is formed.
[0159] Among them, in the embodiment of the present application, for the sake of ease of viewing, the elements are extracted from the comment editing area, which does not mean that the graphics are displayed away from the comment editing area when actually displayed. The selected elements 604B are synthesized to obtain the synthesized element 605B. If the user cancels the selection of some elements and reselects the selected element 606B, after the update, the synthesized element 605B is switched to the synthesized element 607B in the comment editing area 603B.
[0160] In some embodiments, before displaying the composite element, the composite element may be determined by:
[0161] Method 1: when multiple first target elements include a text element, fill the text element with the color of other target elements, and synthesize the graphics of other target elements with the text element to form a synthesized element, wherein the other target elements are first target elements that are not text elements.
[0162] For example, the color types of other target elements can be multiple, and the color of other target elements can be the color of other target elements that accounts for a preset proportion, or the color located at a preset position. Synthesizing the graphics of other target elements with text elements can be achieved through a neural network model, or directly splicing, overlaying, etc. the layers of the graphics of other target elements and text elements.
[0163] Method 2: when the multiple first target elements do not include a text element, each first target element is superimposed to form a composite element.
[0164] For example, the superposition between elements can be achieved by superimposing layers, or by calling a generative pre-trained model, or by processing based on preset superposition rules.
[0165] refer to Figure 4 , Figure 4 It is a schematic diagram of the element synthesis principle provided by the embodiment of the present application. A variety of different expression elements will be displayed in the live broadcast room.Figure 4 It shows the element synthesis principle 401. Element 1 and Element 2 have undergone graphic transformations and combined with each other to obtain Element 3. Element 1 is an eye, Element 2 is a crying face with both eyes, and Element 3 is a crying face with one eye.
[0166] Method 3: When the number of at least one first target element is multiple, splice the first target elements with an associated relationship among the multiple first target elements, and superimpose the splicing result with other target elements to form a synthesized element.
[0167] Here, other target elements are elements among the multiple first target elements other than the first target elements with an associated relationship, and the synthesized element is obtained by calling an image transformation function based on the multiple first target elements.
[0168] In some embodiments, the synthesis order of the synthesis method can also be the order in which the user selects the elements.
[0169] In some embodiments, in response to the existence of associated elements of the selected first target element, display the associated elements of the first target element; in response to the selected multiple first target elements being mutually exclusive elements, hide the associated elements of the mutually exclusive first target elements.
[0170] Exemplarily, taking a sports game live broadcast as an example, when the user selects a team, obtain the slogan and expression corresponding to the team in the dataset, and display the slogan and expression corresponding to the team in the comment editing area; when the user selects the image elements of the emblems of two competing teams, display the general slogan and expression in the comment editing area.
[0171] Reference Figure 6C , Figure 6C is a schematic diagram of element synthesis provided by an embodiment of the present application; the arrow between the comment editing area 601C and the comment editing area 602C represents the sequence. Two elements in the comment editing area 601C are triggered, and the synthesized element 604C and the selected element 603C are displayed in the comment editing area 602C. If Figure 6C the emblem graphics of other teams in the comment editing area are selected, it is presented as Figure 6D , Figure 6D is a schematic diagram of element synthesis provided by an embodiment of the present application; compared with the comment editing area 601C, the emblem 2 in the comment editing area 601D is selected, and the comment editing area 602D displays the updated element 603D associated with the current live content of the team of the emblem 2. The content of the updated element 603D is the text element "LOSE", and the emblem 2 and the updated element 603D are synthesized into the synthesized element 604D. The synthesized element updated in real time in the comment editing area varies according to different user selections.
[0172] If the image elements of the team emblems of both teams are selected, the update element 603D is hidden.
[0173] In the embodiments of the present application, other unselected elements are updated according to the selected elements, which improves the efficiency of displaying elements and the efficiency of users editing interactive information.
[0174] In some embodiments, the first interactive information includes a first time point, and the first time point represents the release time of the first interactive information; after step 304, refer to Figure 3D , Figure 3D is a schematic flowchart of the interactive processing method for the live broadcast room provided by the embodiments of the present application. Execute Figure 3D Steps 3041 to 3042 are specifically described below.
[0175] In step 3041, in response to a playback operation for the live broadcast room, a live broadcast playback video is displayed.
[0176] Exemplarily, when the live broadcast corresponding to the first interactive information ends, step 3041 can be executed.
[0177] In step 3042, at least one identifier of the first time point in the live broadcast playback video is displayed.
[0178] Exemplarily, the identifier of the first time point is an element in the first interactive information.
[0179] In some embodiments, step 3042 can be implemented in the following manner: the identifier of each first time point is displayed in the time axis of the live broadcast playback video, where the position of the identifier of the first time point in the time axis corresponds to the first time point, and the identifier of the first time point is used to trigger the playback progress to jump to the first time point when triggered.
[0180] The identifier can be a simplified version or the original version of the element in the interactive information. Refer to Figure 7A , Figure 7A is a schematic diagram of the display manner of the live broadcast playback video provided by the embodiments of the present application; multiple time point identifiers 703A are displayed in the time axis 702A of the live broadcast playback video 701A, and the icon of each time point identifier 703A is presented as the pattern of the element included in the corresponding interactive information.
[0181] In response to any time point identifier 703A being triggered, the jump is made to the position corresponding to the triggered time point identifier 703A.
[0182] In some embodiments, before step 3042, the following processing is performed:
[0183] Obtain the elements carried in each first interaction message sent in the live broadcast room; obtain the first time point corresponding to each first interaction message; based on each first time point and the elements carried in each first interaction message, annotate the time axis of the recorded screen video of the live broadcast room, and use the annotated recorded screen video as the live broadcast replay video.
[0184] Exemplarily, the recorded screen video of the live broadcast room is generated after the current live broadcast in the live broadcast room. The recorded screen video includes the video content of the entire live broadcast, or the recorded screen video only includes the video content from the moment the user clicks the control to start recording the screen to the end of the live broadcast. Locate the time axis based on the first time point to obtain multiple time point positions that can be used for annotation, and use the elements carried in the interaction message as the visual representation of the label to obtain the live broadcast replay video.
[0185] In some embodiments, after step 304, refer to Figure 3E , Figure 3E which is a schematic flowchart of the interaction processing method for the live broadcast room provided by the embodiments of the present application. Execute Figure 3E step 3043 in
[0186] to display at least one video clip collection corresponding to the live broadcast replay video.
[0187] Here, the video clip collection includes multiple video clips in the live broadcast replay video. The labels of each video clip collection are different, and the label is a common element, and the common element is an element that appears in the first interaction messages of multiple video clips. Figure 7B , Figure 7B is a schematic diagram of the display method of the live broadcast replay video provided by the embodiments of the present application; in the video clip collection 702B, each video clip corresponds to the same type of element, and the pattern of this element is used as the label 701B of the video clip collection 702B. In the video clip collection 704B, each video clip corresponds to the same type of element, and the pattern of this element is used as the label 703B of the video clip collection 704B.
[0188] In some embodiments, before step 3043, refer to Figure 3F , Figure 3F which is a schematic flowchart of the interaction processing method for the live broadcast room provided by the embodiments of the present application. Execute Figure 3F steps 3051 to 3055 in
[0189] In step 3051, obtain the elements carried in each first interaction message sent in the live broadcast room.
[0190] Exemplarily, the first interaction information includes a timestamp (representing the release time) as a time point, an element, and an identifier of the account that sent the first interaction information. The element is extracted from the first interaction information, and the element can be a composite element or a first element and a second element.
[0191] In step 3052, obtain the first time point included in each first interaction information.
[0192] Extract the timestamp data from the first interaction information to obtain the time point.
[0193] In step 3053, perform the following processing for each first time point: Based on the first time point and a preset segment length, perform segment division processing on the recorded screen video of the live broadcast room to obtain a video segment corresponding to each first time point.
[0194] Exemplarily, the position corresponding to the video segment in the time axis of the recorded screen video includes the first time point. The preset segment length can be set according to the actual application scenario.
[0195] In step 3054, classify each video segment according to the element carried in the first interaction information corresponding to each video segment to obtain the type of the video segment.
[0196] Each element corresponds to one type, so the elements of the video segments of the same type are common elements.
[0197] In step 3055, divide the video segments into corresponding video segment collections according to the types to which they belong.
[0198] Exemplarily, different types of video segments are divided into corresponding different types of video segment collections, so that the elements of different types of video segment collections are different, and the elements of the video segments in the same video segment collection are common elements.
[0199] In some embodiments, a browsing diagram showing the time point on the time axis and the playback method of the video collection segment can be displayed in the same video playback interface. If the user clicks any one of them, the clicked playback method starts to play.
[0200] In the embodiments of the present application, multiple playback methods are provided, which can improve the efficiency of the user to find the playback video and save the time required for the user to obtain the playback video.
[0201] In some embodiments, for users who have performed interactions during the live broadcast, corresponding playback videos or playback video collections are generated according to the interaction information already published by the users. Or, on the basis of recommending the playback videos of the interaction information published by themselves, playback videos corresponding to other interacting users can also be recommended to the interacting users.
[0202] In some embodiments, for users who did not watch the live stream or did not perform interactions during the live stream, the time points and tags of similar users are recommended through big data. Similar users can be determined in the following ways: According to the object classification model, similarity calculation methods (such as cosine similarity, Euclidean distance, etc.) are used to achieve the division and classification of user groups, and user groups with preferences similar to those of non-interactive users are found; According to the results of facial expression tagging analysis, the visual expression tags most likely to match the preferences of non-interactive users are recommended for them. The tag recommendation function can be implemented using recommendation system algorithms and machine learning algorithms.
[0203] In some embodiments, the embodiments of the present application also propose an interaction processing method for a live broadcast room, referring to Figure 3G , Figure 3G is a schematic flowchart of the interaction processing method for the live broadcast room provided by the embodiments of the present application.
[0204] In step 311, the live broadcast room is displayed.
[0205] Here, the live broadcast room includes a comment editing entry.
[0206] In step 312, in response to a trigger operation on the comment editing entry, a comment editing area is displayed.
[0207] Here, the comment editing area includes at least one first element, and the first element changes according to the current live broadcast progress of the live broadcast room. The first element is used to generate first interaction information, and the first interaction information is used to be sent into the live broadcast room.
[0208] Exemplarily, the principles of step 311 and step 312 can refer to Figure 3A above, and will not be elaborated here.
[0209] In the embodiments of the present application, at least one first element in the live broadcast room and the comment editing area is displayed, and the first element is related to the current live broadcast progress of the live broadcast room; By displaying elements related to the live broadcast progress, the user experience of the live broadcast progress is improved, and users are promoted to perform interaction operations during the live broadcast. The first interaction information sent is generated based on the first target element, and the finally sent interaction information is related to the live broadcast progress. Compared with the prior art's interaction method of only displaying fixed elements, the elements displayed are changed in real time according to the live broadcast progress, and there is no need for users to actively search for elements that match the live broadcast progress, which can improve the interaction efficiency of users in the live broadcast room.
[0210] Next, an exemplary application of the interaction processing method for the live broadcast room of the embodiments of the present application in an actual application scenario will be described.
[0211] During the live broadcast, users can communicate with the host or other users watching the live broadcast through a variety of different expression elements. In the related technologies, the expression elements provided in the live broadcast room are single, which is difficult to meet the interactive needs of users during the live broadcast. For example, when users watch a text live broadcast without a video screen, they usually lack a real-time sense of participation and interaction. When watching the corresponding video clips after the game, they cannot immediately connect the video content with what they saw during the text live broadcast. Another example is the video playback after the live broadcast. The existing playback videos are usually manually edited into multiple segments, and the video collections seen by all users are the same.
[0212] Using chatting and regular expressions as the user participation method in the live broadcast, this form of user participation and emotional expression is single. The display form of text lacks a vivid sense of picture. At the same time, typing has a higher participation cost. And the fixed expressions cannot accurately express the user's current mental activities, lacking a sense of real-time and interactivity as a whole. In the embodiment of the present application, according to the real-time progress of the game, the corresponding text slogans and expression elements are obtained, which is more real-time. Users can freely create personalized elements by freely combining different elements, improving the sense of real-time and interactivity.
[0213] In the existing schemes of element fusion, only two fixed expression elements can be fused, and the fixed expression elements cannot change in real time according to the usage scenario. In the embodiment of the present application, multiple elements are combined, and the elements will change according to the scenario, which can generate more different element combinations, and the user's creative freedom is higher, combining playability and real-time.
[0214] The manually edited video segments cannot be customized for different users. In the embodiment of the present application, by obtaining the expression output of the user when watching the live broadcast, after the live broadcast ends, a video collection that conforms to the user's viewing mood at that time is sorted out for the user, providing personalized content for the user. For users who do not participate in the expression output, through big data recommendation, a video collection content that conforms to their preferences is recommended for the users. Meeting the personalized information acquisition requirements of different users.
[0215] The embodiment of the present application provides an interactive processing method for a live broadcast room. When users watch a text live broadcast, a real-time participation interactive method is provided for the users. According to the change of the real-time progress of the game field, elements related to the current live broadcast content are displayed, facilitating users to quickly query appropriate expression elements; and, users participate in comment interaction in the form of freely set combined elements. The sent elements contain the user's current emotional information and the time information of the game, which can be used as the editing basis during video playback to help users generate video playback segments that more conform to their personal preferences.
[0216] Hereinafter, the interactive processing method for the live broadcast room provided by the embodiment of the present application will be explained with reference to the drawings. Refer toFigure 9 , Figure 9 is a schematic flowchart of the interactive processing method for a live broadcast room provided by an embodiment of the present application, described with a terminal device as the execution subject.
[0217] In step 901, elements updated according to the live broadcast content are displayed in the live broadcast interface.
[0218] Exemplarily, the elements displayed in the live broadcast interface include two types. One type is fixed elements, that is, elements irrelevant to the live broadcast content; the other type is elements updated according to the live broadcast content, hereinafter referred to as real-time elements, and such elements change with the live broadcast content. The types of elements include but are not limited to: text slogans (text elements), emoji elements (emoji elements), and image elements.
[0219] Exemplarily, in an embodiment of the present application, a sports game is taken as an example for explanation. The elements updated according to the live broadcast content will change with the progress of the game. For example, the real-time elements change with the success or failure of a team's goal and the score situation.
[0220] Exemplarily, when a user browses the live broadcast of a game in an application, an interactive chat area, as well as a comment editing area for element creation and element display, can be triggered in the live broadcast room interface.
[0221] Refer to Figure 5A , Figure 5A is a schematic diagram of the live broadcast room interface provided by an embodiment of the present application; the live broadcast room interface 501A is an interface for sports text live broadcast, and the live broadcast content 502A and the comment interaction area 503A are displayed in the live broadcast room interface 501A. If a user wants to join the interaction after viewing the interaction content or the live broadcast content in the comment interaction area 503A, the comment editing area can be triggered.
[0222] In response to a trigger operation on the live broadcast room interface 501A, it is transferred to Figure 5A the interface of Figure 5B Refer to Figure 5B , Figure 5B is a schematic diagram of the live broadcast room interface provided by an embodiment of the present application; the live broadcast content 502A and the comment editing area 504A are displayed in the live broadcast room interface 501A, where "LETS GO" and "DEFENCE" are elements updated according to the live broadcast content.
[0223] Continue to refer to Figure 9 , in step 902, in response to an element selection operation, the synthesized element of the selected element is displayed, and in response to a sending operation, the interactive information carrying the synthesized element is displayed.
[0224] Exemplarily, the user can click to select one or more elements. After the user clicks, the synthesized element pattern is displayed in real time in the comment editing area. After the user clicks to send the synthesized element pattern, the synthesized element pattern is sent to the chat area.
[0225] For ease of understanding, refer to Figure 4 , Figure 4 which is a schematic diagram of the element synthesis principle provided by an embodiment of the present application. Multiple different expression elements will be displayed in the live broadcast room. Figure 4 shows the element synthesis principle 401, where element 1 and element 2 have graphic transformations and are combined with each other to obtain element 3. Element 1 is an eye, element 2 is a crying expression with both eyes, and element 3 is a crying expression with one eye.
[0226] Exemplarily, the elements used to generate the synthesized elements include, but are not limited to, the following types: image elements, text elements, and expression elements.
[0227] Refer to Figure 6A , Figure 6A which is a schematic diagram of element synthesis provided by an embodiment of the present application. A variety of first elements 602A that change with the live broadcast content are displayed in the comment editing area 601A. The types of the first elements include, but are not limited to: image elements (such as the team logo image), text elements (such as the slogans "DEFENCE" and "LETS GO"), and expression elements. The expression elements can be used to represent the emotions of the user.
[0228] Refer to Figure 5D , Figure 5D which is a schematic diagram of the live broadcast room interface provided by an embodiment of the present application; in the comment editing area 504A, the user selects three different elements, and a third element 507A is generated based on the three elements. Among them, the comment interaction area 505A is Figure 5A the fully expanded comment interaction area 503A in
[0229] Refer to Figure 5E , Figure 5E which is a schematic diagram of the live broadcast room interface provided by an embodiment of the present application; in response to the send operation for the third element 507A, a first interaction message 508A including the third element is displayed in the comment interaction area 505A.
[0230] Exemplarily, based on the above Figure 5D and Figure 5EAs can be seen from the examples, the elements used to generate synthetic elements include, but are not limited to, the following types: the team logo graphics (image elements) of the team, the text slogan graphics (text elements), and the user emoji graphics (emoji elements). Among them, the team logo of the team comes from the two teams in the current game in the live content, and both the text slogan and the user emoji graphics are related to the current game progress. After the user selects multiple elements, several elements are combined and redrawn to generate the synthetic element.
[0231] Reference Figure 5C , Figure 5C is a schematic diagram of the live broadcast room interface provided by the embodiment of the present application; if the user selects the first element and sends the first element, the first interaction information 506A is formed in the comment interaction area 505A, and the first interaction information 506A includes the content "LETS GO" of the first element.
[0232] In some embodiments, the process of element synthesis includes, but is not limited to, the following situations: synthesizing the team logo element into the text slogan, extracting the main color of the team logo as the main color of the text, extracting the main color of the team logo as the unified outline effect of the element, etc.
[0233] For example, the following explains the synthesis process. The user triggers the comment editing area and makes a selection. Suppose the user selects the first team logo, the first slogan, and an emoji. At this time, the main text of the synthetic element uses the main color of the first team logo, places the main part of the first team logo in the text, and uses the secondary color of the first team as the overall outline of the element.
[0234] Reference Figure 6B , Figure 6B is a schematic diagram of element synthesis provided by the embodiment of the present application. The arrows between the comment editing area 601B, the comment editing area 602B, and the comment editing area 603B represent the sequence. When the user selects multiple elements for the comment editing area 601B, the selected elements 604B in the comment editing area 602B are formed. Among them, in the embodiment of the present application, for the convenience of viewing, the elements are extracted from the comment editing area, which does not mean that the graphics are displayed separately from the comment editing area during actual display. The selected elements 604B are synthesized to obtain the synthetic element 605B. If the user cancels the selection of some elements and re-selects them as the selected elements 606B, after the update, the synthetic element 605B is switched to the synthetic element 607B in the comment editing area 603B.
[0235] Reference Figure 6C , Figure 6CIt is a schematic diagram of element synthesis provided by an embodiment of the present application; the arrow between the comment editing areas 601C and 602C represents the sequence. Two elements in the comment editing area 601C are triggered and are shown as the synthesized element 604C and the selected element 603C in the comment editing area 602C.
[0236] In some embodiments, if the user changes the selection of elements, cancels the first team and the first emoji, and changes to the second team. At this time, the element display area shows the changes of the synthesized elements in real time. The main text of the element becomes mainly the main color of the logo of the second team, with the secondary color of the second team as the outline, and combines the main elements of the second team in the middle of the text.
[0237] For example: if Figure 6C the logo graphics of other teams in the comment editing area are selected and presented as Figure 6D , Figure 6D It is a schematic diagram of element synthesis provided by an embodiment of the present application; compared with the comment editing area 601C, the logo 2 in the comment editing area 601D is selected, and the comment editing area 602D shows the updated element 603D associated with the current live content of the team of logo 2. The content of the updated element 603D is the text element "LOSE", and the logo 2 and the updated element 603D are synthesized into the synthesized element 604D. The synthesized elements updated in real time in the comment editing area vary according to different user selections.
[0238] In some embodiments, when the rhythm of the game between the two sides changes during the game, the elements in the elements will change following the game content. Continuing to refer to Figure 6C , Figure 6C the text element "MVP" in is displayed when the corresponding team presents a winning state, and the emoji element part also changes following the game rhythm. The user can form different synthesized elements according to different element selections.
[0239] In the embodiments of the present application, the user can customize and combine synthesizable elements: the generated elements are composed of single or multiple elements, and the user can freely select each element. After the user has made the selection, the selected elements are synthesized into a pattern to form a synthesized element and sent to the comment interaction area. The elements created are different according to the different elements freely selected by each user.
[0240] Continuing to refer to Figure 9 , in step 903, in response to the end of the live broadcast, a live broadcast replay video is generated according to the interaction information in the live broadcast content.
[0241] Exemplarily, for users who have sent elements during a live broadcast and want to view the playback clips of the video clip after the live broadcast ends, a live broadcast playback video that fully conforms to the user's viewing mood at that time can be generated according to the mood elements contained in the elements sent by the user at that time.
[0242] The live broadcast playback video that fully conforms to the user's mood at that time refers to a live broadcast playback video that labels the mood elements corresponding to the interaction information as the labels of the video clips where the interaction information is sent. There are two presentation methods: using the time point when the interaction information is sent as the time node label and the interaction information as the icon of the node to form a live broadcast playback video with node labels; intercepting the video clips associated with the time points when the interaction information is sent to form a collection of different video clips corresponding to the same interaction information.
[0243] There are two forms of presenting the video collection. Form one includes: splicing all the clips into a video in chronological order, using the mood elements as visual guidance (time node labels, clicking on the node labels can jump to the corresponding clips), and displaying them on the timeline. In response to the trigger operation for the node, it is positioned to the moment corresponding to the node to view the corresponding video clip.
[0244] Reference Figure 7A , Figure 7A is a schematic diagram of the live broadcast playback video display method provided by the embodiments of the present application; in the timeline 702A of the live broadcast playback video 701A, multiple time point identifiers 703A are displayed, and the icon of each time point identifier 703A is presented as the pattern of the element included in the corresponding interaction information.
[0245] In some embodiments, the content of the synthetic elements customized by the user is related to the progress of the competition. Each element contains a time stamp and the mood expressions selected by the user. The user sends elements at certain nodes during the process of watching the text live broadcast. After the competition ends, a live broadcast playback video that fully conforms to the user's mood at that time can be generated according to the mood elements contained in the elements sent by the user at that time, so as to help the user efficiently review the competition scenes that accurately touch the user himself.
[0246] Form two includes: after integrating and classifying all the mood elements, forming an aggregated theme of multiple video clips with the clips corresponding to the same mood.
[0247] Reference Figure 7B , Figure 7BIt is a schematic diagram of the live replay video display method provided by an embodiment of the present application; in the video clip collection 702B, each video clip corresponds to the same type of element, and the pattern of this element is used as the first label 701B of the video clip collection 702B. In the video clip collection 704B, each video clip corresponds to the same type of element, and the pattern of this element is used as the second label 703B of the video clip collection 704B.
[0248] In some embodiments, for users who have not watched the live content during the live broadcast, according to the labels of the users, a similar group of users who have watched the live broadcast is found. The emotional labels of the similar group of users when watching the live broadcast are refined, and a video replay collection aggregated by viewing emotions is generated for the users who have not watched the live broadcast. According to the differences of users, the video clip collections of the live replay videos watched by each user are different. If presented in the form of a complete video, the time points in the video are different.
[0249] In some embodiments, when a large number of users participate in the generation and sending of interactive information during the live broadcast, the label data of these users is recorded. When a new user who has not participated in the live broadcast views the live replay video after the live broadcast ends, the label of the new user can be matched with the big data to find a user group similar to the new user. The key emotional labels marked by this part of the user group are mapped to the corresponding video timeline, and the corresponding video clips are refined. A series of video replay clips that meet the user's preferences are generated for the new user.
[0250] In some embodiments, for a live sports game broadcast, the types of elements include: team logo elements, text elements (formed by slogans), and emoji elements.
[0251] For team logo elements, since the two teams in the live broadcast are known in advance, the vector graphics of all the team logos can be stored in a data set. Analyze the graphics to obtain at least one main color and at least one secondary color of each team logo, and store them in the data set. Analyze the graphics, extract each graphic main element, and store it in the database. The graphic main element can be used to be displayed in the comment editing area.
[0252] For text elements, a text collection for describing sports games is obtained in advance. For example, it describes the corresponding text slogans when different teams score or defend at different score states of both sides. There is a mapping relationship between each piece of text in the text collection and the corresponding live content. When the live content includes the preset live content in the mapping relationship, the corresponding text is called according to the mapping relationship to display the text material.
[0253] For expression elements, before the start of the game live broadcast, expression elements used to describe the emotions of the audience watching the game are pre-stored. When the corresponding live content is triggered, the pre-stored expression elements are displayed. During the game live broadcast, text data of user comments related to the game on the Internet is collected and pre-processed. The pre-processing includes removing special characters, punctuation marks, and stop words, and converting the text to lowercase, etc.
[0254] Exemplarily, for the pre-stored expression elements, the emotion type of each element can be determined in the following way: Use sentiment analysis technology to perform sentiment classification on user comments, and classify the comments as positive, negative, or neutral. Sentiment analysis can be trained and predicted using machine learning algorithms (such as Naive Bayes, Support Vector Machine, etc.) or deep learning models (such as Recurrent Neural Network, Convolutional Neural Network, etc.).
[0255] For the expression elements classified by emotion type, an expression element library is constructed, which contains expression elements of different emotion categories. An existing expression element library or a custom expression element library can be used. For each emotion category, according to the comment samples of this category, keywords and phrases related to this emotion are extracted. Techniques such as the bag-of-words model, Term Frequency–Inverse Document Frequency (TF-IDF), etc. can be used to extract keywords and phrases.
[0256] In some embodiments, during the live broadcast, the real-time elements to be displayed can be determined in the following way: Obtain user comments in real time, call the keyword and phrase extraction method to extract the keywords and phrases therein. Use a keyword and phrase matching algorithm (such as cosine similarity, edit distance, etc.) to calculate the similarity between the keywords and phrases in the user comments and the keywords and phrases of each emotion category. According to the similarity score, select the emotion category that best matches the user comments, and select an expression element from the expression element library of this category as the matching result, and display the expression element corresponding to the matching result on the user's terminal device.
[0257] Reference Figure 8 , Figure 8 is the preset element table provided by the embodiments of this application. The matchup situation is different game description texts. According to the mapping relationship between each game description text and the corresponding team's expression elements and copywriting (text elements), when the corresponding game description text is displayed in the live content, or when the situation conforms to the game description text, the corresponding expression element or text element is displayed.
[0258] In some embodiments, among the elements displayed in the comment editing area, the user can select one or more. During the selection process, the displayed elements can be updated according to the user's selection. For example, when the user selects a team, the slogans and expressions in the panel are obtained from the currently corresponding ones in the dataset; refer to Figure 6D , when the user selects the image element of team logo 2, the updated element 603D corresponding to team logo 2 is displayed. Another example: when the user selects two teams, the common slogans and expressions are shown in the panel. The elements selected by the user are combined and synthesized using image synthesis technology in the order of team, slogan, and expression. When implementing image synthesis, the position and size of the elements need to be considered to ensure that they are correctly combined. The image transformation functions in the image processing library can be used to achieve this.
[0259] For example, when the user selects a team, the main color of the team in the database is filled into the vector graph of the text slogan; the secondary color of the team is used as the overall outline of the synthesized graph; the core element of the team is displayed at any position of the text slogan graph.
[0260] In some embodiments, image processing algorithms are used to obtain the main color and secondary color of the team. The following are two common methods: K-means clustering algorithm: The pixel values of the team logo image are divided into multiple color clusters, and then the largest color cluster is selected as the main color of the team logo. The K-means clustering algorithm can be implemented using the clustering function in the image processing library. Color histogram analysis algorithm: Calculate the color histogram of the team logo image, and then select the color with the highest value in the histogram as the main color of the element. The color histogram analysis algorithm can be implemented using the histogram function in the image processing library. An image edge detection algorithm is used to detect the edges of the elements and add outline lines to them. The following are two commonly used image edge detection algorithms: Canny algorithm: The image is smoothed using a Gaussian filter, then the gradient and non-maximum suppression of the image are calculated, and finally a double-threshold algorithm is used to detect the edges. The Canny algorithm can be implemented using the edge detection function in the image processing library. Sobel algorithm: The gradient of the image is calculated using the Sobel operator, and then a threshold algorithm is used to detect the edges. The Sobel algorithm can be implemented using the edge detection function in the image processing library.
[0261] In some embodiments, when the user does not select a team or selects two teams, common expressions are used for the slogans and expressions, and common colors are used for the main color and secondary color. That is, when the user selects elements without associated elements for the synthesis of new elements, or selects mutually exclusive elements (the team logos of two teams) for the synthesis of new elements, common expression elements are displayed.
[0262] In some embodiments, for users who have watched live content, the interaction information already published by the users is extracted. Each piece of interaction information carries a corresponding live broadcast progress timestamp, and video segments are extracted based on the timestamps carried by each piece of interaction information. For example: obtain the video segment of 15s after the timestamp, and use JavaScript to calculate the timestamp and position. When a user who has watched the live broadcast views the video of the complete live broadcast replay, the timestamp data of the interaction information is mapped to the position on the timeline of the video, and the elements carried by the interaction information are displayed at the position corresponding to the timestamp in the timeline of the replay video. Also, when the user clicks on an emoji, the video jumps to the node of the corresponding video.
[0263] In some embodiments, traverse the complete live broadcast replay video, extract the timestamps of each piece of interaction information, and intercept video segments of a preset length near each timestamp. Classify the video segments according to the elements corresponding to each piece of interaction information included in the video segments to obtain multiple different types of video segments, and add the video segments of each type to different sets. Each set of video segments can also be combined into a new video file or playlist respectively. Video editing libraries or frameworks such as MoviePy or FFmpeg can be used to implement the functions of video merging and editing.
[0264] In some embodiments, for users who have not interacted during the live content, collect a large amount of element data published by users during the live broadcast and store it in a database. Use data mining techniques and machine learning algorithms to analyze and process the data, store the emojis corresponding to each user, as well as the timestamps and video segments corresponding to the emojis. According to information such as the behavior, interests, and preferences of the users, use machine learning algorithms and data mining techniques to build an object classification model. Use clustering algorithms (such as K-means, DBSCAN, etc.) to divide users into different groups, or use classification algorithms (such as decision trees, random forests, etc.) to classify users. According to the object classification model, use similarity calculation methods (such as cosine similarity, Euclidean distance, etc.) to achieve the division and classification of user groups, and find user groups with preferences similar to those of the users who have not participated in the interaction. For each user group, analyze its emoji data to find the video tags that are most likely to match the preferences of the users who have not participated in the interaction. Association rule mining algorithms and collaborative filtering algorithms can be used to implement the analysis and mining of emoji data. According to the results of the emoji tag analysis, recommend the emoji tags that are most likely to match the preferences of the users who have not participated in the interaction for the users who have not participated in the interaction. Recommendation system algorithms and machine learning algorithms can be used to implement the tag recommendation function. According to the recommended emojis, display the emojis on the video timeline for the users who have not participated in the interaction, and display the collection of video segments aggregated by the emojis.
[0265] Exemplarily, the method of recommending a video collection and the timeline of the playback video to users who have not participated in the interaction is also applicable to users who have watched the video. For example: In the timeline of the playback video presented to users who participate in the interaction, it includes both the expression elements corresponding to the interaction information sent by the user himself and the expression elements corresponding to the interaction information posted by other users similar to this user.
[0266] The embodiments of the present application have the following effects:
[0267] 1. According to the real-time progress of the live broadcast, the embodiments of the present application obtain the corresponding text elements and emojis, which is more real-time. Users can freely create personalized elements by freely combining different elements, improving the sense of real-time and interactivity.
[0268] 2. The embodiments of the present application realize the combination of multiple elements, and the elements will change according to the scene, which can generate more different combinations of elements. Users have a higher degree of creative freedom, combining playability and real-time.
[0269] 3. By obtaining the expression output of the user himself during the live broadcast, after the live broadcast ends, the embodiments of the present application sort out a video collection that conforms to the user's mood at that time for the user, providing personalized content for the user. For users who do not participate in the expression output, through big data recommendation, video collection content that conforms to their preferences is recommended for the users. It meets the personalized information acquisition requirements of different users.
[0270] 4. Because the expression elements output by the user reflect the real emotions of the user when watching the text live broadcast, based on this, a live broadcast playback video that conforms to the user's mood at that time can be generated for the user after the text live broadcast ends, which can help the user efficiently review the game scene, improve the efficiency of the user's content search, and thus enhance the content viewing experience.
[0271] Next, the implementation of the interactive processing device 455 in the live broadcast room provided by the embodiments of the present application as a software module will be continued. In some embodiments, as Figure 2 shown, the software module in the interactive processing device 455 in the live broadcast room stored in the memory 450 may include: a display module 4551 for displaying the live broadcast room; a display module 4551 for displaying a comment editing area, where the comment editing area includes at least one first element, and the first element changes according to the current live broadcast progress of the live broadcast room; a display module 4551 for displaying at least one first target element in a selected state in response to an element selection operation, where the first target element is any one of the selected first elements; a display module 4551 for displaying first interaction information in the comment interaction area of the live broadcast room in response to a send operation, where the first interaction information includes the first target element.
[0272] In some embodiments, a display module 4551 is configured to, when displaying a comment editing area, display at least one second element to be selected in the comment editing area, where the second element is an element that remains fixed during the live broadcast in the live broadcast room; when displaying at least one first target element in a selected state in response to an element selection operation, the method further includes:
[0273] Display at least one second target element in a selected state in the comment editing area, where the second target element is any one of the selected second elements, and the second target element is used to combine with the first target element to generate first interaction information.
[0274] In some embodiments, the first element is used to characterize the current live broadcast progress feature of the live broadcast room, and the live broadcast progress feature includes at least one of the following:
[0275] The interaction progress among multiple objects in the live broadcast room; keywords that appear in the live broadcast room; objects that appear in the live broadcast room.
[0276] In some embodiments, an acquisition module 4552 is configured to, after displaying the live broadcast room, acquire the live broadcast progress feature of the current live broadcast content in the live broadcast room; the method for determining at least one first element to be selected in the comment editing area includes any one of the following:
[0277] When the live broadcast progress feature is the interaction situation among multiple objects in the live broadcast room, perform sentiment analysis processing on the interaction situation by invoking a neural network model to obtain the current emotion type of the live broadcast room, and use the relevant elements of the current emotion type as at least one first element to be selected; when the live broadcast progress feature is an object that appears in the live broadcast room, perform graphic analysis processing on the object that appears in the live broadcast room, and use the obtained image elements as at least one first element to be selected; when the live broadcast progress feature is a keyword that appears in the live broadcast room, generate text elements based on the keyword; perform search processing based on the live broadcast progress feature to obtain search content, and perform graphic analysis processing on the search content to obtain at least one first element to be selected.
[0278] In some embodiments, the first element is used to characterize the playing unit where the current live broadcast progress of the live broadcast room is located, where the playing unit is obtained by dividing the live broadcast content of the live broadcast room.
[0279] In some embodiments, the types of the playing unit include: a time unit obtained by periodically dividing according to a preset duration; a sub-shot unit obtained by dividing according to different camera shots in the live broadcast room; a plot unit obtained by dividing according to different plots in the live broadcast room.
[0280] In some embodiments, the types of the first element include:
[0281] Image elements, where the image elements are obtained by extracting from the video frames corresponding to the live content, or are obtained by searching according to the live progress characteristics; text elements, where the text elements are obtained by performing speech recognition on the audio data corresponding to the live content, or are obtained by performing text recognition on the video frames corresponding to the live content; expression elements, where the expression elements are elements related to the emotion type represented by the live content, and the emotion type is obtained by performing emotion recognition on at least one of the video frames, audio data, and interaction situation corresponding to the live content.
[0282] In some embodiments, the obtaining module 4552 is configured to obtain the current live progress characteristics of the live broadcast room after the live broadcast room is displayed; query at least one candidate first element from the candidate element set based on the live progress characteristics.
[0283] In some embodiments, the obtaining module 4552 is configured to obtain the preview information of the live broadcast room before querying at least one candidate first element from the candidate element set based on the live progress characteristics, where the preview information is used to be displayed before the start of the live broadcast in the live broadcast room; perform feature extraction processing on the preview information to obtain attribute features, where the attribute features include at least one of the following: image features, text features, and expression features; perform graphic analysis processing based on the image features to obtain image elements, where the image elements are used as candidate elements; perform image conversion processing based on the text features to obtain text elements, where the text elements are used as candidate elements; use the expression elements corresponding to the expression features as candidate elements; form a candidate element set with each candidate element, where each candidate element is used as a candidate first element when being queried based on the live progress characteristics.
[0284] In some embodiments, when there are multiple selected first target elements, the multiple first target elements are carried in the first interaction information in a synthesized manner; the display module 4551 is configured to display a synthesized element, where the synthesized element is obtained by fusing the multiple first target elements, and the synthesized element is used to be carried in the first interaction information.
[0285] In some embodiments, the display module 4551 is configured to, after displaying the synthesized element, the method further includes:
[0286] In response to a cancellation operation on any first target element, switch the cancelled first target element from the selected state to the unselected state; update the synthesized element based on the current first target elements.
[0287] In some embodiments, the display module 4551 is configured to, before displaying the composite element, when multiple first target elements include text elements, fill the text elements with the color of other target elements and composite the graphics of the other target elements with the text elements to form a composite element, where the other target elements are non-text first target elements; when the multiple first target elements do not include text elements, stack each first target element to form a composite element; when the number of at least one first target element is multiple, splice the first target elements having an association relationship among the multiple first target elements and stack the splicing result with the other target elements to form a composite element, where the other target elements are elements other than the first target elements having an association relationship among the multiple first target elements, and the composite element is obtained by calling an image transformation function based on the multiple first target elements.
[0288] In some embodiments, the display module 4551 is configured to display the associated element of the first target element in response to the existence of an associated element for the selected first target element; and hide the associated elements of the mutually exclusive first target elements in response to the selected multiple first target elements being mutually exclusive elements.
[0289] In some embodiments, the first interaction information includes a first time point, and the first time point represents the release time of the first interaction information; the display module 4551 is configured to, after displaying the first interaction information in the comment interaction area of the live broadcast room, in response to a playback operation for the live broadcast room, display the live broadcast playback video, and
[0290] display an identifier of at least one first time point in the live broadcast playback video, where the identifier of the first time point is an element in the first interaction information.
[0291] In some embodiments, the display module 4551 is configured to display an identifier of each first time point in the timeline of the live broadcast playback video, where the position of the identifier of the first time point in the timeline corresponds to the first time point, and the identifier of the first time point is used to trigger the playback progress to jump to the first time point when triggered.
[0292] In some embodiments, the acquisition module 4552 is configured to, before displaying an identifier of each first time point in the timeline of the live broadcast playback video, acquire the elements carried by each first interaction information sent in the live broadcast room; acquire the first time point corresponding to each first interaction information; annotate the timeline of the screen recording video of the live broadcast room based on each first time point and the elements carried by each first interaction information, and use the annotated screen recording video as the live broadcast playback video.
[0293] In some embodiments, the display module 4551 is configured to display at least one video clip collection corresponding to the live replay video after displaying the first interaction information in the comment interaction area of the live broadcast room, where the video clip collection includes multiple video clips in the live replay video, each video clip collection has a different label, and the label is a common element, and the common element is an element that appears in the first interaction information of multiple video clips.
[0294] In some embodiments, the acquisition module 4552 is configured to, before displaying at least one video clip collection corresponding to the live replay video, acquire the elements carried in each first interaction information sent in the live broadcast room; acquire the first time point included in each first interaction information; and perform the following processing for each first time point:
[0295] Based on the first time point and a preset clip length, perform clip division processing on the screen recording video of the live broadcast room to obtain a video clip corresponding to each first time point, where the position of the video clip corresponding to the time axis of the screen recording video includes the first time point; classify each video clip according to the elements carried in the first interaction information corresponding to each video clip to obtain the type of the video clip; and divide the video clips into corresponding video clip collections according to the types to which they belong.
[0296] In some embodiments, the software module in the interaction processing device 455 of the live broadcast room stored in the memory 450 may include: a display module 4551 configured to display the live broadcast room, where the live broadcast room includes a comment editing entry; the display module 4551 is further configured to, in response to a trigger operation on the comment editing entry, display a comment editing area, where the comment editing area includes at least one first element, the first element changes according to the current live broadcast progress of the live broadcast room, and the first element is used to generate a first interaction information for sending to the live broadcast room.
[0297] An embodiment of the present application provides a computer program product, which includes a computer program or computer executable instructions, and the computer program or computer executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or computer executable instructions from the computer-readable storage medium, and the processor executes the computer program or computer executable instructions, so that the electronic device executes the interaction processing method of the live broadcast room in the above embodiments of the present application.
[0298] An embodiment of the present application provides a computer-readable storage medium storing computer executable instructions, where computer executable instructions or a computer program are stored, and when the computer executable instructions or the computer program are executed by a processor, the processor will be caused to execute the interaction processing method of the live broadcast room provided in the embodiments of the present application. For example, as Figure 3AThe interactive processing method of the live broadcast room shown.
[0299] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; it may also be various devices including one or any combination of the above memories.
[0300] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0301] As an example, the computer-executable instructions may or may not correspond to files in the file system, may be stored as part of a file that stores other programs or data, for example, stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or code portions).
[0302] As an example, the executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.
[0303] In summary, at least one first element in the live broadcast room and the comment editing area is displayed through the embodiments of the present application, and the first element is related to the current live broadcast progress of the live broadcast room; by displaying elements related to the live broadcast progress, the user experience of the live broadcast progress is improved, and the user is promoted to perform interactive operations during the live broadcast. The first interactive information sent is generated based on the first target element, and the finally sent interactive information is related to the live broadcast progress. Compared with the prior art that only displays fixed elements in the interactive manner, the elements displayed are changed in real time according to the live broadcast progress, and there is no need for the user to actively search for elements that match the live broadcast progress, which can improve the interactive efficiency of the user in the live broadcast room.
[0304] The above is only the embodiments of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. An interactive processing method for a live broadcast room, characterized in that, The method includes: Displaying a live broadcast room; Displaying a comment editing area, wherein the comment editing area includes at least one first element, and the first element changes according to the current live broadcast progress of the live broadcast room; In response to an element selection operation, displaying at least one first target element in a selected state, wherein the first target element is any one of the first elements that is selected; In response to a sending operation, displaying first interaction information in the comment interaction area of the live broadcast room, wherein the first interaction information includes the first target element.
2. The method according to claim 1, characterized in that, When displaying the comment editing area, the method further includes: Displaying at least one second element to be selected in the comment editing area, wherein the second element is an element that remains unchanged during the live broadcast process of the live broadcast room; When, in response to an element selection operation, displaying at least one first target element in a selected state, the method further includes: Displaying at least one second target element in a selected state in the comment editing area, wherein the second target element is any one of the second elements that is selected, and the second target element is used to combine with the first target element to generate the first interaction information.
3. The method according to claim 1, wherein The first element is used to represent the current live broadcast progress characteristics of the live broadcast room, and the live broadcast progress characteristics include at least one of the following: The interaction progress among multiple objects in the live broadcast room; Keywords that appear in the live broadcast room; Objects that appear in the live broadcast room.
4. The method according to claim 3, wherein After displaying the live broadcast room, the method further includes: Obtaining the live broadcast progress characteristics of the current live broadcast content of the live broadcast room; The method for determining at least one first element to be selected in the comment editing area includes any one of the following: When the live broadcast progress characteristic is the interaction situation among multiple objects in the live broadcast room, performing sentiment analysis processing by calling a neural network model based on the interaction situation to obtain the current emotion type of the live broadcast room, and using the relevant elements of the current emotion type as the at least one first element to be selected; When the live broadcast progress characteristic is the object that appears in the live broadcast room, performing graphic analysis processing on the object that appears in the live broadcast room, and using the obtained image elements as the at least one first element to be selected; When the live broadcast progress characteristic is the keyword that appears in the live broadcast room, generating text elements based on the keyword; Performing a search process based on the live broadcast progress characteristic to obtain search content, and performing graphic analysis processing on the search content to obtain the at least one first element to be selected.
5. The method according to claim 1, characterized in that, The first element is used to represent the playing unit where the current live broadcast progress of the live broadcast room is located, wherein the playing unit is obtained by dividing the live broadcast content of the live broadcast room.
6. The method according to claim 5, wherein The types of the playing unit include: time units periodically divided according to a preset duration; Shot units divided according to different shots of the live broadcast room; Plot units divided according to different plots of the live broadcast room.
7. The method according to any one of claims 3 to 6, characterized in that The types of the first element include: Image elements, where the image elements are obtained by extracting from video frames corresponding to the live content, or are searched according to the live progress characteristics; Text elements, where the text elements are obtained by performing speech recognition on audio data corresponding to the live content, or by performing text recognition on video frames corresponding to the live content; Emotion elements, where the emotion elements are elements related to the emotion type represented by the live content, and the emotion type is obtained by performing emotion recognition on at least one of the video frames corresponding to the live content, the audio data, and the interaction situation; 8. The method according to claim 1, characterized in that, After displaying the live broadcast room, the method further includes: Obtaining the current live progress characteristics of the live broadcast room; Querying at least one first element to be selected from the candidate element set based on the live progress characteristics; 9. The method according to claim 8, wherein Before querying at least one first element to be selected from the candidate element set based on the live progress characteristics, the method further includes: Obtaining the preview information of the live broadcast room, where the preview information is used to be displayed before the live broadcast of the live broadcast room starts; Performing feature extraction processing on the preview information to obtain attribute features, where the attribute features include at least one of the following: image features, text features, and emotion features; Performing graphic analysis processing based on the image features to obtain image elements, where the image elements are used as candidate elements; Performing image conversion processing based on the text features to obtain text elements, where the text elements are used as candidate elements; Using the emotion elements corresponding to the emotion features as candidate elements; Forming the candidate element set with each candidate element, where each candidate element is used as a first element to be selected when queried based on the live progress characteristics; 10. The method according to claim 1, wherein When multiple first target elements are selected, the multiple first target elements are carried in the first interaction information in a synthetic manner; The method further includes: Displaying a synthetic element, where the synthetic element is obtained by fusing multiple first target elements, and the synthetic element is used to be carried in the first interaction information; 11. The method according to claim 10, wherein After displaying the synthetic element, the method further includes: In response to a cancellation operation on any of the first target elements, switching the cancelled first target element from the selected state to the unselected state; Updating the synthetic element based on the current first target elements; 12. The method according to claim 10, wherein Before displaying the synthetic element, the method further includes: When multiple first target elements include text elements, filling the text elements with the color of other target elements and synthesizing the graphics of the other target elements with the text elements to form the synthetic element, where the other target elements are non-text first target elements; When multiple first target elements do not include text elements, stacking each first target element to form the synthetic element; When the number of the at least one first target element is multiple, splice the first target elements having an association relationship among the multiple first target elements, and superimpose the splicing result with other target elements to form the composite element, where the other target elements are elements other than the first target elements having the association relationship among the multiple first target elements, and the composite element is obtained by calling an image transformation function based on the multiple first target elements.
13. The method according to claim 10, wherein The method further includes: In response to the existence of an associated element for the selected first target element, display the associated element of the first target element; In response to the selected multiple first target elements being mutually exclusive elements, hide the associated elements of the mutually exclusive first target elements.
14. The method according to claim 1, characterized in that The first interaction information includes a first time point, and the first time point represents the release time of the first interaction information; After the first interaction information is displayed in the comment interaction area of the live broadcast room, the method further includes: In response to a replay operation for the live broadcast room, display a live broadcast replay video, and Display an identifier of at least one first time point in the live broadcast replay video, where the identifier of the first time point is an element in the first interaction information.
15. The method according to claim 14, wherein The displaying the identifier of at least one first time point in the live broadcast replay video includes: Display an identifier of each first time point in the time axis of the live broadcast replay video, where the position of the identifier of the first time point in the time axis corresponds to the first time point, and the identifier of the first time point is used to trigger the playback progress to jump to the first time point when triggered.
16. The method according to claim 15, characterized in that, Before displaying an identifier of each first time point in the time axis of the live broadcast replay video, the method further includes: Obtain elements carried by each first interaction information sent in the live broadcast room; Obtain the first time point corresponding to each first interaction information; Based on each first time point and the elements carried by each first interaction information, annotate the time axis of the screen recording video of the live broadcast room, and use the annotated screen recording video as the live broadcast replay video.
17. The method according to claim 1, characterized in that After the first interaction information is displayed in the comment interaction area of the live broadcast room, the method further includes: Display at least one video segment collection corresponding to the live broadcast replay video, where the video segment collection includes multiple video segments in the live broadcast replay video, labels of each video segment collection are different, and the label is a common element, and the common element is an element that appears in the first interaction information of the multiple video segments.
18. The method according to claim 17, wherein Before displaying at least one video segment collection corresponding to the live broadcast replay video, the method further includes: Obtain elements carried by each first interaction information sent in the live broadcast room; Obtain the first time point included in each first interaction information; Perform the following processing for each first time point: Based on the first time point and a preset segment length, perform segment division processing on the recorded screen video of the live broadcast room to obtain video segments corresponding to each of the first time points, where the position corresponding to the video segment in the time axis of the recorded screen video includes the first time point; According to the elements carried in the first interaction information corresponding to each video segment, perform classification processing on each video segment to obtain the type of the video segment; Divide the video segments into corresponding video segment collections according to the types to which they belong.
19. An interactive processing method for a live broadcast room, characterized in that, The method includes: Display a live broadcast room, where the live broadcast room includes a comment editing entry; In response to a trigger operation on the comment editing entry, display a comment editing area, where the comment editing area includes at least one first element, the first element changes according to the current live broadcast progress of the live broadcast room, and the first element is used to generate first interaction information for sending into the live broadcast room.
20. An interactive processing device for a live broadcast room, characterized in that, The device includes: A display module for displaying a live broadcast room; The display module for displaying a comment editing area, where the comment editing area includes at least one first element, and the first element changes according to the current live broadcast progress of the live broadcast room; The display module for, in response to an element selection operation, displaying at least one first target element in a selected state, where the first target element is any one of the first elements that is selected; The display module for, in response to a sending operation, displaying first interaction information in the comment interaction area of the live broadcast room, where the first interaction information includes the first target element.
21. An interactive processing device for a live broadcast room, characterized in that, The device includes: A display module for displaying a live broadcast room, where the live broadcast room includes a comment editing entry; The display module is further configured to, in response to a trigger operation on the comment editing entry, display a comment editing area, where the comment editing area includes at least one first element, the first element changes according to the current live broadcast progress of the live broadcast room, and the first element is used to generate first interaction information for sending into the live broadcast room.
22. An electronic device, characterized in that, The electronic device includes: A memory for storing computer-executable instructions or a computer program; A processor for, when executing the computer-executable instructions or the computer program stored in the memory, implementing the interactive processing method of the live broadcast room according to any one of claims 1 to 19.
23. A computer-readable storage medium stores computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the interactive processing method of the live broadcast room according to any one of claims 1 to 19.
24. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or the computer program, when executed by the processor, implement the interactive processing method of the live broadcast room according to any one of claims 1 to 19.