Live streaming processing method and apparatus, device, storage medium, and product
By automatically identifying and displaying the subject being explained on the live stream page, the problem of cumbersome manual operation is solved, and the interactive experience of live streaming is improved.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2025-09-09
- Publication Date
- 2026-04-23
AI Technical Summary
During the live stream, users need to manually identify and change the object being explained, which makes the operation cumbersome and they may forget to do so, affecting the interactive experience of the live stream.
By automatically identifying and displaying the subject being explained in a specific way based on the live video stream, including adding preset icons to the live page, manual operation is reduced and interaction efficiency is improved.
It enables intuitive location of the subject being explained without additional operation during live streaming, improving the live streaming interactive experience and simplifying the operation process.
Smart Images

Figure CN2025120130_23042026_PF_FP_ABST
Abstract
Description
Live streaming processing methods, devices, equipment, media, and products
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411456390.3, filed on October 17, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to a live streaming processing method, apparatus, device, medium, and product. Background Technology
[0004] With the continuous development of internet technology, more and more users are interacting with information through live streaming, such as explaining related target objects in the live stream. As the interactive information in live streams becomes more diverse, users are paying increasing attention to the interactive experience.
[0005] In related technologies, for an object being explained during a live stream, the user needs to manually locate the object in the object display area and then click the corresponding "Start Explanation" control to mark it, allowing viewers to quickly locate the object. To cancel or change the object being explained, the user must first click the corresponding "End Explanation" control to deselect it. Because this is a manual process, it is cumbersome, and there's a possibility that the streamer might forget to do so or need to pause the live stream to click the "Start Explanation" or "End Explanation" controls. In such cases, the object being explained may not be correctly marked, making it difficult for viewers to quickly find it and negatively impacting the live stream's interactive experience. Summary of the Invention
[0006] This disclosure provides a live streaming processing method, apparatus, device, medium, and product.
[0007] In a first aspect, embodiments of this disclosure provide a live streaming processing method, the method comprising:
[0008] The live stream page of the live stream room is displayed; wherein, the live stream room is associated with multiple objects to be explained, which are displayed in a first manner; the live stream page includes a live video stream corresponding to the live stream room;
[0009] The object being explained is determined from among the multiple objects to be explained in the live video stream, and the object being explained is displayed on the live page in a second manner.
[0010] Secondly, embodiments of this disclosure also provide a live streaming processing apparatus, the apparatus comprising:
[0011] A live streaming page display module is used to display the live streaming page of a live streaming room; wherein, the live streaming room is associated with multiple objects to be explained, which are displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room;
[0012] The narration object display module is used to determine the object being narrated among multiple objects to be narrated based on the live video stream, add a preset identifier to the object being narrated, and display the object being narrated on the live page in a second manner.
[0013] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Storage device for storing one or more programs.
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the live streaming processing method as described in any of the embodiments of this disclosure.
[0017] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the live streaming processing method as described in any of the embodiments of this disclosure.
[0018] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the live streaming processing method as described in any of the embodiments of this disclosure. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0020] Figure 1A is a schematic flowchart of a live streaming processing method provided in an embodiment of this disclosure;
[0021] Figure 1B is a schematic diagram of an interface for presenting the object being explained on a live streaming page in a live streaming processing method provided in an embodiment of the present disclosure.
[0022] Figure 1C is a schematic diagram of an interface for presenting the object being explained on a live streaming page in another live streaming processing method provided in this embodiment of the present disclosure.
[0023] Figure 2 is a flowchart illustrating another live streaming processing method provided in this embodiment of the present disclosure;
[0024] Figure 3 is a flowchart illustrating another live streaming processing method provided in an embodiment of this disclosure;
[0025] Figure 4 is a flowchart illustrating another live streaming processing method provided in this embodiment of the present disclosure;
[0026] Figure 5 is a schematic diagram of a live streaming processing device provided in an embodiment of this disclosure; and
[0027] Figure 6 is a schematic diagram of the structure of an electronic device for implementing an embodiment of the present disclosure. Detailed Implementation
[0028] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0029] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0030] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0032] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0033] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0034] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0035] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0036] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0037] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0038] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0039] Figure 1A is a flowchart illustrating a live streaming processing method provided in an embodiment of this disclosure. This embodiment is applicable to scenarios where the subject being explained in a live stream is automatically identified and displayed. The method can be executed by a live streaming processing device, which can be implemented in software and / or hardware, optionally through an electronic device such as a mobile terminal, PC, or server. As shown in Figure 1A, the method of this embodiment specifically includes:
[0040] S110. Display the live streaming page of the live streaming room; wherein, the live streaming room is associated with multiple objects to be explained, which are displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room.
[0041] The live stream room can be understood as a resource-sharing space displaying a live video stream. The live stream room can be determined based on the streaming address of the live video stream. The live page can be understood as a page used to present information related to the live stream room. The object to be explained can be understood as an object associated with and displayed in the live stream room that can be explained within the live stream room. The object to be explained can also be understood as an object to be explained within the live stream room. For example, the object to be explained can include, but is not limited to, objects in the video frame of the live video stream, objects that interact with objects in the live video stream (e.g., special effects or interactive information), and objects associated with the live stream room through a preset association method (e.g., bound object links) (e.g., recommended items in the live stream room), etc., at least one of these objects can be associated with and displayed in the live stream room. The live video stream can be understood as media content to be displayed on the live page. For example, the live video stream can include live audio data and / or live image data. The live audio data is audio data streamed to the live stream room. The live audio data includes at least live voice data. More specifically, the live audio data can be the live voice data of the target user or audio data collected by a preset audio acquisition device. The live image data can be understood as the image data displayed on the live page, used to present the live streaming video on the live page.
[0042] In this embodiment of the disclosure, the live streaming page of the live streaming room may specifically include: at least one live streaming room entry point; and, in response to an entry trigger operation for the live streaming room entry point, displaying the live streaming page of the live streaming room corresponding to the triggered live streaming room entry point. The live streaming room entry point can be understood as a control for entering the live streaming room. For example, the live streaming room entry point may be a link to enter the live streaming room. The link to enter the live streaming room can be presented through images and / or text, or through a live streaming preview window. The live streaming preview window can synchronously play the live streaming content in the live streaming room. The live streaming room entry point can be presented in one or more forms such as a button, link, image, and floating window to guide entry into a specific live streaming room. The live streaming room entry point typically contains some basic information about the live streaming room, such as the name of the live streaming room, cover image, and tags, to provide an initial understanding of the live streaming content. The entry trigger operation can be understood as an operation that triggers the live streaming room entry point to enter the live streaming room. For example, the entry trigger operation may be an operation of clicking the live streaming room entry point.
[0043] Specifically, entry-triggered operations (such as entry-click operations) can be detected. Upon receiving an entry-triggered operation, the live stream entry point affected by the operation can be obtained. Then, based on information contained in the live stream entry point (such as the live stream identifier and / or Uniform Location Resource Identifier), the live stream information related to the live stream can be obtained. The obtained live stream information is then used to load and display the live stream page corresponding to the live stream. The live stream page may include one or more live stream-related components, such as the live video stream, chat room, object display list of live stream objects, and broadcaster information, to provide a rich live stream experience.
[0044] S120. Determine the object being explained among the multiple objects to be explained based on the live video stream, and display the object being explained on the live page in a second manner.
[0045] The object being explained can be understood as the object to be explained in the live video stream. It is understood that the object being explained, determined from the live video stream among multiple objects to be explained, is the object with the highest matching degree or strongest correlation to the live video stream among the multiple objects to be explained. Specifically, the correlation between the target live data in the live video stream and the data of multiple objects to be explained can be determined, and then...; the methods for determining the data correlation between different objects to be explained and the live video stream can be the same or different. In this embodiment, the methods for determining the data correlation between the object to be explained and the live video stream may include, but are not limited to: determining the consistency between the object attribute data of the object to be explained and the target live data extracted from the live video stream; and / or, determining the consistency between the object attribute data corresponding to the object to be explained and the preset matching data corresponding to the target live data, etc.
[0046] Taking the object to be explained as an item to be explained as an example, assuming the target live data in the live video stream is a cup of brand A, then we can determine whether the object attribute data of the object to be explained is consistent with the target live data extracted from the live video stream. That is, we need to find the cup of brand A in the object to be explained as the object being explained. Assuming the target live data in the live video stream is fabric, and fabric is not included in the object to be explained, we can determine whether the object attribute data corresponding to the object to be explained is consistent with the preset matching data corresponding to the target live data. In this case, the preset matching data corresponding to the target live data can be customized. If the object attribute data corresponding to the object to be explained includes customization, then the object to be explained can be determined as the object being explained.
[0047] As mentioned above, the live video stream may include live audio data and / or live image data. Optionally, determining the object being explained from among the multiple objects to be explained based on the live video stream may include: determining the object being explained from among the multiple objects to be explained based on the live audio data and / or live image data in the live video stream. In other words, the object to be explained that matches the live video stream can be determined from among the multiple objects to be explained based on the live audio data and / or live image data in the live video stream, and this object will be the object being explained.
[0048] In this embodiment of the disclosure, the second method, distinct from the first method, is a display method adopted after the object to be explained is determined to be the object being explained. Specifically, the display information corresponding to the second method differs from at least some of the information in the display information corresponding to the first method. For example, the display information includes at least one of the following: display style, display position, display size, display color, and display status (e.g., blinking status).
[0049] Optionally, displaying the object being explained in the live stream page in the second manner includes: displaying the object being explained in the live stream page at a first preset position in the form of a window. Using this technical solution, after opening the live stream page, the object being explained in the live video stream can be viewed intuitively without additional operation, improving the live stream interactive experience.
[0050] The first preset position can be a pre-set location on the live stream page that can be obscured. The target window for presenting the object being explained can be presented on top of the live stream page based on the first preset position. For example, the first preset position can be an edge position on the live stream page, such as the lower left corner, upper right corner, or lower right corner. Generally, the main content of the live stream page is often presented in the center area of the screen. By adopting this technical solution, the obstruction of the main content of the live stream page can be avoided as much as possible.
[0051] As shown in Figure 1B, the live streaming page 10 of the live streaming room can display the live streaming account identifier 101 corresponding to the live streaming room and the live streaming screen 102 of the live video stream. A target window 20 can also be included on top of the live streaming page 10, and the object information of the object being explained can be displayed in the target window 20.
[0052] Based on the above technical solution, in order to avoid obscuring the content of the live page for too long, after the object being explained is displayed in a window at a first preset position on the live page, if the display time of the object being explained reaches the preset display time, the display of the object being explained in the window can be canceled, or the display of the window can be canceled.
[0053] Optionally, displaying the object being explained in the live stream page in the second manner includes: if an object display panel is displayed on the live stream page, displaying the object being explained at a second preset position in the object display panel. The object display panel can be understood as a display panel presented at the top layer of the live stream page for displaying multiple objects to be explained. When multiple objects to be explained are associated with the live stream room and displayed in the first manner, an object display request is received through the live stream page, the object display panel is displayed on the live stream page, and multiple objects to be explained are displayed in the object display panel. For example, the multiple objects to be explained in the object display panel are displayed in the form of a list. In other words, the object display panel can display an object display list including multiple objects to be explained. The second preset position can be a pre-set position in the object display panel. For example, the second preset position can be the top position in the object display panel. More specifically, the second preset position can be the first or second position in the object display list, etc.
[0054] As an optional implementation of this disclosure, displaying the object being explained in the live stream page in a second manner includes: adding a preset identifier to the object being explained, displaying the object being explained in the live stream page in a second manner, and associating the preset identifier with the object being explained. The preset identifier can be a pre-set identifier to indicate that the object to be explained is the object being explained. The display form of the preset identifier can be various, such as being presented as preset characters and / or preset graphics. Using this technical solution, the preset identifier can more intuitively highlight that the object to be explained is the object being explained.
[0055] In this embodiment of the disclosure, associating the preset identifier with the object being explained can be achieved by displaying the preset identifier in a display area adjacent to the display area of the object being explained, or by displaying the preset identifier within the display area of the object being explained, etc. When the preset identifier is displayed in association with the object being explained, the relative display information between the preset identifier and the object remains unchanged. The display information of the preset identifier can change as the display information of the object being explained changes. For example, if the display position of the object being explained changes, the display position of the preset identifier also changes accordingly to maintain the relative display position between the preset identifier and the object being explained.
[0056] As shown in Figure 1C, the live streaming page 10 of the live streaming room can display the live streaming account identifier 101 corresponding to the live streaming room and the live streaming video stream 102. An object display panel 30 is displayed on top of the live streaming page 10. The object display panel 30 includes multiple object display areas. The object being explained can be displayed in the object display area 301 of the object display panel 30, and a preset "Understanding" label is added to the object being explained to intuitively indicate that the object display area 301 displays the object being explained.
[0057] As an optional technical solution of this disclosure embodiment, after determining the object being explained among the multiple objects to be explained based on the live video stream, the method further includes: if it is determined that the duration of the object being explained is greater than or equal to a first preset duration, displaying the object being explained in the first manner. The first preset duration may be a pre-set value used to determine whether the time spent by the object being explained exceeds an expected duration threshold. The specific value of the first preset duration can be set according to actual needs and is not specifically limited here. The first preset duration may be the same as or different from the second preset duration. Considering that the object being explained in the live stream may have changed within the first preset duration, if it is determined that the duration of the object being explained is greater than or equal to the first preset duration, for example, if the time since the beginning of identifying the object being explained exceeds 60 seconds, the identification result can be discarded without processing, and the original display method of the object being explained can be maintained, that is, the identified object being explained can continue to be displayed in the first manner. By adopting this technical solution, a high degree of matching between the identified object being explained and the live video stream can be guaranteed.
[0058] In this embodiment of the disclosure, the object being explained is displayed in a third manner when the conditions for canceling the explanation for the object being explained are met. The conditions for canceling the explanation can be understood as various triggering conditions (cancellation) for canceling the display of the object being explained in the second manner. For example, the conditions for canceling the explanation include, but are not limited to, at least one of the following: receiving a cancellation operation for the object being explained; canceling the association display of the object being explained with the live stream; receiving a new object being explained, etc.
[0059] In this embodiment, the third method differs from the second method. That is, at least some information in the display information of the third method is different from that of the second method. For example, the third method can be the same as the first method, that is, restoring the object being explained to be displayed in the first method. The third method can also be different from the first method, for example, moving the object being explained to a target position. In practical applications, when the object to be explained is set as the object being explained, the object being explained can be inserted above the previous object to be explained that was set as the object being explained. Continuing the previous example, assuming the object being explained is displayed at the first position in the object display list, and the first method corresponding to the object being explained (the display method before it was determined to be the object being explained) is displayed at the fifth position in the object display list, in response to the cancellation operation for the object being explained, the object being explained can be displayed at the fifth position in the object display list, or the object being explained can be moved to the second or last position in the object display list, etc. In this case, the moved object being explained is no longer the object being explained and is restored to the object to be explained.
[0060] As an optional technical solution in this embodiment, after displaying the object being explained in the live stream page in the second manner, it may further include: updating the object being explained to the set object being explained in response to a explanation setting operation for the object to be explained. The explanation setting operation can be understood as setting the object to be explained as the object being explained. When the object being explained is already displayed, in response to the explanation setting operation for the object to be explained, the object being explained will be replaced with the set object being explained; that is, the currently displayed object being explained will no longer be displayed, and the set object being explained will be displayed as the object being explained. This technical solution supports changes to the object being explained through interaction, supporting both automatic recognition and responsiveness to custom settings, achieving flexible display of the object being explained, enriching the interactive methods of the live stream, and improving the live stream interactive experience. This technical solution is particularly suitable for situations where the object being explained is not automatically determined or the determined object is inaccurate.
[0061] Based on the above technical solution, further, after updating the object being explained to the set object to be explained, the method further includes: in response to a cancellation operation on the object being explained, displaying the object being explained in a third manner, and returning to execute the operation of determining the object being explained from among multiple objects to be explained based on the live video stream. This technical solution combines automatic identification of the object being explained with manual setting of the object being explained, enriching the ways to set the object being explained in the live stream and improving the live streaming interactive experience.
[0062] In this embodiment of the disclosure, the cancellation of the explanation operation can be understood as canceling the operation of displaying the object to be explained as the object being explained. After receiving the cancellation of the explanation operation for the object being explained, the display of the object being explained in the second manner can be canceled, and the object being explained can be displayed in the third manner. At this time, it is also possible to return to the operation of determining the object being explained from among the multiple objects to be explained based on the live video stream, so as to continue to automatically identify the object being explained from among the objects to be explained.
[0063] It is understandable that, when the preset identifier is associated with the object being explained, if the conditions for canceling the explanation for the object being explained are met, the association between the preset identifier and the object being explained is canceled, and the object being explained is displayed in a third manner.
[0064] As an optional technical solution in this embodiment, after displaying the object being explained in the live stream page in the second manner, the method further includes: displaying the object being explained based on a preset display duration when it is displayed on the live stream page. The preset display duration can be understood as a pre-set duration for displaying the object being explained on the live stream page in a single instance. The advantage of this setting is that the presentation duration of the object being explained on the live stream page can be customized.
[0065] In one optional implementation, the object being explained is displayed based on a preset display duration. Specifically, this can involve determining the continuous display duration of the object being explained on the live streaming page, and canceling the display of the object being explained on the live streaming page when the continuous display duration reaches the preset display duration. In short, when the object being explained is displayed on the live streaming page, it can be removed from the live streaming page once the preset display duration has been reached.
[0066] Optionally, canceling the display of the object being explained on the live stream page includes: canceling the display window used to present the object being explained on the live stream page. For example, this could involve changing the display status attribute of the display window used to present the object being explained on the live stream page to a hidden attribute, or destroying the display window created on the live stream page to present the object being explained, etc.
[0067] It is understandable that, if the preset display duration exceeds the live stream duration, the object being explained will remain displayed on the live stream page. It should be noted that the object being explained here is not a specific object to be explained. If the display duration of the object being explained does not reach the preset display duration and the object being explained changes, the object being explained displayed on the live stream page will be updated accordingly.
[0068] The technical solution of this disclosure embodiment, by displaying the live broadcast page of the live broadcast room, can intuitively present the live broadcast content related to the live broadcast room. The live broadcast page includes a live video stream corresponding to the live broadcast room, which can improve the audiovisual effect of the live broadcast room. Moreover, multiple objects to be explained are associated with the live broadcast room in a first manner, which can further enrich the presentation content of the live broadcast room. Then, by determining the object being explained among the multiple objects to be explained based on the live video stream, the object being explained is displayed on the live broadcast page in a second manner. This can automatically identify the object being explained among the multiple objects to be explained in the live broadcast room, saving the time of manually searching and setting the object being explained. Furthermore, by adjusting the display mode of the object being explained, the object being explained can be displayed differently, so as to quickly locate the object being explained in the live broadcast room. This solves the technical problems of the relatively complicated operation of setting the object being explained and the low efficiency of locating the object being explained in related technologies, simplifies the interactive operation of the live broadcast room, optimizes the interactive mode of the live broadcast room, and improves the interactive experience of the live broadcast room.
[0069] Figure 2 is a flowchart illustrating another image processing method provided in this embodiment. Based on the above embodiments, this embodiment further refines the method of determining the object being explained using the live video stream. Optionally, determining the object being explained among multiple objects to be explained based on the live video stream includes: determining the explanation object data in the live video stream based on the live audio data in the live video stream, and determining the object being explained based on the explanation object data and object attribute data corresponding to the multiple objects to be explained. For detailed implementation, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. As shown in Figure 2, the method of this embodiment may specifically include:
[0070] S210. Display the live streaming page of the live streaming room; wherein, the live streaming room is associated with multiple objects to be explained, which are displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room.
[0071] S220. Determine the narration object data in the live video stream based on the live audio data in the live video stream, determine the object being narrated based on the narration object data and the object attribute data corresponding to multiple objects to be narrated, and display the object being narrated on the live page in a second manner.
[0072] In this embodiment of the disclosure, the data of the object to be explained can be understood as the data in the live audio data associated with the object to be explained. The live video stream displayed on the live page typically includes live audio data. The live audio data is generally closely related to the live content. Providing audio explanations of the object to be explained in the live room is also one of the more effective and convenient ways to explain an object. Therefore, in this embodiment of the disclosure, the live audio data in the live video stream can be analyzed to determine the explanation data in the live video stream that is related to the object to be explained.
[0073] As an optional implementation of this disclosure, determining the narration object data in the live video stream based on the live audio data in the live video stream includes: extracting audio keywords from the live audio data in the live video stream, and identifying the audio keywords as the narration object data in the live video stream. Specifically, audio keywords can be extracted from the live audio data in the live video stream based on object attribute data corresponding to multiple objects to be narrated. Using this technical solution, keywords related to the object to be narrated in the live audio data can be accurately extracted, and irrelevant information can be effectively filtered out, improving data processing efficiency.
[0074] As another optional implementation of this disclosure, determining the narration object data in the live video stream based on the live audio data in the live video stream includes: performing semantic recognition on the live audio data in the live video stream, and determining the narration object data in the live video stream based on the semantic recognition result. For example, the live audio data in the live video stream can be input into a semantic recognition model to obtain the narration object data in the live video stream. The semantic recognition model can be trained on a pre-established machine learning model based on sample audio data and sample narration data.
[0075] The object attribute data can be understood as data used to characterize the features of the object to be explained and to distinguish different objects. For example, the object attribute data may specifically include object description data and / or object image data. Unlike object image data, the object description data can be data presented in character form to describe the object to be explained. For example, the object description data may include, but is not limited to, at least one of object identification description data, object type description data, object content description data, object feature description data, and object appearance description data. Taking the object to be explained as an item to be explained as an example, the object identification description data can at least be used to describe the item's name and / or brand; the object type description data can be used to describe the category to which the item belongs; the object content description data can be used to describe the components and / or the number of components of the item; the object feature description data includes data used to describe at least one feature of the item, such as its material, use, origin, variety, style, size, and texture; and the object appearance description data includes at least one of information describing the item's shape, color, and taste.
[0076] When the object attribute data includes object description data, determining the object being explained based on the explanation object data and the object attribute data corresponding to the multiple objects to be explained may include: performing data matching between the explanation object data and the object description data corresponding to each of the objects to be explained, and determining the object being explained based on the data matching result corresponding to each object to be explained. In this case, the object being explained is the object to be explained with the highest matching degree among the multiple objects to be explained.
[0077] When the object attribute data includes object image data, determining the object being explained based on the explanation object data and the object attribute data corresponding to multiple objects to be explained may include: performing data matching between the explanation object data and the object image data corresponding to each object to be explained, and determining the object being explained based on the data matching result corresponding to each object to be explained. Specifically, for each object to be explained, the object image data of the object to be explained can be identified to obtain object description data corresponding to the object image data. Then, the explanation object data can be performed data matching between the explanation object data and the object description data corresponding to each object to be explained, and the object being explained can be determined based on the data matching result corresponding to each object to be explained. Alternatively, the explanation object data and the object image data of each object to be explained can be input into a correlation detection model to determine the correlation index between the explanation object data and each object image data, and then the object being explained can be determined based on the correlation index corresponding to each object to be explained.
[0078] When the object attribute data includes object description data and object image data, determining the object being explained based on the explanation object data and the object attribute data corresponding to multiple objects to be explained may include: for each object to be explained, determining a first correlation index between the explanation object data and the object description data corresponding to the object to be explained, determining a second correlation index between the explanation object data and the object image data corresponding to the object to be explained, determining a target correlation index based on the first correlation index and the second correlation index; and then determining the object being explained based on the target correlation index corresponding to each object to be explained.
[0079] The technical solution of this disclosure determines the narration object data in the live video stream based on the live audio data in the live video stream. It can effectively parse the data associated with the narrated object in the live video stream from the live audio data in the live video stream. Then, it determines the narrated object based on the narration object data and the object attribute data corresponding to multiple narrated objects. It can accurately determine the narrated object that matches the narration object data from multiple narrated objects, i.e., the narrated object, by using the object attribute data corresponding to multiple narrated objects. This ensures the matching degree between the live audio data and the narrated object and realizes the automatic and accurate identification of the narrated object.
[0080] Figure 3 is a flowchart illustrating another image processing method provided in this embodiment. Based on the above embodiments, this embodiment further refines the method of determining the object being explained through the live video stream. Optionally, determining the object being explained based on the explanation object data and object attribute data corresponding to multiple objects to be explained includes: acquiring target image data from the live image data in the live video stream; and, upon acquiring the target image data, determining the object being explained based on the explanation object data, the target image data, and object attribute data corresponding to multiple objects to be explained. Specific implementation details can be found in the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. As shown in Figure 3, the method of this embodiment may specifically include:
[0081] S310. Display the live streaming page of the live streaming room; wherein, the live streaming room is associated with multiple objects to be explained, which are displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room.
[0082] S320. Determine the narration object data in the live video stream based on the live audio data in the live video stream, and obtain the target image data in the live image data in the live video stream.
[0083] In this embodiment of the disclosure, the live video stream may include live image data, and the target image data can be understood as image data in the live image data that is associated with the object to be explained.
[0084] As an optional implementation of this disclosure, obtaining target image data from the live image data in the live video stream may include: determining image difference information between multiple frames of live image data in the live video stream, and determining target image data from the live image data based on the image difference information. This technical solution can quickly identify changes in the live image data, thereby capturing the object to be explained, and is particularly suitable for situations where the object to be explained changes during a live broadcast.
[0085] As another optional implementation of this disclosure, obtaining target image data from the live image data in the live video stream may include: obtaining target image data from the live image data in the live video stream based on the narration object data. Using this technical solution, target image data associated with the live audio data can be quickly and effectively determined based on the narration object data in the live audio data, thereby more accurately determining the object being narrated within the narration object.
[0086] Optionally, obtaining the target image data from the live image data in the live video stream based on the narration object data may include: inputting the narration object data and the live image data into an image segmentation model to obtain the target image data. The image segmentation model may be obtained by training a diffusion model using sample narration data and sample image data.
[0087] Optionally, obtaining target image data from the live image data in the live video stream based on the narration object data includes: performing semantic recognition on the live image data in the live video stream to obtain image semantic information; and obtaining target image data from the live image data in the live video stream based on the image semantic information and the narration object data. The image semantic information is used to indicate the live display object contained in the live image data and the object type information corresponding to the live display object.
[0088] As another optional implementation of this disclosure, obtaining target image data from the live image data in the live video stream may include: obtaining target image data from the live image data in the live video stream based on object attribute data corresponding to multiple objects to be explained. Using this technical solution, target image data associated with the objects to be explained in the live image data can be accurately determined based on object attribute data corresponding to multiple objects to be explained, thereby identifying the object being explained more quickly and accurately.
[0089] S330. Upon obtaining the target image data, determine the object being explained based on the explanation object data, the target image data, and the object attribute data corresponding to the multiple objects to be explained, and display the object being explained on the live streaming page in a second manner.
[0090] The object attribute data may include object description data and object image data. As described above, obtaining the target image data indicates that the live image data includes the object being explained. At this point, the object being explained, i.e., the object being taught, can be determined from among the multiple objects by combining the explanation object data, the target image data, and the object attribute data corresponding to the multiple objects being explained.
[0091] As an optional implementation of this disclosure, when the target image data is obtained, for each object to be explained, a description matching result is determined based on the object data and the corresponding object description data; an image matching result is determined based on the target image data and the object image data; and an object matching result is determined based on the description matching result and the image matching result. Then, the object being explained is determined based on the object matching results corresponding to multiple objects to be explained. The advantage of this setup is that it can combine the target image data and the object data to be explained to quickly and easily match the object being explained from multiple objects to be explained, improving the speed of determining the object being explained and thus quickly displaying the object being explained, enhancing the live streaming interactive experience.
[0092] As another optional implementation of this disclosure, when the target image data is obtained, the display object data of the explanation object data in the live video stream can be determined first based on the target image data, and object matching data can be determined based on the explanation object data and the target image data. Then, for each object to be explained, a first matching result is determined based on the object matching data and the object description data corresponding to the object to be explained, a second matching result is determined based on the target image data and the object image data corresponding to the object to be explained, and a target matching result corresponding to the object to be explained is determined based on the first matching result and the second matching result. Finally, the object being explained is determined based on the target matching results corresponding to multiple objects to be explained. The advantage of this setup is that it allows for a deep integration of the explanation object data and the target image data, enabling further analysis of the explanation object data determined from the live audio data using the target image data to obtain more accurate object matching data, thereby further improving the accuracy of the determined object being explained.
[0093] In practical applications, the live stream may or may not capture the object to be explained. That is, the live video stream's image data may or may not include the target image data. In other words, there may be cases where the target image data is not acquired. In this embodiment, if the target image data is not acquired, the object being explained can be determined based on the explanation object data and the object attribute data corresponding to multiple objects to be explained.
[0094] The technical solution of this disclosure, based on determining the data of the object to be explained in the live video stream by means of live audio data in the live video stream, further analyzes the live image data in the live video stream by acquiring target image data from the live image data in the live video stream to obtain target image data related to the object to be explained. This increases the analytical dimensions of the live video stream and provides richer data basis for determining the data of the object being explained. Furthermore, by acquiring the target image data, the object being explained is determined based on the data of the object to be explained, the target image data, and object attribute data corresponding to multiple objects to be explained. This allows for the determination of the object being explained by combining live audio data and live image data, achieving cross-validation of the analysis results of live audio data and live image data, effectively improving the accuracy of identifying the object being explained, and further enhancing the live interactive experience.
[0095] Figure 4 is a flowchart illustrating another image processing method provided in this embodiment. Based on the above embodiments, this embodiment further refines the method for automatically identifying the object being explained under what circumstances. Optionally, determining the object data to be explained in the live video stream based on the live audio data in the live video stream includes: determining the start / stop state of the object recognition mode, wherein the object recognition mode is at least used to automatically identify the object to be explained included in the live video stream, and the start / stop state includes an enabled state and a disabled state; when the object recognition mode is in the enabled state, determining the object data to be explained in the live video stream based on the live audio data in the live video stream. For detailed implementation, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. As shown in Figure 4, the method of this embodiment may specifically include:
[0096] S410. Display the live streaming page of the live streaming room; wherein, the live streaming room is associated with multiple objects to be explained, which are displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room.
[0097] S420. Determine the start / stop status of the object recognition mode, wherein the object recognition mode is at least used to automatically identify the objects to be explained included in the live video stream, and the start / stop status includes an enabled status and a disabled status.
[0098] In this embodiment of the disclosure, the object recognition mode can be understood as a pre-set feature that, when enabled, can automatically identify objects to be explained within the live video stream. Whether the object recognition mode is enabled can be customized. In other words, the start and stop status of the object recognition mode can be customized.
[0099] As an optional implementation of this disclosure, determining the start / stop state of the object recognition mode includes: obtaining the attribute value of the start / stop setting item of the object recognition mode in the target page, and determining the start / stop state of the object recognition mode based on the attribute value. Further, the start / stop state of the object recognition mode can be adjusted by setting the attribute value of the start / stop setting item of the object recognition mode in the target page. Specifically, determining the start / stop state of the object recognition mode may include: displaying the target page, wherein the target page includes the start / stop setting item of the object recognition mode; and, in response to a start / stop setting operation for the start / stop setting item, determining the start / stop state of the object recognition mode based on the start / stop setting operation.
[0100] The start / stop setting item can be understood as including information items with operable controls for setting whether the object recognition mode is enabled. In other words, the object recognition mode can be set to an enabled or disabled state by operating the start / stop setting item. Specifically, whether the object recognition mode is set to an enabled or disabled state can be determined by the attribute value corresponding to the start / stop setting item. The enabled state corresponds to a first attribute value, and the disabled state corresponds to a second attribute value. The start / stop setting operation of the start / stop setting item can be an operation for setting the attribute value corresponding to the start / stop setting item. Based on this, determining the start / stop state of the object recognition mode according to the start / stop setting operation can specifically include: determining the attribute value of the start / stop setting item according to the start / stop setting operation, and determining the start / stop state of the object recognition mode according to the attribute value.
[0101] In this embodiment, the method for setting the attribute value of the start / stop setting item is associated with the display style of the start / stop setting item. For example, the start / stop setting item can be a switch-type control including a set button, and the attribute value of the start / stop setting item can be determined by placing the set button at the position corresponding to the on state or the position corresponding to the off state. As another example, the start / stop setting item can be a mode switching control, and in response to a control trigger operation on the mode switching control, the start / stop state of the object recognition mode is switched.
[0102] S430. When the object recognition mode is in the enabled state, determine the narration object data in the live video stream based on the live audio data in the live video stream.
[0103] When the object recognition mode is in the enabled state, it means that it is necessary to automatically identify the object being explained in the live video stream. At this time, the object data to be explained in the live video stream can be determined based on the live audio data in the live video stream.
[0104] S440. Determine the object being explained based on the data of the object being explained and the object attribute data corresponding to the multiple objects to be explained, and display the object being explained on the live broadcast page in a second manner.
[0105] In some cases, the object being explained may not be identified within the second preset time period. In such cases, the object being explained in the live video stream often cannot be automatically displayed. As an optional implementation of this disclosure, if the object being explained is not identified within the second preset time period, the object recognition mode is switched from the enabled state to the disabled state, and / or, target prompt information is generated and displayed. When the object being explained is not identified, it may be due to a backlog of recognition requests or an execution error in the recognition of the object being explained. In this case, the above technical solution reduces the recognition pressure on the object being explained by promptly adjusting the object recognition mode to the disabled state; by generating and displaying target prompt information, the reason for the lack of feedback from the object being explained can be understood in a timely manner, so as to determine subsequent actions, such as continuing to wait for feedback from the object being explained or disabling the object recognition mode.
[0106] The second preset duration can be a pre-set maximum duration (i.e., an upper limit) for the object being explained. The specific value of the second preset duration can be set according to actual needs and is not specifically limited here. For example, it could be 30 seconds or 60 seconds. The target prompt information can be at least one of the following: the recognition of the object being explained has timed out, the automatic recognition of the object being explained has failed, the object recognition mode has been switched to a disabled state, or the object recognition mode can be switched to a disabled state.
[0107] As an optional implementation of this disclosure, if the object to be explained does not exist in the live broadcast room, the object recognition mode is switched from the enabled state to the disabled state. In this case, it is not necessary to identify the object being explained among the objects to be explained, so the object recognition mode can be disabled to save performance.
[0108] Optionally, switching the object recognition mode from the enabled state to the disabled state may specifically include: changing the attribute value of the start / stop setting item to the attribute value corresponding to the disabled state, thereby switching the object recognition mode from the enabled state to the disabled state.
[0109] The technical solution of this disclosure embodiment, based on the above-mentioned optional technical solutions, adds the determination of the start / stop state of the object recognition mode for automatically identifying the object to be explained included in the live video stream. By determining the start / stop state of the object recognition mode, it is possible to accurately determine whether the object to be explained included in the live video stream is automatically identified. Since the start / stop state includes an enabled state and a disabled state, it can support flexible setting of the start / stop state of the object recognition mode. Then, by determining the explanation object data in the live video stream based on the live audio data in the live video stream when the object recognition mode is in the enabled state, it is possible to automatically identify the object to be explained included in the live video stream and then analyze the live video stream, which can ensure the effective execution of the operation and reduce unnecessary performance consumption.
[0110] Figure 5 is a schematic diagram of a live streaming processing device provided in an embodiment of this disclosure. As shown in Figure 5, the device includes a live streaming page display module 510 and a narration object display module 520. The live streaming page display module 510 is used to display the live streaming page of a live streaming room; wherein the live streaming room is associated with multiple objects to be narrated, displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room; the narration object display module 520 is used to determine the object being narrated among the multiple objects to be narrated based on the live video stream, and to display the object being narrated in a second manner on the live streaming page.
[0111] The technical solution of this embodiment displays the live page of the live room through the live page display module 510, which can intuitively present the live content related to the live room. The live page includes a live video stream corresponding to the live room, which can improve the audiovisual effect of the live room. Moreover, multiple objects to be explained are associated with the live room in a first manner, which can further enrich the content presented in the live room. Then, the explanation object display module 520 determines the object being explained among the multiple objects to be explained according to the live video stream, and displays the object being explained in a second manner on the live page. This can automatically identify the object being explained among the multiple objects to be explained in the live room, saving the time of manually searching and setting the object being explained. Furthermore, by adjusting the display mode of the object being explained, the object being explained can be displayed differently, so as to quickly locate the object being explained in the live room. This solves the technical problems of the relatively complicated operation of setting the object being explained and the low efficiency of locating the object being explained in related technologies, simplifies the interactive operation of the live room, optimizes the interactive mode of the live room, and improves the interactive experience of the live room.
[0112] Based on any of the above optional technical solutions, optionally, the explanation object display module 520 includes an audio data analysis submodule and an explanation object determination submodule, wherein the audio data analysis submodule is used to determine the explanation object data in the live video stream based on the live audio data in the live video stream; the explanation object determination submodule is used to determine the object being explained based on the explanation object data and object attribute data corresponding to multiple objects to be explained, wherein the object attribute data includes object description data and / or object image data.
[0113] Based on any of the above-mentioned optional technical solutions, optionally, the narration object determination submodule includes: an image data analysis unit and an audio-visual combination recognition unit. The image data analysis unit is used to acquire target image data from the live image data in the live video stream; the audio-visual combination recognition unit is used, upon acquiring the target image data, to determine the object being narrated based on the narration object data, the target image data, and object attribute data corresponding to multiple objects to be narrated.
[0114] Optionally, based on any of the above-mentioned optional technical solutions, the audio-visual combination recognition unit includes a matching data determination subunit, a matching result determination subunit, and a subject-to-be-explained object determination subunit. Specifically, the matching data determination subunit is used to determine the display object data of the subject-to-be-explained data in the live video stream based on the target image data, and to determine object matching data based on the subject-to-be-explained data and the target image data; the matching result determination subunit is used, for each subject-to-be-explained object, to determine a first matching result based on the object matching data and the object description data corresponding to the subject-to-be-explained object, to determine a second matching result based on the target image data and the object image data corresponding to the subject-to-be-explained object, and to determine a target matching result corresponding to the subject-to-be-explained object based on the first matching result and the second matching result; the subject-to-be-explained object determination subunit is used to determine the subject being explained based on the target matching results corresponding to multiple subjects.
[0115] Optionally, based on any of the above-mentioned optional technical solutions, the audio data analysis submodule includes a mode start / stop determination unit and a narration data determination unit. The mode start / stop determination unit is used to determine the start / stop state of the object recognition mode, which is used at least to automatically identify objects to be narrated in the live video stream; the start / stop state includes an enabled state and a disabled state. The narration data determination unit is used to determine the narration object data in the live video stream based on the live audio data in the live video stream when the object recognition mode is in the enabled state.
[0116] Optionally, based on any of the above-mentioned optional technical solutions, the mode start / stop determination unit includes a mode setting page display unit and a mode start / stop status determination unit. The mode setting page display unit is used to display a target page, wherein the target page includes start / stop settings for the object recognition mode; the mode start / stop status determination unit is used to determine the start / stop status of the object recognition mode based on a start / stop setting operation performed on the start / stop settings.
[0117] Based on any of the above-mentioned optional technical solutions, the object display module 520 may optionally include a first object display submodule and / or a second object display submodule. The first object display submodule is used to display the object being explained at a first preset position on the live streaming page in the form of a window; the second object display submodule is used to display the object being explained at a second preset position on the object display panel if an object display panel is displayed on the live streaming page.
[0118] Optionally, based on any of the above-mentioned optional technical solutions, the live streaming processing device further includes: a module for updating the object being explained. The module is configured to, after the object being explained is displayed in the live streaming page in the second manner, update the object being explained to the set object being explained in response to a narration setting operation for the object to be explained.
[0119] Optionally, based on any of the above-mentioned optional technical solutions, the live streaming processing device further includes a narration cancellation display module. The narration cancellation display module is configured to, after updating the object being narrated to the set object to be narrated, in response to a cancellation operation on the object being narrated, display the object being narrated in a third manner, and return to perform the operation of determining the object being narrated from among the multiple objects to be narrated based on the live video stream.
[0120] Based on any of the above optional technical solutions, optionally, the object display module 520 is specifically used to add a preset identifier to the object being explained, display the object being explained in the live broadcast page in a second manner, and associate the preset identifier with the object being explained.
[0121] Optionally, based on any of the above-mentioned optional technical solutions, the live streaming processing device further includes a time-limited display unit. The time-limited display unit is configured to display the object being explained based on a preset display duration after the object being explained is displayed on the live streaming page in a second manner.
[0122] The live streaming processing apparatus provided in this disclosure can execute the live streaming processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the live streaming processing method.
[0123] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0124] Referring now to FIG6, a schematic diagram of the structure of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG6 is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present disclosure.
[0125] As shown in Figure 6, the electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0126] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 6 illustrates electronic device 600 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0127] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0128] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0129] The electronic device provided in this disclosure and the live streaming processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this disclosure can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0130] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the live streaming processing method provided in the above embodiments.
[0131] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0132] According to one or more embodiments of this disclosure, Example 1 provides a live streaming processing method, including: displaying a live streaming page of a live streaming room; wherein the live streaming room is associated with a plurality of objects to be explained displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room; determining the object being explained among the plurality of objects to be explained based on the live video stream, and displaying the object being explained in the live streaming page in a second manner.
[0133] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, which further includes: Optionally, determining the object being explained among a plurality of objects to be explained based on the live video stream includes: determining the explanation object data in the live video stream based on the live audio data in the live video stream, and determining the object being explained based on the explanation object data and object attribute data corresponding to the plurality of objects to be explained, wherein the object attribute data includes object description data and / or object image data.
[0134] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, which further includes: Optionally, determining the object being explained based on the explanation object data and the object attribute data corresponding to the plurality of objects to be explained includes: obtaining target image data in the live image data in the live video stream, and, when the target image data is obtained, determining the object being explained based on the explanation object data, the target image data and the object attribute data corresponding to the plurality of objects to be explained.
[0135] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 3, further comprising: optionally, determining the object being explained based on the explanation object data, the target image data, and object attribute data corresponding to multiple objects to be explained, comprising: determining display object data of the explanation object data in the live video stream based on the target image data, and determining object matching data based on the explanation object data and the target image data; for each object to be explained, determining a first matching result based on the object matching data and object description data corresponding to the object to be explained, determining a second matching result based on the target image data and object image data corresponding to the object to be explained, and determining a target matching result corresponding to the object to be explained based on the first matching result and the second matching result; and determining the object being explained based on the target matching results corresponding to multiple objects to be explained.
[0136] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 2, further comprising: optionally, determining the narration object data in the live video stream based on the live audio data in the live video stream includes: determining the start / stop state of an object recognition mode, wherein the object recognition mode is at least used to automatically identify the object to be narrated included in the live video stream, and the start / stop state includes an enabled state and a disabled state; when the object recognition mode is in the enabled state, determining the narration object data in the live video stream based on the live audio data in the live video stream.
[0137] According to one or more embodiments of this disclosure, Example Six provides the method of Example Five, further comprising: optionally, determining the start / stop state of the object recognition mode includes: displaying a target page, wherein the target page includes an object recognition mode start / stop setting item; and determining the start / stop state of the object recognition mode based on the start / stop setting operation in response to a start / stop setting operation for the start / stop setting item.
[0138] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, which further includes: optionally, displaying the object being explained in the live broadcast page in a second manner, including: displaying the object being explained in the live broadcast page at a first preset position in the form of a window; and / or, if an object display panel is displayed in the live broadcast page, displaying the object being explained at a second preset position in the object display panel.
[0139] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 1, which further includes: optionally, after displaying the object being explained in the live streaming page in a second manner, it further includes: updating the object being explained to the set object being explained in response to an explanation setting operation for the object to be explained.
[0140] According to one or more embodiments of this disclosure, Example Nine provides the method of Example Eight, which further includes: optionally, after updating the object being explained to the set object to be explained, it further includes: in response to a cancellation operation for the object being explained, displaying the object being explained in a third manner, and returning to perform the operation of determining the object being explained from a plurality of objects to be explained based on the live video stream.
[0141] According to one or more embodiments of this disclosure, Example 10 provides the method of Example 1, which further includes: optionally, displaying the object being explained in the live broadcast page in a second manner includes: adding a preset identifier to the object being explained, displaying the object being explained in the live broadcast page in a second manner, and displaying the preset identifier in association with the object being explained.
[0142] According to one or more embodiments of this disclosure, Example 11 provides the method of Example 1, which further includes: optionally, after displaying the object being explained in the live broadcast page in a second manner, the method further includes: displaying the object being explained based on a preset display duration when the object being explained is displayed in the live broadcast page.
[0143] According to one or more embodiments of this disclosure, Example Twelve provides a special effects processing device, including: a live streaming page display module for displaying a live streaming page of a live streaming room; wherein the live streaming room is associated with multiple objects to be explained displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room; and an object to be explained display module for determining the object being explained among the multiple objects to be explained based on the live video stream, adding a preset identifier to the object being explained, and displaying the object being explained in the live streaming page in a second manner.
[0144] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0145] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0146] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0147] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0149] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not necessarily limiting in certain circumstances; for example, an image data analysis unit can also be described as "for acquiring target image data from live image data in a live video stream".
[0150] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0151] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0153] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0154] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A live streaming processing method, comprising: The live stream page of the live stream room is displayed; wherein, the live stream room is associated with multiple objects to be explained, which are displayed in a first manner; the live stream page includes a live video stream corresponding to the live stream room; The object being explained is determined from among the multiple objects to be explained in the live video stream, and the object being explained is displayed on the live page in a second manner.
2. The live processing method of claim 1, wherein, The step of determining the object being explained from among multiple objects to be explained based on the live video stream includes: The narration object data in the live video stream is determined based on the live audio data in the live video stream. The object being narrated is determined based on the narration object data and the object attribute data corresponding to multiple objects to be narrated. The object attribute data includes object description data and / or object image data.
3. The live processing method of claim 2, wherein, The step of determining the object being explained based on the data of the object being explained and the object attribute data corresponding to multiple objects to be explained includes: Obtain target image data from the live image data in the live video stream. If the target image data is obtained, determine the object being explained based on the explanation object data, the target image data, and object attribute data corresponding to multiple objects to be explained.
4. The live processing method of claim 3, wherein, The step of determining the object being explained based on the data of the object being explained, the target image data, and object attribute data corresponding to multiple objects to be explained includes: Based on the target image data, determine the display object data of the narration object data in the live video stream, and determine the object matching data based on the narration object data and the target image data; For each of the objects to be explained, a first matching result is determined based on the object matching data and the object description data corresponding to the object to be explained; a second matching result is determined based on the target image data and the object image data corresponding to the object to be explained; and a target matching result corresponding to the object to be explained is determined based on the first matching result and the second matching result. The object being explained is determined based on the target matching results corresponding to multiple objects to be explained.
5. The live processing method of any of claims 2-4, wherein, The step of determining the narration object data in the live video stream based on the live audio data in the live video stream includes: Determine the start / stop status of the object recognition mode, wherein the object recognition mode is at least used to automatically identify the objects to be explained included in the live video stream, and the start / stop status includes an enabled state and a disabled state. When the object recognition mode is in the enabled state, the narration object data in the live video stream is determined based on the live audio data in the live video stream.
6. The live processing method of claim 5, wherein, The start / stop status of the determined object recognition mode includes: Display the target page, which includes a start / stop setting for the object recognition mode; In response to a start / stop setting operation for the start / stop setting item, the start / stop state of the object recognition mode is determined based on the start / stop setting operation.
7. The live processing method of any of claims 1-6, wherein, The step of displaying the object being explained on the live stream page in a second manner includes: The object being explained is displayed in a window at a first preset position on the live stream page; and / or, If an object display panel is displayed on the live streaming page, the object being explained will be displayed at a second preset position in the object display panel.
8. The live processing method of any of claims 1-7, wherein, After displaying the object being explained on the live stream page in the second manner, the method further includes: In response to the explanation setting operation for the object to be explained, the object being explained is updated to the set object to be explained.
9. The live processing method of claim 8, wherein, After updating the object being explained to the set object to be explained, the method further includes: In response to a cancel operation on the object being explained, the object being explained is displayed in a third manner, and the operation of determining the object being explained from among the plurality of objects to be explained based on the live video stream is returned.
10. The live processing method of any of claims 1-9, wherein, The step of displaying the object being explained in the live stream page in a second manner includes: Add a preset identifier to the object being explained, display the object being explained in the live broadcast page in a second manner, and associate the preset identifier with the object being explained.
11. The live processing method of any of claims 1-10, wherein, After displaying the object being explained in the live stream page in a second manner, the method further includes: When the object being explained is displayed on the live streaming page, the object being explained is displayed based on a preset display duration.
12. A live streaming processing device, comprising: The live streaming page display module is configured to display the live streaming page of a live streaming room; wherein, the live streaming room is associated with multiple objects to be explained, which are displayed in a first manner; the live streaming page includes a live video stream corresponding to the live streaming room; The narration object display module is configured to determine the object being narrated among a plurality of objects to be narrated based on the live video stream, and to display the object being narrated on the live page in a second manner.
13. An electronic device, comprising: One or more processors; A storage device is configured to store one or more programs, wherein, When the one or more programs are executed by the one or more processors, the one or more processors implement the live streaming processing method as described in any one of claims 1-11.
14. A storage medium containing computer-executable instructions, wherein, The computer-executable instructions, when executed by a computer processor, are used to perform the live streaming processing method as described in any one of claims 1-11.
15. A computer program product comprising a computer program, wherein, When the computer program is executed by the processor, it implements the live streaming processing method as described in any one of claims 1-11.
Citation Information
Patent Citations
Information display method and device, electronic equipment and storage medium
CN113935813A
Target recommendation object determination method and device, electronic equipment and storage medium
CN114399699A
Information display method and device, electronic equipment and storage medium
CN117093790A
Live broadcast interaction method and device, computer equipment and computer readable storage medium
CN118042236A
Live broadcast processing method and device, equipment, medium and product
CN119342272A