Method and apparatus for processing live video
Real-time object identification and annotation in live streaming videos address the limitation of viewer interaction by allowing dynamic display of associated information, enriching the viewing experience.
Patent Information
- Application Number
- CN202211448956.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-18
AI Technical Summary
In the live video, users can only see the information about one product explained by the current anchor, and cannot see the information about other products at the same time.
By obtaining the binding information between the object in the live video and the information to be marked, the position of the object is identified in real time and marked, and sending the information to be marked of the object for real-time display at the user terminal.
Real-time identification and labeling of objects in live videos is realized, enriching the functions of live videos, allowing users to see information about multiple products at the same time.
Smart Images

Figure CN115761583B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, specifically to the fields of video, live streaming, and artificial intelligence technologies including image processing and deep learning, and particularly to a method and apparatus for processing live video. Background Art
[0002] Live streaming may refer to video live streaming by a live streamer through a live streaming terminal. The live stream signal is often transmitted to the terminals of various users in real time. The live streamer sends the live video stream to the server, and after the server processes the live video stream, it sends the processing result to the terminals of various users. In this way, users can view the live video content in real time.
[0003] Since there are more and more companies and live streamers for live commerce, more and more users choose to purchase goods during live streaming. However, in the current live video, users can often only see the product information of one product being explained at present, such as the product link. Summary of the Invention
[0004] A method, apparatus, electronic device, and storage medium for processing live video are provided.
[0005] According to a first aspect, a method for processing live video is provided, including: obtaining binding information between at least one object to be recognized in the live video and information to be annotated, where the object includes an item or a body part, and the information to be annotated is used to annotate the position of the bound object; in the live video, performing real-time recognition on the object indicated by the binding information to obtain the object position information of the object, where the information to be annotated is displayed in real time at the annotation real-time display position corresponding to the position indicated by the object position information; based on the binding information, determining the information to be annotated and the object position information of each object in at least one object as the information to be sent; and sending the information to be sent of at least one object.
[0006] According to a second aspect, a method for processing live video is provided, the method including: determining binding information between at least one object to be recognized in the live video and information to be annotated, where the object includes an item or a body part, and the information to be annotated is used to annotate the position of the bound object, and is displayed in real time at the annotation real-time display position corresponding to the position of the object; sending the binding information; where the processing process of the binding information includes: in the live video, performing real-time recognition on the object indicated by the binding information to obtain the object position information of the object, where the information to be annotated is displayed in real time at the annotation real-time display position corresponding to the position indicated by the object position information; based on the binding information, determining the information to be annotated and the object position information of each object in at least one object as the information to be sent; and sending the information to be sent of at least one object.
[0007] According to a third aspect, a processing apparatus for live video is provided. The apparatus includes: an acquisition unit configured to acquire binding information between at least one object to be recognized in the live video and annotation information, where the object includes an item or a body part, and the annotation information is used to annotate the position of the bound object; a recognition unit configured to perform real-time recognition on the object indicated by the binding information in the live video to obtain the object position information of the object, where the annotation information is displayed in real time at the real-time display position corresponding to the position indicated by the object position information; a determination unit configured to determine, based on the binding information, the annotation information and the object position information of each object in at least one object as information to be sent; and a sending unit configured to send the information to be sent of at least one object.
[0008] According to a fourth aspect, a processing apparatus for live video for a live terminal is provided. The apparatus includes: a binding information determination unit configured to determine binding information between at least one object to be recognized in the live video and annotation information, where the object includes an item or a body part, the annotation information is used to annotate the position of the bound object, and is displayed in real time at the real-time display position corresponding to the position of the object; and an upload unit configured to send the binding information; where the processing process of the binding information includes: performing real-time recognition on the object indicated by the binding information in the live video to obtain the object position information of the object, where the annotation information is displayed in real time at the real-time display position corresponding to the position indicated by the object position information; determining, based on the binding information, the annotation information and the object position information of each object in at least one object as information to be sent; and sending the information to be sent of at least one object.
[0009] According to a fifth aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of the embodiments of the processing method for live video.
[0010] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method of any one of the embodiments of the processing method for live video.
[0011] According to a seventh aspect, a computer program product is provided, including a computer program, where the computer program, when executed by a processor, implements the method of any one of the embodiments of the processing method for live video.
[0012] According to the solution of the present disclosure, objects in a live video can be recognized in real time, which helps the device for watching the live video to display the annotation content of the objects in real time. Moreover, through the binding information, it is possible to select an object in the live video for annotation. In addition, the device that receives the information to be sent can display the information to be annotated of at least one object in the live video, enriching the functions of the live video. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Other features, objects, and advantages of the present disclosure will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0014] Figure 1 is an exemplary system architecture diagram to which some embodiments of the present disclosure can be applied;
[0015] Figure 2 is a flowchart of an embodiment of the method for processing a live video according to the present disclosure;
[0016] Figure 3 is a schematic diagram of an application scenario of the method for processing a live video according to the present disclosure;
[0017] Figure 4a is a flowchart of another embodiment of the method for processing a live video according to the present disclosure;
[0018] Figure 4b is a timing diagram of another embodiment of the method for processing a live video according to the present disclosure;
[0019] Figure 5 is a schematic structural diagram of an embodiment of the device for processing a live video according to the present disclosure;
[0020] Figure 6 is a block diagram of an electronic device for implementing the method for processing a live video of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Details of the embodiments of the present disclosure are included to help understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.
[0022] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0023] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other. The following will describe the present disclosure in detail with reference to the drawings and in conjunction with the embodiments.
[0024] Figure 1 An exemplary system architecture 100 is shown that can apply embodiments of the method for processing live video or the apparatus for processing live video according to the present disclosure.
[0025] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0026] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. For example, the terminal device 101 may be the terminal used by the live streamer for live streaming, that is, the live terminal, and the terminal devices 102, 103 may be the terminals used by different live stream viewers, that is, user terminals.
[0027] Various communication client applications may be installed on the terminal devices 101, 102, 103, such as live streaming applications, video applications, instant messaging tools, email clients, social platform software, etc.
[0028] The terminal devices 101, 102, 103 here may be hardware or software. When the terminal devices 101, 102, 103 are hardware, they may be various electronic devices with a display screen, including but not limited to smart phones, tablet computers, e-book readers, laptop computers, and desktop computers, etc. When the terminal devices 101, 102, 103 are software, they may be installed in the above-listed electronic devices. It may be implemented as multiple software or software modules (such as multiple software or software modules for providing distributed services), or may be implemented as a single software or software module. No specific limitation is made here.
[0029] The server 105 may be a server that provides various services, such as a background server that provides support for the terminal devices 101, 102, 103. The background server may analyze and process data such as live videos received, and feedback the processing results (such as the object position information and information to be annotated of the objects in the live video) to the terminal devices.
[0030] It should be noted that the method for processing live video provided in the embodiments of the present disclosure can be executed by the server 105 or the terminal devices 101, 102, 103. Correspondingly, the apparatus for processing live video can be disposed in the server 105 or the terminal devices 101, 102, 103.
[0031] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0032] Continuing to refer to Figure 2 , a flowchart 200 of an embodiment of the method for processing live video according to the present disclosure is shown. The method for processing live video can be used in an electronic device such as a server, and the method includes the following steps:
[0033] Step 201, obtain the binding information between at least one object to be recognized in the live video and the information to be annotated, where the object includes an item or a body part, and the information to be annotated is used to annotate the position of the bound object.
[0034] In this embodiment, the execution subject (such as Figure 1 the server shown, that is, the server side) on which the method for processing live video runs can obtain the binding information, and the binding information binds the object and the information to be annotated, that is, there is a corresponding relationship between the object and the information to be annotated. In practice, the object and the information to be annotated are in one-to-one correspondence, and each object has corresponding information to be annotated. The object here is the object to be recognized in the video frame of the live video, such as the upper body, hands, legs, feet, etc. of the body. Or the object can also be an item, such as a mask, a hat, a shoe, etc.
[0035] One or more objects can be presented in the live video, and the object indicated by the binding information can be all or part of the presented objects.
[0036] When a user terminal, a live terminal, or other terminals display the live video or the processed live video, the information to be annotated can be presented in the video in real time. For example, the object to be recognized corresponds to the same position as the object presented in the live video, but may be different. For example, the object to be annotated and recognized is the foot in the video frame, and the position of the foot in the video frame currently presented, that is, the recognition result, is a shoe. Or, the object indicated by the binding information can also be an item, that is, the shoe itself.
[0037] The information to be annotated is used to annotate the object indicated by the binding information, that is, the position of the object in the live video. The content to be annotated is the information to be annotated. For example, the information to be annotated may include the attribute value of the object. If the object is an item (such as a commodity), the attribute value of the object may be item information such as the introduction of the item or the link of the item.
[0038] Step 202, in the live video, perform real-time recognition on the object indicated by the binding information to obtain the object position information of the object. Among them, the information to be annotated is displayed in real time at the real-time display position corresponding to the position indicated by the object position information.
[0039] In this embodiment, the above-mentioned execution entity can perform real-time recognition on the object indicated by the binding information in the live video to obtain the position information of the object, that is, the object position information. In practice, the above-mentioned execution entity can perform real-time recognition on the current video frame in the live video. Then, the video frame finally annotated on the user terminal is the current video frame or the processed current video frame.
[0040] Step 203, based on the binding information, determine the information to be annotated and the object position information of each object in at least one object as the information to be sent.
[0041] In this embodiment, the binding information includes the correspondence relationship between the object and the information to be annotated, that is, the binding relationship. Therefore, the above-mentioned execution entity can determine the information to be annotated for each object in at least one object in the binding information. And, the above-mentioned execution entity needs to determine the information to be annotated and the position information of each object in at least one object as the information to be sent to the user terminal, that is, the information to be sent. The information to be sent includes the position information and the information to be annotated of at least one object, and there is an association relationship between the sent position information and the information to be sent.
[0042] Step 204, send the information to be sent of at least one object.
[0043] In this embodiment, the above-mentioned execution entity can send the information to be sent of at least one object to the user terminal. In this way, the user terminal can display the object and the information to be annotated marked thereon in the displayed live video. The information to be sent is real-time messaging (IM) information.
[0044] Specifically, on the display of the user terminal, the information to be annotated can be displayed in real time at the real-time display position corresponding to the position indicated by the object position information.
[0045] The method provided by the above embodiments of the present disclosure can perform real-time recognition on objects in a live video, thereby helping the device for watching the live video to display the annotation content of the objects in real time. Moreover, through the binding information, it is possible to select an object in the live video for annotation. In addition, the device that receives the information to be sent can display the information to be annotated of at least one object in the live video, enriching the functions of the live video.
[0046] In some alternative implementation manners of any embodiment of the present disclosure, the step of real-time display includes: generating real-time display position information corresponding to the position information of the object for the information to be annotated; sending the real-time display position information to the user terminal, so that the user terminal displays the information to be annotated at the real-time display position indicated by the real-time display position information in the live video, where the real-time display position information is a position adjacent to the position of the object.
[0047] In these implementation manners, the above execution subject can generate the real-time display position information of the information to be annotated and send it to the user terminal. In this way, the user terminal can display the information to be annotated in real time at the position indicated by the real-time display position information, that is, the position adjacent to the object. In addition, the above execution subject can also send the real-time display position information to the live terminal. In practice, this step of real-time display can be performed on terminal devices such as user terminals and live terminals.
[0048] Specifically, the adjacent position indicates that the distance between the object and the corresponding information to be annotated is relatively close. The real-time display position information corresponding to the position information of the object indicates the position where the information to be annotated is to be displayed. The indicated position can be various preset positions in the video frame, such as an adjacent position or the edge of the video frame. If the indicated position is the edge of the video frame, the indicated positions can be sorted in the order from top to bottom or from left to right corresponding to the objects.
[0049] For example, in the current video frame of the live video, the distance between the object and the corresponding information to be annotated can be less than or equal to a preset distance threshold. Or, the distance between the object and the corresponding information to be annotated can be equal to a preset distance, such as a preset proportion of the width of the video frame.
[0050] The devices in these implementation manners can generate the positional relationship between the object and the information to be annotated in the image, so as to control the annotation position of the information to be annotated, which helps to improve the annotation effect.
[0051] For the above-mentioned information to be annotated, generating real-time display position information corresponding to the object position information may include: generating initial information for the information to be annotated, which corresponds to the real-time display position information of the object position information, where there is a first position relationship between the initial information and the object position information; in response to the initial information indicating outside the edge of the video frame, generating real-time display position information indicating a position relationship other than the first position relationship, and the real-time display position information indicates inside the edge of the video frame.
[0052] In practice, the object may be at the edge of the video frame. According to the preset position relationship between the object and the annotation information, if the annotation information is outside the edge of the video frame, then it is necessary to re-determine a position of the annotation information that can be within the video frame range. The re-determined position of the annotation information, that is, the annotation real-time display position, is different from the position indicated by the initial information, and the position relationships between these two and the object are also different.
[0053] These implementation manners can adjust the position relationship between the display position of the annotation information and the object by re-determining the real-time display position information, ensuring that the annotation information can be successfully displayed.
[0054] In some optional implementation manners of any embodiment of the present disclosure, the obtained position information includes the object position information of each object in at least one object; based on the binding information, determining the information to be annotated and the object position information of each object in at least one object as the information to be sent, including: for each object in at least one object, extracting the information to be annotated of the object from the binding information; associating the object position information of the object with the information to be annotated, and using the association result as the information to be sent.
[0055] In these optional implementation manners, the above-mentioned execution subject may associate the object position information and the information to be annotated of each object, so as to obtain the associated object position information and the information to be annotated as the association result, so as to display the information to be annotated at the corresponding position of the object position, that is, the annotation real-time display position.
[0056] These implementation manners can simultaneously display the information to be annotated of each object (such as each of at least two objects) in the live video through information association, thereby overcoming the problem that the live broadcast can only display the information of the currently explained item.
[0057] In some optional implementation manners of any embodiment of the present disclosure, the method further includes: sending recognizable object information to the live broadcast terminal, so that the live broadcast terminal determines the candidate information to be annotated indicated by the selection operation as the information to be annotated for annotation in response to receiving a selection operation on the candidate information to be annotated after displaying the recognizable object information, where the object indicated by the recognizable object information includes the object corresponding to the candidate information to be annotated indicated by the selection operation.
[0058] In these implementation manners, a user of the user live terminal, such as a live streamer, can select the information to be annotated from each of the candidate information to be annotated. The user needs to make the selection with reference to the recognizable object information sent by the server. The recognizable object information refers to the information of the object that the server can recognize, indicating that the server has the recognition ability for the object. The user can select the information to be annotated of the object that the server can recognize with reference to this information.
[0059] The user in these implementation manners can select the information to be annotated of the object to be annotated, and can select the products that the server has the recognition ability through the recognizable objects of the server, so as to ensure the smooth implementation of annotating the object.
[0060] In some optional implementation manners of any embodiment of the present disclosure, the information to be annotated includes object profile information; the display step of the information to be annotated includes: for each object in at least one object, displaying the object profile information of the object at a position adjacent to the object; in response to receiving a selection operation on the object profile information, displaying the detailed object information of the object.
[0061] In these implementation manners, devices such as the user terminal can display the object profile information as the information to be annotated of the object. If a user watching the live stream selects this information (such as clicking), the detailed object information, such as the shopping page of the object, can be displayed. In some cases, the information to be annotated can also include information such as the identification graph (such as a trademark) of the object.
[0062] These implementation manners can further display the detailed information of the object after the profile is selected, so as to further meet the actual needs of the user for the object in the live stream.
[0063] Continue to refer to Figure 3 , Figure 3 is a schematic diagram of an application scenario of the live video processing method according to this embodiment. In Figure 3 the application scenario, the object in the binding information includes the upper body and the feet, and the corresponding information to be annotated is a certain brand of upper garment and a certain brand of shoes respectively.
[0064] Further refer to Figure 4a, which shows a flowchart 400 of another embodiment of the method for processing live video. This flowchart 400 can be used in electronic devices such as live terminals, and includes the following steps: Step 401, determining the binding information between at least one object to be recognized in the live video and the information to be marked, where the object includes an item or a body part, and the information to be marked is used to mark the position of the bound object, and to be displayed in real time at the marked real-time display position corresponding to the position of the object; Step 402, sending the binding information; where the processing process of the binding information includes: in the live video, recognizing the object indicated by the binding information in real time to obtain the object position information of the object, and the information to be marked is displayed in real time at the marked real-time display position corresponding to the position indicated by the object position information; based on the binding information, determining the information to be marked and the object position information of each object in at least one object as the information to be sent; sending the information to be sent of at least one object.
[0065] The method provided by the above embodiments of the present disclosure can recognize objects in the live video in real time, which helps the user terminal and other terminals to display the information to be marked in real time. Moreover, through the binding information, objects in the live video can be selected for marking. In addition, the terminal that receives the information to be sent can display the information to be marked of at least one object in the live video at the same time, enriching the functions of the live video.
[0066] In some optional implementation manners of this embodiment, determining the binding information between the objects presented in the live video and the information to be marked includes: obtaining the information to be marked of at least one object and displaying the information to be marked; in response to receiving a binding operation on at least some of the at least one object and the information to be marked, generating the binding information indicated by the binding operation. In these implementation manners, the binding operation is performed after receiving the recognizable object information sent by the server and displaying the recognizable object information. The user of the live terminal performs a one-to-one binding operation on the object and the information to be marked, so as to achieve the binding.
[0067] The above execution subject can obtain the information to be marked of at least one object in various ways. For example, the above execution subject can have a preset set of information to be marked and obtain the information to be marked from this set.
[0068] The binding operation can be various. For example, when the user performs a binding operation of selecting an object and the information to be marked, and the live terminal detects it, the binding information is generated. For example, the object selected by the user is the upper body, and the information to be marked selected is the item information of a coat, such as a product introduction or a link. For another example, the user can also select the legs and the item information of the pants for binding.
[0069] In addition, the binding operation can also be that the user inputs an object and the information to be marked, so as to generate the binding information.
[0070] These implementation methods enable live users, such as the host, to perform binding operations, thereby flexibly controlling the bound object.
[0071] In some alternative implementation methods of this embodiment, obtaining the information to be annotated for an object includes: in response to receiving a selection operation on candidate information to be annotated after presenting the recognizable object information, determining the candidate information to be annotated indicated by the selection operation as the information to be annotated for annotation, where the object indicated by the recognizable object information includes the object corresponding to the candidate information to be annotated indicated by the selection operation.
[0072] Users in these implementation methods can select the information to be annotated for the object to be annotated, and can select products that the server has the ability to recognize through the recognizable objects on the server, thereby ensuring the smooth implementation of annotation.
[0073] Further referring to Figure 4b , which shows the timing 400 of another embodiment of the method for processing live video. Among them, the server obtains a list of recognizable object information and then sends it to the live terminal. The live terminal presents the list of recognizable object information to the user, and then receives the user's selection operation on the candidate information to be annotated to obtain a list of information to be annotated. The live terminal can upload the list of information to be annotated to the server and, after receiving the feedback from the server, present the list of information to be annotated. The live terminal, in response to receiving a binding operation on the information to be annotated and the object, generates and sends binding information to the server. The server uses artificial intelligence to identify the object indicated by the binding information and issues instant messaging information, which carries the object position information and the information to be annotated of the object. The user terminal parses out the object position information and the information to be annotated and presents the information to be annotated at the real-time display position corresponding to the position indicated by the object position information. The user terminal, in response to receiving the user's click on the information to be annotated, displays the item details.
[0074] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for processing live video. This device embodiment corresponds to the method embodiment shown in Figure 2 . Except for the features described below, this device embodiment may also include the same or corresponding features or effects as the method embodiment shown in Figure 2 . This device can be specifically applied to various electronic devices.
[0075] Such as Figure 5As shown in the figure, the processing device 500 for live video in this embodiment includes: an acquisition unit 501, an identification unit 502, a determination unit 503, and a sending unit 504. Among them, the acquisition unit 501 is configured to acquire the binding information between at least one object to be identified in the live video and the information to be marked, where the object includes an item or a body part, and the information to be marked is used to mark the position of the bound object; the identification unit 502 is configured to perform real-time identification on the object indicated by the binding information in the live video to obtain the object position information of the object, where the information to be marked is displayed in real time at the real-time display position corresponding to the position indicated by the object position information; the determination unit 503 is configured to determine the information to be marked and the object position information of each object in at least one object as the information to be sent based on the binding information; the sending unit 504 is configured to send the information to be sent of at least one object.
[0076] In this embodiment, for the specific processing of the acquisition unit 501, the identification unit 502, the determination unit 503, and the sending unit 504 of the processing device 500 for live video and the technical effects brought by them, reference can be made to Figure 2 the relevant descriptions of steps 201, step 202, step 203, and step 204 in the corresponding embodiments, which will not be elaborated here.
[0077] In some optional implementation manners of this embodiment, the obtained position information includes the object position information of each object in at least one object; the determination unit is further configured to perform the following operations based on the binding information to determine the information to be marked and the object position information of each object in at least one object as the information to be sent: for each object in at least one object, extract the information to be marked of the object from the binding information; associate the object position information of the object with the information to be marked, and use the association result as the information to be sent.
[0078] In some optional implementation manners of this embodiment, the information to be marked includes object profile information; the display step of the information to be marked includes: for each object in at least one object, display the object profile information of the object at a position adjacent to the object; in response to receiving a selection operation on the object profile information, display the detailed object information of the object.
[0079] In some optional implementation manners of this embodiment, the display step of the information to be marked includes: a generation unit configured to generate real-time display position information corresponding to the object position information for the information to be marked; a distribution unit configured to display the information to be marked at the real-time display position indicated by the real-time display position information of the live video, where the real-time display position information is a position adjacent to the position of the object.
[0080] In some alternative implementation manners of this embodiment, for the information to be annotated, generating real-time display position information corresponding to the object position information includes: generating initial information of the real-time display position information corresponding to the object position information for the information to be annotated, where there is a first position relationship between the initial information and the object position information; in response to the initial information indicating outside the edge of the video frame, generating real-time display position information indicating a position relationship other than the first position relationship, and the real-time display position information indicates inside the edge of the video frame.
[0081] This application also provides a processing device for live video, which can be used in electronic devices such as live terminals. The device includes: a binding information determination unit configured to determine the binding information between at least one object to be recognized in the live video and the information to be annotated, where the object includes an item or a body part, and the information to be annotated is used to annotate the position of the bound object and to perform real-time display at the real-time display position corresponding to the position of the object; an upload unit configured to send the binding information; where the processing process of the binding information includes: in the live video, performing real-time recognition on the object indicated by the binding information to obtain the object position information of the object, and the information to be annotated is used to perform real-time display at the real-time display position corresponding to the position indicated by the object position information; based on the binding information, determining the information to be annotated and the object position information of each object in at least one object as the information to be sent; sending the information to be sent of at least one object.
[0082] In some alternative implementation manners of this embodiment, the binding information determination unit is further configured to determine the binding information between at least one object to be recognized in the live video and the information to be annotated in the following manner: obtaining the information to be annotated of at least one object and displaying the information to be annotated; in response to receiving a binding operation on at least some of the at least one object and the information to be annotated, generating the binding information indicated by the binding operation.
[0083] In some alternative implementation manners of this embodiment, the binding information determination unit is further configured to obtain the information to be annotated of at least one object in the following manner: in response to receiving a selection operation on candidate information to be annotated after displaying the recognizable object information, determining the candidate information to be annotated indicated by the selection operation as the information to be annotated for annotation, where the object indicated by the recognizable object information includes the object corresponding to the candidate information to be annotated indicated by the selection operation.
[0084] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0085] Figure 6FIG. 0 shows a schematic block diagram of an exemplary electronic device 600 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0086] As Figure 6 shown, the device 600 includes a computing unit 601 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0087] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0088] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method for processing live video. For example, in some embodiments, the method for processing live video can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method for processing live video described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method for processing live video by any other suitable means (e.g., by means of firmware).
[0089] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0090] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0091] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0092] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0093] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0094] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0095] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is imposed herein.
[0096] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for processing a live video, the method comprising: Obtaining binding information between at least one object to be recognized in the live video and information to be annotated, wherein the object includes an item or a body part, and the information to be annotated is used to annotate the position of the bound object; In the live video, performing real-time recognition on the object indicated by the binding information to obtain the object position information of the object, wherein the information to be annotated is displayed in real time at the real-time display position corresponding to the position indicated by the object position information, including: generating initial information of real-time display position information corresponding to the object position information for the information to be annotated, wherein there is a first position relationship between the initial information and the object position information; in response to the initial information indicating outside the edge of the video frame, generating real-time display position information indicated by a position relationship other than the first position relationship, and the real-time display position information indicates inside the edge of the video frame; displaying the information to be annotated at the real-time display position indicated by the real-time display position information in the live video; Based on the binding information, determining the information to be annotated and the object position information of each object in the at least one object as information to be sent; Sending the information to be sent of the at least one object.
2. The method according to claim 1, wherein, The obtained position information includes the object position information of each object in the at least one object; The determining, based on the binding information, the information to be annotated and the object position information of each object in the at least one object as information to be sent includes: For each object in the at least one object, extracting the information to be annotated of the object from the binding information; Associating the object position information of the object with the information to be annotated, and using the association result as the information to be sent.
3. The method according to claim 1 or 2, wherein, The information to be annotated includes object profile information; The step of displaying the information to be annotated includes: For each object in the at least one object, displaying the object profile information of the object at a position adjacent to the object; In response to receiving a selection operation on the object profile information, displaying the object detailed information of the object.
4. The method according to claim 1, wherein The real-time display position information is a position adjacent to the position of the object.
5. A method for processing a live video, the method comprising: Determining binding information between at least one object to be recognized in the live video and information to be annotated, wherein the object includes an item or a body part, the information to be annotated is used to annotate the position of the bound object, and is displayed in real time at the real-time display position corresponding to the position of the object; Sending the binding information; Among them, the processing process of the binding information includes: in the live video, performing real-time recognition on the object indicated by the binding information to obtain the object position information of the object. Among them, the real-time display of the to-be-annotated information at the corresponding annotation real-time display position indicated by the object position information includes: generating initial information of the real-time display position information corresponding to the object position information for the to-be-annotated information, where there is a first position relationship between the initial information and the object position information; in response to the initial information indicating outside the edge of the video frame, generating real-time display position information indicated by a position relationship other than the first position relationship, and the real-time display position information indicates inside the edge of the video frame; displaying the to-be-annotated information at the annotation real-time display position indicated by the real-time display position information in the live video; based on the binding information, determining the to-be-annotated information and the object position information of each object in the at least one object as the to-be-sent information; sending the to-be-sent information of the at least one object.
6. The method according to claim 5, wherein The determination of the binding information between at least one object to be recognized in the live video and the to-be-annotated information includes: Obtaining the to-be-annotated information of the at least one object and displaying the to-be-annotated information; In response to receiving a binding operation on at least some of the at least one object and the to-be-annotated information, generating the binding information indicated by the binding operation.
7. The method according to claim 6, wherein The obtaining of the to-be-annotated information of the at least one object includes: In response to receiving a selection operation on candidate to-be-annotated information after displaying recognizable object information, determining the candidate to-be-annotated information indicated by the selection operation as the to-be-annotated information to be annotated, where the object indicated by the recognizable object information includes the object corresponding to the candidate to-be-annotated information indicated by the selection operation.
8. A processing device for a live video, the device includes: An obtaining unit configured to obtain the binding information between at least one object to be recognized in the live video and the to-be-annotated information, where the object includes an item or a body part, and the to-be-annotated information is used to annotate the position of the bound object; A recognition unit configured to perform real-time recognition on the object indicated by the binding information in the live video to obtain the object position information of the object. Among them, the real-time display of the to-be-annotated information at the corresponding annotation real-time display position indicated by the object position information includes: a generating unit configured to generate initial information of the real-time display position information corresponding to the object position information for the to-be-annotated information, where there is a first position relationship between the initial information and the object position information; in response to the initial information indicating outside the edge of the video frame, generating real-time display position information indicated by a position relationship other than the first position relationship, and the real-time display position information indicates inside the edge of the video frame; a sending unit configured to display the to-be-annotated information at the annotation real-time display position indicated by the real-time display position information in the live video; A determination unit, configured to determine, based on the binding information, the information to be annotated and the object position information of each object in the at least one object as information to be sent; A sending unit, configured to send the information to be sent of the at least one object.
9. The device according to claim 8, wherein The obtained position information includes the object position information of each object in the at least one object; The determination unit is further configured to perform the determination, based on the binding information, of the information to be annotated and the object position information of each object in the at least one object as information to be sent in the following manner: For each object in the at least one object, extract the information to be annotated of the object from the binding information; Associate the object position information of the object with the information to be annotated, and use the association result as the information to be sent.
10. The device according to one of claims 8-9, wherein, The information to be annotated includes object profile information; The display step of the information to be annotated includes: For each object in the at least one object, display the object profile information of the object at a position adjacent to the object; In response to receiving a selection operation on the object profile information, display the detailed object information of the object.
11. The apparatus according to claim 8, wherein The real-time display position information is a position adjacent to the position of the object.
12. A processing device for a live video, the device includes: A binding information determination unit, configured to determine the binding information between at least one object to be recognized in a live video and the information to be annotated, where the object includes an item or a body part, and the information to be annotated is used to annotate the position of the bound object and to be displayed in real time at a real-time display position corresponding to the position of the object; An upload unit, configured to send the binding information; Wherein, the processing process of the binding information includes: in the live video, perform real-time recognition on the object indicated by the binding information to obtain the object position information of the object, where the real-time display of the information to be annotated at the real-time display position corresponding to the position indicated by the object position information includes: generate initial information of real-time display position information corresponding to the object position information for the information to be annotated, where there is a first position relationship between the initial information and the object position information; in response to the initial information indicating outside the edge of the video frame, generate real-time display position information indicated by a position relationship other than the first position relationship, and the real-time display position information indicates inside the edge of the video frame; display the information to be annotated at the real-time display position indicated by the real-time display position information in the live video; determine, based on the binding information, the information to be annotated and the object position information of each object in the at least one object as information to be sent; send the information to be sent of the at least one object.
13. The apparatus according to claim 12, wherein, The binding information determination unit is further configured to perform the determination of the binding information between at least one object to be recognized in a live video and the information to be annotated in the following manner: Obtain the information to be annotated of the at least one object and display the information to be annotated; In response to receiving a binding operation on at least some of the at least one object and the information to be annotated, generate the binding information indicated by the binding operation.
14. The apparatus according to claim 13, wherein, The binding information determining unit is further configured to obtain the information to be annotated for the at least one object in the following manner: In response to receiving a selection operation on candidate information to be annotated after presenting recognizable object information, determine the candidate information to be annotated indicated by the selection operation as the information to be annotated for annotation, where the object indicated by the recognizable object information includes the object corresponding to the candidate information to be annotated indicated by the selection operation.
15. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
17. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Live broadcast scene recognition method and device
CN110213610A
Live broadcast information pushing method, device and system, electronic equipment and computer medium
CN114928768A